Broadcom has introduced VMware Private AI Cloud, a platform for running AI inference and AI agents on infrastructure a company already owns. Instead of shipping sensitive data out to an external model provider, the design brings the models to wherever the data already sits. The announcement was made at VMware Explore, the company's annual event held in Las Vegas, and the list of validated models includes a Japanese-language model from NEC.
Move the Models, Not the Data
When enterprises try to put AI into production, the first obstacle is rarely performance. It is location. Whether customer records or design documents can be pushed straight into an external API is a question that stops the conversation outright in some industries. VMware Private AI Cloud reverses the direction and pulls the models into a closed internal environment.
Under the hood, this is a repackaging of existing products. VMware Cloud Foundation 9 provides the base, with VMware AI Factory as the AI-ready foundation, VMware Tanzu Platform running applications and agents, VMware vDefend handling network defense, and VMware Avi Load Balancer managing traffic. Inference workloads, agentic applications, containers, and conventional virtual machines all share one operating model. Ram Velaga, who leads Broadcom's Infrastructure Software Group, described the result as the point where private cloud and private AI infrastructure converge.
Broadcom's own research puts the share of enterprises already running production inference on private cloud, or planning to, at 56 percent. Read another way, the organizations that have moved past experimentation are the ones leaning toward keeping workloads in-house.
Putting Token Spend on the Operations Dashboard
Broadcom names three cost areas it is targeting: hardware capital expenditure, operational complexity, and token consumption that grows with every request. The third is a line item traditional virtualization platforms were never built to track, and it starts to bite the moment a pilot becomes production.
VMware Cloud Foundation 9 adds NVMe memory tiering, which treats fast NVMe storage as an extension of memory, along with cluster-wide storage deduplication. Both target efficiency before organizations resort to buying more capacity for models that consume large amounts of memory and storage. The platform also supports heterogeneous infrastructure spanning CPUs, GPUs, and other accelerators from multiple vendors, so hardware can be chosen on availability and cost rather than locked to a single configuration.
On the operations side, there is token monitoring for model usage, multi-tenant model sharing so several teams or applications can draw on the same model service, more detailed GPU and vGPU allocation tracking, and an observability dashboard that pulls those measurements together. Knowing who used which model and how much is what makes internal chargeback and usage caps practical.
For performance, Broadcom points to independent testing against MLPerf Inference v5.1 that showed VCF matching bare-metal infrastructure. The argument is that enterprises no longer need to stand up a separate bare-metal AI stack because virtualization gets in the way.
NEC's Japanese-Optimized cotomi Joins the Validated List
The clearest part of this announcement is that Broadcom named the models it has tested. More than 150 open-source and commercial models run on VMware Cloud Foundation, and models from Google, NVIDIA, NEC, Alibaba Cloud, and Z.ai are now formally validated.
- NVIDIA Nemotron 3: A family of open multimodal models built on a hybrid Mamba-Transformer Mixture-of-Experts architecture, supporting context windows of up to 1 million tokens and aimed at long-running agentic workflows
- Google DeepMind Gemma 4: An open-weight multimodal family, useful for organizations that want to build autonomous agents entirely inside their own infrastructure
- NEC cotomi: A model optimized for Japanese. Broadcom says it delivers a 40 percent improvement in token efficiency
- Alibaba Cloud Qwen 3.7-Max: A proprietary multimodal model with a 1-million-token context window, built for multimodal reasoning and agentic work
- Z.ai GLM 5.2: An open-source model aimed at local coding and reasoning agents
The model runtime is based on vLLM, giving supported models a common serving layer. That makes it realistic to hand models to internal users as a Model-as-a-Service offering. For organizations in Japan, having a Japanese-language model inside the validated set is a concrete decision point.
Agents Start With Nothing Allowed
Autonomous agents interpret instructions, query data on their own, invoke tools, and reach into connected systems. Unlike conventional applications, they have room to act beyond their intended scope, and a mistake in permission design can lead to both data exposure and runaway spending.
Tanzu Platform answers this with a deny-by-default runtime. APIs, networks, MCP servers, and the internet are all invisible to an agent unless explicitly granted. Credentials live in an isolated store and are never handed to the agent itself, which closes off the path where a prompt injection attack coaxes secrets out of the model.
Sitting above that is AgentMinder, a newly announced control plane. It treats each agent as an enterprise identity, in the same way an employee has one, and binds that identity to a mission, an approved set of tools, and authorized resources. Tool invocation is governed at runtime under least-privilege rules, and every action leaves an audit record. Once agents start multiplying across departments, that kind of registry is what keeps the situation legible.
On the network side, vDefend handles microsegmentation and virtual patching, and can spot unsanctioned AI use by watching traffic flows. Avi Load Balancer adds web application firewall and API protection, along with detection for agents reaching toward tools they should not touch. TrueSource, which governs open-source supply chains, has been extended beyond the Java, Python, and Node.js ecosystems to cover data infrastructure components including PostgreSQL, RabbitMQ, MySQL, and Valkey.
Availability
AgentMinder is available now. The Tanzu Platform capabilities announced at VMware Explore are scheduled for general availability in fall 2026.
Summary
VMware Private AI Cloud is not an announcement about a new model. It is about rebuilding the foundation so existing models can run internally at a cost that can actually be calculated. Named validated models, visible token consumption, and deny-by-default agents all point at operations rather than experiments, which makes this easiest to evaluate for enterprises that already run their own infrastructure. How well it works in practice will depend on the deployment reports that follow general availability this fall.
