NVIDIA has laid out in detail its concept of the AI factory, a new class of infrastructure built to produce intelligence[1]. Just as power plants once generated electricity, an AI factory is a facility that converts energy into tokens to manufacture intelligence around the clock[1]. The company argues that performance is measured by tokens per watt and cost per token, and that this efficiency translates directly into revenue[1].
What Is an AI Factory
NVIDIA says AI is no longer simply software but has become essential infrastructure that underpins society[1]. At the center of this shift is the AI factory, a large-scale computing foundation dedicated to producing intelligence[1]. In the same way that power plants converted energy into electricity during the industrial age, NVIDIA frames AI factories as facilities that convert energy into tokens, the unit of production for reasoning models and agents[1].
The metrics that matter here are tokens per second, tokens per watt, cost per token, utilization, and uptime[1]. NVIDIA states that performance per watt feeds directly into revenue and that cost per token determines the economics of every AI factory[1]. Each facility is optimized as a whole, spanning models, compute, networking, memory, software, storage, power, and cooling, to keep intelligence in continuous output[1].
How Agentic AI Changes the Workload
An AI factory is built for always-on inference that does far more than answer a single prompt[1]. Agentic AI, which autonomously plans and carries out tasks, handles not only reasoning and planning but also search, tool use, data retrieval, code writing, and real actions[1]. These systems even spin up their own sub-agents and learn how to use domain-specific tools, making workloads longer, deeper, and far more compute-intensive[1].
Supporting this requires fast memory, storage to hold context, networking for coordination, software for orchestration, and CPUs for execution, all paired with accelerated compute[1]. NVIDIA says keeping the entire workflow moving smoothly is what determines performance, describing inference as a real-time orchestration challenge that now spans the full machine[1].
Competitiveness Is Measured Per Watt
NVIDIA argues that performance per watt has become the ultimate yardstick for an AI factory's competitiveness[1]. Citing the InferenceX benchmarks from research firm SemiAnalysis, it says its Blackwell Ultra generation GPU delivered the lowest cost per token[1].
The company also shared concrete figures. According to NVIDIA, its NVIDIA GB300 NVL72 systems, built on Blackwell Ultra, generate up to 50 times more tokens per megawatt than the prior generation and can cut cost per token to one-thirty-fifth that of the earlier Hopper generation[1]. A software framework called NVIDIA Dynamo orchestrates long-context reasoning and high-volume inference, keeping utilization high[1]. With the next-generation NVIDIA Vera Rubin platform, NVIDIA says this trend goes further, raising performance per watt up to 35 times with LPX and driving token cost lower through deeper full-stack optimization[1].
From Chips to Full-Stack Design, Backed by Partners and Digital Twins
NVIDIA says what began with GPUs has expanded into full-stack AI factories that include accelerated compute, high-speed interconnects, liquid-cooled systems, inference software, autonomous agents, and reference architectures[1]. To bring these to enterprise data centers, it works closely with system partners such as Cisco, Dell, HPE, Lenovo, and Supermicro[1]. NVIDIA notes that it runs its own enterprise AI factory, where hundreds of autonomous agents assist its engineering and operations teams, offering it as a practical proof point[1].
Building large facilities, the company says, takes more than optimized compute[1]. The NVIDIA DSX reference designs aim to build gigawatt-scale AI factories at the lowest token cost per megawatt[1]. The NVIDIA Omniverse DSX Blueprint uses digital twins, virtual replicas of real facilities, that connect facilities, hardware, and software to support everything from pre-build design validation to post-deployment operational improvement[1]. NVIDIA sums it up this way: the last industrial revolution turned energy into work, while this one turns energy into intelligence[1].
The company adds that it will discuss the full picture of these AI factories in a keynote by founder and CEO Jensen Huang at NVIDIA GTC Taipei at COMPUTEX, scheduled for June 1 at 11 a.m. Taipei time[2].
Summary
NVIDIA's message is a reframing of the data center, from a place that stores files into a factory that produces tokens. Measuring competitiveness by tokens per watt and cost per token, through whole-system optimization that includes power and cooling rather than chips alone, looks aimed squarely at the surging compute demand driven by generative AI. How much concrete detail emerges at the June 1 keynote is the next thing to watch.
Source: https://blogs.nvidia.com/blog/ai-factories-the-new-infrastructure-of-intelligence/
Source: https://blogs.nvidia.com/blog/nvidia-gtc-taipei-computex-2026-news/
