On September 15, at the AI Infra Summit in Santa Clara, California, NVIDIA published new efficiency results and partner collaborations for AI factories[1]. Its latest platform, Vera Rubin NVL72, recorded up to 30 times more throughput per megawatt than the previous generation on a benchmark built from real agentic work. The yardstick for AI infrastructure is shifting away from peak compute and toward how many tokens a system can produce from each megawatt of power.
From Peak Performance to Tokens per Megawatt
The speaker was Ian Buck, vice president of hyperscale and high-performance computing at NVIDIA. This year's AI Infra Summit drew more than 8,000 attendees, up from 3,500 a year earlier[1]. The question at the center of the event was how to keep AI infrastructure running inside a hard power budget.
The driver is the spread of agentic AI. Agents chain reasoning steps together and call external tools repeatedly as they assemble an answer. A single session can accumulate hundreds of thousands of input tokens, roughly 15 times the token volume of a simple chat request, according to NVIDIA[1]. More processing translates directly into more power.
NVIDIA's response is to codesign everything from silicon to the grid. Vera Rubin systems, the Dynamo inference software, NeMo libraries, and the networking and storage layers, including NVLink, Spectrum-X Ethernet, ConnectX SuperNICs and BlueField DPUs, are treated as a single AI factory[1].
Lambda Runs 19 Nodes on the Power Budget of 16
NVIDIA DSX MaxLPS is the software that dynamically optimizes power allocation. It continuously monitors consumption across GPUs and racks, shifting available power to where it is needed most and reclaiming capacity that static provisioning leaves idle[1][2]. Because training and inference draw power in different patterns, mixed-workload AI factories see the largest benefit.
AI cloud provider Lambda presented validation results from DSX MaxLPS running on Blackwell-generation servers. Within the power envelope normally allocated to 16 nodes, Lambda ran 19, raising cluster-wide token throughput by 24 percent, from about 4 million to 5 million tokens per second, and improving performance per watt by 23 percent[1]. For next-generation Vera Rubin NVL72 AI factories, NVIDIA expects MaxLPS to enable up to 40 percent more GPU capacity within the same megawatt budget in suitable deployment environments[1][3].
At the rack level, Intelligent Power Smoothing software and expanded energy buffering absorb short power spikes. With less headroom reserved for momentary peaks, systems can run closer to sustained demand[1].
30x the Previous Generation on Real Agentic Workloads
The most striking result came from SemiAnalysis AgentX. AgentX records actual agentic coding sessions and replays them with context growth, tool-call delays and sub-agent spawning preserved, measuring inference the way agents actually behave. It captures the shape of agentic load that single-request benchmarks tend to miss[1].
On AgentX, running the DeepSeek V4 Pro model, Vera Rubin NVL72 delivered up to 30 times higher throughput per megawatt than GB300 NVL72, with cost per million tokens up to 45 times lower[1]. In power-constrained deployments, throughput per megawatt sets the revenue ceiling of an AI factory and cost per million tokens sets the margin on it. NVIDIA attributes the gain to extreme codesign: the NVL72 scale-up domain, sixth-generation NVLink, NVFP4 precision on fifth-generation Tensor Cores, and an inference stack spanning TensorRT LLM and Dynamo[1].
For low-latency inference, NVIDIA Groq 3 LPX complements Vera Rubin. On models above 2 trillion parameters running at long context, the combined platform delivers up to 35 times higher token throughput per megawatt than GB200 NVL72, and on a 100K-context Qwen 3.8 27B workload Groq 3 LPX reached 2,529 output tokens per second per user[1].
AI Factories That Work With the Grid
A second announcement turns AI factories into a flexible grid resource. Emerald AI worked with NVIDIA inside the flexible-load interconnection program run by utility Silicon Valley Power to demonstrate automated load reduction. The system responded to hundreds of demand signals while protecting AI workload performance[1].
DSX Flex makes this possible. It receives load-shedding requests, demand-response events and pricing signals from the grid, then acts automatically within a predefined workload hierarchy. The most critical jobs keep running while lower-priority work pauses temporarily and resumes afterward[1]. Data centers move from being purely consumers of electricity to resources that can modulate their draw according to grid conditions.
NVIDIA also disclosed that Amazon's Annapurna Labs is working with it on NVHBM custom high-bandwidth memory, that d-Matrix is integrating NVLink Fusion with its Raptor XPUs, and that Pinterest is using Blackwell and Dynamo to bring conversational AI to visual discovery[1].
Vera CPU Results and NVLink 6 Resiliency
The CPU side produced numbers as well. Perplexity benchmarked the Vera CPU for SPACE, its secure sandbox platform for agentic AI, and measured 1.9 times faster sandbox starts. Redpanda reported 5.5 times lower latency and 73 percent higher throughput than other CPUs, Starburst cited 3 times faster query throughput, and Kinetica measured 2.7 times faster analytical queries. ClickHouse said Vera was the fastest machine it had measured on ClickBench[1].
The larger a deployment grows, the more staying up becomes a performance question. Across hundreds of thousands of GPUs, transient errors, signal degradation and hardware failures are unavoidable. NVLink 6 answers with a multilayer resiliency architecture: custom forward error correction, physical layer retry and link recovery at the physical layer, plus credit-based flow control, dynamic routing and link rebalancing at the network layer to contain faults locally[1]. Keeping one failure from stalling the whole fabric matters as much to effective AI factory output as power efficiency does.
Summary
What the AI Infra Summit showed is that the measure of AI infrastructure is moving from how fast it runs to how much work it delivers per megawatt. Vera Rubin NVL72 combined with DSX MaxLPS fits more GPUs into the same power envelope and recorded up to 30 times more throughput per megawatt than the previous generation on real agentic work. Together with the DSX Flex grid demonstration, AI factories are being redefined around the question of how to spend every watt.
Source[2]: https://developer.nvidia.com/blog/maximizing-ai-factory-performance-per-watt-with-nvidia-dsx-maxlps
Source[3]: https://blogs.nvidia.com/blog/vera-rubin-nvl72-efficiency-ai-agents/
