NVIDIA argued on its official blog that in an era of tightening power constraints, performance per watt is the single most important metric for the profitability of an AI factory[1]. The idea is that how many tokens a system can generate within a fixed power budget determines its revenue and profit. Across the newest generation of leading open models, NVIDIA says the Blackwell-generation GB300 NVL72 delivered up to 25x the performance per watt of the previous Hopper generation. In production, frontier AI labs such as Anthropic and OpenAI already run inference on Blackwell.

Why Performance per Watt Is the Ultimate Metric

Power is an inescapable constraint for AI infrastructure. NVIDIA explains that the number of tokens an AI factory can generate within a fixed power budget determines its revenue and profit margin. Token price or a chip's theoretical peak alone cannot measure real efficiency, and because performance per watt can only be earned through real-world results rather than gamed, it becomes the foundation for AI factories[1].

As autonomous agentic AI drives token demand even higher, the infrastructure decisions organizations make today will separate those who can scale from those who cannot in a power-constrained world. Virtually every frontier model today uses a mixture-of-experts (MoE) architecture, and serving that efficiently at rack scale requires codesigning every layer of the stack, from silicon to software, NVIDIA stresses[1].

Blackwell's Up-to-25x Efficiency Gain

According to NVIDIA, across the newest generation of leading open models, the GB300 NVL72 delivers up to 25x the performance per watt of the Hopper generation. The company notes this figure reflects where Blackwell stands today, a starting point that continues to improve[1].

Because workloads optimized for latency and those optimized for throughput and cost demand different operating points, NVIDIA presents performance as a Pareto curve for each model rather than a single number. It also offers tools such as DynoSim that let teams find their optimal point before spending a single GPU-hour on validation[1].

This efficiency is the result of extreme codesign across silicon and software. NVLink Switch, which is critical to rack-scale performance, is now in its sixth generation with the Vera Rubin platform, and SHARP performs in-network computing inside the switch to offload work from the GPUs. On the software side, NVIDIA Dynamo, TensorRT LLM, SGLang and vLLM stack optimizations such as NVFP4 quantization, disaggregated serving and large-scale expert parallelism, and software gains keep compounding. On the DeepSeek V4 model, performance per watt improved by up to 5x in a single month[1].

The "Only 60% Usable" Problem and DSX MaxLPS

NVIDIA points out that in AI factories, power lost to cooling and rack-level inefficiencies can mean only about 60 percent of the electricity pulled from the grid turns into useful AI work[1].

Closing that gap is DSX MaxLPS, the power-and-efficiency software in the DSX platform. By shifting power between GPUs and racks in real time and combining warm-water liquid cooling with techniques like power steering, it lets operators run up to 40 percent more GPUs within the same power budget[1].

Production Track Record and Deployments

NVIDIA emphasizes that rack-scale reliability at AI-factory scale is hard-won. Rack-scale systems introduce failure modes single-node deployments never encounter, so time in production becomes an advantage in itself. Its Blackwell NVL72 continues to deliver sustained performance and rack-level reliability across a diverse range of models and use cases, which is why frontier labs such as Anthropic and OpenAI use it for inference[1].

Concrete deployments are expanding. CoreWeave runs the Kimi K2.6 model on GB300 NVL72, combining NVFP4 quantization with EAGLE3 speculative decoding to boost inference performance. Perplexity runs Qwen3 235B and a post-trained Qwen3.5-397B-A17B on GB200 NVL72, serving millions of queries a day. Fireworks AI deploys the GLM 5.2 model on Blackwell, supporting the production environments of customers including Cursor and Factory AI. This accumulated production experience gives the next-generation NVIDIA Vera Rubin a head start[1].

Summary

What NVIDIA is proposing is a shift in how AI infrastructure is judged, moving from raw performance to performance per watt. The Blackwell-generation GB300 NVL72 shows up to 25x the performance per watt of Hopper, and DSX MaxLPS lets operators run up to 40 percent more GPUs within the same power budget. With adoption spreading among Anthropic, OpenAI, CoreWeave, Perplexity and others in production, in an age where power is the biggest constraint, infrastructure efficiency looks set to shape the competitiveness of AI services themselves.

Source: https://blogs.nvidia.com/blog/performance-per-watt-ai-infrastructure-efficiency/