As AI moves from pilots to production AI factories, the yardstick for performance is shifting from a chip's peak specifications to the cost per token. NVIDIA says its full-stack inference software, codesigned with its GPUs, CPUs and networking, has already cut token costs by up to 5x on the DeepSeek V4 model on the Blackwell platform in just one month[1]. Here is why software, not only hardware, increasingly decides the economics of inference.
Cost Per Token Becomes the New Metric
The traditional metric was the peak performance of an individual chip. NVIDIA argues that in production, what matters is how many useful tokens a system can deliver per dollar, per watt and within the required latency targets[1].
Behind this shift is the rise of agentic AI. Conventional web, search and SaaS workloads were predictable: a request might load a page or read from and write to a database, following similar software paths, and they scaled by adding more of the same servers. Agentic AI is different. It reasons, plans, calls tools, spins up specialist subagents and manages massive context across multi-turn workflows[1].
As a result, a single request turns into a distributed computing problem that can span hundreds of subagents, thousands of tasks and multiple large language models, running across GPUs, CPUs, DPUs and storage. NVIDIA frames software as the factor that decides whether that complexity becomes wasted capacity or a lower cost per token[1].
Three Layers That Lift Performance
NVIDIA's inference software is described as connecting three layers so that individual optimizations turn into system-level performance[1].
The first is production operation, which coordinates distributed serving, orchestration, autoscaling and memory management so inference runs across the right compute and storage. The second is application acceleration, which runs models with high performance using runtime optimizations such as overlapping compute and communication and kernel fusion, while still leaving developers room to tune and customize. The third is infrastructure access, which exposes GPU, networking, memory and system capabilities without requiring developers to manage every device instruction set or data-transfer protocol directly[1].
When these layers work as one system, individual optimizations compound. According to NVIDIA, disaggregated serving, large expert parallelism over NVLink, NVFP4 precision and multi-token prediction each deliver gains on their own, but combined they increase throughput by up to 20x[1]. This layering of software is what underpinned the result of cutting DeepSeek V4 token costs by up to 5x on Blackwell within a month[1].
Adopters and the Open Source Flywheel
Real deployments are already underway. Baseten used the open source NVIDIA TensorRT-LLM library to serve DeepSeek V4 Pro on Blackwell for reasoning, coding and long-context workloads, applying proprietary runtime optimizations to deliver up to 50% more tokens per second. Cognition uses the NVIDIA Dynamo inference framework to manage inference GPUs and gain a path to scale reinforcement learning workloads. Deep Infra serves frontier open source models on Blackwell from day zero, and Together AI used TensorRT-LLM to help Cursor move quickly from model optimizations to production for its real-time coding experience[1].
This foundation is amplified by the open source ecosystem. Many of today's most widely used AI frameworks and inference projects are built on NVIDIA CUDA, so new research and optimizations run with leading performance on NVIDIA GPUs from day zero[1]. PyTorch is a prime example: launched in 2016 with native CUDA support, it has coevolved with NVIDIA's architecture and exposes technologies such as Tensor Cores, Transformer Engine and NVFP4[1].
The effect shows up in concrete numbers. NVIDIA notes that DFlash speculative decode delivers up to 15x more throughput on existing hardware, and that FastVideo generates 1080p videos in less than five seconds. When the new open model DeepSeek V4 was released, the vLLM and SGLang inference frameworks had day-zero deployment recipes for Blackwell, and within about a month performance on Blackwell improved by up to 5x, cutting token costs to roughly one-fifth of previous levels[1]. The more developers optimize CUDA-native inference paths, the more those improvements feed back into the ecosystem through production deployments. NVIDIA calls this the open source flywheel[1].
It is worth noting that the multipliers and cost reductions cited here were all published by NVIDIA itself. Readers gauging the real-world impact should weigh third-party verification and the fact that results vary with the model and workload in question.
Summary
NVIDIA positions inference competitiveness as having shifted from raw chip performance to the cost per token delivered by software. A full-stack design that connects three layers, together with a CUDA-centered open source ecosystem, underpins its claim of cutting DeepSeek V4 costs by up to 5x on Blackwell in one month. Beyond hardware generations, continuous software improvement now appears central to the economics of inference.
出典:https://blogs.nvidia.com/blog/inference-software-lowest-token-cost/
