NVIDIA has introduced "intelligence per dollar" as a new yardstick for the agentic AI era in a post on its official blog[1]. The company argues that AI has moved past training a model once and calling it done: post-training, the phase where a model keeps sharpening its skills while in operation, is now the main event. It positions its new Vera Rubin platform as the infrastructure built to handle that continuous load efficiently, claiming it can train the largest models with one-fourth the GPUs of the previous generation and rethinking the cost structure of AI infrastructure in the process.

Post-Training Shifts From a One-Time Step to a Loop That Never Stops

NVIDIA uses a professional athlete as an analogy. What sets elite performers apart is what happens between games: continuous refinement, adjusting to new opponents and sharpening skills based on what the last game exposed. Agentic AI works the same way[1]. Unlike a generative model that simply responds to a prompt, an agent is given a goal and must plan for itself, adapt as environments shift and tools change, and recover from problems it hits mid-run.

For that reason, post-training, the phase that refines a model after its initial training, is no longer a one-time finishing step. The tools an agent relies on can change week to week, and production surfaces edge cases no test set anticipated. Training loops back from production as new problems appear, and the compute footprint grows not because any single run is larger, but because the runs never stop. NVIDIA describes this as the central workload of the agentic era[1].

"Intelligence per Dollar" as a New Yardstick

The goal of post-training is to maximize the yield of every forward pass (inference) and backward pass (weight update) in the learning loop, raising the intelligence gained for every dollar spent. That is where "intelligence per dollar" comes in[1].

NVIDIA places the metric one layer above the familiar "cost per token." Where cost per token measures operating efficiency — what it costs to deliver 1 million tokens — intelligence per dollar asks what it costs to build a model worth serving and keep it worth serving as its environment changes. The two are nested rather than competing: infrastructure that lowers cost per token also lowers the cost of every unit of intelligence built into the model[1].

As a concrete example, NVIDIA points to Nemotron 3 Ultra, an open-weight 550-billion-parameter model with a fully disclosed training recipe. It scored 71.7 percent on SWE-bench verified, a standard benchmark that scores real software bug fixes, producing a fix that passed a project's own tests for roughly seven of every 10 real open-source bugs[1][2].

Vera Rubin Trains Top Models on One-Fourth the GPUs

To make this continuous post-training economically viable, NVIDIA says its current Blackwell platform lowered the cost per run and made frequent post-training practical. The next-generation Vera Rubin platform extends that trajectory further, training the largest models with one-fourth the GPUs of the Blackwell generation[1].

Vera Rubin was codesigned end to end for agentic continuous learning, aiming to run more rollouts per run, keep more environments in play at once and sustain post-training cycles that never stop.

How It Is Used in the Field

NVIDIA also highlights several adopters. Prime Intellect, which continuously post-trains frontier open models, runs its training on today's Blackwell while planning to scale reinforcement learning environments and generate more rollouts per run on Vera Rubin. When the company tested realistic reinforcement learning workloads, it found that Vera CPUs delivered on average 30 percent greater throughput per CPU than x86 architectures[1].

Perplexity, the search AI, runs reinforcement learning post-training asynchronously across hundreds of NVIDIA GPUs. An RDMA-based weight transfer engine syncs trillion-parameter models between training and inference nodes in under two seconds, and the resulting post-trained Qwen3 235B models are served on GB200 NVL72 systems[1].

Together AI, which offers post-training as a service, exposes supervised fine-tuning, reinforcement learning and preference optimization through an API and plans to harness Vera Rubin next. NVIDIA stresses that open libraries such as NeMo Gym for training environments and NeMo RL for distributed post-training turn this work from bespoke research code into repeatable infrastructure[1].

Summary

The "intelligence per dollar" metric NVIDIA is putting forward widens the lens from the cost of a single inference to the cost of building a capable model and keeping it capable. As post-training becomes a loop that never stops, the compute load keeps climbing, yet Vera Rubin is said to handle it on one-fourth the GPUs of the Blackwell generation. Together with field figures from the likes of Prime Intellect and Perplexity, and the question of how broadly they will reproduce, it is an announcement that signals the infrastructure race of the agentic era is changing the very yardstick for cost.

出典:https://blogs.nvidia.com/blog/nvidia-vera-rubin-post-training-intelligence-per-dollar/

出典:https://developer.nvidia.com/blog/nvidia-nemotron-3-ultra-powers-faster-more-efficient-reasoning-for-long-running-agents/