NVIDIA announced that it posted the fastest time to train on every one of the seven benchmarks in the latest MLPerf Training 6.0 round[1]. It was also the only platform to scale a submission to 8,192 GPUs, and its newest GB300 NVL72 system delivered up to 1.6x faster training than the previous generation. As AI models keep growing in size, NVIDIA used this round to push performance, scale, and reliability of the training platform all at once.
What Is MLPerf Training 6.0?
MLPerf Training is a peer-reviewed, industry-standard benchmark that compares AI model training performance under common conditions[1]. Because every vendor races the same workloads to a fixed quality target, it offers a like-for-like view of how capable a given compute platform really is.
Version 6.0 added two new mixture-of-experts (MoE, an approach that activates only a subset of expert networks depending on the input) pretraining workloads: DeepSeek-V3 671B and GPT-OSS-20B. The additions reflect how central MoE architectures have become for large-scale models[1].
Fastest on All Seven, the Only Full Sweep
NVIDIA was the only platform to submit results across every one of the seven benchmarks in the suite, and it delivered the fastest time to train on all of them[1]. This round, the company submitted on two rack-scale systems, GB200 NVL72 and GB300 NVL72.
Inside each rack, fifth-generation NVLink Switches connect 72 GPUs with high bandwidth into a unified pool of compute and memory, letting the 72 GPUs act as one giant GPU[1]. Large-scale MoE training requires all-to-all communication to route tokens to the right expert subnetwork across GPUs, and NVLink's bandwidth is what handles that efficiently. NVIDIA also showcased its NVFP4 low-precision training methods, which raise performance while meeting strict accuracy requirements; most recently it used NVFP4 to pretrain the 550-billion-parameter NVIDIA Nemotron 3 Ultra model[1].
GB300 NVL72: Up to 1.6x Faster Than the Previous Generation
In this round, GB300 NVL72 delivered up to 1.6x faster training than GB200 NVL72 at the same scale[1]. Blackwell Ultra capabilities such as higher compute density with NVFP4, expanded memory capacity, and a higher power ceiling that lets the GPU sustain peak performance drive this gain. NVIDIA has published a separate, deeper technical breakdown of the results[2].
Scaling to 8,192 GPUs and Partner Records
For distributed training at scale, NVIDIA offers two scale-out networking platforms, NVIDIA Quantum InfiniBand and NVIDIA Spectrum-X Ethernet, so data centers can build large clusters tuned to their own infrastructure[1].
On DeepSeek-V3 671B, the largest MoE model in the suite, NVIDIA scaled its submission to 8,192 GPUs using GB200 NVL72 systems, the largest Blackwell-based submission in MLPerf Training to date. It also submitted results at 5,120 GPUs on Llama 3.1 405B, one of the largest dense LLMs in the suite[1].
Partner records stand out as well. Microsoft Azure scaled Llama 3.1 405B training to 8,192 GPUs on GB200 NVL72 and reached the reference quality target in 7.07 minutes, the fastest time to train for that benchmark. CoreWeave delivered the fastest time to train for DeepSeek-V3 671B, hitting the quality target in 2.02 minutes at 8,192-GPU scale using GB300 NVL72 systems connected with Spectrum-X Ethernet[1].
Reliability Built for Large-Scale Operation
In production, training runs can span weeks or months across hundreds of thousands of GPUs. At that scale, effective training throughput depends not only on system performance but also on the resiliency that keeps runs reproducible over time[1].
NVIDIA says it engineers resiliency along two axes. The first is preventing failures: each GPU is screened across more than 30 manufacturing test stages before it reaches a data center, and once deployed, the RAS (reliability, availability, and serviceability) engine monitors nearly the entire chip while self-healing capabilities route around detected faults without stopping the workload. At the network level, Spectrum-X Ethernet reroutes around failed links in milliseconds, keeping the fabric healthy without disrupting the job[1].
The second is faster recovery. The NVIDIA Resiliency Extension (NVRx) detects and manages underperforming nodes before they slow the rest of the cluster, and when an interruption occurs, it resumes from a recent checkpoint (a saved snapshot of the training state) rather than restarting the entire job, minimizing lost time[1].
A Growing Blackwell Ecosystem
Partner participation was robust this round, with 19 organizations submitting results, including Microsoft Azure, CoreWeave, Dell Technologies, Google Cloud, Fujitsu, and Supermicro[1].
Concrete outcomes were shared as well. Cohere achieved 3x faster training on GB200 NVL72 for its North agentic AI platform. Midjourney, which trained its v8 image generation model on a Blackwell cluster, is now scaling a large fleet of Blackwell Ultra GPUs on CoreWeave for upcoming image and video models. Thinking Machines Lab saw 2x faster training and serving on GB300 NVL72 on Google Cloud compared with the prior generation. And Nebius cut Higgsfield's training time by 30% on Blackwell infrastructure, supporting a platform that now serves 22 million users and generates more than 6 million pieces of AI content per day[1].
Summary
In MLPerf Training 6.0, NVIDIA Blackwell posted the fastest time to train on all seven benchmarks and was the only platform to scale to 8,192 GPUs. GB300 NVL72 improved by up to 1.6x over the previous generation, and together with the new MoE workloads, NVFP4 low-precision training, and millisecond-level fault rerouting, the results showcase the platform's all-around strength as a training foundation.
Source: https://blogs.nvidia.com/blog/blackwell-mlperf-training-6-0/
