OpenAI announced on May 5, 2026, that it has released MRC (Multipath Reliable Connection), a new networking protocol for the GPU fabrics inside AI supercomputers, as a contribution to the Open Compute Project (OCP). MRC was developed over two years with AMD, Broadcom, Intel, Microsoft and NVIDIA, and is already deployed across OpenAI's largest NVIDIA GB200 supercomputers, including the Stargate site in Abilene, Texas and Microsoft's Fairwater facilities[1][2].
What MRC is: an AI-training extension on top of RoCE
MRC extends RDMA over Converged Ethernet (RoCE), the standard that enables hardware-accelerated remote direct memory access between GPUs and CPUs. It draws on techniques developed by the Ultra Ethernet Consortium (UEC) and layers SRv6 (IPv6 Segment Routing) source routing on top, so it can support large-scale AI training fabrics[1].
The protocol is built into the latest 800Gb/s network interfaces and lets OpenAI spread a single transfer across hundreds of paths, route around failures in microseconds and run a simpler network control plane[1]. The motivation is the predictable performance required to keep training jobs running for the systems behind ChatGPT, which is now used by more than 900 million people every week[1].
Multi-plane topology and packet spraying to crush congestion
There are three pieces to MRC. The first is a multi-plane topology: instead of treating each interface as a single 800Gb/s link, OpenAI splits it into eight 100Gb/s paths and builds eight independent planes. A switch that previously connected 64 ports at 800Gb/s can now connect 512 ports at 100Gb/s[1].
The result is that roughly 131,000 GPUs can be fully connected with only two tiers of switches, where a conventional 800Gb/s network would have required three or four. OpenAI says this design lowers power draw, the count of components that can fail and the total cost of the network, and details it in the co-authored paper "Resilient AI Supercomputer Networking using MRC and SRv6"[1][3].
The second is adaptive packet spraying. Classic RoCE forces a single transfer down a single path so that packets arrive in order, which leaves a multi-plane fabric underused. MRC deliberately sprays the packets of one transfer across hundreds of paths and lets the destination GPU memory absorb them out of order. When a path congests, the connection swaps it; if a packet is lost, MRC takes the safe option and stops using that path until probe packets confirm it is healthy again[1]. To distinguish congestion from failure, MRC uses packet trimming: a switch can drop only the payload while forwarding the header to the destination, signaling an explicit retransmission request[1].
SRv6 source routing and what it looks like in production
The third pillar replaces dynamic routing with SRv6 source routing. Traditionally, switches run a dynamic protocol such as BGP (Border Gateway Protocol) to compute paths and reroute around failures. With MRC, the sender specifies the path each packet takes, embedding a sequence of switch identifiers into the IPv6 destination address[1].
OpenAI reports that during training they have observed multiple link flaps per minute between Tier 0 and Tier 1 switches, with no measurable impact on synchronous pretraining jobs. A recent frontier model run for ChatGPT and Codex required rebooting four Tier 1 switches, and that maintenance proceeded without coordinating with the training teams[1]. Where a conventional fabric might take seconds to tens of seconds to stabilize after a failure, MRC reroutes on a microsecond timescale, sharply reducing the risk that a single fault stalls a synchronous training run[1].
The flagship deployment site is the Stargate supercomputer that OCI has built in Abilene, Texas, where NVIDIA GB200 GPUs are slated to reach roughly 450,000 units, with MRC underpinning the training jobs that scale[2][4]. Microsoft, for its part, is running MRC across the "AI WAN" that ties together its purpose-built Fairwater data centers, treating the multi-site footprint as a single virtual supercomputer[3].
Summary
MRC reframes AI training networks from "fast as long as nothing breaks" to "fast even though things will break." The combination of multi-plane topology, packet spraying and SRv6 source routing is what makes a 100,000-plus GPU cluster on only two switch tiers practically operable. With the specification now in OCP, the open question is how quickly cloud providers and hardware vendors can ship interoperable implementations.
References: [1] OpenAI "Supercomputer networking to accelerate large scale AI training" https://openai.com/index/mrc-supercomputer-networking/
[2] AMD "AMD and OpenAI Advance AI Networking at Scale with MRC" https://www.amd.com/en/blogs/2026/amd-advances-ai-networking-at-scale-with-mrc.html
[3] SDxCentral "Microsoft details 'AI WAN' connecting distributed Fairwater AI superfactory" https://www.sdxcentral.com/news/microsoft-details-ai-wan-connecting-distributed-fairwater-ai-superfactory/
[4] DataCenterDynamics "OpenAI and Oracle to deploy 450,000 GB200 GPUs at Stargate data center in Abilene, Texas" https://www.datacenterdynamics.com/en/news/openai-and-oracle-to-deploy-450000-gb200-gpus-at-stargate-abilene-data-center/
