MRC introduces a new transport paradigm for AI data centers

MRC: A New Transport Protocol for AI Data Centers

In May 2026, OpenAI — together with AMD, Broadcom, Microsoft, and NVIDIA — published the MRC (Multipath Reliable Connection) protocol specification through the Open Compute Project. This is not an academic paper. It is an engineering result that has been running in production on Stargate for some time.

It addresses a problem that cannot be sidestepped at 100,000+ GPU scale: the network design assumptions behind traditional RoCE are failing systematically at this size.

The Problem: Three Costs That Keep Growing

Cost #1: Flow collisions.

Traditional RoCE RC pins every QP to a single ECMP path. When two AllReduce flows happen to hash onto the same egress link, bandwidth degrades — and the collision does not self-heal. The next iteration hits the same result. The larger the cluster, the higher the probability: collective communication at N-GPU scale generates O(N²) QP pairs, and ECMP hash collisions across that many fixed paths are inevitable.

Cost #2: Failure blast radius.

In a single-plane 800G Clos network, a single T0–T1 link flap triggers BGP reconvergence, which takes seconds to tens of seconds. For a synchronous training job, a second of network interruption means AllReduce stalls and the entire cluster's GPUs are waiting on the fabric. At Stargate's scale, multiple T0–T1 link flaps per minute are the norm, not the exception.

Cost #3: No fault localization.

Dynamic routing via ECMP + BGP makes the actual path of each flow unpredictable. When training step time suddenly spikes, operators cannot tell which switch or which link is the culprit. Replacing a link requires coordinating with the training team, stopping the job, and burning GPU time.

MRC's Design Choices

1. Eight-Plane Architecture

MRC exploits per-lane NIC port breakout: a single OSFP 800G port is physically 8 independent 100G PAM4 lanes. A breakout cable (one OSFP head, eight SFP-DD tails) splits these 8 lanes to 8 different T0 switches, forming 8 independent Clos planes.

The consequence: the same 51.2 Tb/s switch ASIC that has 64 ports at 800G has 512 ports at 100G. Two switch tiers cover 131,000 GPUs — no third tier needed.

More importantly, the failure blast radius changes fundamentally. A single T0–T1 link failure removes ~3% of NIC bandwidth in a single-plane 800G design. In MRC's 8-plane design it removes 1/256 ≈ 0.4%. A full NIC port failure (1 of 8 planes gone) lets the job continue at 7/8 bandwidth without crashing.

This is fundamentally different from Multi-Rail. In Multi-Rail, each server's NIC connects to exactly one rail; cross-rail traffic must traverse the spine. In MRC, every server connects to all 8 planes simultaneously. Any two servers have 8 independent paths between them. Path management lives in the NIC firmware — NCCL sees nothing.

2. Host-Side Packet Spraying: Per-Packet Path Selection

Traditional RoCE RC segments a large WRITE into First/Middle/Last packets. Only the First packet carries the RETH header (remote memory address + rkey); subsequent packets carry no addressing information and depend on in-order arrival to accumulate offsets from the base address. Out-of-order delivery is fatal — a Middle packet arriving without a base address leaves the NIC with nowhere to DMA.

MRC's solution: every data packet carries a complete RETH header (remote virtual address + rkey, 16 bytes of overhead). Each packet is self-contained. Arrival order is irrelevant; each packet DMAs directly to its correct memory location.

An important distinction: MRC's spraying is host-driven out-of-order delivery, not the incidental reordering produced by switch hash algorithms. Correspondingly, the NIC must be able to reassemble out-of-order packets.

The NIC firmware maintains an EV (Entropy Value) table for independent per-packet path selection. Each QP holds hundreds of EVs, each corresponding to a physical path. In plain terms: traditional RoCE uses the same 5-tuple hash for all packets within a QP. In MRC, every packet within a QP can theoretically take a different path. Still manageable by EV table to be with certainty.

3. Drop PFC, Use Multipath Congestion Control

Loss and congestion handling is also upgraded. PFC is disabled; SACK/NACK with selective retransmission replaces it.

When a switch detects congestion, the response depends on severity. Mild congestion — or congestion attributable to hash collisions — triggers an immediate EV switch, steering the packet to a different path. Severe or widespread congestion triggers rate reduction, now controlled by a send window driven by combined ECN and RTT signals (NSCC), not by PFC pause frames.

For unmanageable congestion, switches do not drop packets. Instead they strip the payload and forward only the header (packet trimming); the receiver sends a NACK to trigger precise retransmission, avoiding Go-Back-N.

For incast scenarios, MRC has no definitive answer. Given the inherent latency of ECN and RTT feedback, some retransmissions are likely unavoidable.

4. Replace BGP with SRv6

The fundamental problem with dynamic routing: the path a packet actually takes differs from what operators believe it takes, so fault localization comes down to guesswork.

MRC disables dynamic routing entirely and replaces it with SRv6 static source routing. The sender encodes the full hop sequence (T0 → T1 → T0 → destination) into the IPv6 destination address. Switches look up a local static forwarding table — configured once at fabric bring-up, never updated thereafter.

Two outcomes follow. First, BGP reconvergence disappears — link failures are handled by NIC firmware at EV granularity: mark BAD, stop sending, send probes, recover. The control plane is never touched. Second, observability becomes precise: a clustermapper process runs on each server, sending probes to every T0 every millisecond. These probes take the exact same forwarding path as data packets. T0-loopback succeeds but T1-loopback fails — the fault is on that T0–T1 link. No guessing.

The performance impact of probe traffic — on both NIC and switch — at this frequency warrants further evaluation.

Looking Ahead: MRC and the Test Framework

For validation and benchmarking, this means the evaluation framework needs to change alongside the protocol. Single-hash RDMA QPs cannot reproduce MRC's per-packet path selection behavior. Single-plane test beds do not produce valid results for multi-plane deployments. Link flap injection moves from an edge case to a core validation item.

MRC defines a new performance baseline. It also defines new test requirements.

limit
3