OpenAI teams with AMD, NVIDIA, Intel, Microsoft, and Broadcom on faster AI training networks

news
OpenAI teams with AMD, NVIDIA, Intel, Microsoft, and Broadcom on faster AI training networks

OpenAI teams with AMD, NVIDIA, Intel, Microsoft, and Broadcom on faster AI training networks

OpenAI has worked with AMD, NVIDIA, Intel, Microsoft, and Broadcom on a new networking protocol designed to make large AI training systems faster and more reliable.

The protocol is called MRC, short for Multipath Reliable Connection. It is meant to solve one of the biggest problems in large AI clusters: keeping thousands of GPUs fed with data without delays, failures, or network congestion slowing everything down.

Large AI training runs depend on constant data movement between GPUs, CPUs, and network hardware. If one transfer arrives late, part of the system can sit idle while waiting. That wastes expensive compute time, especially in clusters with tens of thousands or hundreds of thousands of GPUs.

MRC tries to reduce that problem by spreading a single transfer across many paths instead of relying on one large network route. It can also route around failures in microseconds, which helps training continue with fewer interruptions.

OpenAI says the protocol is built into the latest 800 Gb/s network interfaces. Instead of treating one interface as a single 800 Gb/s link, MRC can split it into smaller links across multiple switches. For example, one interface can be divided into eight 100 Gb/s paths, creating several parallel network planes.

That design can make very large GPU clusters easier to build. OpenAI says a network using this approach can fully connect about 131,000 GPUs with only two switch tiers. A more traditional 800 Gb/s network may need three or four tiers, which adds complexity.

Here is a quick look at MRC:

AreaDetails
Full nameMultipath Reliable Connection
Main goalImprove AI training network speed and resilience
Companies involvedOpenAI, AMD, NVIDIA, Intel, Microsoft, and Broadcom
Main problem addressedCongestion, late transfers, and hardware failures
Network supportLatest 800 Gb/s interfaces
Key ideaSplit traffic across many smaller paths
Major benefitKeeps large GPU clusters running with fewer delays
AvailabilityReleased through the Open Compute Project

MRC extends RDMA over Converged Ethernet, also known as RoCE. That allows faster memory access between GPUs and CPUs over Ethernet based networks, which is important for AI training systems.

OpenAI has already used MRC across supercomputers with NVIDIA GB200 Blackwell GPUs. These include systems used for training frontier AI models, including Oracle Cloud Infrastructure in Abilene, Texas, and Microsoft’s Fairwater supercomputers.

The protocol is also expected to play a role in OpenAI’s Stargate supercomputer plans. That project is targeting 10 GW of AI compute by 2029, with several gigawatts already deployed.

The larger point is that AI progress now depends on more than faster GPUs. At this scale, networking is just as important. If data cannot move quickly and reliably between chips, much of the available compute is wasted. MRC is OpenAI’s attempt to make those massive AI systems more efficient, while giving the wider industry a shared protocol to build on.

Discover: News

Discussion (0)

Be the first to comment.