MetaRoCE: Evolving Network Transports for AI
Aug 25, 2026
Addressing the Networking Demands of AI Training
As AI clusters continue to scale, the communication patterns generated by these environments are placing new demands on network infrastructure. Traditional Ethernet transport protocols were designed for cloud, storage, and HPC workloads, but large-scale AI training requires thousands of accelerators to communicate and synchronize continuously, creating different demands on the network.
Developed after Meta identified limitations in existing transport protocols at AI training scale, MetaRoCE introduces a more intelligent, multipath approach to moving data across the network, helping large AI environments make more effective use of available network resources.
AI architectures are also evolving beyond traditional scale-out clusters toward scale-across environments that span multiple data center sites. Supporting these deployments requires networking technologies that can efficiently utilize available paths, adapt to changing network conditions, and maintain performance across increasingly complex infrastructure. MetaRoCE was implemented on the AMD Pensando™ Pollara 400 AI NIC before being deployed on the AMD Pensando™ Vulcano 800 AI NIC, showing how programmable networking can help move a new transport from testing, validation, and toward to production as requirements evolve.
What MetaRoCE Does Differently
MetaRoCE changes how traffic moves across the network. Rather than assigning a connection to a single network path, MetaRoCE can use multiple available paths simultaneously, helping improve bandwidth utilization and reduce the impact of congestion on any individual link.
That multipath capability is especially relevant to large AI training clusters where bandwidth utilization and congestion management across many parallel paths directly impact job completion time. In these environments, paths can vary in utilization and congestion, and MetaRoCE can dynamically distribute traffic and adapt to changing network conditions, helping maintain throughput and reduce the impact of congestion on individual links.
MetaRoCE also moves more transport intelligence to the endpoint. Rather than relying solely on the network to manage congestion, the receiving NIC signals a target rate, which the sender uses to cap the transmission window and schedule traffic across available paths. If a packet is lost, only the missing packet is retransmitted rather than everything that followed. Together, endpoint control, multipathing, and efficient packet recovery help maintain throughput and resiliency while the underlying fabric remains standard Ethernet.
Accelerating Proof of Concept to Production
The programmable AMD AI NIC architecture played a key role in the development of MetaRoCE. Using the AMD Pensando™ Pollara 400 AI NIC as the initial development and validation platform, Meta and AMD were able to implement protocol changes in software, evaluate their impact on network behavior, and refine the transport through real-world testing. The teams did not have to wait for a future silicon generation each time the protocol changed.
That same programmability is simplifying the transition from proof of concept to production. MetaRoCE could be carried forward from the AMD Pensando™ Pollara 400 AI NIC to the higher-bandwidth AMD Pensando™ Vulcano 800 AI NIC without restarting the transport effort on a new fixed-function implementation. This preserved the work already invested in developing and validating the protocol while providing a path to 800G deployment.
The progression from AMD Pensando Pollara 400 AI NIC to AMD Pensando Vulcano 800 AI NIC illustrates an important benefit of programmable networking: innovation can move forward without starting over with each hardware generation.
Programmability for an Evolving AI Ecosystem
New workload patterns, deployment models, and transport innovations often emerge faster than traditional hardware refresh cycles. Programmable networking allows new capabilities and protocol enhancements to be introduced without requiring a completely new networking architecture.
AMD programmable AI NICs provide the hardware foundation to implement and evolve advanced transport technologies on standard Ethernet. MetaRoCE is one example. Rather than fixing transport behavior when the chip is designed, the programmable architecture from AMD allows protocol logic to be developed, tested, and refined as networking requirements change.
That flexibility extends beyond the initial deployment. If the MetaRoCE specification is updated, or future protocol enhancements are required, those changes can be implemented through the programmable data path rather than waiting for an entirely new hardware platform. This gives infrastructure teams a way to introduce updates while preserving existing hardware investments.
The same approach supports work by AMD with open standards and industry collaboration, including Multipath Reliable Connection (MRC) and the Ultra Ethernet Consortium (UEC). Programmability provides a practical path for adopting new transport technologies while preserving interoperability and the openness of Ethernet-based infrastructure.
Evolving at the Pace of AI
MetaRoCE shows what becomes possible when networking infrastructure can evolve alongside changing requirements. As networking requirements continue to change, transport protocols and specifications will change with them. Moving intelligence to the endpoint and implementing transport logic in software provides a path to introduce enhancements, refine behavior, and add capabilities without waiting for the next hardware generation. MetaRoCE puts that model into practice while preserving the openness and flexibility of Ethernet infrastructure.