AI Networking Built for Scale
Jul 23, 2026
Networking Is Becoming a Key Determinant of AI Performance and Economics
The AI bottleneck is shifting. While compute remains essential, increasingly it is the network that determines system performance. The speed and efficiency of data movement across GPUs, racks, and clusters now influences how fast models train, how effectively they serve users, and the overall cost of delivering AI. As training, distributed inference, and agentic AI scale to tens of thousands of synchronized GPUs, the networking traffic between these GPUs becomes increasingly important in system performance. These workloads compound, and each one raises the demand on the network beneath it.
Meeting that demand takes an architecture that treats networking as a system across three layers, spanning the front-end that feeds the AI infrastructure, the scale-up interconnect that binds GPUs into one compute resource, and the scale-out network that extends AI across racks and data centers. At Advancing AI 2026, AMD is introducing new networking capabilities across all three layers, integrated in AMD Helios™ as one rack-scale system under a single programmable software platform, spanning the AMD Pensando™ Salina DPU for front-end agentic AI acceleration, UALoE for open, resilient scale-up networking, and the AMD Pensando™ Vulcano 800 AI NIC for scale-out and scale-across.
Keeping pace with these evolving workloads requires more than added bandwidth. Network bandwidth has roughly doubled every two years as the industry has moved from 400G to 800G and now toward 1.6T, but speed alone is not sufficient. AI factories must support training, distributed inference, and agentic workloads concurrently, each placing fundamentally different demands on the network: training requires tight synchronization across thousands of GPUs, inference requires consistent low latency as demand fluctuates, and agentic AI requires persistent memory access and coordination across multiple services simultaneously. Addressing all three at once, within a single data center, is the infrastructure challenge hyperscalers face today. The response has been architectural: intelligence is moving into the network itself, with programmable AI NICs and DPUs distributing awareness and decision-making to the endpoints where AI traffic originates.
Front-End: Free Compute to Feed the AI Factory
Front-end networking connects users and data to the AI infrastructure, handling ingestion, authentication, storage access, and traffic security. Normally, each of those functions runs on the host CPU, tying up cores that would otherwise serve inference and reasoning, and the overhead only grows as the cluster scales.
As AI shifts toward agentic workloads, longer context windows, higher concurrency, and persistent memory access have expanded the DPU's mandate well beyond networking offload. A DPU is now central to how agentic systems access data, coordinate work, and sustain performance under load, and programmability is what elevates it from a fixed-function chip to a platform. Infrastructure services can be updated and differentiated through software, compounding innovation across every deployment and generation. The AMD Pensando™ Salina DPU provides network traffic security, accelerates software-defined networking, and speeds storage access, while extending GPU memory capacity through DPU-managed NVMe so KV-cache can support million-token contexts, delivering up to 1.4x the performance of NVIDIA BlueField-31 and returning up to 22 CPU cores per server to AI work.2
Scale-Up: Unite Multiple GPUs Into One System
The largest models exceed the memory and compute of any single GPU and must be distributed across many. To run efficiently, those GPUs cannot behave as separate machines exchanging messages, but must operate as one system. Techniques such as tensor parallelism require continuous exchange of intermediate results across GPUs, placing strict demands on bandwidth, latency, and consistency. Any degradation in the interconnect slows the entire group to its weakest link, at which point additional GPUs stop contributing.
Inside AMD Helios, UALoE connects all 72 AMD Instinct™ MI455X GPUs as a single compute resource on a single-hop topology, delivering up to 260 TB/s of aggregate bandwidth and access to 31 TB of HBM4 memory within a single rack, without committing customers to a proprietary interconnect.
At the core of the AMD scale up solution is resilient design, which is required for scale-up to be used in production. Each GPU connects through 18 independent UALoE stations across separate switch trays, so when a hardware event affects one path, the system maintains the unified GPU domain automatically without terminating the workload. AMD Fabric Manager and AMD Fabric OS together provide the operational layer, handling provisioning, workload placement, telemetry, fault recovery, and real-time switch-layer visibility, transforming the scale-up network into a managed, observable platform. vPods extend that further, allowing customers to partition the domain into software-defined GPU environments so concurrent workloads coexist in isolation on the same rack without sacrificing performance.
Scale-Out: Distribute AI Across Data Centers
Rarely is a single rack large enough for AI at full scale, so training frontier models and serving millions of users distributes work across racks and data centers. At this scale, gradients must move without stalling, latency must stay predictable as demand fluctuates, and reliability, availability, and serviceability (RAS) matter as much as bandwidth alone. The network must recover quickly from congestion, packet loss, and path disruptions so a single event doesn’t stall a job.
Designed for next-generation AI infrastructure, the AMD PensandoTM Vulcano 800 AI NIC delivers programmable, high-performance scale-out networking for Helios, enabling up to 2.4 Tbps of scale-out bandwidth per GPU. The additional bandwidth helps accelerate collective communication and reduce training bottlenecks. The next-generation AMD AI NIC™ further delivers up to 13% improvement in AI job completion times3 and up to 33% lower switching costs.4 The AMD Pensando™ Pollara 400 AI NIC extends these capabilities to existing AI clusters, bringing open, Ethernet-based AI networking to deployments already operating at scale.
Open and Programmable Networking Is Built to Keep Evolving
Built on third-generation programmable P4 engines, AMD networking absorbs new transports and optimizations through software, letting infrastructure evolve without costly hardware replacement. The result is a programmable Ethernet fabric that keeps pace with AI rather than holding it back.
Proprietary interconnects can increase dependency on a single vendor ecosystem, while open standards are intended to provide choice and speed industry-wide innovation. Grounded in open Ethernet standards and ecosystems including the Ultra Ethernet Consortium (UEC), Ultra Accelerator Link (UAL), and ESUN, AMD delivers one coherent system across front-end, scale-up, and scale-out, built to carry AI infrastructure through what comes next.
Related Blogs
Footnotes
- PEN-017: Testing conducted by AMD Performance Labs as of 15th April 2025 on the AMD Pensando Salina DPU, on a test system comprising of 2x Dual socket Xeon Dell power edge XE9680 function; Cisco 64x400G Switch; IXIA 2x400G tester as a traffic generator from Keysight; 2xAMD Pensando™ Salina DPU; 5th gen Xeon 8568 - 48 core CPU with PCIe Gen-5; Operating System Version : Ubuntu® 22.04.5 LTS; Kernel Version: 5.15.0-139-generic; BIOS version: 1.3.6 [
Mitigation: Off (default)
System profile setting: Performance (default)
SMT: enabled (default)]
AMD Pensando performs an average of 117 MPPS (millions packets per second) in AMD testing while Nvidia Bluefield-3 published performance is 80 MPPS. https://hc33.hotchips.org/assets/program/conference/day1/HC2021.NVIDIA.IdanBurstein.v08.norecording.pdf Slide 6 for ~1.45x the performance with AMD Pensando Salina DPU.
Results may vary based on factors including but not limited to system configuration and software settings.
- 22 CPU Core Savings with Accelerated Connections running on AMD DPU at Microsoft. Core Savings for Customer SLB Appliances running on Sirius Appliance at Microsoft as part of Accelerated Connections Service vs Running SLB on Standard Server
- Based on AMD Engineering silicon modeling and AMD synthetic benchmark simulation, using the MOE_4p5T hi_sparsity_ea9 benchmark test to project time- to- solution speed up in days using the FP8 training datatype on a simulated LLM system modeled with 8000 AMD Instinct MI455X GPUs and (3) versus (2) AMD Pensando Vulcano NICs.
Results reflect analysis of pure training compute (decode layers only) Configuration evaluated with a global batch size of 4096 and a sequence length of 4K, using SwiGLU activation. Results exclude evaluation and checkpointing overhead. FlashAttention v3 with matmul–softmax overlap is assumed. AllReduce and All2All communication costs are fully accounted for and not hidden via tiled compute–communication overlap. Gradient synchronization and FSDP weight prefetch are evaluated across varying levels of overlap to assess scale-out sensitivity. Results assume ideal and not fully optimized real-world behavior and may vary when actual product(s) are released in market.
- PEN-022: AMD comparison and pricing as of May 18, 2026, for network fabric costs to support 32,000 GPUs. Comparison of a Vulcano-based NIC (VULCANO-CUSTOM-2.4T) deployed as part of a Helios rackscale system with a network based on 1.6T Tomahawk 6 switching with 200G SerDes versus using a competitor 800G NIC with 800G Tomahawk 6 switching with 100G SerDes. Both fabrics were fat-tree topologies built on Tomahawk 5 800G switching platforms, with NIC costs considered comparable. The Vulcano-based design is estimated to deliver up to 33% savings in network switching costs by enabling a more cost-effective architecture with fewer switching platforms, more bandwidth per port on the network, and reduced transceiver cables/optics.
TH6-100G Serdes Fat-Tree (Competition):
Switching (TH6C BCM78914 - 128x800G):
• Leaf Units 1,000
• Spine Units 500
• Total Switches 1,500
• TH6 100G Unit Price $79,587
• Total Switching Cost $60M
Cables/Optics:
• NIC Transceivers (800G-DR8) 64,000 @ $500
• Leaf/Spine Transceivers (800G-DR8) 192,000 @ $500
• MPO Cables 256,000 @ $89
• Total Optics/Cables $64M
Total Fabric Cost TH6-100G (Switches + Cables/Optics): $124M
TH6-200G Serdes Fat-Tree (Vulcano-Custom2.4T / AMD Solution):
Switching (TH6P BCM78910 - 64x1.6T):
• Leaf Units 1,000
• Spine Units 500
• Total Switches 1,500
• TH6 200G Unit Price $66,336
• Total Switching Cost $50M
Cables/Optics:
NIC Transceivers (1.6T-DR8) 32,000 @ $900 (50% fewer vs. competition)
Leaf/Spine Transceivers (1.6T-DR8) 64,000 @ $900
MPO Cables 128,000 @ $89 (50% fewer vs. competition)
Total Optics/Cables $43M
Total Fabric Cost TH6-200G (Switches + Cables/Optics): $93M
Capex Savings (Fabric only):
Savings $: $30.7M
Savings %: 33.1%
Pricing sources: SemiAnalysis Hyperscaler Networking Model data used with permission; full analysis available via SemiAnalysis subscription. Edgecore switch pricing as of May 17, 2026. Results may vary based on system configuration.
- PEN-017: Testing conducted by AMD Performance Labs as of 15th April 2025 on the AMD Pensando Salina DPU, on a test system comprising of 2x Dual socket Xeon Dell power edge XE9680 function; Cisco 64x400G Switch; IXIA 2x400G tester as a traffic generator from Keysight; 2xAMD Pensando™ Salina DPU; 5th gen Xeon 8568 - 48 core CPU with PCIe Gen-5; Operating System Version : Ubuntu® 22.04.5 LTS; Kernel Version: 5.15.0-139-generic; BIOS version: 1.3.6 [
Mitigation: Off (default)
System profile setting: Performance (default)
SMT: enabled (default)]
AMD Pensando performs an average of 117 MPPS (millions packets per second) in AMD testing while Nvidia Bluefield-3 published performance is 80 MPPS. https://hc33.hotchips.org/assets/program/conference/day1/HC2021.NVIDIA.IdanBurstein.v08.norecording.pdf Slide 6 for ~1.45x the performance with AMD Pensando Salina DPU.
Results may vary based on factors including but not limited to system configuration and software settings. - 22 CPU Core Savings with Accelerated Connections running on AMD DPU at Microsoft. Core Savings for Customer SLB Appliances running on Sirius Appliance at Microsoft as part of Accelerated Connections Service vs Running SLB on Standard Server
- Based on AMD Engineering silicon modeling and AMD synthetic benchmark simulation, using the MOE_4p5T hi_sparsity_ea9 benchmark test to project time- to- solution speed up in days using the FP8 training datatype on a simulated LLM system modeled with 8000 AMD Instinct MI455X GPUs and (3) versus (2) AMD Pensando Vulcano NICs.
Results reflect analysis of pure training compute (decode layers only) Configuration evaluated with a global batch size of 4096 and a sequence length of 4K, using SwiGLU activation. Results exclude evaluation and checkpointing overhead. FlashAttention v3 with matmul–softmax overlap is assumed. AllReduce and All2All communication costs are fully accounted for and not hidden via tiled compute–communication overlap. Gradient synchronization and FSDP weight prefetch are evaluated across varying levels of overlap to assess scale-out sensitivity. Results assume ideal and not fully optimized real-world behavior and may vary when actual product(s) are released in market. - PEN-022: AMD comparison and pricing as of May 18, 2026, for network fabric costs to support 32,000 GPUs. Comparison of a Vulcano-based NIC (VULCANO-CUSTOM-2.4T) deployed as part of a Helios rackscale system with a network based on 1.6T Tomahawk 6 switching with 200G SerDes versus using a competitor 800G NIC with 800G Tomahawk 6 switching with 100G SerDes. Both fabrics were fat-tree topologies built on Tomahawk 5 800G switching platforms, with NIC costs considered comparable. The Vulcano-based design is estimated to deliver up to 33% savings in network switching costs by enabling a more cost-effective architecture with fewer switching platforms, more bandwidth per port on the network, and reduced transceiver cables/optics.
TH6-100G Serdes Fat-Tree (Competition):
Switching (TH6C BCM78914 - 128x800G):
• Leaf Units 1,000
• Spine Units 500
• Total Switches 1,500
• TH6 100G Unit Price $79,587
• Total Switching Cost $60M
Cables/Optics:
• NIC Transceivers (800G-DR8) 64,000 @ $500
• Leaf/Spine Transceivers (800G-DR8) 192,000 @ $500
• MPO Cables 256,000 @ $89
• Total Optics/Cables $64M
Total Fabric Cost TH6-100G (Switches + Cables/Optics): $124M
TH6-200G Serdes Fat-Tree (Vulcano-Custom2.4T / AMD Solution):
Switching (TH6P BCM78910 - 64x1.6T):
• Leaf Units 1,000
• Spine Units 500
• Total Switches 1,500
• TH6 200G Unit Price $66,336
• Total Switching Cost $50M
Cables/Optics:
NIC Transceivers (1.6T-DR8) 32,000 @ $900 (50% fewer vs. competition)
Leaf/Spine Transceivers (1.6T-DR8) 64,000 @ $900
MPO Cables 128,000 @ $89 (50% fewer vs. competition)
Total Optics/Cables $43M
Total Fabric Cost TH6-200G (Switches + Cables/Optics): $93M
Capex Savings (Fabric only):
Savings $: $30.7M
Savings %: 33.1%
Pricing sources: SemiAnalysis Hyperscaler Networking Model data used with permission; full analysis available via SemiAnalysis subscription. Edgecore switch pricing as of May 17, 2026. Results may vary based on system configuration.