AI Networking Built for Scale

Jul 23, 2026

Dark data center aisle with server racks, glowing blue lines. Text: 'NETWORKING BUILT FOR AI', 'AMD together we advance_'.

Networking Is Becoming a Key Determinant of AI Performance and Economics 

The AI bottleneck is shifting. While compute remains essential, increasingly it is the network that determines system performance. The speed and efficiency of data movement across GPUs, racks, and clusters now influences how fast models train, how effectively they serve users, and the overall cost of delivering AI. As training, distributed inference, and agentic AI scale to tens of thousands of synchronized GPUs, the networking traffic between these GPUs becomes increasingly important in system performance. These workloads compound, and each one raises the demand on the network beneath it.

Meeting that demand takes an architecture that treats networking as a system across three layers, spanning the front-end that feeds the AI infrastructure, the scale-up interconnect that binds GPUs into one compute resource, and the scale-out network that extends AI across racks and data centers. At Advancing AI 2026, AMD is introducing new networking capabilities across all three layers, integrated in AMD Helios™ as one rack-scale system under a single programmable software platform, spanning the AMD Pensando™ Salina DPU for front-end agentic AI acceleration, UALoE for open, resilient scale-up networking, and the AMD Pensando™ Vulcano 800 AI NIC for scale-out and scale-across.

Keeping pace with these evolving workloads requires more than added bandwidth. Network bandwidth has roughly doubled every two years as the industry has moved from 400G to 800G and now toward 1.6T, but speed alone is not sufficient. AI factories must support training, distributed inference, and agentic workloads concurrently, each placing fundamentally different demands on the network: training requires tight synchronization across thousands of GPUs, inference requires consistent low latency as demand fluctuates, and agentic AI requires persistent memory access and coordination across multiple services simultaneously. Addressing all three at once, within a single data center, is the infrastructure challenge hyperscalers face today. The response has been architectural: intelligence is moving into the network itself, with programmable AI NICs and DPUs distributing awareness and decision-making to the endpoints where AI traffic originates.

Front-End: Free Compute to Feed the AI Factory

Front-end networking connects users and data to the AI infrastructure, handling ingestion, authentication, storage access, and traffic security. Normally, each of those functions runs on the host CPU, tying up cores that would otherwise serve inference and reasoning, and the overhead only grows as the cluster scales.

As AI shifts toward agentic workloads, longer context windows, higher concurrency, and persistent memory access have expanded the DPU's mandate well beyond networking offload. A DPU is now central to how agentic systems access data, coordinate work, and sustain performance under load, and programmability is what elevates it from a fixed-function chip to a platform. Infrastructure services can be updated and differentiated through software, compounding innovation across every deployment and generation. The AMD Pensando™ Salina DPU provides network traffic security, accelerates software-defined networking, and speeds storage access, while extending GPU memory capacity through DPU-managed NVMe so KV-cache can support million-token contexts, delivering up to 1.4x the performance of NVIDIA BlueField-31 and returning up to 22 CPU cores per server to AI work.2

Scale-Up: Unite Multiple GPUs Into One System

The largest models exceed the memory and compute of any single GPU and must be distributed across many. To run efficiently, those GPUs cannot behave as separate machines exchanging messages, but must operate as one system. Techniques such as tensor parallelism require continuous exchange of intermediate results across GPUs, placing strict demands on bandwidth, latency, and consistency. Any degradation in the interconnect slows the entire group to its weakest link, at which point additional GPUs stop contributing.

Inside AMD Helios, UALoE connects all 72 AMD Instinct™ MI455X GPUs as a single compute resource on a single-hop topology, delivering up to 260 TB/s of aggregate bandwidth and access to 31 TB of HBM4 memory within a single rack, without committing customers to a proprietary interconnect.

At the core of the AMD scale up solution is resilient design, which is required for scale-up to be used in production. Each GPU connects through 18 independent UALoE stations across separate switch trays, so when a hardware event affects one path, the system maintains the unified GPU domain automatically without terminating the workload. AMD Fabric Manager and AMD Fabric OS together provide the operational layer, handling provisioning, workload placement, telemetry, fault recovery, and real-time switch-layer visibility, transforming the scale-up network into a managed, observable platform. vPods extend that further, allowing customers to partition the domain into software-defined GPU environments so concurrent workloads coexist in isolation on the same rack without sacrificing performance.

Scale-Out: Distribute AI Across Data Centers

Rarely is a single rack large enough for AI at full scale, so training frontier models and serving millions of users distributes work across racks and data centers. At this scale, gradients must move without stalling, latency must stay predictable as demand fluctuates, and reliability, availability, and serviceability (RAS) matter as much as bandwidth alone. The network must recover quickly from congestion, packet loss, and path disruptions so a single event doesn’t stall a job.

Designed for next-generation AI infrastructure, the AMD PensandoTM Vulcano 800 AI NIC delivers programmable, high-performance scale-out networking for Helios, enabling up to 2.4 Tbps of scale-out bandwidth per GPU. The additional bandwidth helps accelerate collective communication and reduce training bottlenecks. The next-generation AMD AI NIC™ further delivers up to 13% improvement in AI job completion times3 and up to 33% lower switching costs.4 The AMD Pensando™ Pollara 400 AI NIC extends these capabilities to existing AI clusters, bringing open, Ethernet-based AI networking to deployments already operating at scale.

Open and Programmable Networking Is Built to Keep Evolving

Built on third-generation programmable P4 engines, AMD networking absorbs new transports and optimizations through software, letting infrastructure evolve without costly hardware replacement. The result is a programmable Ethernet fabric that keeps pace with AI rather than holding it back.

Proprietary interconnects can increase dependency on a single vendor ecosystem, while open standards are intended to provide choice and speed industry-wide innovation. Grounded in open Ethernet standards and ecosystems including the Ultra Ethernet Consortium (UEC), Ultra Accelerator Link (UAL), and ESUN, AMD delivers one coherent system across front-end, scale-up, and scale-out, built to carry AI infrastructure through what comes next.

Share:

Contributors


Related Blogs