What Is AI Networking?

AI networking is the high-performance fabric that powers AI at scale. It connects GPUs, accelerators, and systems, delivering the throughput, synchronization, and performance consistency that training, inference, and agentic AI workloads demand as they grow.

The Critical Reality

As AI workloads grow in scale and complexity, system performance becomes increasingly dependent on networking. From front-end networking to scale-up and scale-out fabrics, each layer shapes how efficiently AI infrastructure performs.

Performance at Scale Starts with Networking

Data center with glowing lights

Performant

Frontier AI models require a network that moves data to and between accelerators with the high bandwidth and low latency to keep them fully utilized.

Abstract teal and orange light streaks overlaid with scrolling source code, suggesting data flow or AI processing

Programmable

AI protocols move fast. A programmable network keeps infrastructure flexible and ready for what comes next.

Teal and orange network lines and nodes connected across a dark textured surface, symbolizing data connectivity

Open

Open systems put choice in organizations' hands, enabling flexible infrastructure builds and better cost optimization at every scale.

The 7 Architectural Principles of Scalable AI Networking

Discover how strategic networking choices impact overall AI performance, infrastructure costs, scalability, and future growth.

needed

Front-End: Feeds the AI Factory

Front-end networking connects users and data to AI infrastructure for reliable ingestion at scale.

Delivers concurrent acceleration of storage, security, and networking services while preserving compute resources for AI workloads.

Key Technology:

Scale-Up: Unified AI Systems

UALoE and tightly coupled software enables a scale-up domain where many GPUs act as one compute resource, with high-bandwidth networking supporting tensor parallelism, large models, and resilient, efficient workloads.

Unifies 72 GPUs and provides resilient networking with alternate connections, automated fault response, and workload isolation for continuous operation.

Key Technology:

Scale-Out & Across: Distributed AI

Scale-out networking delivers the bandwidth, efficiency, and resiliency for training and inference at the largest scale to be distributed across multiple data centers.

Accelerates distributed AI across racks, data centers, and regions with predictable performance, rapid recovery, and lower infrastructure cost.

Key Technology:

Use Cases

Model Training

Challenge

Training billion-parameter models demands precise, low-latency synchronization across hundreds of thousands of GPUs.

Solution

Ultra-low latency collective communications accelerate training and improve cluster utilization.

Radiating teal and orange light streaks with sparkling particles, suggesting high-speed data transfer

Distributed Inference

Challenge

Responding to millions of inference requests in real time demands consistent, low-latency performance under unpredictable demand.

Solution

Low-latency, deterministic networking keeps inference responses fast and consistent, even as demand fluctuates.

Abstract background

Agentic AI

Challenge

Agentic AI workloads rely on continuous coordination between multiple complex services.

Solution

Accelerated infrastructure services enable efficient execution, low-latency communication, and improved utilization across large-scale AI workflows.

Teal and orange particle mesh forming a flowing wave over blurred cityscape lights

Importance of Scale-Up Systems for Production AI

As AI scales beyond individual nodes, scale-up networking ties many GPUs into one compute resource and provides the high availability that helps keep training and agentic workloads running through network failures.

These AI workloads are often sensitive to restarts, so scale-up networking must provide the resilience to quickly recover the network from disruptions without stalling active jobs or losing accumulated progress.

AMD Portfolio for AI Networking

AMD delivers a broad portfolio of AI networking technologies designed to meet the performance, scalability, and resiliency demands of modern AI infrastructure.

AMD AI Networking Software Stack

With a multi-generational mature software stack, the Helios Management Software enables intelligent fabric health operations through health monitoring, telemetry, RAS and automation.

Abstract blue digital circuit pathway with glowing light source and pixel-like data patterns

AMD Pensando™ DPUs

Networking, security, and storage services run concurrently to enable deterministic performance and efficient resource utilization across cloud and AI deployments.

AMD Helios Rackscale

AMD Helios Rackscale Solution: Leadership Rack Performance for Hyperscale AI1

The AMD Helios Rackscale solution design is a fully integrated AI infrastructure, combining the latest AMD Instinct™ GPUs, AMD EPYC™ Server CPUs, and AMD Pensando™ networking, designed using open industry standards enabling large-scale inference, frontier model training and fine-tuning.

Open Ecosystem & Standards

AMD collaborates with leading industry organizations to help advance AI.

Open Compute Project logo
Ultra Accelerator Link logo
Ultra Ethernet Consortium logo

FAQ

Frequently Asked Questions

AI networking is optimized for the synchronized, high-volume GPU communication that runs across scale-up and scale-out architectures, and provides additional intelligent congestion management, fast recovery, and deep observability that traditional workloads don't require.

An AI NIC adds programmable logic and dedicated acceleration that optimizes GPU-to-GPU communication, helps reduce latency, and enhances reliability at scale.

Scale-up refers to maximizing performance within nodes or tightly coupled systems. Scale-out involves expanding across clusters and pods within a data center. Scale-across enables unified management and orchestration across multiple, geographically dispersed data centers so they operate as one unified system.

AMD builds on proven Ethernet standards and maintains extensive partner ecosystem relationships to help ensure interoperability while allowing vendor flexibility.

UALoE (Ultra Accelerator Link over Ethernet) is used in rack-scale systems for scale-up network fabrics. It enables many GPUs to operate as one unified system. It delivers the performance of tightly coupled GPUs while keeping the reliability and openness of production AI requires, all built on Ethernet standards for future-ready flexibility.

The AMD software stack includes AMD Cluster Controller (ACC) for automated Day 0/Day 2 operations, AMD Fabric Manager (AFM) for GPU partitioning and isolation, and AMD Fabric Operating System (AFOS) for reliable switch operations. Together, they help eliminate manual configuration and make large-scale AI clusters practical to operate in production.

Contact Us

Contact an AMD sales representative.

Footnotes
  1. Calculations by AMD Performance Labs in June 2025, based on the projected memory capacity/ bandwidth and scale up/out bandwidth specifications of AMD Instinct™ MI455X 72xGPU “Helios” AI Rack vs. the publicly announced NVIDIA “Vera Rubin” 72xGPU “Oberon” Rack. Server manufacturers may vary configurations, yielding different results. MI350-045A
    Calculations by AMD Performance Labs in September 2025, based on the FP8/FP4 datatypes and the projected specifications for AMD Instinct™ MI455X 72xGPU “Helios” AI Rack vs. Publicly announced specs for the NVIDIA “Vera Rubin” 72xGPU “Oberon” AI Rack. Actual results based on production silicon may vary. Server manufacturers may vary configurations, yielding different results. MI350-046B