AMD Pensando™ DPUs
Networking, security, and storage services run concurrently to enable deterministic performance and efficient resource utilization across cloud and AI deployments.
Explore how AMD networking technology boosts AI performance
AI networking is the high-performance fabric that powers AI at scale. It connects GPUs, accelerators, and systems, delivering the throughput, synchronization, and performance consistency that training, inference, and agentic AI workloads demand as they grow.
Frontier AI models require a network that moves data to and between accelerators with the high bandwidth and low latency to keep them fully utilized.
AI protocols move fast. A programmable network keeps infrastructure flexible and ready for what comes next.
Open systems put choice in organizations' hands, enabling flexible infrastructure builds and better cost optimization at every scale.
Discover how strategic networking choices impact overall AI performance, infrastructure costs, scalability, and future growth.
Front-end networking connects users and data to AI infrastructure for reliable ingestion at scale.
Delivers concurrent acceleration of storage, security, and networking services while preserving compute resources for AI workloads.
Key Technology:
UALoE and tightly coupled software enables a scale-up domain where many GPUs act as one compute resource, with high-bandwidth networking supporting tensor parallelism, large models, and resilient, efficient workloads.
Unifies 72 GPUs and provides resilient networking with alternate connections, automated fault response, and workload isolation for continuous operation.
Key Technology:
Scale-out networking delivers the bandwidth, efficiency, and resiliency for training and inference at the largest scale to be distributed across multiple data centers.
Accelerates distributed AI across racks, data centers, and regions with predictable performance, rapid recovery, and lower infrastructure cost.
Key Technology:
Model Training
Training billion-parameter models demands precise, low-latency synchronization across hundreds of thousands of GPUs.
Ultra-low latency collective communications accelerate training and improve cluster utilization.
Distributed Inference
Responding to millions of inference requests in real time demands consistent, low-latency performance under unpredictable demand.
Low-latency, deterministic networking keeps inference responses fast and consistent, even as demand fluctuates.
Agentic AI
Agentic AI workloads rely on continuous coordination between multiple complex services.
Accelerated infrastructure services enable efficient execution, low-latency communication, and improved utilization across large-scale AI workflows.
As AI scales beyond individual nodes, scale-up networking ties many GPUs into one compute resource and provides the high availability that helps keep training and agentic workloads running through network failures.
These AI workloads are often sensitive to restarts, so scale-up networking must provide the resilience to quickly recover the network from disruptions without stalling active jobs or losing accumulated progress.
AMD delivers a broad portfolio of AI networking technologies designed to meet the performance, scalability, and resiliency demands of modern AI infrastructure.
With a multi-generational mature software stack, the Helios Management Software enables intelligent fabric health operations through health monitoring, telemetry, RAS and automation.
The AMD Helios Rackscale solution design is a fully integrated AI infrastructure, combining the latest AMD Instinct™ GPUs, AMD EPYC™ Server CPUs, and AMD Pensando™ networking, designed using open industry standards enabling large-scale inference, frontier model training and fine-tuning.
AMD collaborates with leading industry organizations to help advance AI.
FAQ
AI networking is optimized for the synchronized, high-volume GPU communication that runs across scale-up and scale-out architectures, and provides additional intelligent congestion management, fast recovery, and deep observability that traditional workloads don't require.
An AI NIC adds programmable logic and dedicated acceleration that optimizes GPU-to-GPU communication, helps reduce latency, and enhances reliability at scale.
Scale-up refers to maximizing performance within nodes or tightly coupled systems. Scale-out involves expanding across clusters and pods within a data center. Scale-across enables unified management and orchestration across multiple, geographically dispersed data centers so they operate as one unified system.
AMD builds on proven Ethernet standards and maintains extensive partner ecosystem relationships to help ensure interoperability while allowing vendor flexibility.
UALoE (Ultra Accelerator Link over Ethernet) is used in rack-scale systems for scale-up network fabrics. It enables many GPUs to operate as one unified system. It delivers the performance of tightly coupled GPUs while keeping the reliability and openness of production AI requires, all built on Ethernet standards for future-ready flexibility.
The AMD software stack includes AMD Cluster Controller (ACC) for automated Day 0/Day 2 operations, AMD Fabric Manager (AFM) for GPU partitioning and isolation, and AMD Fabric Operating System (AFOS) for reliable switch operations. Together, they help eliminate manual configuration and make large-scale AI clusters practical to operate in production.
Contact an AMD sales representative.