AMD Pensando™ Vulcano 800 AI NIC: Built to Scale-Out and Across
Jul 23, 2026
Changing Workload Demands
As AI models grow in scale and complexity, networking plays an increasingly larger role in determining whether a cluster delivers its theoretical compute capacity. Multi-node training runs and distributed inference deployments are now standard practice, yet the networking fabrics underpinning most clusters were not designed for this reality. The result is significant idle compute time as GPUs stall waiting on data, directly translating to higher operational cost and slower iteration cycles. Addressing this requires rethinking how clusters are connected, with the AI NIC becoming a first-order consideration in infrastructure decision-making.
The AMD Pensando™ Vulcano 800 AI NIC is purpose-built for clusters operating at this scale. It delivers up to 2.4 Tbps of scale-out bandwidth per GPU, via 800 Gbps connectivity per AI NIC and a unique 3 NICs-per-GPU configuration capability, providing the throughput headroom that training and distributed inference workloads require. Rather than treating high bandwidth as a single-dimension metric, the AI NIC is designed around two distinct networking constructs that together determine real-world cluster efficiency: scale-out and scale-across.
Scaling Out and Across
Scale-out and scale-across describe two distinct but complementary connectivity requirements that together define how well a cluster handles modern AI workloads.
Scale-out refers to the ability to distribute a workload across multiple nodes, expanding the effective GPU pool beyond what a single server can provide. For large model training, this means coordinating gradient synchronization across dozens or hundreds of nodes with minimal stall time. For inference, it means routing requests across a fleet of serving instances without the network becoming the limiting factor. At this layer, bandwidth and latency between nodes directly determine whether adding more hardware translates to proportional throughput gains, or whether returns diminish as communication overhead grows.
Scale-across extends beyond the rack: as AI deployments grow, clusters increasingly span multiple data centers. In these architectures, traffic between data centers carries the same performance-critical communication that intra-rack traffic does within a single facility. Latency and bandwidth between sites therefore have the same consequence as within them: degraded throughput, longer job completion times, and underutilized GPU capacity. The AMD Pensando™ Vulcano 800 AI NIC is designed to maintain high-bandwidth, low-latency connectivity across this topology, enabling scale-across networking to operate effectively at inter-data center distances without requiring separate, purpose-built interconnects for each layer of the cluster hierarchy.
Together, these two patterns cover the full topology of a modern AI cluster. The AMD Pensando Vulcano 800 AI NIC is designed to address both simultaneously, which is the architectural requirement that disaggregated inference and multi-node training have made non-negotiable.
Performance
Performance in AI clusters is not determined by peak throughput alone — job completion rates, fault tolerance, and sustained utilization under real-world conditions are equally consequential. The AMD Pensando™ Vulcano 800 AI NIC is designed to improve across all three.
The AMD Pensando Vulcano 800 AI NIC delivers up to 13% faster AI job completion times.1 This figure reflects the potential for reduction in retransmits, stalls, and recovery cycles that degrade job completion times in less capable fabrics. At scale, that improvement compounds: faster job completion can translate directly to higher cluster utilization and more training iterations per unit of operational cost.
Underpinning this is a set of RAS (Reliability, Availability, Serviceability) capabilities designed to keep clusters productive under real-world conditions. Multi-plane architecture distributes traffic across independent, separate data paths, so a single link or component failure degrades gracefully rather than disrupting the full job. Fault isolation, in-service diagnostics, and rapid recovery mechanisms reduce mean time to repair without requiring cluster-wide maintenance windows. The cumulative effect is to improve cluster uptime and maintenance of AI infrastructure’s production-readiness.
Programmability
Workload requirements in AI are shifting faster than traditional silicon development cycles allow, with new model architectures and communication primitives emerging within months. The AMD Pensando™ Vulcano 800 AI NIC addresses this directly through a fully programmable software and hardware stack, enabling protocol behavior, congestion control, and transport logic to be updated in software.
That programmability is grounded in a deliberate architectural commitment to open-standards based networking. AMD was first to market with Ultra Ethernet Consortium (UEC) support in its previous generation of AI NIC, establishing early leadership in defining next-generation Ethernet for AI workloads. The AMD Pensando Vulcano 800 AI NIC carries that forward with leadership support for MRC (Multipath Reliable Connection), a transport protocol co-developed for OpenAI to address the reliability and throughput requirements of large-scale training environments.
Openness
Clusters built on open, interoperable standards retain the ability to preserve vendor choice and maintain flexibility as the market evolves. The AMD Pensando™ Vulcano 800 AI NIC is designed for this environment, operating within the broader Ethernet ecosystem rather than requiring a vertically integrated stack to deliver its full capability.
AI workload requirements are not stabilizing — distributed inference, multi-node training, and new architectures have each become standard practice within a short window, and the trajectory points to continued evolution. For clusters built on proprietary networking fabrics, that pace of change creates a compounding problem: every new requirement either fits within what a single vendor supports, or it forces a difficult choice between operational workarounds and full fabric replacement. Open, standards-based infrastructure avoids that constraint entirely. Because the AMD Pensando Vulcano 800 AI NIC operates natively within the Ethernet ecosystem, operators can integrate new hardware generations and reconfigure cluster topology as demands shift, without being gated by a vendor's roadmap. The window to build infrastructure that can keep pace with workload change is narrowing, and the choice of networking fabric is central to it.
That baseline efficiency extends further through specific design choices in the AI NIC. Its multi-plane architecture distributes traffic across independent data paths, delivering fault tolerance and throughput with less infrastructure than a conventional single-plane fabric would demand. The AMD Pensando Vulcano 800 AI NIC achieves up to 33% lower switching costs2 through reduced cable and transceiver usage.
The AMD AI NIC™: Built for the Pace of AI
The constraints shaping AI infrastructure today have put networking at the forefront, from a supporting consideration to a foundational one. Decisions made at the fabric layer now have direct consequences for cluster utilization, operational cost, and the ability to adapt as requirements shift.
The AMD Pensando™ Vulcano 800 AI NIC is built with that reality in mind. Across scale-out and scale-across connectivity, job completion performance, operational resilience, programmability, and total cost of ownership, it is designed to address the specific demands that large-scale AI deployments place on networking infrastructure. As the default networking of choice for the AMD Helios™ rack-scale solution, the AMD AI NIC technology represents a purpose-built answer to the connectivity challenges that modern AI workloads have made unavoidable.
Explore AMD Pensando™ AI NIC technology.
Related Blogs
Footnotes
- MI400-018: Based on AMD Engineering silicon modeling and AMD synthetic benchmark simulation, using the MOE_4p5T hi_sparsity_ea9 benchmark test to project time- to- solution speed up in days using the FP8 training datatype on a simulated LLM system modeled with 8000 AMD Instinct MI455X GPUs and (3) versus (2) AMD Pensando Vulcano NICs.
Results reflect analysis of pure training compute (decode layers only) Configuration evaluated with a global batch size of 4096 and a sequence length of 4K, using SwiGLU activation. Results exclude evaluation and checkpointing overhead. FlashAttention v3 with matmul–softmax overlap is assumed. AllReduce and All2All communication costs are fully accounted for and not hidden via tiled compute–communication overlap. Gradient synchronization and FSDP weight prefetch are evaluated across varying levels of overlap to assess scale-out sensitivity. Results assume ideal and not fully optimized real-world behavior and may vary when actual product(s) are released in market. MI400-018
- PEN-022: AMD comparison and pricing as of May 18, 2026, for network fabric costs to support 32,000 GPUs. Comparison of a Vulcano-based NIC (VULCANO-CUSTOM-2.4T) deployed as part of a Helios rackscale system with a network based on 1.6T Tomahawk 6 switching with 200G SerDes versus using a competitor 800G NIC with 800G Tomahawk 6 switching with 100G SerDes. Both fabrics were fat-tree topologies built on Tomahawk 5 800G switching platforms, with NIC costs considered comparable. The Vulcano-based design is estimated to deliver up to 33% savings in network switching costs by enabling a more cost-effective architecture with fewer switching platforms, more bandwidth per port on the network, and reduced transceiver cables/optics.
TH6-100G Serdes Fat-Tree (Competition):
Switching (TH6C BCM78914 - 128x800G):
• Leaf Units 1,000
• Spine Units 500
• Total Switches 1,500
• TH6 100G Unit Price $79,587
• Total Switching Cost $60M
Cables/Optics:
• NIC Transceivers (800G-DR8) 64,000 @ $500
• Leaf/Spine Transceivers (800G-DR8) 192,000 @ $500
• MPO Cables 256,000 @ $89
• Total Optics/Cables $64M
Total Fabric Cost TH6-100G (Switches + Cables/Optics): $124M
TH6-200G Serdes Fat-Tree (Vulcano-Custom2.4T / AMD Solution):
Switching (TH6P BCM78910 - 64x1.6T):
• Leaf Units 1,000
• Spine Units 500
• Total Switches 1,500
• TH6 200G Unit Price $66,336
• Total Switching Cost $50M
Cables/Optics:
NIC Transceivers (1.6T-DR8) 32,000 @ $900 (50% fewer vs. competition)
Leaf/Spine Transceivers (1.6T-DR8) 64,000 @ $900
MPO Cables 128,000 @ $89 (50% fewer vs. competition)
Total Optics/Cables $43M
Total Fabric Cost TH6-200G (Switches + Cables/Optics): $93M
Capex Savings (Fabric only):
Savings $: $30.7M
Savings %: 33.1%
Pricing sources: SemiAnalysis Hyperscaler Networking Model data used with permission; full analysis available via SemiAnalysis subscription. Edgecore switch pricing as of May 17, 2026. Results may vary based on system configuration.
- MI400-018: Based on AMD Engineering silicon modeling and AMD synthetic benchmark simulation, using the MOE_4p5T hi_sparsity_ea9 benchmark test to project time- to- solution speed up in days using the FP8 training datatype on a simulated LLM system modeled with 8000 AMD Instinct MI455X GPUs and (3) versus (2) AMD Pensando Vulcano NICs.
Results reflect analysis of pure training compute (decode layers only) Configuration evaluated with a global batch size of 4096 and a sequence length of 4K, using SwiGLU activation. Results exclude evaluation and checkpointing overhead. FlashAttention v3 with matmul–softmax overlap is assumed. AllReduce and All2All communication costs are fully accounted for and not hidden via tiled compute–communication overlap. Gradient synchronization and FSDP weight prefetch are evaluated across varying levels of overlap to assess scale-out sensitivity. Results assume ideal and not fully optimized real-world behavior and may vary when actual product(s) are released in market. MI400-018 - PEN-022: AMD comparison and pricing as of May 18, 2026, for network fabric costs to support 32,000 GPUs. Comparison of a Vulcano-based NIC (VULCANO-CUSTOM-2.4T) deployed as part of a Helios rackscale system with a network based on 1.6T Tomahawk 6 switching with 200G SerDes versus using a competitor 800G NIC with 800G Tomahawk 6 switching with 100G SerDes. Both fabrics were fat-tree topologies built on Tomahawk 5 800G switching platforms, with NIC costs considered comparable. The Vulcano-based design is estimated to deliver up to 33% savings in network switching costs by enabling a more cost-effective architecture with fewer switching platforms, more bandwidth per port on the network, and reduced transceiver cables/optics.
TH6-100G Serdes Fat-Tree (Competition):
Switching (TH6C BCM78914 - 128x800G):
• Leaf Units 1,000
• Spine Units 500
• Total Switches 1,500
• TH6 100G Unit Price $79,587
• Total Switching Cost $60M
Cables/Optics:
• NIC Transceivers (800G-DR8) 64,000 @ $500
• Leaf/Spine Transceivers (800G-DR8) 192,000 @ $500
• MPO Cables 256,000 @ $89
• Total Optics/Cables $64M
Total Fabric Cost TH6-100G (Switches + Cables/Optics): $124M
TH6-200G Serdes Fat-Tree (Vulcano-Custom2.4T / AMD Solution):
Switching (TH6P BCM78910 - 64x1.6T):
• Leaf Units 1,000
• Spine Units 500
• Total Switches 1,500
• TH6 200G Unit Price $66,336
• Total Switching Cost $50M
Cables/Optics:
NIC Transceivers (1.6T-DR8) 32,000 @ $900 (50% fewer vs. competition)
Leaf/Spine Transceivers (1.6T-DR8) 64,000 @ $900
MPO Cables 128,000 @ $89 (50% fewer vs. competition)
Total Optics/Cables $43M
Total Fabric Cost TH6-200G (Switches + Cables/Optics): $93M
Capex Savings (Fabric only):
Savings $: $30.7M
Savings %: 33.1%
Pricing sources: SemiAnalysis Hyperscaler Networking Model data used with permission; full analysis available via SemiAnalysis subscription. Edgecore switch pricing as of May 17, 2026. Results may vary based on system configuration.