AI Is No Longer Just a Compute Problem
Jul 23, 2026
IDC projects the scale-out networking market will reach $79 billion by 2030 (Datacenter AI Infrastructure Spend for AI Networking, December 2025) with the total value of AI infrastructure spending projected to surge from $486 billion in 2025 to $1.38 trillion by 2030 (Worldwide Quarterly AI Infrastructure Tracker, April 2026). As AI NIC unit shipments grow at a 40% compound annual growth rate from 2025 to 2030 (Datacenter AI Infrastructure Spend for AI Networking, December 2025), one can see that the network layer is attracting investment at a pace that tracks closely with compute.
These numbers reflect a shift in how the industry is thinking about AI infrastructure. This blog explores five learnings that examine what that shift means in practice, and what organizations need to understand to make infrastructure decisions that hold up at scale.
Learning #1: Network infrastructure decisions should be strategically made alongside compute
AI performance is co-determined by compute and networking. However, GPU utilization may degrade when the network cannot fully sustain the volume and velocity of communication those GPUs require, and at cluster scale, that degradation is not marginal. Organizations that sequence networking decisions after compute procurement tend to find that the inefficiencies introduced early are difficult to address without significant re-engineering downstream.
Teams that treat networking as a secondary procurement decision are not simply leaving performance on the table — they are building a constraint into their infrastructure that compounds as workloads grow and cluster sizes increase.
Learning #2: Legacy Ethernet architectures were designed for a different era of workloads
Traditional three-tier datacenter networks were built around north-south traffic: traffic between users/applications/storage and servers. AI workloads operate differently. Large-scale training often requires tightly synchronized communication between thousands of accelerators simultaneously, generating east-west traffic at a volume and frequency that conventional architectures were never designed to handle.
Fat-tree topologies and RoCEv2 deployments have served as the default approaches for many organizations, but their limitations become increasingly apparent as cluster sizes grow. Congestion, packet loss, and latency are structural outcomes of applying a general-purpose network architecture to a workload with fundamentally different requirements. The inefficiencies compound as scale increases.
IDC identifies a clear industry shift towards next-gen open Ethernet solutions that reduce reliance on single-vendor architectures and address certain scalability and operational challenges with some traditional solutions. However, not all Ethernet approaches are equivalently suited to AI workloads, and the architectural decisions made within that shift carry significant performance implications.
Learning #3: Networking decisions underlie the hidden costs of AI infrastructure.
Hardware costs are often the most visible line items in AI infrastructure planning — compute, networking equipment, and facility buildout are relatively straightforward to model. The costs that are harder to quantify upfront are operational: how long it takes to bring a cluster online, how much manual tuning is required before the network performs reliably, and how much engineering time is consumed managing infrastructure failures and maintaining consistency across thousands of nodes.
Cluster downtime carries a cost that often scales directly with cluster size. Upgrades, maintenance windows, and fault recovery events that are manageable at small scale may become significant issues impacting productivity as deployments grow. Architectural decisions made at the outset — around resiliency, traffic distribution, and job completion efficiency — determine how often those events occur and how quickly the network recovers from them.
Multiplane architectures offer a concrete illustration of how architectural choice translates into cost outcomes. IDC estimates that multiplane designs can reduce networking infrastructure costs primarily through reduced optics requirements. The organizations best positioned for AI at scale are those that account for total cost of ownership, not just initial procurement, when evaluating their networking architecture.
Learning #4: Proprietary networking stacks may introduce long-term risk that open, standards-based architectures are designed to avoid
AI networking standards are still actively maturing. Transport protocols, congestion control mechanisms, and lossless fabric specifications are areas of ongoing development across the industry, and the landscape today is not where it will be in the future. Organizations making architectural commitments now are doing so in an environment where the standards themselves have not yet fully settled.
Committing to a fully proprietary stack in this environment transfers a meaningful degree of infrastructure control to a single vendor's roadmap. The ability to adopt best-of-breed components as the market evolves becomes constrained, and switching costs increase as requirements change and the gap between what the stack supports and what the workload needs widens. These constraints carry strategic implications beyond performance, affecting long-term infrastructure flexibility.
Open networking approaches address this directly by enabling interoperability across vendors and infrastructure components. Emerging specifications such as UEC-ready RDMA and MRC (Multipath Reliable Connection) are beginning to address the limitations of legacy protocols at scale, and standards-based solutions are positioned to adopt them as they mature. Programmable architectures extend this further, allowing transport protocols and congestion control mechanisms to be updated via software rather than requiring hardware replacement as the industry evolves.
Learning #5: Operational excellence and time to production are emerging as the defining selection criteria for AI networking
A network's peak performance specifications are a starting point, not a complete picture. At scale, what determines infrastructure value is often how the network behaves across the full operational lifecycle — through upgrades, failures, configuration changes, and the routine maintenance that large cluster deployments make unavoidable. Networks that require significant intervention to remain performant may impose a continuous operational cost that rarely surfaces in pre-deployment planning.
High-resolution telemetry and granular observability change the nature of that operational burden. When issues can be identified faster, across logical interfaces and transport layers, teams move from reactive troubleshooting to proactive management. The difference in mean time to resolution at cluster scale is not incremental — it has direct implications for GPU utilization and job completion rates.
Out-of-box performance is increasingly a selection criterion for the same reason. Infrastructure that delivers reliable baseline performance from deployment can shorten the path to business value. Architectures that require extensive manual tuning before they perform predictably delay it — and at the pace AI infrastructure is scaling, that delay is a cost few organizations are positioned to absorb.
The Network Is Where AI Strategy Gets Decided
The organizations that get AI infrastructure right in the upcoming years will not simply be the ones that procured the most compute. They will be the ones who understood, early, that compute without the right network is underutilized — and made infrastructure decisions accordingly.
The networking layer is where performance, cost, adaptability, and operational efficiency converge — and the decisions made at that layer will define the ceiling of what AI deployments can achieve. AMD AI NIC™ technology, featured as a vendor spotlight, is built around the architectural principles this research identifies as decisive: programmability, open standards, and operational resilience.
Read the full IDC spotlight paper, Enabling AI-Ready Scale-Out Networking for AI Workloads, for a deeper analysis of the trends shaping AI network design and a detailed look at how leading organizations are building infrastructure that scales without compromise.
FAQs
What is a proprietary networking stack?
A proprietary networking stack is a closed set of networking hardware, software, and protocols controlled by a single vendor. Unlike open, standards-based approaches, proprietary stacks limit interoperability with other vendors.
Why does programmability matter?
Programmable architectures allow transport protocols and congestion control to be updated via software as standards evolve, without requiring hardware replacement.
What are Ultra Ethernet Consortium (UEC) and Multipath Reliable Connection (MRC)?
UEC and MRC are emerging open networking specifications, with both being part of a broader industry effort to establish standards for high-performance Ethernet transport in large GPU cluster deployments.
What does "lossless networking" mean in the context of AI?
Lossless networking helps ensure that packets are not dropped during transmission, which is critical for AI training workloads where packet loss causes GPU stalls and directly impacts job completion times.