How Agentic AI Is Expanding the Role of the DPU

Oct 05, 2026

With every generation of AI workloads, there has come new challenges, requirements and standards setting the stage for new infrastructure requirements. Training demands large scale-out GPU clusters with hundreds of thousands of GPUs and the high-bandwidth networks to connect them. Inference adds latency sensitivity with high request volumes, placing a large amount of pressure on the front-end network services required to route and balance traffic across model deployments. Now, agentic AI is now driving another shift, introducing a new level of complexity across infrastructure.

While inference is largely stateless, agentic workflows are iterative and persistent. A single user request can trigger a chain of model calls, tool invocations, and retrieval operations. As agentic deployments scale, the volume of state (in the form of context) that must be created, stored, and moved grows with the workloads.

The result is more front-end infrastructure traffic, greater data movement between accelerator memory and storage, and growing pressure on CPU resources. At the same time those resources are in higher demand than ever to orchestrate agentic AI workflows. This is where the data processing unit, or DPU, becomes more consequential as it offloads critical infrastructure services from the CPU. Hyperscalers deploy DPUs at the front-end of the network to create a dedicated infrastructure processing domain, separate from application workloads, for networking, security, storage, and telemetry. This maintains tenant isolation, improves efficiency, and preserves CPU resources for the workloads that need them.

Agentic AI is expanding the DPU’s role in two concrete ways: 

1. The front-end network demand is growing faster than organizations can afford to service it with CPUs; and 

2. DPUs are becoming central to how AI workload data moves between accelerator memory and disaggregated storage across the network.

1. The Infrastructure Tax on Agentic AI

According to Gartner, agentic AI consumes 5 to 30 times more tokens per task than a standard chatbot query, and that increase in demand flows directly into the network.1 The front-end networking environment this creates looks fundamentally different from what inference alone required: more connections, more load balancing decisions, and more security policy enforcement. All of which requires infrastructure processing that has traditionally consumed CPU cycles. The same cores are now in high demand for agent orchestration, application logic, and data retrieval. Diverting those cores to service networking overhead is a tradeoff that grows more costly as agent deployments scale.

The value of the DPU for agentic AI has already been established at hyperscale. In one customer deployment, DPU offloading reclaimed 22 CPU cores per server previously consumed by a single infrastructure service.2 As agentic workloads scale and front-end traffic grows, that reclaimed capacity becomes more consequential, not just for efficiency, but for the ability to serve more agents and users at production scale.

2. When Context Becomes the Bottleneck

Agentic workloads generate and depend on a growing amount of context and serving that context efficiently has become one of the harder infrastructure problems to solve. Recent research points to a striking consequence: in agentic inference, as context grows, the amount of KV-cache storage in memory increases, which becomes a big constraint on GPUs. The bottleneck has moved.

GPU memory is fast, but finite when trying to keep every agent's complete context permanently in local HBM, which is not practical at scale. The industry is moving toward hierarchical memory architectures, with active context in HBM and colder KV cache in lower-cost memory and flash storage. NVMe-over-Fabrics makes this possible by allowing storage to become a disaggregated, shared pool accessible across the network with low latency. The DPU sits in the middle of that, handling transport, data movement, encryption, and tenant isolation, so the CPU and GPU do not have to. The compute stays focused on running agents and inference, and the DPU handles the inbound data ingestion, offload and server load balancing. 

This is where the DPU's role expands beyond offload. In an agentic AI system, the DPU effectively becomes the infrastructure processor responsible for managing context across that hierarchy — offloading KV-cache orchestration, storage protocol management, and data movement roles from the CPU, so it can stay focused on the inference and agent workloads that actually matter.

The larger opportunity is not simply lower storage latency; it is an entirely different scale of AI infrastructure. By storing, sharing, and retrieving context from a distributed storage tier, AI providers can support far more concurrent agents than local GPU memory alone would allow. Context can be reused across sessions and shared across inference engines, which means less recomputation and better use of the GPUs running the work. 

amd-salina-dpu

Figure 1 - AMD Pensando™ Salina DPU

The AMD Pensando™ Salina DPU is built for exactly this role. It supports KV-cache storage for contexts exceeding one million tokens across hundreds of thousands of concurrent sessions. It's one of the only fully P4-programmable DPU on the market to deliver concurrent networking, security, and storage services3. — This allows the Salina DPU to uniquely give cloud operators the flexibility to adapt as context requirements and agentic architectures continue to evolve.

Building for What Comes Next

Agentic AI has brought about massive industry shifts in a short period of time; however, these architectures are still evolving. As a result, traffic patterns, storage requirements, and security policies will continue to shift as deployments mature. A fully programmable DPU lets operators introduce new services, adapt newer data paths, and implement new protocols as workloads change. That programmability extends the useful life of the infrastructure without replacing it.

What started as a critical cloud networking component has grown into something more fundamental for AI infrastructure in the agentic AI era. In a rack-scale deployment, a single programmable DPU can handle front-end network traffic, the connections between GPU nodes, and the storage path that extends context beyond GPU memory. One device, managing the flow of information across the whole system. This is the direction AMD is building toward.

As AI becomes more distributed, stateful, and multi-agent, the infrastructure that manages information flow will have as much influence on system performance as the compute that processes it. The DPU's programmable architecture puts it at the center of that infrastructure — not as a peripheral, but as a principal.

Footnotes
  1. 22 CPU - 22 CPU Core Savings with Accelerated Connections running on AMD DPU at Microsoft. Core Savings for Customer SLB Appliances running on Sirius Appliance at Microsoft as part of Accelerated Connections Service vs Running SLB on Standard Server 
    https://www.ciscolive.com/c/dam/r/ciscolive/global-event/docs/2024/pdf/CSSSPG-1015.pdf

  2. Gartner Predicts That by 2030, Performing Inference on an LLM With 1 Trillion Parameters Will Cost …

  3. https://www.amd.com/content/dam/amd/en/documents/pensando-technical-docs/product-briefs/pensando-salina-product-brief.pdf

Share:

Article By


Corporate Vice President, Business and Product Management Network Technology Solutions Group

Related Blogs