MoRI UMBP Empowers AMD Instinct™ MI355X GPU to Demonstrate Token-per-Dollar TCO Leadership on the Public AgentX Leaderboard
Oct 01, 2026
Agentic applications run long multi-turn sessions, which makes the cost of serving them depend overwhelmingly on how much KV cache the server can reuse instead of recomputing.
In July, together with Moonshot AI, we introduced SGLang + MoRI UMBP (Unified Memory & Bandwidth Pool) in Rebuilding Agentic AI from First Principles for AMD GPU. MoRI UMBP is a KV cache infrastructure built by the AMD MoRI team from first principles, starting from the agentic workload itself and purpose-built for the AMD platform. Over the past month we have contributed that work to the SGLang community, where it lands in the open source ecosystem with the new KVCache Store Linker, so the whole community benefits. On the public SemiAnalysis AgentX benchmark with DeepSeek-V4-Pro-0813 1.6T, AMD Instinct™ MI355X GPU - running MoRI disaggregation with MoRI UMBP as the unified KV cache pool, on top of the ongoing optimization work by AMD in the SGLang community - now delivers, at its peak, 69M total tokens per $1 TCO, against 46M for NVIDIA B200 running Dynamo SGLang: a 1.5x advantage in tokens per dollar, and outperforms NVIDIA GB200, B300 and GB300 at some interactivity levels (SemiAnalysis, updated 2026-09-25).
The rest of this post covers in detail how UMBP produces that result.
A Quick Recap of the Public AgentX Benchmark
AgentX is a public benchmark built by SemiAnalysis from real-world agentic coding traces, served through the public InferenceX harness and dashboard. It is fast becoming an industry standard for agentic inference evaluation, adopted worldwide by leading teams including OpenAI, Meta, Inferact, RadixArk, MiniMax, Alibaba Qwen, Moonshot AI, Zhipu GLM and Oracle. Its methodology, results and CI are all public, so every number in this post can be checked independently, including by NVIDIA.
AgentX replays whole agent sessions: each turn appends tool results to the accumulated context and re-queries the model. That traffic has four properties chat benchmarks do not produce.
- Long-running, multi-turn sessions - roughly 43 turns per session, around 1M tokens of context traffic over a session's life.
- Long contexts, short outputs - median 142K input tokens against 444 output tokens per turn.
- Extensive prefix reuse - above 96% of prompt tokens are repeats of a prefix the server has already seen.
- Subagents and tool calls, which fan a single user task out into several concurrent sessions sharing a prefix.
It reports TTFT, P90 interactivity (per-user tok/s), TPGS (total tokens per GPU-second, counting cached tokens), and TCO (infrastructure cost per token) — the last of which is what Figure 1 plots.
At 96% prefix reuse, cache management determines prefill cost. And because the agent blocks on the full response, the latency that matters is end-to-end, dominated by decode. AgentX therefore presents three challenges we had to solve:
- The KV cache does not fit in HBM: With many long sessions live at once and subagents fanning out, most reusable KV has already been evicted from GPU memory at moderate concurrency.
- The KV cache is not sharded across TP ranks: Under DeepSeek-V4 sparse attention, the MLA latent and DSA indexer KV are replicated in full on every rank.
- Cache warmth is not free: A host cache that lives inside the engine process is lost on every restart, and on every rolling upgrade.
MoRI UMBP + the SGLang KVCache Store Linker
We addressed these three challenges by pairing MoRI UMBP with the SGLang KVCache Store Linker. This section walks through what we observed on AgentX, how the linker was designed in response, and the results it delivered.
Observations on AgentX
MoRI UMBP was initially integrated into SGLang as a HiCache L3 storage backend. Evaluating this configuration on AgentX, we made six observations:
- Wasted shareable DRAM: The per-rank L2 host cache occupies DRAM that could otherwise serve the shareable L3 backend, shrinking effective shareable capacity.
- Indirect data path: KV moves between L1 HBM and the L3 backend through the L2 host tier, which adds load/offload overhead and rules out a direct L1 ⇔ L3 data path.
- Per-rank local management: Each rank manages its own host cache and makes all KV decisions on local information only, which is globally suboptimal.
- Redundant replication: Under MLA + TP, each TP rank replicates the KV cache and loads/offloads independently, wasting memory and PCIe bandwidth.
- No layer-wise pipelining: The L2→L1 load overlaps with compute layer by layer, but the fetch from the external L3 backend into L2 is not pipelined — it must complete before compute starts.
- Cache dies with the engine: The host cache lives in the engine process, so any restart discards it.
The KVCache Store Linker
We proposed to the SGLang maintainers an option to bypass the L2 host tier entirely with a direct data path between L1 HBM and external KV cache stores — a direction that turned out to align closely with the community's own plans. The MoRI team then co-designed the KVCache Store Linker with the SGLang community, integrating MoRI UMBP as a first-class backend. The linker connects SGLang's unified radix tree straight to the distributed DRAM pool. On a prefix match, prefill pulls KV pages from DRAM instead of recomputing them.
Linker + MoRI UMBP resolves all six issues above:
- Fully shareable DRAM across DP ranks and model instances — enabling DP + round-robin deployments via cross-DP-rank KV cache sharing.
- Direct L1 ⇔ L3 path with no intermediate overhead; on its own, this improves TTFT by up to 13% over running without MoRI UMBP.
- Global KV management — MoRI UMBP places and evicts KV based on global information, more effective than per-rank local policies.
- Deduplication + split load/offload by rank, alleviating memory and PCIe bandwidth pressure. A TP-N prefill now stores and fetches one copy of the replicated MLA/DSA KV instead of N; at TP8, eight keys become one. That multiplies effective DRAM capacity and cuts host traffic by the same factor.
- Layer-wise pipelined loading, with MoRI UMBP hiding the added per-layer request overhead via batching, layer grouping, a ranged API, and an optimized GPU gather kernel for host-to-device KV loading.
- Cache survives engine restarts — in MoRI UMBP standalone mode the KV pool lives in a separate per-node process, so restarts and upgrades reuse it with no warm-up.
Results
These fixes pay off even when the DRAM tier is barely used. In an on/off run with the same recipe (1P1D TP8 + TP8, 16 MI355X GPUs, a 600 GB DRAM KV budget, and the consistent_hashing router for KV cache affinity), HBM already serves about 95.6% of prompt tokens at concurrency 128–256 and the DRAM tier less than 1%. Turning MoRI UMBP on still gives +8.3% throughput per GPU and –35% P90 TTFT at concurrency 256, and +2.7% and –34% at 128 (Figure 2). This gain is architectural, not from offloading, and it maps directly to two of the observations above:
- Wasted shareable DRAM: With MoRI UMBP off, the host KV pool still fills up (72% at concurrency 128, 100% at 256) while serving at most 0.1% of prompt tokens. MoRI UMBP on turns that DRAM into one shareable pool instead.
- Indirect data path: With the DRAM tier nearly idle in both runs, the gain comes from the linker's direct HBM ⇔ DRAM path, which removes the intermediate staging layer and its load/offload overhead.
A large enough cache changes the topology. Once the deduplicated DRAM tier is in place, prefill no longer needs TP8 simply to hold KV. At concurrency 16–48 we run a TP4 prefill with a TP8 decode (12 GPUs) instead of TP8 + TP8 (16 GPUs). MoRI UMBP serves 30%, 52% and 75% of prompt tokens at concurrency 16, 32 and 48, while less than 3% are recomputed. Throughput per GPU rises 24–34% on 25% fewer GPUs.
Further Optimizations
Alongside this, the AMD SGLang team continues to deliver optimizations for DeepSeek-V4-Pro-0813, covering the following areas.
- FP4 sparse-attention indexer: DeepSeek-V4's DSA indexer scores the whole context for every query token at every layer, and carries its own per-token KV. Running it on AITER FP4 kernels on gfx950 cuts that KV from 132 B to 68 B per token — more concurrent sequences in HBM, less indexer bandwidth per decode step, and fewer bytes over the MoRI link.
- Optimistic prefill with request-owned speculative KV : In PD disaggregation a request normally waits for decode to bootstrap it before prefill can start; at high concurrency that handshake is pure queueing time. Letting prefill start optimistically, with the speculative KV owned by the request rather than a pre-reserved decode slot, cut P90 TTFT by 27.7% at concurrency 256. Removing a host sync from DSpark prefill slot expansion cut P90 TTFT a further 13–16% at concurrency 128–256.
- Per-stream split-K for MLA decode picks the split-K factor per index stream instead of applying one setting to layers whose KV lengths differ by orders of magnitude.
Summary
This work produced two results.
- First, against NVIDIA Blackwell systems: On the public AgentX leaderboard, MI355X surpasses GB200 NVL72, B300 and GB300 NVL72 at selected operating points, by up to 5.6× in tokens per dollar (see the table after Figure 1).
- Second, against ourselves a month earlier: Same benchmark, same concurrency of 192:
Peak throughput per GPU also rose from 22.9k to 55.8k (2.4x), at concurrency 256.
In an agentic workload, 96% of prompt tokens have already been computed, so the cost per token is set by the system that manages where those tokens live. MoRI UMBP is that system: it turns DRAM into a deduplicated KV pool that is shareable across instances and survives engine restarts, wired straight into SGLang's radix tree through the KVCache Store Linker.
The first item on the July roadmap was completing the MoRI UMBP integration; that is now delivered. The next is bringing MoRI UMBP to the broader ecosystem e.g. ATOM, vLLM, llm-d.
Acknowledgements
We thank the SGLang community for design reviews and fast upstreaming, SemiAnalysis for the AgentX benchmark and InferenceX CI infrastructure, and the AMD MoRI, SGLang, and AITER teams.
References
- SemiAnalysis - AgentX benchmark and InferenceX dashboard
- vLLM - AgentX
- AMD - Rebuilding Agentic AI from First Principles for AMD GPU, together with Moonshot AI ("What is UMBP")