Agentic AI: AMD EPYC™ 9005 CPUs Wins Today, EPYC 9006, formerly code-named “Venice”, Takes It to a New Level

Jul 23, 2026

Abstract background and Data Center

6th Generation AMD EPYC™ “Venice” Server CPUs mark another major step forward for data center CPU architecture and the next phase of enterprise AI infrastructure. Built on the “Zen 6” core architecture and TSMC’s advanced 2 nm process technology, AMD EPYC 9006 Series Server CPUs are designed to deliver greater compute density, expanded memory bandwidth, and next-generation I/O for cloud, enterprise, HPC, database, and AI workloads. With up to 256 cores and 512 threads per socket, up to 16 channels of DDR5 memory, JEDEC-standard MRDIMM support up to 12,800 MT/s, and PCIe® Gen 6 connectivity, AMD EPYC 9006 Series Server CPUs provides a powerful platform foundation for the next generation of AI workloads, especially emerging agentic AI systems.

As shown in Figure 1, 5th Gen AMD EPYC Server CPUs already deliver leadership performance across the diverse workloads that define modern agentic AI deployments.1,2,3,4,5 AMD EPYC 9006 Series Server CPUs builds on that leadership with additional performance, scale, and platform capability, enabling enterprises to support increasingly complex AI pipelines.1,2,3,4,5 As agentic AI continues to evolve, infrastructure demands are expected to grow rapidly with increased demands for concurrency, orchestration complexity, retrieval volume, and tool execution. With 6th Gen AMD EPYC Server CPUs designed for agentic AI, industry leading memory capacity, and platform throughput, AMD EPYC 9006 Series Server CPUs helps future-ready AI infrastructure investments by providing the performance headroom and scalability needed for the next generation of enterprise agentic AI applications. 

Figure 1: Agentic AI Workload Performance (CPU Centric)
Figure 1: Agentic AI Workload Performance (CPU Centric)

To understand why this matters, it is useful to examine how agentic AI changes the performance equation. Unlike traditional single-shot inference workloads, agentic AI systems execute complex end-to-end workflows: characterizing user requests, assembling context, planning and routing actions, retrieving enterprise knowledge, querying databases, invoking APIs, parsing documents, executing transient tools, verifying results, and streaming grounded responses. Overall performance therefore depends on the combined efficiency of computing, memory, storage, networking, orchestration software, and runtime isolation. In this environment, the CPU plays a pivotal role, not merely as a supporting component for accelerators, but as the engine that drives orchestration, retrieval, tool execution, and response generation across the production agentic AI pipeline.

To evaluate this emerging class of workloads, we worked closely with customers and industry partners to identify representative applications and develop benchmarks that reflect real-world agentic AI deployments. As shown in Figure 2, the framework decomposes the agentic AI workflow into distinct stages: Gateway, Assemble Context, Plan and Route, Retrieve Context and Similarity Search, Reasoning, Enterprise Tools, Ephemeral Tools, Verification, and Stream Response. Together, these stages capture the end-to-end execution path of modern AI agents and enable a comprehensive assessment of platform performance across orchestration, retrieval, reasoning, tool execution, and response generation as shown in Figure 3.

Agentic AI continuously moves through compute, memory, I/O, security, and acceleration-intensive stages.
Figure 2. Agentic AI distinct compute stages
Figure 2. Agentic AI distinct compute stages
Figure 3. Agentic AI execution pipeline
Figure 3. Agentic AI execution pipeline

Table 1 summarizes the benchmark workloads and their corresponding characteristics that are representative of the major stages of the agentic AI pipeline as shown in Figure 3. The selected workloads were chosen to capture the architectural behaviors most commonly observed in production agentic AI deployments, including high-concurrency request processing, dynamic control flow, memory-intensive execution, synchronization overhead, I/O-intensive behavior, and heterogeneous runtime environments. Because the Gateway and Stream Response tiers exhibit similar computational characteristics, they are evaluated as a single category. Likewise, the Assemble Context, Plan and Route, and Verification tiers share comparable resource requirements and are consolidated into a common evaluation group. While the Reasoning tier is central to agentic AI and is predominantly powered by GPU-accelerated inference, the AMD Instinct™ platform has already established leadership in performance and scalability for these workloads. Accordingly, this blog concentrates on the CPU infrastructure that powers the remainder of the agentic AI pipeline and does not cover GPU inference performance.

Table 1. Agentic AI Workloads and Characteristics

Pipeline Tier

Workloads and Characteristics

Gateway;
Stream Response

NGINX with WRK for high-concurrency gateway and streaming request handling. Captures high-concurrency networking, event-driven scheduling, irregular control flow, and long-lived session management.


Assemble Context, Plan and Route; Verification

IREE tokenizer for prompt preprocessing and tokenization throughput. Captures text preprocessing, tokenization, and lookup-intensive execution.


Assemble Context, Plan and Route; Verification

Small-LLM inference for routing, planning, and critique workloads. Captures lightweight inference, prompt processing, and moderate compute intensity.


Assemble Context, Plan and Route; Verification

TPCx-AI for end-to-end AI pipeline orchestration. Captures multi-stage AI pipelines and orchestration-heavy workflows.

Retrieve Context & Similarity Search

FAISS for embedding similarity search and vector-index traversal. Captures high-dimensional embedding search, sparse index traversal, and vector similarity matching.

Reasoning

LLM inference on GPU nodes of both frontier models as well as smaller models depending on the use case, optimizing time to first token and tokens per second per dollar.

Enterprise Tools

TPC-H (or derivative) on MySQL for decision-support analytics queries. Captures OLAP analytics, join-intensive processing, and large-scale data scanning.

Enterprise Tools

TPC-C (or derivative) on MySQL for transactional database operations. Captures high-concurrency OLTP execution with short critical sections.

Enterprise Tools

Redis benchmark for low-latency in-memory key-value access. Captures low-latency, in-memory key-value access.

Enterprise Tools

YCSB with MongoDB for document-oriented NoSQL access patterns. Captures semi-structured data access and document-oriented processing.

Ephemeral Tools

Multi-persona agentic workflow replay across development, RAG ingestion, security, and media-processing scenarios. Captures concurrent agent execution, mixed AI/crypto/compression workloads, document ingestion, embedding generation, vector indexing, and resource sharing across co-resident agents. See the appendix for complete examples of the source workloads.

Several important insights emerge from this study. Most notably, agentic AI performance is not constrained by any single bottleneck. Instead, these workloads continuously transition among networking, orchestration, retrieval, tool execution, data processing, and inference-adjacent activities, making overall performance dependent on a balanced combination of compute throughput, memory bandwidth, cache efficiency, storage performance, synchronization mechanisms, and vector-processing capabilities. Concurrency further amplifies these demands. As AI agents generate parallel requests, retrieval operations, and tool invocations, the overhead associated with scheduling, context switching, locking, and coordination becomes an increasingly important factor in determining scalability. In addition, I/O and data movement often play a role that is equally as important as computation. Many stages of the agentic AI pipeline spend significant time interacting with APIs, databases, distributed services, storage systems, and networks, making end-to-end responsiveness highly dependent on platform latency and throughput. Taken together, these findings reinforce a central conclusion: agentic AI is fundamentally a systems-level workload, requiring a platform capable of efficiently balancing compute, memory, storage, networking, and orchestration resources across the entire execution pipeline with dynamically changing phases.

Using the methodology described above, AMD performance analysis engineers evaluated the CPU-centric stages of the agentic AI pipeline on platforms powered by Intel® Xeon® 6980P (128 cores / 256 threads), AWS Graviton5 (192 cores / 192 threads), AMD EPYC 9965 (192 cores / 384 threads), and AMD EPYC 9996 (256 cores / 512 threads) processors. Figure 1 presents the relative performance results across the evaluated workload tiers and Figure 4 shows the geomean across all the stages. Additional details on the benchmark methodology and system configurations are provided in the Appendix.

Figure 4: Agentic AI Workload Performance Geomean (CPU Centric)
Figure 4: Agentic AI Workload Performance Geomean (CPU Centric)

As shown in Figure 4, AMD EPYC 9005 Series Server CPUs delivers a strong performance uplift across the CPU-centric stages of the agentic AI pipeline, improving by 82% versus Intel® Xeon® 6980P and leadership versus AWS Graviton5 when taking geomean across all agentic AI pipeline execution stages6. This demonstrates that AMD EPYC 9005 Series Server CPUs provides a stronger current-generation platform for orchestration, retrieval, database, tool-execution, and response-handling stages that define production agentic AI infrastructure. AMD EPYC 9006 Series Server CPUs raise the bar further, delivering geomean uplift of 174% versus Intel® Xeon® 6980P and strong gains versus AWS Graviton56. Together, these results show a clear progression: AMD EPYC 9005 Series Server CPUs establishes a leadership baseline for today’s agentic AI deployments, while AMD EPYC 9006 Series Server CPUs extends that advantage with additional performance headroom for higher concurrency, larger retrieval workloads, heavier enterprise-tool execution, and more complex end-to-end agentic workflows.

AMD EPYC Server CPUs provide generational leadership in performance for agentic AI workloads, delivering strong performance across the diverse stages of the agentic AI pipeline. Compared with competing Intel Xeon and AWS Graviton platforms, AMD EPYC 9005 Series Server CPUs demonstrate significant advantages across a broad set of representative workloads, making it exceptionally well suited for agentic AI deployments. AMD EPYC 9006 Series Server builds on this foundation with further gains in performance, scalability, and efficiency, helping meet the increasing demands of next-generation agentic AI infrastructure.

Appendix A

Testing Methodology

Testing methodology is defined for each pipeline stage so that the benchmark reflects deployed behavior while maintaining comparability across systems. To preserve fairness across platforms, software configuration, benchmark harnesses, dataset scale, concurrency settings, and runtime environments were kept consistent while allowing platform-specific best practices for each execution stage and workload. Optimal NUMA, SMT, and other platform settings were applied for workloads representing AI pipeline stages. All workload configurations used self-contained modules so that network and disk I/O bandwidth did not affect measured throughput or latency. All systems were configured with Ubuntu 24.04.4 LTS, kernel 6.17.0-29-generic, and the CPUfreq governor set to “performance.”

Gateway and Stream Response

This tier uses NGINX Web Server with the WRK load generator to simulate high-concurrency network throughput. Measurements capture platform performance limits for Gateway and Stream Response pipeline characterization. NGINX version 1.24.0-2ubuntu7.9 and WRK version 4.2.0 were used across all platforms.

Plan & Route and Verification & Critique

This tier was evaluated using the IREE tokenizer, small language models, and an end-to-end AI and machine-learning benchmark based on the industry-standard TPCx-AI kit 2.0.0. Tokenization runs multiple instances of the single-threaded workload, with one instance pinned per core and SMT enabled. Small-LLM inference is evaluated by running vLLM instances of Llama-3.1-8B-Instruct with SMT disabled. End-to-end AI at SF30 is evaluated in a multi-instance scenario, with SMT/HT enabled on x86 platforms. Instances are scaled out until all cores are occupied and the system is fully saturated.

Retrieve Context and Similarity Search

This tier is represented by FAISS using the industry-accepted IVF4096,PQ128x4fs index to capture system throughput. Multiple pinned FAISS instances were run on each platform, and the overall metric was calculated by aggregating throughput across all instances.

Enterprise Tools

This tier was evaluated using a decision-support workload derived from the TPC-H kit, a transaction-processing workload derived from the TPC-C kit, Redis-benchmark, and MongoDB-YCSB. Workloads were deployed in 32-vCPU virtual machines and executed sequentially in randomized orders within each virtual machine to reduce order-specific bias. The overall metric was calculated using a geometric mean. Redis Server version 8.6.3, MySQL version 8.0.46, and MongoDB version 7.0.34 were used across all platforms.

 VM1: TPC-C(MySQL), TPC-H(MySQL), Redis(Get),  Redis(Set),MongoDB: Geomean(5x subsets)
 VM2: MongoDB,TPC-C(MySQL), TPC-H(MySQL), Redis(Get),  Redis(Set): Geomean(5x subsets)
 VM3: TPC-H(MySQL), Redis(Get),  Redis(Set),MongoDB, TPC-C(MySQL): Geomean(5x subsets)
 VM4: Redis(Get),  Redis(Set),MongoDB,TPC-C(MySQL), TPC-H(MySQL): Geomean(5x subsets)
 VMn: TPC-C(MySQL), TPC-H(MySQL), Redis(Get),  Redis(Set),MongoDB: Geomean(5x subsets)

Ephemeral Tools

This tier replays short-lived tool and command execution from captured agent sessions. A frontier model is given a tool-agnostic task, its completed run is traced, and the issued commands and tool calls are parsed into an ordered script that is replayed verbatim. This approach measures the tool-execution footprint of an agent session using fixed, pinned traces for reproducibility. The five scenarios are Development, Context Compaction, Security Signing/Attestation Broker, Media Manipulation, and RAG Ingestion.

The five personas invoke tools and libraries representative of common agentic workflows:
Development: grep, git, find, sed, GNU coreutils, GNU bash
Context Compaction: xz, zcat, cmp, grep, gzip, bzip2, zstd, awk
Security: cryptography (pyca), hashlib (stdlib), OpenSSL, GNU bash
Media: Pillow, numpy, GNU bash
RAG Ingestion: pymupdf, sentence-transformers, torch, zentorch, faiss-cpu, numpy
The main characteristics of these tools and libraries are mapped below.

To emulate ephemeral tools, all five personas run concurrently on the same platform and are scaled until resources saturate and additional sessions can no longer be supported. There is no fixed thread limit: each persona runs as a single-threaded worker process oversubscribed to a multiple of the core count and increased until throughput plateaus. All five personas run with an even split on every platform, and each level runs for 180 seconds. The overall metric is the raw geometric mean of the five per-persona throughput rates. Development, Context Compaction, Security, and Media report sessions per second, while RAG Ingestion reports documents per second.

Table 2. Ephemeral Tools — Workloads, Tools, and Shared Characteristics

Workload

Tools

Shared characteristic

openssl-sha256

OpenSSL, cryptography (pyca), hashlib, gzip, bzip2, xz, zcat

CPU-bound integer/bitwise sequential byte-crunch

simdjson

zstd, cmp

SIMD, memory-bandwidth-bound linear scan

ripgrep

grep, find, sed, coreutils, bash

Regex/string scan + traversal + I/O + orchestration

python-pandas

awk, numpy

Columnar/tabular numeric, memory-bound

opencv-template-match

Pillow, torch, zentorch, sentence-transformers, faiss-cpu

Dense FP/SIMD MAC / similarity compute

trafilatura

pymupdf

Parse messy docs, extract clean content

python-pptx

git

Serialize/parse compressed structured container

 

Footnotes
  1. 9xx6-042: Assemble Context, Plan and Route, and Verification tier comparison based on AMD internal testing as of 07/11/2026. TPCx-AI kit 2.0.0 end-to-end AI pipeline at Scale Factor 30, multi-instance, SMT/HT enabled on x86, median AIUCpm.
    Configurations Intel Xeon 6980P (128C/256T, 1P), SMT ON, 64GB DDR5-6400 RDIMM, SNC3, SuperMicro SYS-222HA-TN, BIOS v1.5, Ubuntu 24.04.4 LTS kernel 6.17.0-29-generic, performance governor. AWS Graviton5 (192C/192T, 1P), DDR5-8800, AWS m9gd.metal-48xl, Ubuntu 24.04.4 LTS kernel 6.17.0-29-generic, performance governor. Since cloud, all default BIOS and system config. AMD EPYC 9755 (128C/256T, 1P), SMT ON, 64GB DDR5-6400 RDIMM, AMD reference platform, BIOS RVOT1006C, Ubuntu 24.04.4 LTS kernel 6.17.0-29-generic, performance governor. AMD EPYC 9965 (192C/384T, 1P), SMT ON, 64GB DDR5-6400 RDIMM, AMD reference platform, BIOS RVOT1006C, Ubuntu 24.04.4 LTS kernel 6.17.0-29-generic, performance governor. AMD EPYC 9996 (256C/512T, 1P), SMT ON, 64GB DDR5-8000 RDIMM, AMD reference platform, BKC10.3 90RC4, Default CPU Power 600W cppc on power determinism, SMT ON, Ubuntu 24.04.4 LTS kernel 6.17.0-29-generic, performance governor.
    Summary results (AIUCpm (median)): Intel Xeon 6980P = 1750.36; AWS Graviton5 = 2444.80; AMD EPYC 9755 = 2704.19; AMD EPYC 9965 = 3458.79; AMD EPYC 9996 5982.91.
    Results for AWS Graviton5 were obtained from a publicly available cloud instance. Results compare AMD internal testing of the listed AMD and Intel systems against testing performed on the listed AWS EC2 Graviton5 bare-metal instance. Cloud instance results may be affected by cloud service configuration, instance availability, regional deployment, hypervisor/Nitro behavior, storage/network configuration, and other cloud provider variables.
    Starting with the 6th Gen AMD EPYC™ server processor family, AMD uses Default CPU Power to describe processor power consumption, succeeding AMD's historical TDP reference. Default CPU Power reflects total power consumed across the processor's compute and I/O dies for the stated performance target. Default CPU Power and TDP may both serve as processor power references for product comparison, platform planning, and performance-per-watt analysis. Results may not be directly comparable to physical server configurations due to differences in platform implementation, firmware, operating environment, memory configuration, and system tuning.
  2. 9xx6-043: Retrieve Context and Similarity Search tier comparison based on AMD internal testing as of 07/11/2026. FAISS IVF4096,PQ128x4fs index, sift1m dataset, k=10, aggregate QPS across pinned instances.
    Configurations Intel Xeon 6980P (128C/256T, 1P), SMT ON, 64GB DDR5-6400 RDIMM, SNC3, SuperMicro SYS-222HA-TN, BIOS v1.5, Ubuntu 24.04.4 LTS kernel 6.17.0-29-generic, performance governor. AWS Graviton5 (192C/192T, 1P), DDR5-8800, AWS m9gd.metal-48xl, Ubuntu 24.04.4 LTS kernel 6.17.0-29-generic, performance governor. Since cloud, all default BIOS and system config. AMD EPYC 9755 (128C/256T, 1P), SMT ON, 64GB DDR5-6400 RDIMM, AMD reference platform, BIOS RVOT1006C, Ubuntu 24.04.4 LTS kernel 6.17.0-29-generic, performance governor. AMD EPYC 9965 (192C/384T, 1P), SMT ON, 64GB DDR5-6400 RDIMM, AMD reference platform, BIOS RVOT1006C, Ubuntu 24.04.4 LTS kernel 6.17.0-29-generic, performance governor. AMD EPYC 9996 (256C/512T, 1P), SMT ON, 64GB DDR5-8000 RDIMM, AMD reference platform, BKC10.3 90RC4, Default CPU Power 600W cppc on power determinism, SMT ON, Ubuntu 24.04.4 LTS kernel 6.17.0-29-generic, performance governor.
    Summary results (aggregate QPS (mean)): Intel Xeon 6980P = 316069; AWS Graviton5 = 119179; AMD EPYC 9755 = 369252; AMD EPYC 9965 = 472079; AMD EPYC 9996 = 751453.
    Results for AWS Graviton5 were obtained from a publicly available cloud instance. Results compare AMD internal testing of the listed AMD and Intel systems against testing performed on the listed AWS EC2 Graviton5 bare-metal instance. Cloud instance results may be affected by cloud service configuration, instance availability, regional deployment, hypervisor/Nitro behavior, storage/network configuration, and other cloud provider variables.
    Starting with the 6th Gen AMD EPYC™ server processor family, AMD uses Default CPU Power to describe processor power consumption, succeeding AMD's historical TDP reference. Default CPU Power reflects total power consumed across the processor's compute and I/O dies for the stated performance target. Default CPU Power and TDP may both serve as processor power references for product comparison, platform planning, and performance-per-watt analysis. Results may not be directly comparable to physical server configurations due to differences in platform implementation, firmware, operating environment, memory configuration, and system tuning.
  3. 9xx6-044: Enterprise Tools tier comparison based on AMD internal testing as of 07/11/2026. TPC-H and TPC-C derivatives on MySQL, Redis benchmark, MongoDB-YCSB in 32-vCPU VMs; per-VM geometric mean of 5x subsets. Results are based on TPC-H and TPC-C derived workloads and are not comparable to published TPC benchmark results.
    System Configuration: Intel Xeon 6980P (128C/256T, 1P), SMT ON, 64GB DDR5-6400 RDIMM, SNC3, SuperMicro SYS-222HA-TN, BIOS v1.5, Ubuntu 24.04.4 LTS kernel 6.17.0-29-generic, performance governor. AWS Graviton5 (192C/192T, 1P), DDR5-8800, AWS m9gd.metal-48xl, Ubuntu 24.04.4 LTS kernel 6.17.0-29-generic, performance governor. Since cloud, all default BIOS and system config. AMD EPYC 9755 (128C/256T, 1P), SMT ON, 64GB DDR5-6400 RDIMM, AMD reference platform, BIOS RVOT1006C, Ubuntu 24.04.4 LTS kernel 6.17.0-29-generic, performance governor. AMD EPYC 9965 (192C/384T, 1P), SMT ON, 64GB DDR5-6400 RDIMM, AMD reference platform, BIOS RVOT1006C, Ubuntu 24.04.4 LTS kernel 6.17.0-29-generic, performance governor. AMD EPYC 9996 (256C/512T, 1P), SMT ON, 64GB DDR5-8000 RDIMM, AMD reference platform, BKC10.3 90RC4, Default[BP1.1][ZP1.2] CPU Power=600w cppc on power determinism, SMT ON, Ubuntu 24.04.4 LTS kernel 6.17.0-29-generic, performance governor.
    Summary results (geomean throughput): Intel Xeon 6980P = 2284701; AWS Graviton5 = 2982203; AMD EPYC 9755 = 2546290; AMD EPYC 9965 = 3867149; AMD EPYC 9996 = 6054748.
    Results for AWS Graviton5 were obtained from a publicly available cloud instance. Results compare AMD internal testing of the listed AMD and Intel systems against testing performed on the listed AWS EC2 Graviton5 bare-metal instance. Cloud instance results may be affected by cloud service configuration, instance availability, regional deployment, hypervisor/Nitro behavior, storage/network configuration, and other cloud provider variables. Starting with the 6th Gen AMD EPYC™ server processor family, AMD uses Default CPU Power to describe processor power consumption, succeeding AMD's historical TDP reference. Default CPU Power reflects total power consumed across the processor's compute and I/O dies for the stated performance target. Default CPU Power and TDP may both serve as processor power references for product comparison, platform planning, and performance-per-watt analysis. Results may not be directly comparable to physical server configurations due to differences in platform implementation, firmware, operating environment, memory configuration, and system tuning.
  4. 9xx6-045: Ephemeral Tools tier comparison based on AMD internal testing as of 07/11/2026. Multi-persona literal replay (Development, Context Compaction, Security, Media, RAG Ingestion); raw geometric mean of five per-persona throughput rates.
    System configurations:
    Intel Xeon 6980P (128C/256T, 1P), SMT ON, 64GB DDR5-6400 RDIMM, SNC3, SuperMicro SYS-222HA-TN, BIOS v1.5, Ubuntu 24.04.4 LTS kernel 6.17.0-29-generic, performance governor. AWS Graviton5 (192C/192T, 1P), DDR5-8800, AWS m9gd.metal-48xl, Ubuntu 24.04.4 LTS kernel 6.17.0-29-generic, performance governor. Since cloud, all default BIOS and system config. AMD EPYC 9755 (128C/256T, 1P), SMT ON, 64GB DDR5-6400 RDIMM, AMD reference platform, BIOS RVOT1006C, Ubuntu 24.04.4 LTS kernel 6.17.0-29-generic, performance governor. AMD EPYC 9965 (192C/384T, 1P), SMT ON, 64GB DDR5-6400 RDIMM, AMD reference platform, BIOS RVOT1006C, Ubuntu 24.04.4 LTS kernel 6.17.0-29-generic, performance governor. AMD EPYC 9996 (256C/512T, 1P), SMT ON, 64GB DDR5-8000 RDIMM, AMD reference platform, BKC10.3 90RC4, Default CPU Power 600W cppc on power determinism, SMT ON, Ubuntu 24.04.4 LTS kernel 6.17.0-29-generic, performance governor.
    Summary results (geomean throughput): Intel Xeon 6980P = 1.779; AWS Graviton5 = 2.505; AMD EPYC 9755 = 2.317; AMD EPYC 9965 = 2.970; AMD EPYC 9996 = 4.451.
    Results for AWS Graviton5 were obtained from a publicly available cloud instance. Results compare AMD internal testing of the listed AMD and Intel systems against testing performed on the listed AWS EC2 Graviton5 bare-metal instance. Cloud instance results may be affected by cloud service configuration, instance availability, regional deployment, hypervisor/Nitro behavior, storage/network configuration, and other cloud provider variables. Starting with the 6th Gen AMD EPYC™ server processor family, AMD uses Default CPU Power to describe processor power consumption, succeeding AMD's historical TDP reference. Default CPU Power reflects total power consumed across the processor's compute and I/O dies for the stated performance target. Default CPU Power and TDP may both serve as processor power references for product comparison, platform planning, and performance-per-watt analysis. Results may not be directly comparable to physical server configurations due to differences in platform implementation, firmware, operating environment, memory configuration, and system tuning.
  5. 9xx6-046: Gateway and Stream Response tier comparison based on AMD internal testing as of 07/11/2026. NGINX 1.24.0 with the WRK 4.2.0 load generator, aggregate max requests/sec across pinned instances.
    System configurations:
    Intel Xeon 6980P (128C/256T, 1P), SMT ON, 64GB DDR5-6400 RDIMM, SNC3, SuperMicro SYS-222HA-TN, BIOS v1.5, Ubuntu 24.04.4 LTS kernel 6.17.0-29-generic, performance governor. AWS Graviton5 (192C/192T, 1P), DDR5-8800, AWS m9gd.metal-48xl, Ubuntu 24.04.4 LTS kernel 6.17.0-29-generic, performance governor. Since cloud, all default BIOS and system config. AMD EPYC 9755 (128C/256T, 1P), SMT ON, 64GB DDR5-6400 RDIMM, AMD reference platform, BIOS RVOT1006C, Ubuntu 24.04.4 LTS kernel 6.17.0-29-generic, performance governor. AMD EPYC 9965 (192C/384T, 1P), SMT ON, 64GB DDR5-6400 RDIMM, AMD reference platform, BIOS RVOT1006C, Ubuntu 24.04.4 LTS kernel 6.17.0-29-generic, performance governor. AMD EPYC 9996 (256C/512T, 1P), SMT ON, 64GB DDR5-8000 RDIMM, AMD reference platform, BKC10.3 90RC4, Default CPU Power 600W cppc on power determinism, SMT ON, Ubuntu 24.04.4 LTS kernel 6.17.0-29-generic, performance governor.
    Summary results (max requests/sec): Intel Xeon 6980P = 10162179; AWS Graviton5 = 15331108; AMD EPYC 9755 = 17906196; AMD EPYC 9965 = 24320476; AMD EPYC 9996 = 28789170.
    Results for AWS Graviton5 were obtained from a publicly available cloud instance. Results compare AMD internal testing of the listed AMD and Intel systems against testing performed on the listed AWS EC2 Graviton5 bare-metal instance. Cloud instance results may be affected by cloud service configuration, instance availability, regional deployment, hypervisor/Nitro behavior, storage/network configuration, and other cloud provider variables. Starting with the 6th Gen AMD EPYC™ server processor family, AMD uses Default CPU Power to describe processor power consumption, succeeding AMD's historical TDP reference. Default CPU Power reflects total power consumed across the processor's compute and I/O dies for the stated performance target. Default CPU Power and TDP may both serve as processor power references for product comparison, platform planning, and performance-per-watt analysis. Results may not be directly comparable to physical server configurations due to differences in platform implementation, firmware, operating environment, memory configuration, and system tuning.
  6. 9xx6-048: Geometric mean across the five CPU-centric Agentic AI pipeline tiers (Gateway and Stream Response; Assemble Context, Plan and Route, and Verification; Retrieve Context and Similarity Search; Enterprise Tools; Ephemeral Tools) based on AMD internal testing as of 07/11/2026.
    Intel Xeon 6980P (128C/256T, 1P), SMT ON, 64GB DDR5-6400 RDIMM, SNC3, SuperMicro SYS-222HA-TN, Ubuntu 24.04.4 LTS kernel 6.17.0-29-generic, performance governor. AWS Graviton5 (192C/192T, 1P), DDR5-8800, AWS m9gd.metal-48xl, Ubuntu 24.04.4 LTS kernel 6.17.0-29-generic, performance governor. AMD EPYC 9755 (128C/256T, 1P), SMT ON, 64GB DDR5-6400 RDIMM, AMD reference platform, Ubuntu 24.04.4 LTS kernel 6.17.0-29-generic, performance governor. AMD EPYC 9965 (192C/384T, 1P), SMT ON, 64GB DDR5-6400 RDIMM, AMD reference platform, Ubuntu 24.04.4 LTS kernel 6.17.0-29-generic, performance governor. AMD EPYC 9996 (256C/512T, 1P), SMT ON, 64GB DDR5-8000 RDIMM, AMD reference platform, Ubuntu 24.04.4 LTS kernel 6.17.0-29-generic, performance governor.
    Per-tier ratios vs Intel Xeon 6980P = 1.00: Gateway and Stream Response (AMD EPYC 9755 = 1.76x, AMD EPYC 9965 = 2.39x, AMD EPYC 9996 = 2.83x, AWS Graviton5 = 1.51x); Assemble Context, Plan and Route, and Verification (AMD EPYC 9755 = 1.54x, AMD EPYC 9965 = 2.01x, AMD EPYC 9996 = 3.45x, AWS Graviton5 = 1.41x); Retrieve Context and Similarity Search (AMD EPYC 9755 = 1.17x, AMD EPYC 9965 = 1.49x, AMD EPYC 9996 = 2.38x, AWS Graviton5 = 0.38x); Enterprise Tools (AMD EPYC 9755 = 1.11x, AMD EPYC 9965 = 1.69x, AMD EPYC 9996 = 2.65x, AWS Graviton5 = 1.31x); Ephemeral Tools (AMD EPYC 9755 = 1.30x, AMD EPYC 9965 = 1.67x, AMD EPYC 9996 = 2.50x, AWS Graviton5 = 1.41x).
    Consolidated geometric mean normalized to Intel Xeon 6980P = 1.00: AMD EPYC 9755 = 1.36x, AMD EPYC 9965 = 1.82x, AMD EPYC 9996 = 2.74x, AWS Graviton5 = 1.08x.
    See 9xx6-042, 9xx6-043, 9xx6-044, 9xx6-045, 9xx6-046 for more info.
    Results compare AMD internal testing of the listed AMD and Intel systems against testing performed on the listed AWS EC2 Graviton5 bare-metal instance. Cloud instance results may be affected by cloud service configuration, instance availability, regional deployment, hypervisor/Nitro behavior, storage/network configuration, and other cloud provider variables. Results may vary due to factors including system configurations, software versions and BIOS settings.
Share:

Article By


Corp VP, Datacenter Ecosystems and Application Engineering, Server BU

Related Blogs