7 Takeaways from AMD Advancing AI 2026 to Guide Your 2027 Infrastructure Plan
Aug 04, 2026
Enterprise AI infrastructure planning is changing shape as organizations start shifting more from chatbot-style AI to agentic AI. An AI agent requires a new conversation because it doesn’t just generate a response — it plans, calls tools, queries databases, checks permissions, retrieves memory, and loops, often triggering other agents along the way. Each of those steps still runs on general-purpose compute, which shifts demand across CPU, GPU, networking, and orchestration software at the same time. This forces IT leaders to plan against a new set of criteria. Advancing AI 2026 was where AMD laid out its answer to this shift in detail.
Below are the 7 key takeaways from the event that enterprises should keep in mind when having 2027 infrastructure planning discussions.
Start With Your Usage Model, Not a Spec Sheet
The first question to answer is not which GPU has the best peak performance, it’s “how does my organization actually use AI?” How many users, how many agents, which models, and how much data needs to stay on-premises for latency or sovereignty reasons? That usage model determines which compute tier you need, in what ratio, and where it should sit.
Consider the CPU-GPU Ratio Before You Plan the GPU Count
The CPU isn't a supporting actor to the GPU story anymore. In an agentic workload, it's doing real work of its own, and that changes the order in which you should plan. For years, GPU capacity dominated AI infrastructure planning, with the CPU cast as a supporting host node. An agent's workflow breaks that assumption: every plan-call-query-check-retrieve-loop cycle still runs on general-purpose compute, and that workflow splits across distinct profiles that each lean on the CPU differently.
Agent orchestration and sandboxing benefit from CPUs built for that pattern. Host nodes that keep GPU clusters fed during reasoning and inference need a different profile. However, general-purpose enterprise applications (ERP, CRM, the databases agents query to get work done) still need to run somewhere. Planning around a single CPU standard risks under-provisioning at least one of those needs.
At Advancing AI 2026, AMD announced the latest 6th Generation AMD EPYC™ 9006 Series Server CPUs, the broadest server CPU portfolio built for agentic AI1, spanning cloud, enterprise, general-purpose and HPC workloads. As agent sandboxes, the generation runs the agents, coordinating concurrent agents and executing tasks; as AI host nodes, it keeps accelerators fed with speed and memory bandwidth; as general-purpose servers, it runs business-critical applications and AI support tasks enterprises depend on every day, with enhanced energy efficiency.
The practical implication: size your CPU layer to the workflow first, then size the GPU layer to the model. Getting that order backward is how enterprises end up with GPU clusters that are starved by their own host nodes.
Match the GPU to Your Buying Scenario
Once you know your usage model and CPU ratio, the GPU decision narrows to the question of which buying scenario you're actually in — not which spec sheet looks best.
To help customers right-size their infrastructure, the new AMD Instinct™ MI400 Series GPUs are split into three products, positioned for three different buying scenarios, rather than one GPU sold for every use case. The AMD Instinct MI455X GPU is built for frontier-scale training and inference, backed by 432 GB of HBM4 memory and 3D chip-stack packaging. The AMD Instinct MI430X GPU (already running in exascale supercomputers at Oak Ridge National Laboratory and France's Jean Zay), targets HPC and sovereign AI workloads that need high-precision scientific computing rather than raw AI throughput. The AMD Instinct MI350P GPU targets enterprises upgrading existing data centers without a full build-out: it's air-cooled, fits inside the power and cooling envelope of standard enterprise servers, and delivers more than 4x the tokens per second per dollar versus the competition.2
The takeaway for procurement: start by asking which of these three profiles matches your actual deployment scenario.
Treat the Rack as the Procurement Unit
For organizations that land on the frontier-scale tier, the planning unit also changes from individual components to the rack as a whole.
The headline announcement at Advancing AI was AMD Helios — the first rack-scale AI solution from AMD. As AI workloads scale, the rack itself becomes the unit of engineering. The premise is straightforward: every layer: compute, networking, cooling, and software performs best when designed as one system from the start. Helios pairs 72-Instinct MI455X GPUs with 18-6th Generation AMD EPYC server CPUs as the host compute layer, AMD Pensando™ networking spanning front-end, scale-up and scale-out connectivity, and AMD ROCm™ open software as the programming and deployment layer.
This tier is not just an edge case reserved for hyperscalers. Anthropic, Cerebras, Meta and OpenAI are among the early adopters, and systems will be available through OEMs like Dell, HPE, Lenovo and Supermicro, and infrastructure partners Sanmina and Wiwynn.
With all these components working together, Helios delivers up to 30% more tokens per dollar than the leading competitive solution.3 For planning purposes, that translates to less time spent validating components and more confidence that the rack performs as advertised, because it was engineered as one system. And because ROCm stays open underneath it, that performance doesn't come at the cost of flexibility.
Factor in Software Maturity, Not Just Silicon
None of the above matters if your teams can't build on it quickly. Developer productivity, not just raw silicon, increasingly determines how fast organizations deploy agentic workloads.
AMD introduced ROCm.ai, an AI-driven development platform built to help teams build, optimize and deploy GPU software faster. It also enables popular coding agents, including Claude, Codex and Cursor, to understand AMD platforms and ROCm natively. Leading open-source frameworks including PyTorch, Hugging Face, vLLM and SGLang are already enabled on MI455X.
When you evaluate a platform, ask how mature its software ecosystem is under real workloads. Day-0 model support, framework breadth, and the ability to use the coding agents your teams already know will all help shorten the time from purchase to production.
Plan for Cost and Governance at Scale
Two challenges tend to surface only after AI moves from pilot to production: cost, once usage moves from pilot to real volume, and control, once AI spreads from the data center to every employee's hands. Two customer examples from Advancing AI 2026 illustrate how to plan for both in advance.
AT&T is proof that scaling AI doesn't mean scaling your token bill. Instead of running every task through the biggest available model, AT&T routes each one through a cache-aware gateway to the model suited for that job, helping reduce AI costs by as much as 80%. The system has processed more than 1 trillion tokens through Microsoft Azure-hosted models on AMD Instinct GPUs, part of the Open Telco AI initiative launched with AMD, Microsoft and GSMA. AT&T uses its most valuable tokens from that work to train OTel 2.0, now the largest, best-performing open-source model built for telecom.4
As AI moves beyond the data center, Cisco shows that local AI and centralized control don't have to be a trade-off. Cisco pairs AMD Ryzen™ AI Halo systems with its own networking, observability, governance and security technology, so agentic AI runs closer to employees and their devices without IT losing visibility or control. That layer of control becomes critical as agents act on enterprise data continuously, not just when prompted. For risk and compliance stakeholders, that governance layer may end up mattering more than the silicon underneath it.
Track Physical AI as the Next Procurement Category
The same loop that defines agentic AI- perceive, reason, act, observe the result, loop again, doesn't stop at the edge of the data center. Physical AI is that loop extended into a machine: AI that perceives its surroundings, reasons about them, and acts on the physical world instead of a screen.
AMD Kria™ AI system-on-modules bring the CPU, GPU, NPU, and FPGA needed for that loop onto one integrated platform for robotics and industrial automation, instead of assembling it from four separate parts. Enterprises in manufacturing, logistics and field operations should treat physical AI not as a distant category but one to start tracking now.
The Framework
Put together, these announcements point to the same conclusion for 2027 planning: AI infrastructure decisions become portfolio decisions, not single purchases. The criteria that mattered five years ago, things like peak FLOPS, list price per GPU, no longer tell the whole story on their own. The more durable questions now are about total cost per useful token over the life of a deployment, how much systems-integration effort an architecture demands, and how much flexibility it preserves.
The question for enterprises is not whether to make this shift, it's when and how they align their infrastructure roadmap to it. Essentially, this means treating rack, CPU, GPU, software and edge deployment as one connected decision rather than five separate ones. Dr. Lisa Su’s keynote at Advancing AI 2026 displayed what that looks like in practice, from full-stack systems to real customer deployments proving the approach at scale.
To hear more about the importance of each of these decisions, customer deployments, and product architecture, watch the full Advancing AI 2026 keynote on demand.
Notas al pie
- EPYC-068 - The AMD EPYC server CPU portfolio spans the industry’s broadest ranges of data center deployments, from general-purpose enterprise, cloud, telecom, SMB, and HPC systems to emerging AI environments including sandboxed agentic AI deployments and GPU head node servers. AMD EPYC 6th Generation platforms extend this breadth by uniquely combining high core and thread density of up to 512 threads, advanced memory bandwidth of up to 16 channels of 12.8 GT/s MRDIMM support, next-generation PCIe® Gen 6 connectivity, and select SKUs with boost frequencies up to 5 GHz.
- MI350P-007: Based on AMD internal testing (July 2026), on a (1x) AMD Instinct MI350P GPU vs (1x) NVIDIA H200 NVL GPU on the Llama 3.3 70B Instruct (FP8) online serving output-throughput per dollar (tok/s/USD) comparison at ISL/OSL 1024/1024 across concurrency levels 1, 4, 8, 16, 32, 64, 128, 256, 512; median of 3 runs per point. MI350P based server internal AMD estimated pricing as $327,238.40 USD. RTX_PRO_6000 based server public list price reported on OEM website as $265,928.24 USD as of 7/16/2026. Stated results are the peak per-concurrency ratios: MI350P served via AIMS silogenai/aim-instinct-meta-llama-llama-3-3-70b-instruct:0.12.0-rc6; H200 NVL via NVIDIA NIM nvcr.io/nim/meta/llama-3.3-70b-instruct:2.0.6; RTX PRO 6000 via NVIDIA NIM nvcr.io/nim/meta/llama-3.3-70b-instruct:2.0.6. Configuration: 8x AMD Instinct MI350P PCIe Card (CDNA4, gfx950, 128 CUs, 144 GB HBM3E, SPX compute / NPS1), vBIOS 113-350P-01-1K1-000A, GPU driver 6.19.13-2353916.24.04, ROCm 7.14.0 (AMD-SMI 26.5.0); host 2P AMD EPYC 9455 (48-core), Dell PowerEdge XE7745, BIOS 1.7.6, microcode 0xb002162, SMT Enabled, Ubuntu 24.04.4 LTS, Linux 6.8.0-124-generic || NVIDIA RTX PRO 6000: 8x NVIDIA RTX PRO 6000 Blackwell Server Edition, vBIOS 98.02.8D.00.01, GPU driver 595.45.04, CUDA 13.2; host 2P AMD EPYC 9455 (48-core), Dell PowerEdge XE7745, BIOS 1.6.4, microcode 0xb00215a, SMT Enabled, Ubuntu 24.04.4 LTS, Linux 6.8.0-124-generic. Sever manufacturers may vary configurations, yielding different results. Results may vary due to factors including system configurations, software versions and BIOS settings.
- MI400-025: Based on AMD Performance Labs estimates as of July 2026, tokens-per-dollar performance was calculated using the Kimi K2 Thinking workload (32K input / 8K output) on an AMD Helios rackscale solution compared to an NVIDIA Vera Rubin NVL72 rack. Results reflect estimated aggregate throughput across low, medium, and high-interactivity operating points and hourly pricing projection of system GPUs based on market conditions. System configurations may vary by manufacturer and may produce different results.
- All performance and/or cost savings claims are provided by the 3rd party organization featured herein and have not been independently verified by AMD. Performance and cost benefits are impacted by a variety of variables. Results herein are specific to such 3rd party organization and may not be typical. GD-181a.
- EPYC-068 - The AMD EPYC server CPU portfolio spans the industry’s broadest ranges of data center deployments, from general-purpose enterprise, cloud, telecom, SMB, and HPC systems to emerging AI environments including sandboxed agentic AI deployments and GPU head node servers. AMD EPYC 6th Generation platforms extend this breadth by uniquely combining high core and thread density of up to 512 threads, advanced memory bandwidth of up to 16 channels of 12.8 GT/s MRDIMM support, next-generation PCIe® Gen 6 connectivity, and select SKUs with boost frequencies up to 5 GHz.
- MI350P-007: Based on AMD internal testing (July 2026), on a (1x) AMD Instinct MI350P GPU vs (1x) NVIDIA H200 NVL GPU on the Llama 3.3 70B Instruct (FP8) online serving output-throughput per dollar (tok/s/USD) comparison at ISL/OSL 1024/1024 across concurrency levels 1, 4, 8, 16, 32, 64, 128, 256, 512; median of 3 runs per point. MI350P based server internal AMD estimated pricing as $327,238.40 USD. RTX_PRO_6000 based server public list price reported on OEM website as $265,928.24 USD as of 7/16/2026. Stated results are the peak per-concurrency ratios: MI350P served via AIMS silogenai/aim-instinct-meta-llama-llama-3-3-70b-instruct:0.12.0-rc6; H200 NVL via NVIDIA NIM nvcr.io/nim/meta/llama-3.3-70b-instruct:2.0.6; RTX PRO 6000 via NVIDIA NIM nvcr.io/nim/meta/llama-3.3-70b-instruct:2.0.6. Configuration: 8x AMD Instinct MI350P PCIe Card (CDNA4, gfx950, 128 CUs, 144 GB HBM3E, SPX compute / NPS1), vBIOS 113-350P-01-1K1-000A, GPU driver 6.19.13-2353916.24.04, ROCm 7.14.0 (AMD-SMI 26.5.0); host 2P AMD EPYC 9455 (48-core), Dell PowerEdge XE7745, BIOS 1.7.6, microcode 0xb002162, SMT Enabled, Ubuntu 24.04.4 LTS, Linux 6.8.0-124-generic || NVIDIA RTX PRO 6000: 8x NVIDIA RTX PRO 6000 Blackwell Server Edition, vBIOS 98.02.8D.00.01, GPU driver 595.45.04, CUDA 13.2; host 2P AMD EPYC 9455 (48-core), Dell PowerEdge XE7745, BIOS 1.6.4, microcode 0xb00215a, SMT Enabled, Ubuntu 24.04.4 LTS, Linux 6.8.0-124-generic. Sever manufacturers may vary configurations, yielding different results. Results may vary due to factors including system configurations, software versions and BIOS settings.
- MI400-025: Based on AMD Performance Labs estimates as of July 2026, tokens-per-dollar performance was calculated using the Kimi K2 Thinking workload (32K input / 8K output) on an AMD Helios rackscale solution compared to an NVIDIA Vera Rubin NVL72 rack. Results reflect estimated aggregate throughput across low, medium, and high-interactivity operating points and hourly pricing projection of system GPUs based on market conditions. System configurations may vary by manufacturer and may produce different results.
- All performance and/or cost savings claims are provided by the 3rd party organization featured herein and have not been independently verified by AMD. Performance and cost benefits are impacted by a variety of variables. Results herein are specific to such 3rd party organization and may not be typical. GD-181a.