How AMD Ryzen AI Max PRO 400 Series Processors Bring Local Agentic AI to Business Customers

Sep 28, 2026

The advent of generative artificial intelligence has raised questions around commercial use cases, adoption, and ROI since the technology entered mainstream discourse in late 2022. As AI agents move towards broad deployment, companies face even greater pressure to answer them.

Unlike genAI chatbots, which typically rely on a question-and-answer model, AI agents can carry out complex, multi-step tasks across different applications. When configured to do so, they can also take certain actions autonomously and work on complex projects for extended periods while the person who initiated the task focuses elsewhere.

Harnessing agents’ ability to analyze information and act upon it is an attractive prospect for companies looking to capture more value from AI. At the same time, rapid advances in both cloud and local AI have left IT departments confronting practical questions about where agents should run, what company data they should be allowed to access, and how their capabilities can be made available to employees cost-effectively.

Today, new systems featuring AMD Ryzen™ AI Max and Ryzen AI Max PRO 400 Series processors are available from AMD partners, giving businesses new options for developing and deploying enterprise AI locally. With up to 16 “Zen 5” CPU cores, a 55 TOPS NPU1, 40 GPU compute units, and a massive, unified pool of DRAM, these new systems are purpose-built for local agentic AI.

An image of tasks where Gorgon Halo excels.

These new Ryzen AI Max and Max PRO Series processors build on the same unified memory architecture AMD debuted with the Ryzen AI Max and Max PRO 300 Series, while offering substantially increased memory capacity. Where previous-generation systems topped out at 128GB of total memory with up to 96GB reserved for GPU workloads, the Ryzen AI Max 400 Series offers up to 192GB, with up to 160GB dedicated to the GPU — an increase of 1.67x. The Ryzen AI Max PRO 400 Series is the world's first x86 client processor capable of running models with more than 300B parameters at 4-bit quantization2, with no need for cloud offload.

Agentic workloads can place heavy demands on memory because they may combine multiple models, large context windows, enterprise data, and several applications or tools within the same workflow. The additional memory capacity of Ryzen AI Max 400 Series systems gives businesses more room to develop and run those workloads locally.

 What Does It Take to Run an Enterprise AI Agent Locally?

An enterprise agent needs more than a foundation model. Development teams may need to connect models to internal documents and codebases, build retrieval-augmented generation applications, evaluate different models, create tools and connectors, test agent behavior, and determine which workloads should remain local and which should use cloud services.

Ryzen AI Max PRO 400 Series processors provide a local platform for that development cycle. Their combination of high-performance CPU cores, integrated graphics, a powerful NPU and a large, unified memory pool supports model development, orchestration, prototyping and execution, all on the same system. The platform supports Windows and Linux environments, while the broader Ryzen AI software ecosystem includes widely used tools and frameworks such as PyTorch, vLLM, llama.cpp, Ollama, ComfyUI and LM Studio. 

 When Should an AI Agent Run Local AI vs in the Cloud?

The question of whether AI workloads should run locally or in the cloud is sometimes treated as an all-or-nothing problem, but Ryzen AI Max PRO 400 Series processors exemplifies a different philosophy – one where cloud and local AI resources coexist within the same deployment strategy.

Frontier cloud models remain valuable when a workload requires capabilities that local models cannot provide. In other cases, however, local systems can handle some or all of the workload. 

HP, for example, is pairing the new AMD Ryzen AI Max+ PRO 495 processor with Perplexity Portable Computer in the HP ZBook Ultra G3a 16. This system can run local models and agents alongside professional applications, while users can tap a cloud model if specific steps require frontier intelligence.

Local execution can also give organizations more control over where proprietary data is processed. An agent working against internal documents, engineering files, source code or other sensitive information can perform suitable operations on the system instead of sending every request and working document to an external inference service.

There are also economic reasons to consider a hybrid division of labor.

Agentic AI can consume enormous numbers of tokens — far more than chatbots — because an agent may reason, revise, retrieve information and invoke models repeatedly during a single job. Cloud inference turns each of those operations into a recurring expense.

An image of tokenomics spending. It projects a break-even point for Strix Halo within six months.
See SHO-49

The AMD Tokenomics Calculator lets organizations explore the potential cost of cloud-only, local-only, and hybrid AI deployments using their expected user counts, token consumption, hardware requirements and cloud-model pricing. As the image above shows, the cost of using cloud services can significantly outstrip the cost of local AI in a relatively short period.    

A hybrid deployment lets businesses reserve cloud inference for workloads where its capabilities justify the expense while running other work locally. With the AMD Tokenomics Calculator, IT can evaluate the impact of varying that mix rather than treating local and cloud AI as mutually exclusive choices.

Enterprise agents are moving from experiments into production. With Ryzen AI Max 400 Series and Ryzen AI Max PRO 400 Series systems now reaching customers, AMD is putting substantially more local compute and memory behind that transition, giving businesses another place to build AI, another place to deploy it, and more control over where the work gets done.

See how Agentic PCs powered by AMD support enterprise workflows and explore available commercial systems. Explore Agentic PCs for Business.

1: GD-243: Trillions of Operations per Second (TOPS) for an AMD Ryzen processor is the maximum number of operations per second that can be executed in an optimal scenario and may not be typical. TOPS may vary based on several factors, including the specific system configuration, AI model, and software version. GD-243.

2: GRHP-01: As of 5/11/2026, the AMD Ryzen AI Max+ 495 PRO processor supports up to 160GB dedicated graphics memory, which is capable of running 300 billion+ parameters at 4-bit quantization. GRHP-01.

3: SHO-49 in the endnotes for this image: Specimen use-case using simplified assumptions. Cloud comparison assumes Claude Sonnet 4.5 standard API pricing of $3/M input tokens and $15/M output tokens, with a 10:1 input-to-output token utilization mix. Local scenario assumes AMD Ryzen™ AI Max+ sustained throughput at 128K context of 36 output tokens/sec and 446 input tokens/sec, measured on Pre-Production AMD Ryzen™ AI Halo developer platform using LM Studio 0.4.12 and vulkan llama.cpp, AMD Software: Adrenalin Edition™ Drivers 26.10.08. Scenario consumption is modeled at 8 hours/day effective utilization, equivalent to approximately 0.573M output tokens/day and 5.73M input tokens/day, or approximately 6.3M total tokens/day. Electricity assumes 150W sustained “nightmare case” draw at $0.15/kWh for 24 hours/day, or $16.20/month. The $0.15/kWh assumption is a rounded estimate broadly aligned with U.S. Energy Information Administration February 2026 average retail electricity pricing , including 14.36¢/kWh for U.S. total all-sector pricing and 17.65¢/kWh for U.S. total residential pricing Source: https://www.eia.gov/electricity/monthly/current_month/april2026.pdf. Three-year local total includes upfront hardware cost plus 36 months of electricity; cloud total assumes equivalent token consumption over 36 months, pricing for Claude Sonnet 4.5, source: https://platform.claude.com/docs/en/about-claude/pricing. Pricing information fetched as of May 2026. Hardware price is based on AMD Ryzen™ AI Halo Box 128GB, as of 5/10/2026. Actual results may vary with workload, context length, caching, batching, model choice, electricity rate, hardware configuration, utilization, and real-world agent behavior. SHO-49.

Share:

Article By


Related Blogs