AMD ROCm™ 10: A Simpler Path to Production AI on AMD Instinct™ GPUs

Aug 27, 2026

ROCm.AI GPU software development graphic highlights Open Ecosystem, Higher-Level Abstractions, AI-Assisted Development, and ROCm Everywhere.

Bring the models, frameworks, containers, and AI assistants your teams already know. ROCm 10 makes it easier to move from AMD Instinct GPU powered infrastructure access to production inference and training.

Software decisions are a big part of AI infrastructure decisions.

Customers evaluating AI infrastructure need to know that their models will run, their developers can get productive quickly, their platform teams can operate the environment at scale, and performance can continue improving after the hardware is deployed.

With ROCm 10, AMD is making that path simpler than ever.

ROCm 10 brings together a more consistent software foundation, validated AI frameworks and containers, scalable communication libraries, training software and an AI-native developer experience through ROCm.AI. The goal is straightforward: reduce the work between gaining access to AMD Instinct GPU infrastructure and putting that infrastructure to productive use.

Rather than requiring customers to assemble the software stack themselves, ROCm 10 combines validated AI software, modular delivery, scalable communications, and AI-assisted workflows. Teams can start with the outcome they need and not create a separate integration project.

Get Your AI Workloads Running Faster

Most organizations adopting new AI infrastructure do not want to rebuild their software environment around the hardware. They want to bring the models, frameworks and serving engines their teams already use.

ROCm 10 makes that easier on AMD Instinct GPUs.

Validated vLLM and SGLang containers, Python wheels and modular software packages give AI teams tested paths for running priority models without having to build the entire inference environment from source.

A common build and validation foundation across ROCm packages, framework wheels and containers also helps improve consistency as workloads move from development and evaluation into production.

ROCm 10 is designed to help teams start with the workload they want to run rather than the software components they need to assemble.

A More AI-Native Developer Experience with ROCm.AI

ROCm.AI provides a simpler way for developers to interact with the AMD software stack.

Developers can use AI assistants they already work with, including Claude, Cursor and Codex, and access AMD-validated workflows through AMD Skills.

Instead of starting by searching documentation and manually determining every command, developers can start with an outcome:

Install and validate my ROCm environment.

Serve this model.

Diagnose a system issue.

Help optimize this workload.

AMD Skills provide guided workflows, while ROCm CLI provides a repeatable execution layer for installation, environment inspection, model serving, updates and diagnostics.

ROCm Console adds local visibility into telemetry, logs, runtime status and diagnostic information.

Together, these capabilities are designed to shorten the path from gaining access to AMD Instinct infrastructure to running a validated AI workload.

Train and Serve with Software Built for AMD Instinct

ROCm 10 supports both sides of modern AI infrastructure: training models and serving them at scale.

For inference, teams can continue working with familiar open-source serving engines such as vLLM and SGLang while taking advantage of AMD-optimized libraries, kernels and validated containers.

For training, AMD Primus provides an integrated environment spanning experiment configuration, infrastructure validation, pre-training, monitoring and optimization.

Primus helps teams identify memory-efficient configurations, validate infrastructure before large training runs and reduce manual trial-and-error as workloads scale from individual systems to multi-node clusters.

This gives organizations a more complete software path across model development, training, post-training and inference on AMD Instinct infrastructure.

Move from the First Workload to Production at Scale

Getting one workload running is only the beginning. AI platform and infrastructure teams need need consistent environments, predictable upgrades, visible system state, and reliable scale-out.

ROCm 10 introduces a more modular foundation through the ROCm Core SDK, allowing teams to install the software components required for a workload without necessarily deploying the entire development stack.

This can reduce software footprint while giving platform teams greater control over qualification, deployment and lifecycle management.

ROCm 10 is also built through a common multi-architecture build system, TheRock, helping create greater consistency across packages, wheels and containers. For customers, that means fewer differences between evaluation, development and production environments.

At cluster scale, ROCm 10 advances RCCL with improvements designed to increase communication efficiency and resiliency across distributed AI workloads, including enhancements for large-scale initialization, fault tolerance, GPU-initiated networking, symmetric memory and multi-node collective operations.

These capabilities become increasingly important as training and inference workloads scale across GPUs, nodes and racks.

Operate with Greater Visibility and Control

Production AI also requires the ability to understand what is happening inside the system without disrupting workloads or moving sensitive information outside the environment.

ROCm 10 expands the tools available to platform and performance teams.

Live AMD Thread Trace attachment enables teams to profile a running workload without restarting it, helping reduce disruption during performance analysis.

Local hipBLASLt optimization can tune GEMM kernel selection for specific workload shapes while keeping workload information and model assets within the customer environment.

ROCm CLI dependencies can also be packaged for disconnected environments, while telemetry, diagnostics and optimization can remain local.

For sovereign, regulated and security-sensitive environments, this allows teams to simplify operation while maintaining greater control over where their software and data are managed.

ROCm's open foundation provides additional visibility and flexibility across the software stack, while the same core software technologies support systems ranging from individual AMD Instinct servers to large-scale supercomputers such as Frontier and El Capitan.

Economics: Reach Useful AI Output Faster and Get More from the GPUs You Already Deploy With ROCm.AI Optimization

For executives and model-service operators, the software objective is straightforward: activate GPU capacity quickly, maintain utilization, and improve performance without excessive specialist effort.

ROCm.AI reduces the number of disconnected steps between hardware access and a working workload. AMD Skills provide guided workflows, ROCm CLI standardizes execution, validated containers reduce environment assembly, and ROCm Console improves visibility. This helps engineering and platform teams spend less time configuring the software path and more time running models. 

ROCm Hyperloom extends that workflow into optimization. It profiles inference workloads, identifies bottlenecks across host code and GPU kernels, evaluates changes, benchmarks the result, and validates performance and correctness. This turns optimization into a repeatable workflow rather than a lengthy manual exercise. 

Recent AMD Performance Labs testing shows how full-stack software optimization can improve inference performance. On an 8x AMD Instinct™ MI355X GPU platform, a preview of ROCm.AI based on ROCm 7.2.2, with optimizations including optimized kernels, parallelism, and scheduling, delivered an average 3.3x higher inference throughput across GLM-5, Kimi-K2.5, and DeepSeek-R1-0528 compared with ROCm 7.0. Performance was measured in tokens per second (TPS), with the uplift reported as the combined average across the three models tested.1

ROCm.AI inference chart shows 3.3x average improvement for DeepSeek-R1, GLM-5, and Kimi-K2.5 vs. ROCm 7.

For training, AMD Performance Labs testing showed an average 2.4x higher training throughput with a preview of ROCm.AI compared with ROCm 7.0 across DeepSeek-V2-Lite, DeepSeek-V3-16B, and Qwen3-30B-A3B using Megatron-LM on an 8x AMD Instinct™ MI355X GPU platform. Performance was measured in tokens per second (TPS), with the uplift reported as the combined average across the three models tested. These results are workload- and configuration-specific and demonstrate the impact of coordinated optimization across serving software, communication, memory management, kernels, and workload execution. They should not be presented as universal ROCm 10 performance uplifts.

ROCm.AI training chart shows 2.4x average improvement for DeepSeek-V3, DeepSeek-V2-Lite, and Qwen3-30B-A3B vs. ROCm 7.

A Simpler Software Decision for AMD Instinct Infrastructure

ROCm 10 makes the path to AMD Instinct infrastructure more direct.

AI teams can retain familiar models and frameworks. Platform teams gain repeatable installation, serving, diagnostics, and observability. Infrastructure teams gain modular software delivery and stronger distributed communication. Security teams retain local control. Executives gain a shorter path from deployed capacity to useful AI output.

Organizations do not need to begin by mastering every ROCm component. They can begin with a priority workload, use AMD Skills through an existing AI assistant, establish and validate the environment with ROCm CLI, deploy a validated training or inference stack, and use Hyperloom and ROCm profiling tools to optimize the result.

ROCm 10 is designed to make AMD Instinct software easier to adopt, operate, and scale - so customers can focus on the AI and HPC outcomes the infrastructure is intended to deliver.

Performance Disclosures

1. Inference on ROCm.ai (MI350-081)

MI350-81- Testing by AMD Performance Labs as of July 7, 2026, measuring the inference performance in tokens per second (TPS) of a system configured with an AMD Instinct MI355x 8x GPU platform and AMD ROCm 7.0 software vs a similarly configured system using a preview version of AMD ROCm.ai (ROCm 7.2.2 with optimizations such as Optimized Kernels, Parallelism and Scheduling) running GLM-5, Kimi-K2.5, and DeepSeekk-R1-0528 models.


Stated performance uplift is expressed as a combined average TPS over across the (3) models tested.


Hardware Configuration


Supermicro AS -4126GS-NMR-LCC (board H14DSG-OD)


8x AMD Instinct MI355X. BIOS AMI v1.4a (2025-04-16), GPU firmware SMC 04.86.11.02, TA RAS 27.69.00.10, TA XGMI 32.00.00.20, RLC43, MEC36, SDMA12, Ubuntu 22.04.2 LTS, kernel 5.15.0-70-generic, amdgpu driver 6.16.6, HOST ROCm 7.1.0


Software Configuration(s)


GLM-5 ROCm Docker Image: rocm/sgl-dev:v0.5.8.post1-rocm700-mi35x-20260219


PYTorch Version 2.8.0, SGLang v0.5.8

 

Kimi ROCm Docker Image: vllm/vllm-openai-rocm:v0.16.0, vLLM version 0.16.0


DeepSeek-R1 Docker Image: rocm/7.0:...sgl-dev-v0.5.2-rocm7.0-mi35x-20250915, SGLang version V0.5.13


vs


GLM-5 ROCm Docker Image:


rocm/atom:rocm7.2.2_ubuntu24.04_py3.12_pytorch_release_2.10.0_atom0.1.2.post, ATOM v0.1.2.post


Kimi ROCm Docker Image: vllm/vllm-openai-rocm:v0.22.0, vLLM version V0.22.0


DeepSeek-R1 Docker Image: lmsysorg/sglang-rocm:v0.5.13-rocm720-mi35x-20260612, SGLang version V0.5.13


Server manufacturers may vary configurations, yielding different results. Performance may vary based on configuration, software, vLLM version, and the use of the latest drivers and optimizations. (MI350-81). 

2. Training On ROCm.ai – (MI350-082)

(MI350-82) Testing by AMD Performance Labs as of July 7, 2026, measuring the training performance in tokens per second (TPS) of AMD ROCm 7.0 software vs a preview version of AMD ROCm.ai (ROCm 7.2.2 with optimizations such as Optimized Kernels, Parallelism and Scheduling),) using Megatron -LM on a system with 8x AMD Instinct MI355x 8x GPUs GPU platform running DeepSeek-V2-Lite, DeepSeek-V3-16B, and Qwen3-30B-A3B models.


Stated performance uplift is expressed as the combined average TPS over across the (3) models tested.


Hardware Configuration


Supermicro AS -4126GS-NMR-LCC (board H14DSG-OD)

 

8x AMD Instinct MI355. BIOS AMI v1.4a (2025-04-16), GPU firmware SMC 04.86.11.02, TA RAS 27.69.00.10, TA XGMI 32.00.00.20, RLC 43, MEC 36, SDMA 12, Ubuntu 22.04.2 LTS, kernel 5.15.0-70-generic, amdgpu driver 6.16.6, HOST ROCm 7.1.0


Software Configuration(s)   


DeepSeek-V2-Lite, ROCm 7.2.1 + Primus v26.3


DeepSeek-V3-16B, ROCm 7.2.1 + Primus v26.3

 

Qwen3-30B-A3B, ROCm 7.2.1 + Primus v26.3

 

Server manufacturers may vary configurations, yielding different results. Performance may vary based on configuration, software, and the use of the latest drivers and optimizations. (MI350-82)

Share:

Article By


Senior Manager, Product Marketing

Related Blogs