AMD Delivers Its Broadest MLPerf® Inference 6.1 Submission
Sep 16, 2026
Why MLPerf Inference 6.1 Marks a New Stage for AMD Instinct
The AMD MLPerf Inference 6.1 submission is the broadest AMD entry to date, expanding from three model families in the prior MLPerf 6.0 Inference round, to now six, across AMD Instinct™ MI355X and MI350X GPUs, as well as our brand new MI350P PCIe card. The submission spans language, reasoning, text-to-video, and recommendation inference, giving readers a broader view of how the platform performs across different models and deployment formats.
Breadth is only the beginning of the story. The same AMD Instinct MI355X GPU hardware produced substantially more performance than it did one MLPerf cycle earlier. Separately, Crusoe’s 512-GPU GPT-OSS-120B submission reached the highest aggregate token throughput ever submitted in MLPerf history. Seven partners closely reproduced AMD results, adding evidence that the performance can move beyond a reference system.
Three technical milestones help explain what changed in this round:
- More performance from the same AMD Instinct MI355X GPU hardware: Within one MLPerf cycle, continued AMD ROCm™ software optimization increased 8-GPU GPT-OSS-120B throughput, reduced Wan-2.2 latency, and raised cluster throughput with fewer GPUs.
- Leading performance across AMD Instinct MI355X and MI350P GPUs: AMD Instinct MI355X GPUs led NVIDIA B200 and B300 GPT-OSS-120B results at 8 GPUs and the NVIDIA GB200 result at 72 GPUs. In its first MLPerf round, AMD Instinct MI350P GPUs led selected NVIDIA RTX PRO 6000 Server Edition and H200 NVL results.
- Record token throughput and production scale inference with efficient scaling: At 72 GPUs, the AMD Instinct MI355X GPU achieved 95% scale efficiency on GPT-OSS-120B. Separately, Crusoe delivered record aggregate token throughput at 512 GPUs, reaching 5.75 million Offline tokens per second on GPT-OSS-120B and 2.90 million on DeepSeek-R1.
Taken together, the submission shows a platform moving from early workload enablement toward broader, repeatable, and production-scale inference.
One AMD CDNA 4 Foundation for Two Deployment Paths
AMD Instinct™ MI355X GPUs and MI350P PCIe® cards are built on AMD CDNA™ 4 architecture, but they address different infrastructure requirements. The AMD MI355X GPU provides the scale-up foundation for the highest-throughput and largest-cluster results in this submission. The AMD MI350P GPU brings the same architectural generation into a dual-slot PCIe card designed for deployment in existing data center environments.
The AMD Instinct MI355X GPU combines 256 compute units, 288 GB of HBM3E memory, and 8 TB/s of memory bandwidth. It delivers up to 10.1 PFLOPS of peak theoretical MXFP4 and MXFP6 matrix performance through an OAM form factor and UBB 2.0 platform. This configuration is designed for dense systems where maximum inference throughput and scale-up performance are the primary requirements.
The AMD Instinct MI350P PCIe card provides 128 compute units, 144 GB of HBM3E memory, and 4 TB/s of memory bandwidth. It delivers up to 4.6 PFLOPS of peak theoretical MXFP4 and MXFP6 matrix performance in a PCIe 5.0 dual-slot form factor. This is an important portfolio choice because customers do not all build AI systems the same way. The AMD Instinct MI355X GPU supports environments that prioritize maximum scale-up performance, while the AMD Instinct MI350P PCIe card provides a more direct path to AMD CDNA™ 4 inference performance in existing data center infrastructure. Together, the AMD Instinct MI350 Series supports both deployment models.
That choice sets up the rest of the submission. The hardware defines the available compute, memory, and deployment format. The results that follow show how continued work across AMD ROCm™ software and the open-source inference ecosystem produced more performance from that foundation, including measurable gains on the same AMD Instinct MI355X GPUs between MLPerf rounds.
Defining Moments from the AMD MLPerf Inference 6.1 Submission
The AMD MLPerf® Inference 6.1 submission comes together through three proof points: more performance from the same hardware, broad leadership across AMD Instinct™ MI355X and MI350P deployment paths, and production-scale inference with record aggregate token throughput.
1. The Same AMD Instinct™ MI355X GPUs Deliver More Performance
In less than six months, continued software optimization increased performance on AMD Instinct™ MI355X GPUs. On the same 8 AMD Instinct™ MI355X GPUs, continued work across AMD ROCm™ software and the open-source inference ecosystem increased GPT-OSS-120B throughput by 28% in Offline and 38% in Server, while Wan-2.2 SingleStream performance increased 70%.
The impact becomes even clearer at cluster scale. 72 AMD Instinct MI355X GPUs in MLPerf Inference round 6.1 delivered more GPT-OSS-120B throughput than 94 GPUs deliveredin previous round 6.0. The GPU generation stayed the same, but more of its capacity translated into useful inference work. This gives infrastructure teams more headroom from deployed systems before expanding the cluster.
Wan-2.2 turns that platform story into a workload-maturity story. The workload moved from a first-time Open submission in MLPerf Inference round 6.0 to now a leading Closed submission in 6.1. The AMD Instinct MI355X GPU then delivered 18% higher Offline and 11% higher SingleStream performance than the selected NVIDIA B300 submission.
The significance is the speed of that progression. A new workload moved from initial enablement to a standards-based leading result within one MLPerf cycle. That is the value of a software platform that continues maturing after hardware ships: new workloads advance, serving paths improve, and deployed systems keep completing more useful work.
2. AMD Instinct™ MI355X and MI350P GPUs Delivers Broad Leadership Across AI Inference Workloads
MLPerf Inference Round 6.1 demonstrates the breadth of AMD performance leading across language, text-to-video, and recommendation workloads, as well as across several serving scenarios.
GPT-OSS-120B provides the clearest starting point. On a standard AMD Instinct MI355X 8-GPU system, AMD led the selected NVIDIA B200 and B300 results in both Offline and Server. That matters because the eight-GPU node is a practical building block for inference infrastructure.
The other workloads broaden the evidence. Wan-2.2 shows rapid software maturity, Llama 2 70B adds coverage across batch, server, and interactive inference, and DLRM v3 extends the platform into recommendation workloads used for ranking and personalization.
Performance leadership also extends into a different infrastructure model. In its first MLPerf round, the AMD Instinct MI350P PCIe® card submitted five Closed workloads and delivered leading results against selected NVIDIA RTX PRO 6000 Server Edition and H200 NVL submissions.
That changes the deployment conversation. The AMD Instinct MI355X GPU addresses high-density scale-up systems, while the AMD Instinct MI350P GPU gives data centers built around dual-slot PCIe servers another path to AMD CDNA™ 4 inference performance. Teams can evaluate the platform that fits their existing power, cooling, and server architecture rather than treating one form factor as the answer for every environment.
What the Performance Can Mean in Practice
A Signal65 study commissioned by AMD helps translate benchmark performance into operational terms. Across the workloads tested, the study reported lower cost per document, more tokens per dollar, and more users served within a defined latency target.
Those findings explain why throughput matters. The business question comes down to how much useful work the infrastructure can complete within a budget and how much demand it can serve without missing an application’s response-time target.
3. Record token throughput and production scale inference with efficient scaling
Strong performance at one node matters only if additional GPUs translate into useful throughput. AMD expanded GPT-OSS-120B from 8 to 72 AMD Instinct MI355X GPUs, a ninefold increase in GPU count, while retaining 95% scale efficiency in both Offline and Server.
This result shows the platform can add substantial inference capacity without giving much of the expected performance back to communication, scheduling, and workload-distribution overhead. AMD Instinct MI355X GPUs are not only fast at one node; they continue delivering as deployments grow into real multi-node inference clusters. At 512 GPUs, the story shifts from scaling efficiency to aggregate throughput. Through Crusoe, GPT-OSS-120B reached 5.75 million tokens per second in Offline and 5.39 million tokens per second in Server, the highest aggregate token throughput ever submitted in MLPerf history.
The DeepSeek-R1 submission reached 2.90 million Offline and 2.41 million Server tokens per second, establishing the highest aggregate token throughput on DeepSeek-R1 in MLPerf history.
Taken together, the three defining moments show a platform that improves on the same hardware, delivers broad leadership across scale-up and PCIe® deployment paths, scales efficiently from eight to 72 GPUs, and reaches record aggregate token throughput at 512 GPUs.
AMD Instinct™ Performance Reproduced Across Partner Systems
Dell Technologies and MangoBoost submitted the first 32-GPU heterogeneous inference result, combining 16 AMD Instinct™ MI300X GPUs in Korea with 16 AMD Instinct MI355X GPUs in the United States as one GPT-OSS-120B serving endpoint. The system delivered 285,454 Offline tokens per second and 253,501 Server tokens per second.
The significance is how that capacity came together. Existing AMD Instinct MI300X GPU systems and newer AMD Instinct MI355X GPU systems contributed to the same inference service, showing a path to adding capacity across GPU system generations and locations rather than operating each fleet as a separate environment.
That flexibility sits within a broader partner story. Dell Technologies, Oracle, Hewlett Packard Enterprise, Supermicro, MangoBoost, Crusoe, and MiTAC submitted results across AMD Instinct MI350X, MI350P, and MI355X GPU systems. The average of comparable AMD MI355X GPU partner results landed within 4% of AMD, and some matched or slightly exceeded the AMD reference result.
A benchmark result becomes more useful when it can be repeated outside the original lab. These partner submissions show that consistent AMD Instinct performance can carry across different systems, giving infrastructure teams more confidence when evaluating OEM, cloud, and software-partner deployments.
Choosing the Right Performance Profile for the Data Center
Not every inference deployment is designed around maximum throughput. Power, cooling, and existing infrastructure can be just as important when teams decide how much GPU performance to place in each server.
Across selected GPT-OSS, Llama, Wan, and DLRM workloads, the eight-GPU AMD Instinct™ MI350X GPU platform retained approximately 80% of the AMD Instinct MI355X platform performance. The AMD Instinct MI355X GPU has a 40% higher rated TDP in this comparison, reflecting its position as the maximum-performance option in the portfolio.
That creates two performance profiles:
- AMD Instinct MI350X GPU: A lower-TDP option for deployments where power, cooling, and infrastructure flexibility shape the system design.
- AMD Instinct MI355X GPU: The maximum-performance option for deployments where inference throughput is the primary requirement.
How AMD ROCm™ Software Connects Benchmark Progress to Production AI
AMD ROCm™ software connects the full MLPerf® Inference 6.1 story. AMD ROCm software v7 powered every result in the submission, including higher performance on the same AMD Instinct™ MI355X GPUs, the move from initial Wan-2.2 enablement to a Closed submission, 95% scale efficiency at 72 GPUs, and partner results that closely tracked AMD.
Those milestones represent different stages of the same workload journey. A model must first run correctly, then its attention operations, GEMMs, KV-cache use, scheduling, and communication paths must be tuned for the target system. The result is not simply a higher benchmark score. It is a workload that can make better use of deployed hardware and produce comparable results across AMD and partner systems.
ROCm 10.0 builds on that foundation with a clearer path into production. Teams can start with familiar vLLM and SGLang software, use ROCm.AI tools to diagnose performance, and carry the same software foundation from one system into multi-node infrastructure. Local profiling and tuning also give developers a more direct view of where time is being spent across the serving path.
The story is one of continuity: bring up the workload, improve the serving path, expand the deployment, and reproduce the result. ROCm provides the software foundation across each step.
Annual Cadence Builds the Next Chapter of AMD Instinct™ Inference
The value of the AMD Instinct roadmap is not simply a sequence of product names. It is the cadence behind the platform progress shown throughout MLPerf® Inference round 6.1.
AMD Instinct MI300X GPUs established the foundation in 2023. AMD Instinct MI325X GPUs extended that platform with HBM3E memory and increased compute in 2024, followed by the AMD Instinct MI350 Series in 2025. The current submission shows how that foundation continues to gain value through software optimization, broader workload support, and more deployment choices.
The roadmap continues with the AMD Instinct MI400 Series in 2026, positioned with HBM4 memory and increased compute, followed by the MI500 Series and a next-generation architecture in 2027. For infrastructure teams, that cadence provides a clearer path from the systems being deployed today to the compute and memory requirements of future inference workloads.
Final Takeaway
The AMD MLPerf® Inference round 6.1 submission marks a major step forward for AMD Instinct AI inference. AMD delivered more performance from the same AMD Instinct™ MI355X GPU hardware, leadership across selected AMD Instinct MI355X and MI350P GPU comparisons, 95% scale efficiency at 72 GPUs, record aggregate token throughput through 512 GPU submission by Crusoe, and results from seven ecosystem partners.
The results also show a broader platform coming together. AMD Instinct GPUs provide the hardware foundation across scale-up, lower-TDP, and PCIe® deployment paths. AMD ROCm™ software connects workload enablement, continued optimization, multi-node execution, and reproducibility across partner systems.
END Notes
END Notes
GENERAL DISCLAIMER The information contained herein is for informational purposes only and is subject to change without notice. While every precaution has been taken in the preparation of this document, it may contain technical inaccuracies, omissions and typographical errors, and AMD is under no obligation to update or otherwise correct this information. Advanced Micro Devices, Inc. makes no representations or warranties with respect to the accuracy or completeness of the contents of this document, and assumes no liability of any kind, including the implied warranties of noninfringement, merchantability or fitness for particular purposes, with respect to the operation or use of AMD hardware, software or other products described herein. No license, including implied or arising by estoppel, to any intellectual property rights is granted by this document. Terms and limitations applicable to the purchase or use of AMD products are as set forth in a signed agreement between the parties or in AMD's Standard Terms and Conditions of Sale. GD-18u. © 2026 Advanced Micro Devices, Inc. All rights reserved. AMD, the AMD Arrow logo, AMD Instinct, and combinations thereof are trademarks of Advanced Micro Devices, Inc. The MLPerf name and logo are registered and unregistered trademarks of MLCommons Association in the United States and other countries. All rights reserved. Unauthorized use is strictly prohibited. See www.mlcommons.org for more information. Other product names used in this publication are for identification purposes only and may be trademarks of their respective owners. Certain AMD technologies may require third-party enablement or activation. Supported features may vary by operating system. Please confirm with the system manufacturer for specific features. No technology or product can be completely secure. |