Overview

Optimized GPU Software Stack

AMD ROCm™ is an open software platform that helps data center teams quickly set up, deploy, and scale AI and HPC workloads across GPUs and nodes. Developers can work with familiar frameworks, tune performance-critical code, and use HIP to adapt existing GPU applications and frameworks for supported AMD hardware.

Image Zoom
ROCm AI stack diagram showing Silo AI, Primus, ROCm Inference, open ecosystem tools, Core SDK, expansions, and deployment.

← Watch Video

ROCm.AI: Accelerating Development Velocity and Performance Optimization

An AI-native developer experience that accelerates ROCm development, simplify AI workload development and optimization, and maximize performance on AMD hardware

Development

Build, Train, and Deploy AI at Scale with AMD ROCm Software

AMD ROCm™ software brings together open frameworks, optimized libraries, model tools, and scale-out infrastructure to support AI development from experimentation to production.

Build

Simplified Model Development

Develop with PyTorch, TensorFlow, and JAX using AMD-optimized containers and workflows for popular open models, more than 3M+ from Hugging Face.

Train

AMD Primus - Train AI Models at Scale 

Use AMD Primus and ROCm communication libraries for pre-training, fine-tuning, and distributed training across AMD Instinct™ GPUs.

Optimize

Optimize and Serve AI Models

Use AMD Quark, AITER, vLLM, and SGLang to optimize models, accelerate critical operations, and deploy high-throughput inference.

Deploy

Deploy Across AI Infrastructure

Scale from individual GPUs to multi-node environments with containers, Kubernetes, Slurm, telemetry, and lifecycle-management tools.

Build, Optimize, Scale, and Run High Performance Computing with AMD ROCm

AMD ROCm™ software helps HPC developers build or port scientific applications, profile and tune performance, and move from a single GPU to multi-node runs.

Build

Build and Port Scientific Applications

Develop GPU-accelerated code with HIP or OpenMP® offload, work with HPC frameworks such as Kokkos and RAJA, and use HIPIFY to help translate existing CUDA® code.

Optimize

Profile and Tune Performance

Profile kernels and data movement with ROCm tools, then use optimized libraries for linear algebra, FFTs, solvers, and sparse operations. Validate results as you tune.

Scale

Scale Across GPUs and Nodes

Use GPU-aware MPI for distributed applications and rocSHMEM for GPU-initiated communication, then measure performance as GPU and node counts grow.

Run

Run on HPC Infrastructure

Package applications in HPC containers and submit cluster jobs with Slurm. ROCm is used on leadership-class systems including El Capitan, Frontier, and LUMI.

What’s New

ROCm 10.0: Built for the Age of Agentic AI

Icon
Build with AMD Expertise

AMD Skills brings curated AMD knowledge and validated workflows into AI coding agents, helping developers build AMD-optimized applications with guidance in the tools they already use.

Icon
Run with Simpler Workflows

The ROCm CLI provides a unified interface to set up, manage and operate AI workloads on AMD hardware.

Icon
Optimize End-to-End Inference

ROCm Hyperloom uses agentic AI to profile workloads, identify bottlenecks, implement optimizations and validate performance across host code and GPU kernels.

Icon
Build on a Modular ROCm Core SDK

The ROCm Core SDK provides a modular foundation for building, deploying and optimizing AI workloads across AMD platforms.

Ecosystem and Partners

A broad, open AI ecosystem built with leading frameworks, technology partners, and developers to accelerate innovation, enable portability, and scale AI from development to production.

Supported Hardware

Developer Resources

ROCm Developer screen shot

Frequently Asked Questions

AMD ROCm™ is an open software stack for developing AI and high-performance computing applications on supported AMD GPUs. Teams use it to train and serve AI models, build GPU-accelerated software, and run scientific workloads across GPUs and nodes.

ROCm and CUDA are separate GPU computing platforms. ROCm provides an open software stack for supported AMD GPUs, and its HIP programming model offers a path for porting CUDA source code. The tools and APIs are not directly interchangeable, so migration effort depends on the application and its dependencies.

Not always. Applications built on supported AI frameworks may be able to start with a ROCm-enabled version of the framework. Code that calls CUDA APIs or relies on CUDA-specific libraries may need changes. HIPIFY helps translate supported CUDA code to HIP, after which developers should test correctness and tune performance as needed.

ROCm supports PyTorch, TensorFlow, and JAX for AI development, as well as vLLM and SGLang for model serving. Supported versions depend on the ROCm release and system configuration. Check the current compatibility matrix before selecting a framework version.

First, check the compatibility matrix for your GPU and operating system. Then follow AMD’s installation guide for that configuration. AMD Infinity Hub offers preconfigured containers, while ROCm CLI can assist with setup and diagnostics on supported systems. Run a sample workload to verify your environment before scaling.

ROCm Newsletter

Receive the latest ROCm news.

Footnotes

©2024 Advanced Micro Devices, Inc. All rights reserved. AMD, the AMD Arrow logo, AMD ROCm, AMD Instinct, EPYC, Radeon Instinct, and combinations thereof are trademarks of Advanced Micro Devices, Inc. PyTorch is a trademark or registered trademark of PyTorch. Other product names used in this publication are for identification purposes only and may be trademarks of their respective companies.

  1. For a full list of Radeon parts supported by ROCm, go to https://rocm.docs.amd.com/en/latest/reference/gpu-arch-specs.html