AMD Primus

Train AI Models at Scale — Without the Complexity

Primus brings training setup, optimized transformers, and production operations into one modular experience. Move from first run to fleet-scale training quickly with reproducible workflows, kernel acceleration, cluster validation, and monitoring baked in.

With Primus, teams can:
Software icon

Monitor distributed jobs and debug issues faster

Gear icon

Launch reproducible training experiments with a single CLI

Fast icon

Accelerate transformer workloads using optimized operators

Accomplishments icon

Validate clusters before large training runs

Black fade

The Primus Platform

Primus provides modular components that support the entire training lifecycle — from experiment setup to cluster-scale operations.

Frameworks

Support for Leading Frameworks

Dense LLMs, MoEs, Diffusion, Multimodal, RnR

  • Megatron-LM
  • Pytorch TouchTitan
  • JAX MaxText

AMD Primus

Training Workflows On AMD ROCm™

Reference Architecture For Training At Scale

  • Configure
  • Validate
  • Accelerate
  • Observe
  • Optimize

AI Developers


none

Validate + Observe

  • Primus Tuning Agent & Primus Projection
  • Primus-Bench & Primus Preflight
  • Primus-Lens
none

Production pre-training

  • Primus-LM
  • Primus-Turbo
  • Primus-SaFE
none

Production Scale RL - Coming Soon

RL Workflow Path

  • OSS RL Stack
  • veRL, Miles & Slime Support
  • AMD Validated Dockers & Upstream CI/CD

Built On

AMD ROCm™ Core SDK

Libraries | Compilers | Tools | Runtimes

Performance Benchmarks

Primus benchmarks show how customers can turn AMD Instinct™ GPU capacity into fast model development — scaling from a single node to multi-node clusters while maximizing GPU utilization, shortening time-to-train, and improving training efficiency across diverse model architectures.

ROCm.AI vs. ROCm 7

Accelerating Training Improvement

DeepSeek-V3-16B

DeepSeek-V2-Lite

Qwen3-30B-A3B

Release-to-Release: Performance Improvement
2.01x
2.54x
2.76x
On Average 2.4X
faster than previous release across models

Parallelization & Scheduling

Pipeline parallelism, sharded data parallel, gradient all-reduce overlap

Memory Management

Activation checkpointing, FP8 mixed precision, fused gradient accumulation

Optimized Kernels

Fused FlashAttention (fwd + bwd), grouped GEMM • fused optimizer step

Tested at Scale, Proven in Production

Discover how Zyphra built ZAYA1, from scratch on a full-stack AMD platform. This collaboration shows a competitive model, trained on over 1K GPUs, converging in production settings on an E2E Instinct stack.

Black fade