AMD Primus

Train AI Models at Scale — Without the Complexity

Primus brings training setup, optimized transformers, and production operations into one modular experience. Move from first run to fleet-scale training quickly with reproducible workflows, kernel acceleration, cluster validation, and monitoring baked in.

With Primus, teams can:
Software icon

Monitor distributed jobs and debug issues faster

Gear icon

Launch reproducible training experiments with a single CLI

Fast icon

Accelerate transformer workloads using optimized operators

Accomplishments icon

Validate clusters before large training runs

Black fade

The Primus Platform

Primus provides modular components that support the entire training lifecycle — from experiment setup to cluster-scale operations.

Frameworks

Support for Leading Frameworks

  • Megatron-LM
  • PyTorch TorchTitan
  • JAX MaxText

AMD Primus

Training Workflows On AMD ROCm™

Reference Architecture For Training At Scale

  • Configure
  • Validate
  • Accelerate
  • Observe
  • Optimize

AI Developers


none

Validate + Observe

  • Primus Tuning Agent & Primus Projection
  • Primus-Bench & Primus Preflight
  • Primus-Lens
none

Production pre-training

  • Primus-LM
  • Primus-Turbo
  • Primus-SaFE
none

Production Scale RL - Coming Soon

RL Workflow Path

  • OSS RL Stack
  • veRL, Miles & Slime Support
  • AMD Validated Dockers & Upstream CI/CD

Built On

AMD ROCm™ Core SDK

Libraries | Compilers | Tools | Runtimes

Performance Benchmarks

Primus benchmarks show how customers can turn AMD Instinct™ GPU capacity into fast model development — scaling from a single node to multi-node clusters while maximizing GPU utilization, shortening time-to-train, and improving training efficiency across diverse model architectures.

ROCm.AI vs. ROCm 7

Accelerating Training Improvement

DeepSeek-V3-16B

DeepSeek-V2-Lite

Qwen3-30B-A3B

Release-to-Release: Performance Improvement
2.01x
2.54x
2.76x
On Average 2.4X
faster than previous release across models

Parallelization & Scheduling

Pipeline parallelism, sharded data parallel, gradient all-reduce overlap

Memory Management

Activation checkpointing, FP8 mixed precision, fused gradient accumulation

Optimized Kernels

Fused FlashAttention (fwd + bwd), grouped GEMM • fused optimizer step

Tested at Scale, Proven in Production

Discover how Zyphra built ZAYA1, from scratch on a full-stack AMD platform. This collaboration shows a competitive model, trained on over 1K GPUs, converging in production settings on an E2E Instinct stack.

In terms of MoE models, ZAYA1-base outperforms the recently released MoE models of similar scales on mathematics and general knowledge evaluations, while lagging slightly in coding. ZAYA1-base dramatically outperforms prior open MoE models such as OLMoE (muennighoff2024olmoe) demonstrating both to the strength of the ZAYA1 architecture as well as to improvements in broader pretraining recipes and datasets that have occurred recently.

Model MMLU(0) Acc(%) MMLU-Pro(5) Acc(%) GPQA(0) Acc(%) MATH-hard(4) Exact-Match(%) MBPP+(3) Pass@1(%)
ZAYA1-base 67.01 40.43 30.70 54.15 75.40
Qwen3-1.7B 54.10 32.4 28.9 33.2 55.82
Qwen3-4B 68.31 41.92 33.72 47.05 76.46
OLMoE-1b-7b 53.40 19.71 26.34 4.68 29.10
Qwen3-8b 74.62 47.24 36.07 28.17 81.00
Gemma3-12b-pt 71.44 42.25 35.15 17.98 75.40
Llama3.1-8b 63.29 32.71 31.12 6.50 62.70

TABLE IV: Comparison of ZAYA1-base on core general knowledge, mathematics, and coding evaluations. ZAYA1-base performs extremely strongly considering its active parameter count, significantly outperforming similar MoE models such as OLMoE and going head to head against extremely strong much larger dense models such as Qwen3-4b and Gemma3-12b.

Black fade