AMD Primus
Train AI Models at Scale — Without the Complexity
Primus brings training setup, optimized transformers, and production operations into one modular experience. Move from first run to fleet-scale training quickly with reproducible workflows, kernel acceleration, cluster validation, and monitoring baked in.
With Primus, teams can:
Monitor distributed jobs and debug issues faster
Launch reproducible training experiments with a single CLI
Accelerate transformer workloads using optimized operators
Validate clusters before large training runs
The Primus Platform
Primus provides modular components that support the entire training lifecycle — from experiment setup to cluster-scale operations.
Frameworks
Support for Leading Frameworks
Dense LLMs, MoEs, Diffusion, Multimodal, RnR
- Megatron-LM
- Pytorch TouchTitan
- JAX MaxText
AMD Primus
Training Workflows On AMD ROCm™
Reference Architecture For Training At Scale
- Configure
- Validate
- Accelerate
- Observe
- Optimize
AI Developers
Validate + Observe
- Primus Tuning Agent & Primus Projection
- Primus-Bench & Primus Preflight
- Primus-Lens
Production Scale RL - Coming Soon
RL Workflow Path
- OSS RL Stack
- veRL, Miles & Slime Support
- AMD Validated Dockers & Upstream CI/CD
Built On
AMD ROCm™ Core SDK
Libraries | Compilers | Tools | Runtimes
Performance Benchmarks
Primus benchmarks show how customers can turn AMD Instinct™ GPU capacity into fast model development — scaling from a single node to multi-node clusters while maximizing GPU utilization, shortening time-to-train, and improving training efficiency across diverse model architectures.
ROCm.AI vs. ROCm 7
Accelerating Training Improvement
DeepSeek-V3-16B
DeepSeek-V2-Lite
Qwen3-30B-A3B
Parallelization & Scheduling
Pipeline parallelism, sharded data parallel, gradient all-reduce overlap
Memory Management
Activation checkpointing, FP8 mixed precision, fused gradient accumulation
Optimized Kernels
Fused FlashAttention (fwd + bwd), grouped GEMM • fused optimizer step
Tested at Scale, Proven in Production
Discover how Zyphra built ZAYA1, from scratch on a full-stack AMD platform. This collaboration shows a competitive model, trained on over 1K GPUs, converging in production settings on an E2E Instinct stack.
Get Started with Primus
The Journey from Model Selection to Training at Scale
Experience the production tested, reference architecture for building frontier AI models from scratch on Instinct!
Resources