Skip to main content

Persistent KV Cache for Continuous Inference with VAST Data

abstract background

Abstract

Modern agentic AI systems require persistent context and high-throughput inference infrastructure that scales efficiently. This session explores the role of KV cache on inference workloads on AMD Instinct GPUs, highlighting the advantages of AMD memory architecture for long-context and continuous inference systems. Learn how the VAST AI OS enables persistent KV cache and context-aware inference pipelines, reducing recomputation while improving performance, efficiency, and scalability.

July 22, 2026 3:00 PM - 3:25 PM PDT

Speakers


Presented By