From Sensor to Insight: Production-Ready Single-Chip Edge AI. Proven.
Sep 22, 2026
Great platforms are not defined by silicon alone, but also by the software ecosystem that unlocks their full potential. When we launched the AMD Versal™ AI Edge Series Gen 2, we set out to redefine edge and physical AI with a unified AI platform encompassing preprocessing, AI inference, and postprocessing. Today, the production-ready capabilities of Vitis™ AI, the AMD unified AI toolchain for adaptive SoCs, bring that vision to life. With all Versal AI Edge Series Gen 2 devices in production, developers can move beyond multi-chip complexity to true single-chip intelligence at the edge. Combining real-time sensor processing, power-efficient AI inference, and high-performance embedded compute in one platform, the Versal AI Edge Series Gen 2 simplifies system design while enabling high performance and efficiency. Paired with Vitis AI, it transforms the vision of ‘single-chip end-to-end AI acceleration’ into a production-ready reality, delivering the model support, software maturity, and proven inference performance needed to deploy edge and physical AI at scale.
Where Limits Define Innovation: The Edge and Physical AI Challenge
Edge and physical AI deployments must operate within tight power, thermal, memory, and compute limits. They also demand deterministic latency and high reliability across the full pipeline, not just the AI accelerators. Across applications ranging from advanced driver assistance systems and autonomous robotics to medical imaging, AI systems are increasingly expected to process multiple sensor streams, execute AI inference in real time, and deliver actionable insights within demanding operating environments. In physical AI, closed-loop perception, processing, and action add even more system complexity. Success, therefore, depends not on the performance of a single processing element, but on how efficiently the entire AI pipeline operates as a unified system. In addition, long deployment lifecycles require such systems to adapt to evolving sensors, interfaces, and AI workloads, without costly redesigns. Together, these requirements drive the need for tightly integrated platforms that deliver consistent, predictable performance across the full AI pipeline while retaining flexibility to support evolving application requirements.
AMD Versal AI Edge Series Gen 2: Heterogeneous Compute for End-to-End Acceleration on a Single Chip
Purpose-built for edge and physical AI, the Versal AI Edge Series Gen 2 meets these demands by bringing heterogeneous compute together in a single adaptive SoC, matching each stage of the AI pipeline to the processing engine best suited for the task.
- Programmable logic handles sensor connectivity, flexible I/O, and real-time preprocessing.
- AI Engines deliver power-efficient, high-performance inference without consuming programmable logic resources.
- High-performance scalar compute supports postprocessing and time-sensitive decision making.
- Hard IP (ISP, GPU) helps improve efficiency and leaves more programmable logic for differentiation.
- Single-chip integration reduces system footprint, communication overhead, and thermal burden.
- Integrated functional safety and portfolio scalability support long-life, real-world deployments.
Collectively, these silicon capabilities provide a versatile foundation for accelerating every stage of an embedded AI pipeline. Fully realizing their system-level benefits, however, requires a unified software stack that complements the architecture and coordinates workloads across it. This is where Vitis AI plays a critical role.
AMD Vitis AI: Production-Ready Inference Unlocked for Versal AI Edge Series Gen 2 Devices
Vitis AI provides a unified environment for implementing complete AI pipelines on Versal AI Edge Series Gen 2, helping accelerate time to market while supporting flexible deployment configurations. Developers can combine optimized AI inference with programmable-logic-accelerated preprocessing, while efficiently allocating device resources across concurrent workloads. Learn more about how Vitis AI seamlessly integrates inference into edge AI pipelines here.
Vitis AI 6.2, launched earlier this year, establishes the deployment-ready software foundation, helping developers move from evaluation to real-world deployment. It introduces out-of-the-box support for modern AI models, including convolutional neural networks and vision transformers, along with INT8, BF16, FP16, and mixed-precision implementations. This flexibility allows developers to select the performance and accuracy trade-offs best suited to their application.
Multi-Camera Performance through AI Engine Array Partitioning
One key capability in Vitis AI is its ability to assign individual compute processes to physical partitions of the AI Engine array. This feature provides designers with options when optimizing for latency, throughput, and power consumption. To demonstrate its system-level benefits, the Versal AI Edge Series Gen 2 was evaluated using a multi-camera object detection and classification workload that combined real-time video processing, AI inference, and decision-making on a single device.
In this representative deployment scenario, a 2VE3858 device processed four simultaneous 1080p camera streams running YOLOv8m object detection. Programmable logic and image processing hard IP accelerated sensor connectivity and preprocessing, the AI Engine array executed inference, and the integrated Arm® processing system handled postprocessing and decision-making. The AI Engine array in the 2VE3858 device is a high-performance inference accelerator capable of sustaining an aggregate throughput of 300 fps1 in batch 4 mode. For this application, the system's target throughput was set at 30 fps per stream to match the camera rate, yielding an aggregate target of 120 fps. By leveraging the resource-allocation flexibility enabled by physical partitioning, the YOLOv8m inference target for the four camera streams was met using only 64 of the 144 AI Engine tiles, freeing the remaining tiles for other compute needs or simply for reducing system power.
Next Up: Vitis AI 6.3
Vitis AI 6.3, releasing shortly after this blog is published, extends the momentum of the 6.2 release with improvements in performance, ease of use, and deployment optimization. New agentic workflows help automate custom operator development and provide quantization and mixed-precision guidance, while performance enhancements for CNN and vision transformer workloads further improve inference efficiency. Additional device support across the Versal AI Edge Series Gen 2 portfolio and enhanced AI Analyzer capabilities, including deeper workload insights and an integrated chatbot experience, further simplify development, performance analysis, and system optimization.
Start Your Versal AI Edge Series Gen 2 Inference Design Today
- Download the latest Vitis AI release and evaluate out-of-the-box performance on supported ML models.
- Compile and optimize custom models using available tutorials and example designs.
- Build custom AI solutions:
- Integrate your ML model into an AI pipeline using ecosystem-aligned tools, including ONNX Runtime and GStreamer plugins available through VVAS.
- Reduce time to market by leveraging bundled AI Skills.
- Order VEK385 boards for prototyping and development of your custom designs, and build complete sensor-to-inference pipelines using familiar frameworks.
- Target the Versal AI Edge Series Gen 2 device that best meets your application requirements using AMD Vivado™ 2026.1. Download the latest Vivado tools here.
Related Blogs
Footnotes
- AI Engine latency measurements were performed on an AMD Versal AI Edge Series Gen 2 VEK385 Evaluation Kit by AMD Engineering using a pre-release version of AMD Vitis AI 6.3. YOLOv8m 3x640x640 resolution and INT8/BF16 mixed precision quantization, configured with tp_size = 2 and dp_size = 4. Final Vitis AI 6.3 release numbers may vary. The camera-based design with complete pipeline was implemented with Vitis AI 6.2, using four independent batch 1 NPUs, each assigned to a unique thread and camera stream. Each NPU configured as tp_size = 1 and dp_size = 1. (VAI-001)
- AI Engine latency measurements were performed on an AMD Versal AI Edge Series Gen 2 VEK385 Evaluation Kit by AMD Engineering using a pre-release version of AMD Vitis AI 6.3. YOLOv8m 3x640x640 resolution and INT8/BF16 mixed precision quantization, configured with tp_size = 2 and dp_size = 4. Final Vitis AI 6.3 release numbers may vary. The camera-based design with complete pipeline was implemented with Vitis AI 6.2, using four independent batch 1 NPUs, each assigned to a unique thread and camera stream. Each NPU configured as tp_size = 1 and dp_size = 1. (VAI-001)