MangoBoost Delivers Open AI Models Powered by AMD

Oct 09, 2026

Blue-lit data center with server racks and glowing network paths connecting maps of the US and South Korea.

Introducing Mango Inference, a new serverless inference service that provides production-ready access to leading open models through familiar APIs, and usage-based pricing, with infrastructure powered by AMD Instinct™ GPUs and AMD ROCm™ software.

Executive summary

Open models are approaching the quality of closed frontier models across many enterprise workloads. As the performance gap narrows, organizations are placing greater emphasis on serving efficiency, deployment speed, cost control and infrastructure choice.

Mango Inference is a new serverless inference service (LLM-as-a-service) from MangoBoost that gives developers access to a growing catalog of leading open models. The service combines high-performance inference with usage-based pricing and an intuitive management console with OpenAI and Anthropic-compatible APIs. It runs on AMD Instinct™ GPUs and AMD ROCm™ software.

MangoBoost operates AMD Instinct GPU compute in Korea and the United States as a unified serving environment. This approach gives customers access to open models through one endpoint, while supporting regional data residency and lower-latency inference for users in Korea and across APAC.

Mango Inference stack shows apps and APIs over Mango Inference, ROCm software, AMD Instinct MI355X and MI350P GPUs, and datacenters.
Figure 1. Mango Inference architecture, endpoint, and AMD Instinct serving capacity in Korea and the United States

Production ready open model inference

Mango Inference is available as a serverless, pay-per-token service. MangoBoost manages the models, accelerators and data center infrastructure, allowing development teams to focus on applications rather than model-serving operations.

Developers can access multiple models through a single endpoint and change models with a one-line configuration update. Standard APIs also make it easier to adopt the service or move workloads without redesigning the application integration. Figure 2 summarizes what developers get with Mango Inference.

MangoBoost developer features include API compatibility, coding agents, tool calling, prepaid billing, analytics, TLS 1.3 and retries.
Figure 2. What developers get with Mango Inference

Simple integration and usage-based economics

Familiar APIs. The service supports both OpenAI- and Anthropic-compatible clients, as well as coding agents including Claude Code, OpenCode and OpenHands.

Application-ready capabilities. Mango Inference supports streaming, tool calling, structured outputs and reasoning. Standard HTTPS responses, status codes, Retry-After headers and request IDs integrate with existing retry logic and observability tools.

Transparent consumption. Customers fund prepaid accounts and pay by token usage. Each response reports token consumption, giving teams direct visibility into spend. Cached input is priced separately at a lower rate, which can reduce costs for agentic workloads that repeatedly send long system prompts.

Streamlined onboarding. Developers can sign in, generate an API key, test models and connect applications through a unified management console. The console provides access to the model catalog, context-length information, availability and usage controls. (Figure 3 below)

Mango Inference setup: sign in, add funds, create an API key, then call the API or run a coding agent.
Figure 3. Mango Inference offers seamless flow for users to get started and deploy on users’ applications.

Mango Inference also offers a unified user-friendly UI admin console (Figure 4) for users to do everything needed to manage, track, test, deploy inference services. 

Mango Inference console Models page lists DeepSeek-V4-Pro-0813, GLM-5.2 and MiniMax-M3 as available, with Kimi-K3 coming soon.
Figure 4. The Mango Inference console, Models page, with model catalog, context length and availability.

Performance proven on AMD Instinct GPUs

MangoBoost applies full-stack optimization across model serving, system software, networking and distributed infrastructure. Its work with AMD demonstrates this expertise across four consecutive rounds of MLPerf® Inference, the industry-standard benchmark suite from MLCommons.

  • MLPerf Inference v5.0, April 2025. Four AMD Instinct MI300X nodes delivered the first multi-node MLPerf Inference result on AMD Instinct and achieved approximately 103,000 tokens per second on Llama 2 70B Offline, the highest recorded result at that time. (AMD, ROCm)
  • MLPerf Inference v5.1, September 2025. Four MI300X nodes and two MI325X nodes served Llama 2 70B as one system at approximately 160,000 tokens per second and broke the previous record set in v5.0, producing the first MLPerf cluster to combine GPU generations. (AMD)
  • MLPerf Inference v6.0, April 2026. MangoBoost connected MI355X capacity in the United States with MI300X and MI325X systems in Korea to create the first multi-region MLPerf cluster spanning three GPU generations, achieving 95.5% scaling efficiency on Llama 2 70B across two continents. (AMD, ROCm)
  • MLPerf Inference v6.1, September 2026. A 32-GPU deployment across four sites and two continents served GPT-OSS-120B through one endpoint and achieved 97% scaling efficiency, representing the largest multi-region cluster ever submitted to MLPerf. (AMD, ROCm)

For additional information about these milestones, refer to https://mlcommons.org/benchmarks/. 

MangoBoost infographic highlights four MLPerf Inference rounds on AMD Instinct, from v5.0 multi-node to v6.1 multi-region results.
Figure 5. MangoBoost MLPerf Inference results on AMD Instinct GPUs. Quotes from the MLPerf Inference v5.1 announcement (MLCommons, Dell Technologies) and the MLPerf Inference v6.1 announcement (AMD).

The serving optimizations behind these results, including prefill/decode disaggregation and heterogeneous scheduling, are built into Mango Inference. Detailed write-ups of each MLPerf submission, including system configuration and methodology, as of September 2026, are available on the MangoBoost blog.

Model choice with sovereign AI options

Mango Inference currently includes GLM-5.2, DeepSeek-V4 and MiniMax-M3. Its internal automation and optimization processes are designed to bring new open models onto AMD Instinct GPU infrastructure quickly as the market evolves.

For sovereign AI programs, model choice is only part of the requirement. Organizations also need clarity about where inference runs, who operates the infrastructure and how sensitive data moves. MangoBoost operates its own AI data center in Seoul, with AMD Instinct MI355X and MI350P GPU capacity. Korean customers can therefore run workloads on infrastructure located and operated within Korea, alongside access to global serving capacity in the United States.

A common endpoint provides access to every model in the catalog. Teams can select the appropriate model for each workload without changing the application integration, supporting data residency, language requirements and broader model choice.

An open production stack powered by AMD

Mango Inference is built on AMD Instinct GPUs and the open-source AMD ROCm software platform. ROCm software allows access to the software stack required to evaluate, optimize and deploy AI workloads on AMD infrastructure. MangoBoost combines this open foundation with its system-level engineering to operate a production inference service optimized for performance and scale.

The result is an open stack from model to infrastructure: open-weight models, standard APIs, AMD Instinct GPUs, and ROCm software. This flexibility can help enterprises and sovereign AI programs reduce dependence on a single closed model or infrastructure provider.

Get started with Mango Inference

Developers can begin at inference.mangoboost.io by creating an account and API key, selecting a model and connecting an OpenAI- or Anthropic-compatible application. To celebrate the launch, new accounts receive free credits to explore the model catalog and run their first workloads. MangoBoost will also demonstrate the service at SC26.

For support or model requests, email support@mangoboost.io or use the Support menu in the Mango Inference console.

Additional resources

Share:

Article By


CVP, Enterprise AI

Related Blogs