Advancing Kimi K3 inference on AMD Instinct MI355X GPUs with TokenSpeed
Oct 09, 2026
TL;DR
Multi-vendor agentic inference, with AMD at the core: TokenSpeed combines a C++ scheduler, Python execution plane, and modular kernel APIs, with unified cache management and execution across accelerators.
Leading Kimi K3 performance on 8x AMD Instinct™ MI355X GPU: 216 output tokens/s per user at concurrency 1, 1.4x of ATOM, on par performance at higher concurrency, with higher precision.
Leveraging Triton foundation, agent-assisted Gluon optimization: Specialized kernels, fusion, and Iris communication deliver up to 5.1x faster MLA prefill, 2.1× faster speculative attention, and 4.4× faster MoE all-reduce against the baselines detailed below.
TokenSpeed: Agentic Inference with AMD at its core
TokenSpeed is a recent open-source LLM inference engine designed for low-latency generation on agentic workloads. Coding agents repeatedly return to long conversations, process new tool results, and generate the next action while the user waits. Serving them well requires low decode latency, efficient prefix reuse, and predictable scheduling.
TokenSpeed gives each part of that problem a clear home. A C++ control plane manages request lifecycles and cache ownership as a finite-state machine, with the type system enforcing safe resource reuse. A Python execution plane supports rapid model development. A layered kernel subsystem connects portable model code to specialized implementations through shared APIs, a central registry, and an accelerator plugin mechanism. Prefix caching, speculative decoding, graph execution, and disaggregated serving build on these common mechanisms.
AMD support is fundamental to this design. TokenSpeed enabled Kimi K3 on CDNA4 at Day 0, with prefix caching, speculation, disaggregated serving, and graph-captured decode. TML Inkling also arrived at Day 0 with AMD support, an AMD Quark MXFP4 checkpoint, and Gluon specialized kernels. The same kernel system now covers MHA, MLA, DSA, KDA, MoE, and GEMM on AMD. These results show how common model and scheduling infrastructure can support new architectures while specialized kernels exploit the hardware.
Kimi K3 is a frontier agentic model, and this post shows how TokenSpeed serves its long conversations on MI355X, especially for high interactivity. The multi-turn measurements below report 216 output tokens/s per user at concurrency 1 on eight MI355X GPUs.
Our optimization process starts with portable Triton implementations as a functional baseline and numerical reference. Profiling then guides agent-assisted development of specialized Gluon kernels that implement the same operator interfaces and meet the same numerical contracts. The approach extends across the stack: combining projections, retaining intermediates on chip, and fusing computation into Iris kernels, a Triton-extension that enables multi-GPU kernel programming.
Kimi K3 Performance on Long Agentic Conversations
Agentic coding workloads repeatedly revisit a long conversation while the user waits for the next tool call. To measure this pattern, we replayed multi-turn SWE-smith coding sessions with TokenSpeed’s open-source agentic benchmark. Each conversation opens with a ~50K-token prompt, adds ~800 tokens per turn over 10–15 turns, and generates 500 tokens per reply. The existing run uses eight AMD Instinct MI355X GPUs, TP8, EAGLE3, FP8 KV, and an approximately 90% cache-hit rate.
With one active session, this run reports 216 output tokens/s per user, or 4.62 ms per output token (TPOT), and a 0.62-second average time to first token (TTFT). At this high interactivity situation, which the industry is heavily pushing towards, it achieved 1.4× compared to ATOM. Performance is on par with or slightly faster than ATOM as we scale up to concurrency 2, 4, and 8. At concurrency 16 (C16), each agent receives 39 output tokens/s.
Figure 1: Kimi K3 agentic serving on eight MI355X GPUs. Per-user decode speed is 1000 divided by time per output token in milliseconds; it excludes time to first token and queuing. Total-token throughput counts input and output tokens, including cached prompts. Each point is a single run at one concurrency setting; serving conditions are described in the configuration section.
ATOM leads at C16. At C16 ATOM switches to a different scheduling approach termed prefill coalescing. This delays prefill requests, allowing more decode requests through, until enough prefill requests are available to be batched. This trades TTFT (1.44 -> 2.90s mean) for TPOT (25.17 -> 19.28ms mean) compared to default split scheduling and results in an overall throughput win, however p95 user request TTFT grows 3.54x (2.95 -> 10.46s) in our measurements. ATOM also conducts more aggressive quantization as shown in the following table:
Table 1: Key-layer precisions differ in the TokenSpeed-ATOM comparison above. The table shows the precision used by each benchmarked serving recipe; these are not matched-precision configurations. Projection and expert pairs are activation/weight; KDA state and KV cache entries show storage precision.
Model Component
TokenSpeed
ATOM
Attention Q/K/V and output projections
BF16/BF16
FP8/FP8
Shared-expert projections
BF16/BF16
FP8/FP8
Routed-expert GEMMs
FP8/MXFP4 (small decode); MXFP4/MXFP4 (prefill)
MXFP4/MXFP4 (prefill)
KDA recurrent state
FP32
FP16
MLA KV cache
FP8
FP8
Kimi K3 Architecture and the Eight GPU Execution Path
Large language models generate text one token at a time. Each new token flows through attention, which reads the conversation, and a feed-forward network, which transforms the current representation. Kimi K3 has 93 decoder layers: 69 use Kimi Delta Attention (KDA), and 24 use gated Multi-head Latent Attention (MLA). Its first feed-forward layer is dense; the remaining 92 use Latent MoE. Attention Residuals (AttnRes) connect representations across layer blocks (Figure 2).
Figure 2: Kimi K3 architecture and the eight-GPU execution path. The section numbers show where each part of the layer is covered.
We use eight-way tensor parallelism for attention and MoE, with no expert parallelism (attention TP8; MoE TP8/EP1). Attention heads and each expert’s intermediate dimension are split across eight GPUs. Each GPU computes a partial result, and collectives combine the parts before dependent work begins. All benchmarks in this post were run on MI355X (CDNA4, gfx950).
Component Optimizations and Performance
Each component below begins with its role in the model, then covers kernel optimization, fusion, and distributed work where they apply. Performance numbers name their baseline and hardware.
1. KDA: Keeping Recurrent State on Chip
Standard attention stores information for every earlier token and reads that history for each new token, so its work grows with conversation length. KDA carries a fixed-size recurrent state instead. Each token updates the state and reads from it. The KDA recurrence therefore has constant work per decode step as context grows; the model’s periodic MLA layers still read the full history (Figure 3).
Figure 3: Standard attention reads a growing history; KDA updates a fixed-size recurrent state.
1.1 State layout and prefill data movement
KDA does little arithmetic per byte, so speed comes down to data movement. The recurrence computes one output value by reducing its state row across the key channels. We therefore store the state V-major as [value, key], making each reduction row contiguous. This benefits both the prefill scan and decode. The decode state uses the order its kernel reads, halving the single-user state-update time from 10.5 to 5.2us.
For prefill, the prompt is divided into 64-token chunks. Most chunk-local matrix work runs in parallel, while a scan carries the recurrent state from one chunk to the next. Our Gluon scan keeps that state in registers and produces the output in the same pass, rather than writing a checkpoint after every chunk.
Separately from the V-major state layout, we reorganize the temporary gated-key buffer used by prefill. We change it from token-major to chunk-major order, placing the scan’s reduction dimension contiguously in memory. A device-side chunk index removes per-chunk address work, and the scheduled load/MFMA pipeline overlaps memory movement with computation. These reduce the complete 8K prefill call from 659 to 595µs, a 9.7% reduction.
1.2 Fused decode and speculative state updates
A decode step originally launched separate convolution, recurrence, and normalization kernels, each writing data for the next to read. The fused Gluon kernel keeps those intermediates on chip. It also consumes packed and token-strided projection outputs directly, avoiding extra QKV and gate copies around the decode kernel. The f_b decay projection is also computed inside the recurrence kernel, removing one more launch while retaining the required BF16 rounding before gate processing. Speculative verification avoids saving a state for every rejected position; a replay step commits the accepted prefix.
1.3 KDA prefill, decode, and verification gains
1.6x faster Gluon chunked prefill than the open-source Triton KDA kernel (8K tokens, 1055 -> 659us)
1.11x faster KDA prefill call time at 8K after further reorganizing the gated-key buffer (8K tokens, 659- > 595us)
1.63x faster decode step after fusing four kernels into one (16 users; 25.4 → 15.6 µs per layer)
1.16 - 1.59x faster ordinary KDA decode including f_b projection at B1–64.
1.66x less time checking speculative guesses across all 69 KDA layers (16 users, 4-token window, 2.75 → 1.66ms)
2. MLA: reusing compressed history
K3’s remaining 24 attention layers use MLA. Instead of keeping separate keys and values for 96 attention heads, MLA stores a shared 512-wide latent and a 64-wide key component per token, far less data than an expanded per-head cache. Decode can transform the query into the compressed space and project the result afterwards. Prefill may expand cached latents for its selected attention kernel. K3 uses NoPE (no rotary position embedding), and FP8 cache storage halves the bytes relative to BF16 (Figure 4).
Figure 4: MLA stores a compressed representation of the attention history. Bars drawn to scale. A smaller cache means more users and longer contexts fit in GPU memory, and less data to read per step.
2.1 Pipelined prefill and shared KV reads
For long prompts, MLA is a big compute job. Our FP8 Gluon kernel runs two groups of threads on each compute unit, offset by one step: while one group multiplies, the other loads the next KV tile in parallel.
2.2 Fused cache writes and output processing
Before an MLA kernel runs, the runtime prepares the cache rows and the query. In the prefill path, one kernel packs and converts the K/V operands and handles the cache preparation, replacing eight small launches with one. The same fused query-preparation path is used across MLA phases for normalization and projection. On pure decode, it also prepares the absorbed query. The decode epilogue then combines attention-result reduction, latent-to-value projection, output gating and the final write, avoiding a write and reread of the latent intermediate.
2.3 MLA prefill and verification gains
Figure 5: MLA prompt processing (FP8, 12 heads per GPU). time per call in µs, lower is better · measured on MI355X
3.8–5.1x faster than the portable Triton MLA kernel on full 4K and 8K chunks, and 3.2× on an 8K query with a 32K cached prefix. The 8-wave Gluon implementation is 1.13–1.18× faster than the previous Gluon kernel in these historical measurements (Figure 5).
Figure 6: Checking speculative guesses: reading the history once vs once per position. speedup, higher is better, by number of users (C) · measured on MI355X
1.86x faster at 50K context and 2.14x at 100K context for 16 users, while meeting the numerical comparison used in the kernel tests. Very short histories are slower in this experiment (0.99–1.20x at 1K/8K for 1–16 users). These are historical kernel timings, not a new end-to-end result at the cutoff (Figure 6).
3. AttnRes: fusing representation mixing
In most models, each layer’s output is added to a running residual stream. Kimi K3 groups layers in blocks of 12 and retains completed block representations. Before each attention and feed-forward step, it scores those representations together with the current accumulated stream and feeds their weighted blend into the next operation. Deep layers can select information from earlier blocks (Figure 7).
Figure 7: Attention Residuals blend saved block representations with the current residual stream.
The catch: done naively, this means rereading up to nine large tensors twice per layer. And it sits right after the tensor-parallel all-reduce, so every microsecond it takes adds directly to the step time.
3.1 Single-pass mixing with bounded register use
Our AttnRes kernel scores, blends and normalizes in a single pass over the snapshots, accumulating in 32-bit floats. We also fixed a subtle problem. The kernel used to treat one input size as a constant baked into the compiled program, so every new prompt length triggered a fresh compile in the middle of serving, at 100+ ms each. Making it a runtime value cut one long-prompt test from 46s to 16s.
The cutoff also changes the snapshot loop. Reading one snapshot at a time prevents the compiler from keeping all candidates live together, reducing register pressure. This makes an 8K-token mix of 8 snapshots 1.10× faster (283 → 257 µs). Token counts and batch-dependent strides remain runtime inputs so new serving shapes do not force compilation.
3.2 Sharing AttnRes work across GPUs
For eligible TP8 batches, we distribute the post-attention mix across token rows. Pull reduce-scatter gives each GPU one eighth of the reduced attention output. It updates the residual, mixes all saved snapshots for its assigned tokens, and applies output RMSNorm. Push all-gather then assembles the normalized MoE input on every GPU, while the updated residual shard stays local until the MoE tail consumes it (see Section 6).
This removes seven eighths of the duplicated post-attention mixing and normalization work. The same row-count and compatibility checks apply to prefill, decode, and mixed batches. For example, a compatible 64-row EAGLE3 verification batch gives each GPU eight rows to mix.
3.3 Preparing scores and fusing the final mix
This is the clearest example of why owning every kernel matters. Snapshots only change every 12 layers, so most of the blend can be computed early, as a by-product of an earlier matrix multiply. Only the last step, folding in the newest total, has to wait. We then fold that last step into the all-reduce itself (Figure 8):
Figure 8: Fusing the all-reduce, residual update, AttnRes blend, and normalization removes intermediate memory traffic.
3.4 AttnRes kernel and serving gains
Figure 9: AttnRes kernel vs a straightforward PyTorch version. time per call in µs, log scales, lower is better · measured on MI355X
4.1–5.9x faster at decode sizes and 23x faster on an 8K-token prompt chunk (Figure 9).
−27% time for the fused all-reduce + AttnRes step at 16 tokens (21.7 → 15.8 µs; 8 GPUs)
+7–14% end-to-end tokens/s for 1–4 users from that fusion and related changes
4. Latent MoE: reducing expert and projection overhead
The feed-forward network uses 896 routed experts in 92 of K3’s 93 decoder layers; the first layer is dense. A router chooses 16 experts per token, and two shared experts runs on every token. Routed experts operate in a 3,584-wide latent space rather than the 7,168-wide hidden space, then project their combined result back to full width. MXFP4 expert weights reduce storage and weight traffic; the routed and shared branches retain their distinct projections (Figure 10).
Figure 10: Latent MoE routes each token to selected experts. The grid is a sketch (224 of the 896 experts); the 16 highlighted cells are the experts chosen for one token.
4.1 Expert-weight reuse and smaller compute tiles
With 896 experts, most experts receive only a handful of tokens in any step. GPU matrix-multiply units like tidy blocks of 64 or 128 rows, so the old kernels padded each expert's few tokens up to a full block and wasted most of the work. We attacked this from both ends:
Decode: for small batches, walk each token’s selected routes directly instead of building sorted groups. Activations use FP8 with MXFP4 weights. At 32–64 rows in the supported K3 configuration (attention TP8; MoE TP8/EP1), we introduce moe_sorting for warp-decode MoE that sorts routes by expert so several routes reuse each loaded weight panel.
Prefill: use 16- and 32-row expert tiles where sparse routing would waste a larger tile, and fuse the expert output combine when the activation format and batch size make it beneficial.
4.2 Joining routed and shared expert reductions
Each GPU ends the MoE expert computation step with partial routed and shared expert results that must be summed across all eight GPUs. Our kernels write routed and shared BF16 partials directly into one reusable Iris buffer. One collective reduces both branches across eight GPUs, keeping their results separate and avoiding a concatenation copy. Tiny batches use a push protocol whose arriving data also signals readiness (Section 6).
For eligible larger batches, reduce-scatter gives each GPU both reduced results for one eighth of the token rows. It applies routed RMSNorm and up-projection locally, then adds the shared result and residual during push all-gather. Each GPU performs only one eighth of the normalization and projection work, and gathering only the combined output cuts all-gather traffic by a third.
4.3 Combining input projections and output additions
Each MoE layer applies three projections to the same input: router scores, routed latent inputs, and shared-expert gate/up values. We store their weight matrices consecutively and compute them with one matrix multiply. Where beneficial, its epilogue also applies the shared experts’ activation, avoiding an intermediate write and read.
Output fusion follows the selected expert computation. Some small-batch kernels combine the selected experts’ contributions directly; the expert-sorted 32–64-row path writes route partials and sums them in FP32. At the MoE tail, the supported 2–32-row up-projection kernels also add the shared-expert output and running residual, saving a separate addition kernel.
4.4 Expert, projection, and communication gains
Figure 11: Expert computation before and after optimization, by number of tokens. Time per call in µs, lower is better · measured on MI355X at 7e2f1deb
Together, these changes make expert computation on one GPU 2.2–2.3x faster at 32–64 tokens and 1.5–2.0x faster on 128–2,048-token prompt chunks. At 4,096–8,192 tokens, the gain is 3–4%. The comparison switches the chapter’s expert-computation optimizations off on the same code revision, isolating their combined effect (Figure 11).
Figure 12: MoE input step: one packed kernel vs three GEMMs plus SiTU. Time per call in µs, lower is better · measured on MI355X at 7e2f1deb
2.0–3.0x faster from 1 to 8,192 tokens, including 2.26x at one token and 2.95× at 8,192 tokens. The previously reported end-to-end result took single-user generation from 64.1 to 77.6 tokens/s (+21%, reported) (Figure 12).
Projection work also benefits from the Gluon GEMM additions. Bucketed Gluon decode GEMM work reports up to 3.02× faster MLA kv_b decode projection than torch.mm at a measured TP8 shape. The large-M Gluon GEMM implementation extends the large-M GEMM to ragged prefill widths and routes the measured K3 projection set. Across 160 held-out projection cases, dispatch reduces total GEMM time by about 4%. Latent up-projection fusion reduces latent up-projection plus residual-add time by 11–16% at 2–32 rows.
5. EAGLE3: verifying several tokens per pass
Speculative decoding uses a small, fast draft model to propose several tokens, then asks the target model to verify them in one pass. In the three-step EAGLE3 configuration illustrated in Figure 13, Kimi K3 checks four positions, accepts the matching prefix, and supplies a target token at the rejection point or after all proposals are accepted. Under greedy sampling, acceptance checks the target's greedy choice. Exact token equality across different batching or numerical configurations requires a separate end-to-end check.
Figure 13: EAGLE3 checks several draft tokens in one target-model pass.
Speculation changes the shape of every component's work. For example, with 16 concurrent requests and four verify positions, the target receives 64 rows. The corresponding optimizations are described where they happen: KV reuse across query positions (Section 2.1), expert-weight reuse and top-k reduction (Section 4.1– Section 4.2), and attention row sharding (Section 6.2). The three-step configuration described here uses a chain (topk=1); tree support merged by the cutoff does not change this configuration.
5.1 Performance
We evaluate EAGLE3 performance on a synthetic random dataset with 50,000 input tokens and 500 output tokens per request, comparing execution with and without EAGLE3 (Figure 14). This is a separate evaluation from the multi-turn SWE-smith workload used in the end-to-end serving comparison. EAGLE3 reduces time per output token at every tested concurrency: from 12.9 to 4.2 ms at one concurrent user, and from 28.8 to 11.4 ms at 16 concurrent users.
Figure 14: Time per output token (TPOT) with and without EAGLE3 on eight MI355X GPUs, using a synthetic random dataset with 50K input tokens and 500 output tokens per request. Lower is better; C1–C16 denote concurrent users.
6. Iris: combining communication and computation
Tensor parallelism splits each layer across eight MI355X GPUs, so their partial outputs must be summed before dependent work begins. Our Kimi K3 implementation uses Iris to build communication protocols in native Gluon and fuse computation where relevant. The symmetric heap provides buffers that kernels can directly read or write on peer GPUs. Across low latency decoding and high throughput prefill, we implemented optimized kernels based on message size and the K3 specific computation surrounding each reduction.
Each layer has two all-reduces: one after the attention output projection and one after expert computation. On the attention side, attention residuals mixes the current residual with saved states from earlier layer blocks using learned, token-dependent weights, followed by RMSNorm. The MoE all-reduce combines two partials. The routed experts produce a 3,584-wide latent vector, which we normalize and project to the model's 7,168-wide hidden dimension. The shared expert already produces a full-width output. We place both partials consecutively in one input buffer, giving a combined BF16 workload of 21 KiB per row.
6.1 Small/Medium messages
On the attention side, eligible layers fuse one-shot push all-reduce with the residual update, AttnRes combine, and final RMSNorm. Fusing these steps removes separate epilogue launches and intermediate memory traffic, reducing the fixed overhead that matters most for small messages. On the MoE side, we implemented a barrier-free all-reduce that is 1.2-3× faster than RCCL across the message sizes for which it is in use. Arriving payloads signal readiness, allowing each tile to reduce once its inputs arrive without a separate barrier. Medium-sized messages use an optimized two-shot pull all-reduce, which also handles larger messages ineligible for the row-sharded computation discussed next (Figure 15).
6.2 Large messages
Figure 15: Replicated and row-sharded communication tails: where each GPU computes and exchanges results.
Normally, each TP rank repeats the computation after all-reduce over all M. For larger messages (M >= 40 for MoE, M >= 56 for attention), we use pull reduce-scatter, assign M/8 complete rows to each of the eight GPUs, and delay push all-gather until after the attention or MoE tail. These tails communicate BF16 activations; reductions and normalization use FP32 arithmetic internally.
After attention, each GPU updates the residual and applies AttnRes mixing and output RMSNorm to a [M/8, 7,168] shard instead of [M, 7,168], removing seven eighths of this duplicated row-wise work. We gather the normalized activation for the following MoE, while the updated residual shard stays local. On the MoE side, each GPU reduces [M/8, 3,584] routed values and [M/8, 7,168] shared values. It normalizes the routed shard and projects it back to the hidden width using the replicated [7,168, 3,584] weight. Sharding removes 87.5% of the routed normalization and up-projection work per GPU. The up projection alone drops from 51.38 million to 6.42 million FLOPs per input token per GPU in each MoE block.
We then add the projected routed output, shared output, and residual in push all-gather. Gathering the combined hidden-width output reduces the MoE all-gather tensor from 21 KiB to 14 KiB per row: one third less all-gather payload, or one sixth less across reduce-scatter and all-gather together. Accounting for transfers between the eight GPUs, this saves 6.125 KiB per input token per GPU per MoE block, or 49 MiB per GPU per block for an 8,192-token prefill chunk. Attention gathers the same-width activation as before, so its benefit here is the sharded computation and smaller local intermediates.
K3's 93 decoder layers contain one dense FFN layer and 92 MoE layers. For batches using both row-sharded paths, the optimization applies to all 92 MoE tails and their corresponding 92 post-attention tails. Across those layers, the up-projections alone save 4.136 billion FLOPs per input token per GPU, and the MoE collectives save 563.5 KiB of inter-GPU payload per input token per GPU. For an 8,192-token prefill chunk, that is 4.40 GiB less payload per GPU, or 35.22 GiB across all eight GPUs.
6.3 Collective latency and serving gains
Figures 16 and 17 compare Iris with RCCL baselines for the MoE and attention communication tails on eight MI355X GPUs.
Figure 16: MoE communication-tail latency: Iris and RCCL baselines on eight MI355X GPUs.
Figure 17: Attention communication-tail latency: Iris and RCCL baselines on eight MI355X GPUs.
Conclusion
On eight AMD Instinct MI355X GPUs, TokenSpeed reaches 216 output tokens/s per user at concurrency one in the measured multi-turn SWE-smith workload. Specialized kernels, EAGLE3 verification, and Iris communication fusion reduce data movement, launch overhead, and duplicated work across the serving path. The results show how coordinated kernel and runtime optimization supports interactive generation on long conversations.
Acknowledgements
We thank the TokenSpeed team and the LightSeek Foundation for their close collaboration and help in bringing Kimi K3 to AMD Instinct GPUs. We thank OpenAI and the Triton/Gluon contributors for the compiler and kernel programming tools that underpin this work. We are also grateful to the broader open-source community, including the PyTorch ecosystem and the projects whose libraries, tools, and reference implementations made this work possible.
Footnotes
System configuration
AMD Instinct™ MI355X GPU node
System Model: Supermicro AS-4126GS-NMR-LCC CPU: 2× AMD EPYC 9575F 64-Core Processor (256 threads total) NUMA: 2 NUMA nodes (1 per socket); NUMA auto-balancing disabled Memory: 3072 GiB (24×128 GiB) Micron DDR5-6400, configured at 6000 MT/s Disk: 2×3.84 TB Micron 7450 NVMe SSDs (7.68 TB total raw capacity) GPU: 8× AMD Instinct™ MI355X, 288 GB HBM3E each, 256 CUs per GPU Host OS: Ubuntu 22.04.5 LTS Host Kernel: 6.8.0-84-generic System BIOS: 1.7 System BIOS Vendor: American Megatrends International, LLC. Host GPU Driver: amdgpu 6.16.13 Host ROCm: 7.2.2
Current host configuration verified on October 7, 2026. Benchmark containers may use a different ROCm version.
Benchmark Methodology
We ran Kimi K3 on eight AMD Instinct MI355X GPUs. The multi-turn SWE-smith conversations came from the linked agentic dataset and were replayed with TokenSpeed’s agentic benchmark. We measured 1, 2, 4, 8, and 16 concurrent sessions, using separate corpus slices at each level. Each turn requested up to 500 output tokens with greedy sampling and EOS ignored. The first prompts matched the corresponding TokenSpeed requests; later turns could diverge because each engine’s reply became part of its next prompt.
Each engine was evaluated separately, serving concurrent conversations on the same eight-GPU server. The scheduling distinction is how prompt processing (prefill) competes with ongoing token generation (decode). TokenSpeed uses separate prefill and decode execution batches: a prefill batch temporarily interrupts ongoing decoding. ATOM also keeps the phases separate, but its concurrency-16 recipe enables prefill coalescing: pending prompt work is held while decoding continues, then admitted together into a larger prefill batch. This is designed to reduce small prefill forwards and decode interruptions, at the cost of extra waiting before the first token. Both approaches share the same GPU pool; none merges independent requests or requires separate prefill/decode servers. Figures 18-19 show selected requests with each outlined box representing one execution batch.
Figure 18: Separate batches. Prefill and decode run in different execution batches on the same GPU pool; one prefill batch may contain multiple prompt requests.
Figure 19: Prefill coalescing. Pending prompts accumulate while decoding continues, then run together in a larger prefill-only batch.
Serving Commands
The following Bash commands reproduce the supplied server settings for the multi-turn SWE-smith comparison. Launch one engine at a time on port 21080.
TokenSpeed Serving Configuration
Runtime: built from source in a Python 3.10.12 virtual environment for the Figure 1 benchmark; no Docker image was used for that run.
Commit:7e2f1deb3cf8561f754a1cae46b98c3e2bc91eda. This source-code snapshot identifies the implementation discussed in the article and the measurements in Figures 11, 12, 16, and 17.
Parallelism and decoding: attention TP8 and MoE TP8/EP1, with four-position EAGLE3 verification.
Cache settings: FP8 KV storage and GPU prefix caching enabled.
Serving command
Set MODEL_PATH to the Kimi K3 model directory and TS_DRAFT to the EAGLE3 draft-model directory before running this command.TokenSpeed Serving Command
Concurrency 1, 2, and 4: TP8/DCP1 with seven DSpark draft tokens.
Concurrency 8 and 16: TP8/DCP8 with three DSpark draft tokens and ReplaySSM. Concurrency 16 also enables prefill coalescing with a decode interval of 4 and a maximum queue delay of 5000 ms.
Evaluation settings: FP8 KV storage and GPU prefix caching enabled; synthetic acceptance override and CPU KV offload disabled.
Serving Command
The shared script below preserves the supplied C1, C2, C4, C8, and C16 settings. C denotes client concurrency; the server capacity remains 32 sequences. Save the next three code blocks together as serve_atom.sh.
1. Model paths and quantization
#!/usr/bin/env bash
# Save this block and the next two ATOM blocks as serve_atom.sh.
HF_CACHE=/data/models/hf/hub
MODEL_REV=f831ab66814297da540d832a5235f8e904f29d06
DRAFT_REV=cf6b8244620e7ea4b0651d214f28e89eac75bed6
ATOM_MODEL="$HF_CACHE/models--moonshotai--Kimi-K3/snapshots/$MODEL_REV"
ATOM_DRAFT="$HF_CACHE/models--Inferact--Kimi-K3-DSpark/snapshots/$DRAFT_REV"
ONLINE_QUANT_CONFIG='{
"global_quant_config": "ptpc_fp8",
"exclude_layer": [
"lm_head", "model.embed_tokens", "*self_attn.[qkv]_conv1d*",
"*block_sparse_moe.experts*", "*block_sparse_moe.routed_expert_*",
"*vision_tower*", "*mm_projector*"
]
}
C1–C8 unset the prefill interval and queue-delay variables. C16 sets the interval to 4 and the maximum queue delay to 5000 ms. ATOM_ENABLE_PREFILL_DELAYER remains 1 in all five supplied configurations.