Benchmark notes
All results below are local measurements on one Ryzen AI MAX+ 395 / Radeon 8060S Strix Halo system with 128 GB unified memory. They compare complete model-format and runtime stacks; they are not model-quality rankings.
60K prompt plus 512-token decode
The frozen workload used exactly 60,000 direct prompt tokens, 512 forced output tokens, temperature 0, EOS ignored, a unique cache salt, and no reusable prefix. Temperature 0 was used only to make this cross-runtime throughput measurement deterministic; it is not InclusionAI's recommended live-generation preset. For live use, follow the official model card: temperature=0.6, top_p=0.95, top_k=20, and enable_thinking=true for better performance.
| Stack | Format/runtime | PP | TG | Request wall |
|---|---|---|---|---|
| CIRU native K1 MTP | official InclusionAI packed INT4, vLLM/ROCm | 419.33 tok/s | 19.92 tok/s | 168.75 s |
| AtomicChat AD-IQ4_XXS | mixed GGUF, llama.cpp Vulkan | 265.63 tok/s | 25.11 tok/s | 246.28 s |
| raulvidis ROCmFP4 STRIX MTP | ROCmFP4 GGUF, llama.cpp Vulkan MTP | 249.23 tok/s | 19.00 tok/s | 267.84 s |
At this workload, native vLLM trails Atomic's decode by 20.66% but leads its prefill by 57.87%, completing the combined request 31.48% sooner. Against ROCmFP4, native leads prefill by 68.25%, decode by 4.88%, and combined request time by 37.00%.
The native row used the then-current gpu_memory_utilization=0.75, 262,144-context envelope and K1 MTP. It accepted 254 of 258 draft tokens (98.45%) and completed without cache hits, preemption, JIT during measurement, or device faults. The public launcher now defaults to 0.72 for more system-memory headroom; this retained row has not been relabeled as a 0.72 measurement.
Update status
The native 60K row was measured at CIRU runtime commit fabbc7e928dfa64c29dcdf0f795b05246555f4ae, before the final exact-shape SiLU kernel and ROCm MLA LSE-layout commits. The release branch now ends at 838616875c5dd4913d9f753221bf05f64cb7a7ed; its last performance-affecting runtime commit is 91fab4e6dcaa683911655630dae8a5b140a7cb91. That branch has not yet been rerun on this 60K protocol. The table is therefore a valid retained comparison, but not a claim that 19.92 tok/s is the release branch's maximum 60K speed. Replace the native row only after a matched rerun of all three stacks or a clearly labeled native-only refresh.
Short decode and concurrent serving
- Strict final-commit single-request K1 TG32 median: 26.2343 tok/s.
- Highest individual strict final-commit TG32 row: 26.2445 tok/s.
The released 256K/0.72/six-slot profile was measured with exact 2,048-token prompts, exactly 256 forced output tokens, K1 MTP, temperature 0 as a benchmark control, EOS ignored, unique cache salts, no prefix reuse, and max_num_batched_tokens=8192. Values are medians of three repetitions.
| Concurrency | PP | Aggregate TG | Per-request TG | Wall time |
|---|---|---|---|---|
| C1 | 523.62 tok/s | 26.79 tok/s | 26.77 tok/s | 13.45 s |
| C2 | 520.98 tok/s | 38.23 tok/s | 19.30 tok/s | 20.86 s |
| C3 | 514.93 tok/s | 46.17 tok/s | 15.79 tok/s | 28.08 s |
| C4 | 523.95 tok/s | 57.12 tok/s | 14.74 tok/s | 32.91 s |
| C5 | 510.72 tok/s | 60.55 tok/s | 12.94 tok/s | 39.67 s |
| C6 | 520.49 tok/s | 63.51 tok/s | 13.01 tok/s | 42.60 s |
All 18 runs completed with exact token counts and zero request failures, cache hits, preemptions, post-warmup JIT, engine restarts, or device faults. C1/C2/C4/C5 stayed within a 5% PP/TG/wall regression guard against the retained matched sweep. Use 26.23 tok/s as the strict single-request optimization headline and 63.51 tok/s only when explicitly labeled six-request aggregate throughput.