# Runtime change manifest ## Immutable identities - Original model: `inclusionAI/Ling-3.0-flash-int4` - Model revision: `ca3ea63b0255d212c4fe6020db9e0a51ce136006` - Model format: 24 safetensors shards, 77,012,299,464 bytes total - vLLM upstream base: `d35eb6c44071ea806018841c490f0d2f3219c485` - CIRU runtime head: `838616875c5dd4913d9f753221bf05f64cb7a7ed` - Net vLLM delta: 10 files, 477 insertions, 28 deletions, including regression tests ## Weight changes None. No tensor, quantization scale, tokenizer, template, configuration, or MTP weight was modified. The official packed-INT4 checkpoint is consumed directly. ## Cold-start diagnostics and environment registration — August 10, 2026 - Register the required `VLLM_ROCM_SAFE_MERGE_ATTN_STATES` switch in vLLM's environment registry and read it through that registry. The fork no longer warns that its own packaged setting is unknown. - Make the Inductor max-autotune defaults explicit in the launcher and reject malformed values before model loading. - Explain before launch that the pinned Triton `make_block_ptr`, PyTorch `TypedStorage`, and Inductor `AUTOTUNE` output is expected during a cold compile. Readiness remains `Application startup complete` plus `/health`. - Print the package's expected fork head beside the installed editable-source head. A mismatch now produces the exact model-preserving CIRU rebuild command; the ROCr-only `install.sh --upgrade` path is no longer easy to mistake for a vLLM source update. ## Clean NixOS installer correction — August 10, 2026 - Restore the host-toolchain bridge used by the validated NixOS build. The installer now derives the active GCC, glibc, libstdc++, and NUMA paths instead of embedding machine-specific Nix store hashes. - Supply the pinned ROCr rebuild's `pkg-config`, `xxd`, DRM, ELF, NUMA, and OpenSSL inputs through an ephemeral `nix-shell`, then translate its search flags for CMake. No package is added to the user's Nix profile. - Make the existing-install command update both the CIRU vLLM fork and the corrected ROCr runtime while preserving the 77 GB checkpoint. Keep `install.sh --upgrade` explicitly ROCr-only. ## Installer release-gate hardening — August 10, 2026 - Reconstruct the pinned vLLM commit through a dedicated remote ref and a detached checkout. Re-running the build no longer asks Git to fetch into the branch that the first build left checked out. - Resolve literal `$HOME`/`$PWD` spellings in every direct path-taking helper, quote virtual-environment executables, and reject whitespace or `:` in the runtime root before Linux `LD_PRELOAD` can split it. - Reject empty, unknown, and surplus script arguments instead of silently ignoring them or handing option-like paths to `mkdir`. - Verify every allowlisted checkpoint file after `hf download` and validate the checkpoint config/index before launch. - Support host-dependency installation when already running as root, while retaining `sudo` for normal users. - Replace the WSL Git-clone path—which could pull the full 77 GB checkpoint—with the checksummed 115 MB runtime archive path. ## Installer path and bootstrap correction — August 10, 2026 - Resolve literal agent-supplied `'$HOME'`, `'${HOME}'`, `'$PWD'`, and `'${PWD}'` path prefixes before canonicalization. This prevents creation of a directory literally named `$HOME` when a caller quotes a shell placeholder. - Reject any other unresolved `$...` expression in installer paths instead of silently creating a misleading directory. - Write `ling3-runtime.env` atomically immediately after path resolution and before host dependencies, downloads, ROCr work, or the vLLM build. The env file therefore survives and records the intended paths if a later step fails. ## Source changes ### `csrc/libtorch_stable/sampler.cu` Uses 512 rather than 1024 merge threads on ROCm so Wave32 hardware does not exceed the 64 KB LDS limit. This is a general ROCm correctness change. ### `vllm/v1/attention/ops/merge_attn_states.py` Adds an opt-in Torch implementation selected by `VLLM_ROCM_SAFE_MERGE_ATTN_STATES=1`. It handles token-first and head-first LSE layouts and avoids an observed `gfx1151` HSA fault in the Triton merge kernel. ### `vllm/v1/attention/backends/mla/prefill/flash_attn.py` Normalizes the LSE returned by upstream ROCm FlashAttention variable-length prefill from token-first `[tokens, heads]` to vLLM's documented head-first `[heads, tokens]` adapter contract. This prevents heterogeneous chunked-prefill batches from copying a head extent into a token extent in the MLA context accumulator. ### `vllm/v1/attention/backends/mla/triton_mla.py` Declares uniform query-length support and converts causal multi-token verifier blocks into per-token Triton MLA decode rows with correct causal KV-prefix lengths. This is relevant to MLA speculative verification, not W4A16 specifically. ### `vllm/model_executor/layers/fused_moe/fused_moe.py` Adds a naive small-decode block assignment specialization that bypasses sorting/alignment only under a narrow validated guard. ### `vllm/model_executor/layers/fused_moe/moe_fused_mul_sum.py` Adds a low-launch-overhead Triton reduction for one to three tokens, top-k 8, and hidden size 2560. ### `vllm/model_executor/layers/fused_moe/experts/triton_moe.py` Wires the guarded WNA16 paths together and adds the exact small-shape `gfx1151` SiLU-and-multiply kernel. Every adjacent shape or unsupported configuration retains the upstream route. ### `tests/v1/attention/test_mla_prefill_quant_output.py` Adds adapter-layout coverage for both upstream ROCm FlashAttention and vLLM FlashAttention, plus an adapter-to-accumulator regression that exercises the formerly failing token/head extent mismatch. ## Runtime configuration - Triton MLA attention - Triton MoE backend - native checkpoint MTP, K1 (`num_speculative_tokens=1`) - compile sizes `[1,2]` - CUDA graphs disabled - chunked prefill and prefix caching enabled - AITER and ROCm skinny GEMM disabled - safe attention merge enabled - native context 262,144 tokens - six active sequences for 256K; two for experimental 1M ## ROCr installer recovery correction The first ROCr idle-fix installer used `git clone --no-checkout` and then mistook the intentionally empty worktree for deleted user files. Fresh installs could therefore stop with `refusing to replace a modified ROCr source checkout` before `ling3-runtime.env` was generated. The installer now recognizes and repairs only a newly created or previously failed empty checkout. A non-empty checkout with real local changes remains protected and is still refused. ## ROCr idle-runtime correction The original TheRock ROCm 7.15 wheel's `libhsa-runtime64` kept two native wait threads busy on the validated `gfx1151` host even with no requests in flight. Profiling attributed approximately 50% of one CPU core to `Runtime::AsyncEventsLoop` and 44% to `InterruptSignal::WaitRelaxed`, matching the reported continuous 98 C idle temperature. The package now rebuilds the ABI-matched ROCr library from the wheel's exact `rocm-systems` source pin (`44be71b52284948e58c93f65f46910399773fdcd`) with the host GCC and preloads the checksummed result. No functional ROCr source change is carried. On the validation host, the hottest idle runtime thread fell to 1.40% CPU or less and CPU temperature returned to a 38.8-54.2 C idle range; a 256-token response still measured 26.307 tok/s and passed exact API-output validation. Existing users can apply only this correction with `bash install.sh --upgrade --install-root PATH` without touching model files. ## Performance context The runnable public/upstream-compatible starting path measured 0.268944 tok/s in a warmed deterministic target-only decode test. The released runtime measures 21.4407 tok/s target-only and 26.2343 effective tok/s with native K1 MTP: 79.72x and 97.55x the starting throughput, respectively. The final exact-shape SiLU candidate passed bitwise source fixtures and exact output checks. In the matched final-step comparison against the already-optimized CIRU parent, it improved target-only decode by 7.7457% and K1 effective decode by 4.3743%. Those percentages describe only the last runtime optimization, not the public-to-release gain. Only clean, resource-gated measurements were retained as promotion evidence. ## Public release packaging - Reduced the native 256K default from `gpu_memory_utilization=0.75` to `0.72` while retaining `max_model_len=262144`. The profile was subsequently promoted to six active slots and `max_num_batched_tokens=8192`; the validated configuration reports a 1,457,313-token KV pool and 5.56x full-context capacity. Added an environment override for users who prefer lower memory use, plus prominent copyable instructions for the experimental 1M/two-slot profile. - Corrected the live API example to use InclusionAI's recommended Ling 3.0 Flash sampling preset: `temperature=0.6`, `top_p=0.95`, `top_k=20`, and thinking enabled. Temperature 0 is documented only as a deterministic historical benchmark condition, not a serving recommendation. - Added the CIRU model-card artwork at `assets/ling30int4.png`. - Added distro-aware host dependencies and portable Python 3.12 provisioning through `uv`. - Added native-Linux instructions for Ubuntu/Debian, Fedora, Arch, and the validated NixOS boundary. - Added an explicitly experimental Windows 11 WSL2/AMD ROCDXG path; native Windows vLLM is not claimed. - Initially changed the public native-256K profile from `gpu_memory_utilization=0.82` to the measured `0.75` envelope used by the retained 60K benchmark; the release default was subsequently reduced to `0.72` as recorded above. - Added the retained 60K PP/TG comparison against AtomicChat AD-IQ4_XXS and ROCmFP4 STRIX MTP, with the native row marked for a final-commit refresh. - Fixed the ROCm MLA chunked-prefill LSE layout mismatch in runtime commit `91fab4e6d`. Focused attention tests passed 31/31 (with one unrelated CUDA-only FA4 test deselected), and the exact five-request tool-calling reproducer completed 5/5 scenarios with zero errors while the original server PID remained alive. - Promoted the native 256K profile from five to six active sequences and fixed its scheduler budget at 8,192 batched tokens after a three-repetition C1-C6 sweep. All 18 exact-2K/256-output runs passed with zero request failures, cache reuse, preemptions, post-warmup JIT, engine restarts, or device faults. Median C6 reached PP 520.49 tok/s, aggregate TG 63.51 tok/s, and 42.60 seconds wall time; all comparable retained rows passed the 5% speed-regression guard. - Pinned `max_num_batched_tokens=8192` in both packaged MTP profiles and made the common launcher default explicit. This prevents vLLM's speculative scheduler from silently choosing 2,048 and changing the measured prefill/decode overlap for users who follow the release launchers. - Added `install.sh --upgrade` for a model-preserving engine update, automatic ROCr checksum verification in both launchers, and `CIRU_DISABLE_ROCR_IDLE_FIX=1` as the explicit rollback switch.