jcbtc's picture
Release validated Ling runtime and installer fixes
9d29a15 verified
|
Raw
History Blame Contribute Delete
11.1 kB
# Runtime change manifest
## Immutable identities
- Original model: `inclusionAI/Ling-3.0-flash-int4`
- Model revision: `ca3ea63b0255d212c4fe6020db9e0a51ce136006`
- Model format: 24 safetensors shards, 77,012,299,464 bytes total
- vLLM upstream base: `d35eb6c44071ea806018841c490f0d2f3219c485`
- CIRU runtime head: `838616875c5dd4913d9f753221bf05f64cb7a7ed`
- Net vLLM delta: 10 files, 477 insertions, 28 deletions, including regression tests
## Weight changes
None. No tensor, quantization scale, tokenizer, template, configuration, or MTP weight was modified. The official packed-INT4 checkpoint is consumed directly.
## Cold-start diagnostics and environment registration — August 10, 2026
- Register the required `VLLM_ROCM_SAFE_MERGE_ATTN_STATES` switch in vLLM's
environment registry and read it through that registry. The fork no longer
warns that its own packaged setting is unknown.
- Make the Inductor max-autotune defaults explicit in the launcher and reject
malformed values before model loading.
- Explain before launch that the pinned Triton `make_block_ptr`, PyTorch
`TypedStorage`, and Inductor `AUTOTUNE` output is expected during a cold
compile. Readiness remains `Application startup complete` plus `/health`.
- Print the package's expected fork head beside the installed editable-source
head. A mismatch now produces the exact model-preserving CIRU rebuild command;
the ROCr-only `install.sh --upgrade` path is no longer easy to mistake for a
vLLM source update.
## Clean NixOS installer correction — August 10, 2026
- Restore the host-toolchain bridge used by the validated NixOS build. The
installer now derives the active GCC, glibc, libstdc++, and NUMA paths instead
of embedding machine-specific Nix store hashes.
- Supply the pinned ROCr rebuild's `pkg-config`, `xxd`, DRM, ELF, NUMA, and
OpenSSL inputs through an ephemeral `nix-shell`, then translate its search
flags for CMake. No package is added to the user's Nix profile.
- Make the existing-install command update both the CIRU vLLM fork and the
corrected ROCr runtime while preserving the 77 GB checkpoint. Keep
`install.sh --upgrade` explicitly ROCr-only.
## Installer release-gate hardening — August 10, 2026
- Reconstruct the pinned vLLM commit through a dedicated remote ref and a
detached checkout. Re-running the build no longer asks Git to fetch into the
branch that the first build left checked out.
- Resolve literal `$HOME`/`$PWD` spellings in every direct path-taking helper,
quote virtual-environment executables, and reject whitespace or `:` in the
runtime root before Linux `LD_PRELOAD` can split it.
- Reject empty, unknown, and surplus script arguments instead of silently
ignoring them or handing option-like paths to `mkdir`.
- Verify every allowlisted checkpoint file after `hf download` and validate the
checkpoint config/index before launch.
- Support host-dependency installation when already running as root, while
retaining `sudo` for normal users.
- Replace the WSL Git-clone path—which could pull the full 77 GB checkpoint—with
the checksummed 115 MB runtime archive path.
## Installer path and bootstrap correction — August 10, 2026
- Resolve literal agent-supplied `'$HOME'`, `'${HOME}'`, `'$PWD'`, and
`'${PWD}'` path prefixes before canonicalization. This prevents creation of
a directory literally named `$HOME` when a caller quotes a shell placeholder.
- Reject any other unresolved `$...` expression in installer paths instead of
silently creating a misleading directory.
- Write `ling3-runtime.env` atomically immediately after path resolution and
before host dependencies, downloads, ROCr work, or the vLLM build. The env
file therefore survives and records the intended paths if a later step
fails.
## Source changes
### `csrc/libtorch_stable/sampler.cu`
Uses 512 rather than 1024 merge threads on ROCm so Wave32 hardware does not exceed the 64 KB LDS limit. This is a general ROCm correctness change.
### `vllm/v1/attention/ops/merge_attn_states.py`
Adds an opt-in Torch implementation selected by `VLLM_ROCM_SAFE_MERGE_ATTN_STATES=1`. It handles token-first and head-first LSE layouts and avoids an observed `gfx1151` HSA fault in the Triton merge kernel.
### `vllm/v1/attention/backends/mla/prefill/flash_attn.py`
Normalizes the LSE returned by upstream ROCm FlashAttention variable-length prefill from token-first `[tokens, heads]` to vLLM's documented head-first `[heads, tokens]` adapter contract. This prevents heterogeneous chunked-prefill batches from copying a head extent into a token extent in the MLA context accumulator.
### `vllm/v1/attention/backends/mla/triton_mla.py`
Declares uniform query-length support and converts causal multi-token verifier blocks into per-token Triton MLA decode rows with correct causal KV-prefix lengths. This is relevant to MLA speculative verification, not W4A16 specifically.
### `vllm/model_executor/layers/fused_moe/fused_moe.py`
Adds a naive small-decode block assignment specialization that bypasses sorting/alignment only under a narrow validated guard.
### `vllm/model_executor/layers/fused_moe/moe_fused_mul_sum.py`
Adds a low-launch-overhead Triton reduction for one to three tokens, top-k 8, and hidden size 2560.
### `vllm/model_executor/layers/fused_moe/experts/triton_moe.py`
Wires the guarded WNA16 paths together and adds the exact small-shape `gfx1151` SiLU-and-multiply kernel. Every adjacent shape or unsupported configuration retains the upstream route.
### `tests/v1/attention/test_mla_prefill_quant_output.py`
Adds adapter-layout coverage for both upstream ROCm FlashAttention and vLLM FlashAttention, plus an adapter-to-accumulator regression that exercises the formerly failing token/head extent mismatch.
## Runtime configuration
- Triton MLA attention
- Triton MoE backend
- native checkpoint MTP, K1 (`num_speculative_tokens=1`)
- compile sizes `[1,2]`
- CUDA graphs disabled
- chunked prefill and prefix caching enabled
- AITER and ROCm skinny GEMM disabled
- safe attention merge enabled
- native context 262,144 tokens
- six active sequences for 256K; two for experimental 1M
## ROCr installer recovery correction
The first ROCr idle-fix installer used `git clone --no-checkout` and then
mistook the intentionally empty worktree for deleted user files. Fresh installs
could therefore stop with `refusing to replace a modified ROCr source checkout`
before `ling3-runtime.env` was generated. The installer now recognizes and
repairs only a newly created or previously failed empty checkout. A non-empty
checkout with real local changes remains protected and is still refused.
## ROCr idle-runtime correction
The original TheRock ROCm 7.15 wheel's `libhsa-runtime64` kept two native wait
threads busy on the validated `gfx1151` host even with no requests in flight.
Profiling attributed approximately 50% of one CPU core to
`Runtime::AsyncEventsLoop` and 44% to `InterruptSignal::WaitRelaxed`, matching
the reported continuous 98 C idle temperature.
The package now rebuilds the ABI-matched ROCr library from the wheel's exact
`rocm-systems` source pin (`44be71b52284948e58c93f65f46910399773fdcd`)
with the host GCC and preloads the checksummed result. No functional ROCr source
change is carried. On the validation host, the hottest idle runtime thread fell
to 1.40% CPU or less and CPU temperature returned to a 38.8-54.2 C idle range;
a 256-token response still measured 26.307 tok/s and passed exact API-output
validation. Existing users can apply only this correction with
`bash install.sh --upgrade --install-root PATH` without touching model files.
## Performance context
The runnable public/upstream-compatible starting path measured 0.268944 tok/s in a warmed deterministic target-only decode test. The released runtime measures 21.4407 tok/s target-only and 26.2343 effective tok/s with native K1 MTP: 79.72x and 97.55x the starting throughput, respectively.
The final exact-shape SiLU candidate passed bitwise source fixtures and exact output checks. In the matched final-step comparison against the already-optimized CIRU parent, it improved target-only decode by 7.7457% and K1 effective decode by 4.3743%. Those percentages describe only the last runtime optimization, not the public-to-release gain. Only clean, resource-gated measurements were retained as promotion evidence.
## Public release packaging
- Reduced the native 256K default from `gpu_memory_utilization=0.75` to `0.72` while retaining `max_model_len=262144`. The profile was subsequently promoted to six active slots and `max_num_batched_tokens=8192`; the validated configuration reports a 1,457,313-token KV pool and 5.56x full-context capacity. Added an environment override for users who prefer lower memory use, plus prominent copyable instructions for the experimental 1M/two-slot profile.
- Corrected the live API example to use InclusionAI's recommended Ling 3.0 Flash sampling preset: `temperature=0.6`, `top_p=0.95`, `top_k=20`, and thinking enabled. Temperature 0 is documented only as a deterministic historical benchmark condition, not a serving recommendation.
- Added the CIRU model-card artwork at `assets/ling30int4.png`.
- Added distro-aware host dependencies and portable Python 3.12 provisioning through `uv`.
- Added native-Linux instructions for Ubuntu/Debian, Fedora, Arch, and the validated NixOS boundary.
- Added an explicitly experimental Windows 11 WSL2/AMD ROCDXG path; native Windows vLLM is not claimed.
- Initially changed the public native-256K profile from `gpu_memory_utilization=0.82` to the measured `0.75` envelope used by the retained 60K benchmark; the release default was subsequently reduced to `0.72` as recorded above.
- Added the retained 60K PP/TG comparison against AtomicChat AD-IQ4_XXS and ROCmFP4 STRIX MTP, with the native row marked for a final-commit refresh.
- Fixed the ROCm MLA chunked-prefill LSE layout mismatch in runtime commit `91fab4e6d`. Focused attention tests passed 31/31 (with one unrelated CUDA-only FA4 test deselected), and the exact five-request tool-calling reproducer completed 5/5 scenarios with zero errors while the original server PID remained alive.
- Promoted the native 256K profile from five to six active sequences and fixed its scheduler budget at 8,192 batched tokens after a three-repetition C1-C6 sweep. All 18 exact-2K/256-output runs passed with zero request failures, cache reuse, preemptions, post-warmup JIT, engine restarts, or device faults. Median C6 reached PP 520.49 tok/s, aggregate TG 63.51 tok/s, and 42.60 seconds wall time; all comparable retained rows passed the 5% speed-regression guard.
- Pinned `max_num_batched_tokens=8192` in both packaged MTP profiles and made the common launcher default explicit. This prevents vLLM's speculative scheduler from silently choosing 2,048 and changing the measured prefill/decode overlap for users who follow the release launchers.
- Added `install.sh --upgrade` for a model-preserving engine update, automatic ROCr checksum verification in both launchers, and `CIRU_DISABLE_ROCR_IDLE_FIX=1` as the explicit rollback switch.