Qwen3.5-9B-NVFP4-QAD-LR1e-5-s4000-weight-only
Weight-only (W4A16 serving contract) export of the same trained checkpoint as Qwen3.5-9B-NVFP4-QAD-LR1e-5-s4000 — see that card for the full training provenance (recipe, dataset, W&B) and the measured 2x2 tables. The A16 rows in those tables were measured on this artifact.
Serving
compressed-tensors NVFP4 weight-only (W4A16) artifact: same trained weights as
Qwen3.5-9B-NVFP4-QAD-LR1e-5-s4000, exported with
scripts/export_nvfp4_vllm.py --weight-only — input_activations is null and no
input_global_scale tensors are shipped, so engines serve it W4A16 with no activation
quantization anywhere. This is the exact artifact behind the A16 rows in the tables.
vllm serve weili-0234/Qwen3.5-9B-NVFP4-QAD-LR1e-5-s4000-weight-only --max-model-len 24576
How this artifact was produced
Exported from the arm's final training checkpoint with
scripts/export_nvfp4_vllm.py --weight-only (QATFactory weili/w4a4 @ a9315e4,
PR). Training provenance is identical to
Qwen3.5-9B-NVFP4-QAD-LR1e-5-s4000; the BF16 master weights live at
Qwen3.5-9B-NVFP4-QAD-LR1e-5-s4000-BF16.
Part of a monitored QAD experiment series with full bookkeeping (pre-registered predictions, exact SHAs/configs/seeds per run). Produced with AI assistance (Claude).
- Downloads last month
- -