Qwen3.8-27B-FP8 / recipe.yaml
huginnfork's picture
Upload folder using huggingface_hub
ebafbed verified
Raw
History Blame Contribute Delete
1.87 kB
name: fp8_dynamic_attnbf16
scheme: FP8_DYNAMIC # FP8 weights, FP8 dynamic per-token activations — data-free, vLLM-native
engine: llmcompressor
# Variant of `fp8_dynamic.yaml` that additionally keeps the ENTIRE self-attention
# block in bf16, leaving only the MLPs quantised.
#
# Rationale: on Qwen3.5/3.6 only a quarter of the layers are `full_attention`
# (16 of 64 on Qwen3.6-27B — the rest are `linear_attn`), so the whole self_attn
# block is just ~1.68 B params and FP8-ing it saves only ~1.56 GiB. The MLPs are
# ~17.1 B params and deliver ~15.9 GiB of the savings. Giving up 4.6% of on-disk
# size buys a completely bf16 attention path.
#
# The specific thing this protects: `attn_output_gate: true` means the attention
# OUTPUT GATE is fused into `q_proj`, which is why q_proj is [2*heads*head_dim,
# hidden] = [12288, 5120] rather than [6144, 5120]. Half that tensor is a
# multiplicative per-head gate on what attention writes into the residual stream,
# and the 16 full-attention layers carry the long-range retrieval. Quantisation
# error on a multiplicative gate is qualitatively worse than on an additive
# projection. Everyone (upstream, Qwen official) quantises q_proj; this recipe
# does not.
#
# Use `fp8_dynamic.yaml` for the standard build; use this one when multi-turn /
# long-context stability matters more than 1.5 GiB.
calibration:
dataset: neuralmagic/calibration
config: LLM
split: train
num_samples: 4
max_seq_length: 512
ignore:
- lm_head
- "re:.*visual.*"
- "re:.*linear_attn.*" # Mamba/SSM block stays in bf16 — same rationale as the NVFP4A16 build
- "re:.*self_attn.*" # THE VARIANT: q/k/v/o_proj too, incl. the output gate fused into q_proj
- "re:.*mtp.*"
# Note: Qwen3.6-27B is dense; on MoE bases also include
# "re:.*mlp.gate$" and "re:.*mlp.shared_expert_gate$".
export:
save_compressed: true