Qwen3.8-27B-DFlash2-W4A16

incoai/Qwen3.8-27B-DFlash2 โ€” the DFlash2 block drafter for Qwen3.8-27B (5 Qwen3-style layers, 1.92B parameters, 3.85 GB in bf16) โ€” requantized to W4A16 compressed-tensors (pack-quantized, group 128, symmetric, Marlin) with GPTQ, so it fits next to a 27B target on a 24 GB card: 1.19 GB.

Built by and for syv-ai/qwen38-27b-rtx3090 (Qwen3.8-27B on one RTX 3090 with vLLM): single-user/start_qwen.sh with SPEC=dflash2.

What is quantized

tensors precision note
layers.N.self_attn.{q,k,v,o}_proj, layers.N.mlp.{gate,up,down}_proj (1.61B params) int4 GPTQ, g128 Hessians from the drafter's own inputs on 400 real prompts (~290k rows per layer)
fc (5120 ร— 25600, projects the target's layer 5/19/33/47/61 hidden states) int4 GPTQ, g128 keeping it bf16 measured no acceptance difference
grouped-conv kernel_projections, candidate_selector, norms bf16 tiny; vLLM builds them unquantized

Measured on the W4A16 target (RTX 3090, 8 realistic chat prompts ร— 1,024 tokens, vLLM 0.27.1 with the repo's patches): same acceptance as the bf16 drafter at greedy (3.34-3.65 vs 3.54 tokens per step), about 5% lower at the model's default sampling; per decode step it reads 2.7 GB less, which is what makes it a net win on this card. Details and the pipeline (drafter/capture_dflash2.py, drafter/quant_dflash2.py) in the repo's drafter/README.md.

A variant whose k/v Hessians also blended the context-KV precompute's input distribution โ€” which is the theoretically tidier calibration โ€” measured 7% worse greedy acceptance (3.12 vs 3.34 tokens per step, 118 vs 126 tok/s end to end) and is not what ships here.

Use

With vLLM 0.27.1 + the repo's patches/dflash2-backport.patch (backport of vLLM PR #52816), or vLLM main with PR #52816:

vllm serve <Qwen3.8-27B target> \
  --speculative-config '{"method":"dflash","model":"syvai/Qwen3.8-27B-DFlash2-W4A16","num_speculative_tokens":7}' \
  --attention-backend FLASH_ATTN --kv-cache-dtype bfloat16

The drafter shares the target's embeddings and lm_head (it ships neither); with a quantized target lm_head the repo's patch is needed (upstream refuses a non-bf16 lm_head for the candidate top-k).

Apache-2.0, like the original drafter by Inco.

Downloads last month
5,949
Safetensors
Model size
2B params
Tensor type
I32
ยท
BF16
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for syvai/Qwen3.8-27B-DFlash2-W4A16

Base model

Qwen/Qwen3.8-27B
Quantized
(5)
this model