Qwen3.8-27B-MTPLX-Bare-Speed-FP16 / mtplx_runtime.json

Commit History

Quantized MTP draft head (INT4/g64 affine, all 8 head matrices)
a6f76e2
verified

Youssofal commited on

restamp: recommended draft temperature 0.6 -> 1.0 (drop-day max-fan A/B: 46.05 vs 42.79 tok/s, higher D2/D3 acceptance)
f61818c
verified

Youssofal commited on

Qwen3.8-27B MTPLX Bare-Speed FP16 sibling: mtplx_runtime.json
4b9b838
verified

Youssofal commited on