This is Qwen/Qwen3.8-27B quantized with AutoRound. Model packed by llm-compressor.

Quantization ("compact" settings):

  • INT data types
  • ~2.8 bits for linear layers
  • 6-bit LM head
  • BF16 MTP layers
  • 8-bit activations

The model is compatible with vLLM (tested: v0.26+). Made and Tested with an RTX Pro 6000 (Google Colab's G4). Currently under evaluation: Subscribe to be notified when I publish the results.

Instructions

vllm serve [this_model_id] \
  --tensor-parallel-size 1 \
  --enable-auto-tool-choice \
  --tool-call-parser qwen3_coder \
  --reasoning-parser qwen3 

Add --speculative-config '{"method":"mtp","num_speculative_tokens":X}' for MTP, where good values for X are between 2 and 6 (depending on your tasks and concurrency).

Downloads last month
50
Safetensors
Model size
4B params
Tensor type
I32
·
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for kaitchup/Qwen3.8-27B-hum-compact-MTP-14.1GB

Base model

Qwen/Qwen3.8-27B
Quantized
(708)
this model

Collection including kaitchup/Qwen3.8-27B-hum-compact-MTP-14.1GB