Qwen3.5-9b-4bit-MTPLX

This is a quantized version of Qwen/Qwen3.5-9B for Apple Silicon (MLX) utilizing Multi-Token Prediction via MTPLX.

Details

  • Original Model: Qwen/Qwen3.5-9B
  • Quantization: 4-bit
  • Group Size: 64
  • Framework: MTPLX

Verification Stats

  • Best depth: D2
  • Multiplier vs autoregressive baseline: 1.97脳
  • Verified on: Apple M5 Pro
  • Sampler: temperature 0.6 路 top_p 0.95 路 top_k 20

See mtplx_runtime.json for the full verification record.

Conversion

Built with MTPLX Forge

Using the App

Grab the Latest Release and read the App Instructions

Using the CLI

Install MTPLX

# Homebrew
brew install youssofal/mtplx/mtplx
# Pip
python3 -m pip install mtplx

Pull down the model:

# MTPLX picks this model up automatically when downloaded
mtplx pull banburist/Qwen-Qwen3.5-9b-4bit-MTPLX

Tune the model:

# Tune immediately after it is pulled
mtplx tune --model <model-or-path> --retune

Start MTPLX:

# for start chat
mtplx start
# for API server only
mtplx serve --port 8000

See the MTPLX repo for more information on configuring the model.

About

This was built to meet the hardware limitations of a 16gb machine without sacrificing MTPLX's performance improvements.

Use Youssof Altouhki's original 6-bit model if you have the hardware: Youssofal/Qwen3.5-9B-MTPLX-Optimized-Speed

Credits

Visit Youssof Altoukhi's profile to see all the official MTPLX builds.

License

This model inherits the Apache 2.0 license from the original Qwen3.5-9B model.

Downloads last month
95
Safetensors
Model size
1B params
Tensor type
BF16
U32
F32
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 馃檵 Ask for provider support

Model tree for banburist/Qwen3.5-9b-4bit-MTPLX

Finetuned
Qwen/Qwen3.5-9B
Quantized
(472)
this model