Instructions to use mlx-community/Qwen3.8-27B-OBLITERATED-OptiQ-4bit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use mlx-community/Qwen3.8-27B-OBLITERATED-OptiQ-4bit with MLX:
# Make sure mlx-vlm is installed # pip install --upgrade mlx-vlm from mlx_vlm import load, generate from mlx_vlm.prompt_utils import apply_chat_template from mlx_vlm.utils import load_config # Load the model model, processor = load("mlx-community/Qwen3.8-27B-OBLITERATED-OptiQ-4bit") config = load_config("mlx-community/Qwen3.8-27B-OBLITERATED-OptiQ-4bit") # Prepare input image = ["http://images.cocodataset.org/val2017/000000039769.jpg"] prompt = "Describe this image." # Apply chat template formatted_prompt = apply_chat_template( processor, config, prompt, num_images=1 ) # Generate output output = generate(model, processor, formatted_prompt, image) print(output) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
mlx-community/Qwen3.8-27B-OBLITERATED-OptiQ-4bit
Built with mlx-optiq, the MLX-native toolkit to quantize, fine-tune, and serve LLMs locally on Apple Silicon, no PyTorch and no cloud. All OptiQ quants · Docs
OptiQ mixed-precision quant of OBLITERATUS/Qwen3.8-27B-OBLITERATED, a Qwen3.8-family vision-language model with a bundled MTP speculation head. 21 GB on disk.
What it is
| Property | Value |
|---|---|
| Base | OBLITERATUS/Qwen3.8-27B-OBLITERATED (Qwen3.8, 27B) |
| Method | OptiQ mixed-precision, per-layer 4/8-bit |
| Bit allocation | Reused from the Qwen3.8-27B OptiQ recipe: the architecture is identical, so the per-layer sensitivity ranking transfers directly and no per-model sweep is needed |
| Layer split | 237 components at 4-bit, 261 at 8-bit |
| Group size | 64 |
| On disk | 21 GB |
| MTP | Speculation head preserved in optiq/mtp.safetensors |
| Vision | bf16 vision tower kept in optiq/optiq_vision.safetensors for image input |
Following the naming llama.cpp uses for its mixed quants, the "4bit" label denotes the family, not the weighted average.
Run it
Qwen3.8 and the MTP/vision sidecars register through OptiQ, so import optiq once before loading:
pip install "mlx-optiq>=0.4.27"
import optiq # registers the arch + MTP/vision sidecars
from mlx_lm import load, generate
model, tok = load("mlx-community/Qwen3.8-27B-OBLITERATED-OptiQ-4bit")
prompt = tok.apply_chat_template(
[{"role": "user", "content": "Explain mixed-precision quantization in two sentences."}],
tokenize=False, add_generation_prompt=True,
)
print(generate(model, tok, prompt=prompt, max_tokens=400))
For image input plus an OpenAI- and Anthropic-compatible endpoint with mixed-precision KV cache:
optiq serve --model mlx-community/Qwen3.8-27B-OBLITERATED-OptiQ-4bit
This is a reasoning model, so give it a generous token budget.
Links
- Project website: mlx-optiq.com
- All OptiQ quants: mlx-optiq.com/models
- Base model: OBLITERATUS/Qwen3.8-27B-OBLITERATED
- Downloads last month
- 2,074
4-bit