Qwen3.8 · 27B · heretic · ara

EXL3  ·  4.5 bpw  ·  18.7 GB  ·  Dense


format bpw size codebook arch

base model quantized by collection


An ExLlamaV3 build of trohrbaugh/Qwen3.8-27B-heretic-ara at 4.5 bits per weight. See Quants for sibling repos at other bit‑widths or browse the collection.

Quants

BPW     Head bits     Calibration rows     Size     Status
4.0 8 250 17.2 GB link
4.5 8 250 18.7 GB this repo
5.0 8 250 20.3 GB link

Inference

Loader Use it for
TabbyAPI OpenAI‑compatible HTTP server. Drop‑in for OpenAI clients.
text‑generation‑webui Local chat UI. Pick the ExLlamaV3 loader from the model dropdown.
ExLlamaV3 Direct Python API for embedding the model in your own code or pipeline.

Download

pip install -U huggingface_hub

hf download \
  Honkware/Qwen3.8-27B-heretic-ara-exl3-4.5bpw \
  --local-dir ./Qwen3.8-27B-heretic-ara-exl3-4.5bpw
Quantization recipe  (advanced, embedded in quantization_config.json)
Setting Value
Format EXL3
Bits per weight 4.5
Head bits 8
Calibration rows 250
Codebook mul1
Out‑scales always
Parallel mode enabled

The codebook is recorded in the weights, so loaders pick it up with no configuration. mul1 needs ExLlamaV3 v0.0.3 or newer; an older build ignores the marker and decodes the weights with the wrong codebook.

Loaded automatically by every ExLlamaV3 loader; reproduced here for searchability.

License & use

Use and license follow the base model. Quantization adds no additional restrictions. Refer to the upstream repository for terms, citation, and safety documentation.


Quantized with BlockQuant  ·  convention {org}/{model}-exl3-{bpw}bpw
Downloads last month
180
Safetensors
Model size
9B params
Tensor type
BF16
·
F16
·
I16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Honkware/Qwen3.8-27B-heretic-ara-exl3-4.5bpw

Quantized
(34)
this model

Collection including Honkware/Qwen3.8-27B-heretic-ara-exl3-4.5bpw