Qwen3-0.6B-3bit-g256

Qwen3-0.6B quantized to 3-bit (group_size 256, symmetric) with an 8-bit embedding; โ‰ˆ4.40 bits/weight.

Weights are provided dequantized in fp16 for direct loading and evaluation โ€” the quantization is already baked into the weights, so evaluate the model as-is (no further quantization).

Downloads last month
7
Safetensors
Model size
0.6B params
Tensor type
F16
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for gitarist/Qwen3-0.6B-3bit-g256

Finetuned
Qwen/Qwen3-0.6B
Finetuned
(1173)
this model