gemma-4-12b-it-uncensored-W4A16-AutoRound

4-bit weight / 16-bit activation (W4A16) quantization of zaakirio/gemma-4-12b-it-uncensored.

Attribution

Quantization

  • Tool: Intel AutoRound (v0.14.0)
  • Scheme: W4A16 — 4-bit weights, group size 128, sym=True
  • Mode: RTN (iters=0) — required for Gemma 4 numerical stability
  • Format: native auto_round (quant_method=auto-round)
  • Multimodal projections (vision/audio embedders) kept unquantized (quant_nontext_module=False) to preserve the encoder-free multimodal path.

Serving (vLLM)

Encoder-free Gemma 4 Unified; serve with a vLLM build that supports Gemma4UnifiedForConditionalGeneration. On Turing (Tesla T4) the Triton attention TILE_SIZE patch (vllm-project/vllm#39018) is required.

Downloads last month
510
Safetensors
Model size
3B params
Tensor type
I32
·
BF16
·
F16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for hfacesp/gemma-4-12b-it-uncensored-W4A16-AutoRound

Quantized
(7)
this model