gemma-4-12b-it-uncensored-W4A16-AutoRound
4-bit weight / 16-bit activation (W4A16) quantization of
zaakirio/gemma-4-12b-it-uncensored.
Attribution
- Original model:
google/gemma-4-12B-it— © Google, released under the Gemma license. - Immediate base (quantized here):
zaakirio/gemma-4-12b-it-uncensored - This repository only quantizes that base; all model capabilities and weights originate upstream. License inherits from the original Gemma model.
Quantization
- Tool: Intel AutoRound (
v0.14.0) - Scheme: W4A16 — 4-bit weights, group size 128,
sym=True - Mode: RTN (
iters=0) — required for Gemma 4 numerical stability - Format: native
auto_round(quant_method=auto-round) - Multimodal projections (vision/audio embedders) kept unquantized
(
quant_nontext_module=False) to preserve the encoder-free multimodal path.
Serving (vLLM)
Encoder-free Gemma 4 Unified; serve with a vLLM build that supports
Gemma4UnifiedForConditionalGeneration. On Turing (Tesla T4) the Triton
attention TILE_SIZE patch (vllm-project/vllm#39018) is required.
- Downloads last month
- 510
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support
Model tree for hfacesp/gemma-4-12b-it-uncensored-W4A16-AutoRound
Base model
google/gemma-4-12B Finetuned
google/gemma-4-12B-it Finetuned
zaakirio/gemma-4-12b-it-uncensored