Near-lossless GGUF quants of:

Uses the chat template from google/gemma-4-E4B-it (2026-05-18), replacing the outdated chat template that was originally bundled with the base QAT model.


Original QAT Specifications

The base QAT model was trained with the following specifications:

  • Text & Draft Models: 4-bit
  • Multimodal Projector (mmproj):
    • Audio Encoder: Mostly 2-bit (with some tensors in 4-bit or BF16)
    • Vision Encoder: 8-bit
    • Other parts: BF16

Quantization Error Evaluation

  • Dequantization: Quantized weights were dequantized to F32 for evaluation.
  • Error Calculation: Quantization error metrics (compared to the BF16 baseline) were computed and accumulated in F64 precision to prevent numerical underflow and precision loss.

Metrics:

  • MAE (Mean Absolute Error)
  • RMSE (Root Mean Squared Error)
  • Max Error

Interpretation of Metrics: These error metrics do not directly reflect actual model performance. This is because imatrix optimization intentionally increases measured error by downweighting less important parameters to improve overall performance. However, these metrics remain useful for estimating how close the quantized model is to the original, lossless quality.

gemma-4-E4B-it-qat

Model MAE RMSE Max Error
JMingo/gemma-4-E4B-it-qat-BF16.gguf (Baseline) 0.00000000 0.00000000 0.00000000
JMingo/gemma-4-E4B-it-qat-Q4_0.gguf 0.00002867 0.00005571 0.00256348
idkwhattoputherenow/gemma-4-E4B-it-qat-q4_0-unquantized-q4_0-maxerr.gguf 0.00003258 0.00006382 0.00299072
unsloth/gemma-4-E4B-it-qat-UD-Q4_K_XL.gguf 0.00004080 0.00007184 0.00341797
google/gemma-4-E4B_q4_0-it.gguf 0.00046331 0.00084654 0.07275391
lmstudio-community/gemma-4-E4B-it-QAT-Q4_0.gguf 0.00046331 0.00084654 0.07275391
mradermacher/gemma-4-E4B-it-qat-q4_0-unquantized.i1-Q4_0.gguf 0.00024689 0.00053073 0.03417969
mradermacher/gemma-4-E4B-it-qat-q4_0-unquantized.Q4_K_S.gguf 0.00048946 0.00078925 0.04432869
mradermacher/gemma-4-E4B-it-qat-q4_0-unquantized.Q4_K_M.gguf 0.00047553 0.00078097 0.04432869
mradermacher/gemma-4-E4B-it-qat-q4_0-unquantized.i1-Q4_1.gguf 0.00042801 0.00076768 0.21069336
mradermacher/gemma-4-E4B-it-qat-q4_0-unquantized.i1-Q4_K_S.gguf 0.00048423 0.00078383 0.21218872
mradermacher/gemma-4-E4B-it-qat-q4_0-unquantized.i1-Q4_K_M.gguf 0.00047063 0.00077560 0.21218872

mtp-gemma-4-E4B-it-qat

Model MAE RMSE Max Error
JMingo/mtp-gemma-4-E4B-it-qat-BF16.gguf (Baseline) 0.00000000 0.00000000 0.00000000
JMingo/mtp-gemma-4-E4B-it-qat-Q4_0.gguf 0.00005023 0.00008237 0.00329590
unsloth/mtp-gemma-4-E4B-it.gguf 0.00007141 0.00010650 0.00854492

mmproj-gemma-4-E4B-it-qat

Model MAE RMSE Max Error
JMingo/mmproj-gemma-4-E4B-it-qat-BF16.gguf (Baseline) 0.00000000 0.00000000 0.00000000
JMingo/mmproj-gemma-4-E4B-it-qat-Q2_K_XL.gguf 0.00001205 0.00003343 0.00271225
Downloads last month
775
GGUF
Model size
7B params
Architecture
gemma4
Hardware compatibility
Log In to add your hardware

4-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for JMingo/gemma-4-E4B-it-qat-GGUF

Quantized
(36)
this model