How to use from
OpenClaw
Start the MLX server
# Install MLX LM:
uv tool install mlx-lm
# Start a local OpenAI-compatible server:
mlx_lm.server --model "ToPo-ToPo/gemma-4-E2B-it-qat-mlx-4bit"
Configure OpenClaw
# Install OpenClaw:
npm install -g openclaw@latest
# Register the local server and set it as the default model:
openclaw onboard --non-interactive --mode local \
  --auth-choice custom-api-key \
  --custom-base-url http://127.0.0.1:8080/v1 \
  --custom-model-id "ToPo-ToPo/gemma-4-E2B-it-qat-mlx-4bit" \
  --custom-provider-id mlx-lm \
  --custom-compatibility openai \
  --custom-text-input \
  --accept-risk \
  --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Quick Links

ToPo-ToPo/gemma-4-E2B-it-qat-mlx-4bit

MLX 4-bit conversion of google/gemma-4-E2B-it-qat-q4_0-unquantized, made with mlx-vlm 0.6.3.

Provenance (self-converted)

  • Source: google/gemma-4-E2B-it-qat-q4_0-unquantized (license: gemma)
  • Tool: mlx-vlm 0.6.3 mlx_vlm.convert (q_bits=4, group_size=64, affine), ~5.6 bpw
  • Validated: loads and translates correctly under mlx-vlm 0.6.3.

Usage

from mlx_vlm import load
model, processor = load("ToPo-ToPo/gemma-4-E2B-it-qat-mlx-4bit")

License

Derivative of Google Gemma; governed by the Gemma Terms of Use (https://ai.google.dev/gemma/terms) and Prohibited Use Policy. Converted to MLX.

⚡ MTP drafter (speculative decoding)

Use google/gemma-4-E2B-it-assistant — Google's official MTP drafter for this model. It loads directly in mlx-vlm (>= 0.6.3), needs no conversion, and speculative decoding is lossless. Drafters are size-specific and not interchangeable across Gemma 4 variants.

🔧 Patched chat template

chat_template.jinja differs from Google's Gemma 4 Canonical Chat Template (2026-07-09) by one intentional change; everything else is untouched.

The canonical template suppresses the thinking channel at the start of a normal model turn, but emits nothing after a tool response when enable_thinking is false. The model may then open a thinking channel on its own, and a quantized model sometimes writes the literal word thought into the answer. This patch gives the tool_response branch the same suppression:

{%- elif ns.prev_message_type == 'tool_response' -%}
    {%- if enable_thinking -%}
        {{- '<|channel>thought\n' -}}
    {%- else -%}
        {{- '<|channel>thought\n<channel|>' -}}
    {%- endif -%}
{%- endif -%}

Only that case changes — the other prompt paths render byte-identical to the canonical template. To get stock behaviour, replace chat_template.jinja with the one from the base model repo; the weights are unaffected.

Downloads last month
61
Safetensors
Model size
1B params
Tensor type
BF16
·
U32
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ToPo-ToPo/gemma-4-E2B-it-qat-mlx-4bit

Quantized
(48)
this model

Collection including ToPo-ToPo/gemma-4-E2B-it-qat-mlx-4bit