Instructions to use AMR5210/qwen3-0.6b-qat-q2k with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use AMR5210/qwen3-0.6b-qat-q2k with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf AMR5210/qwen3-0.6b-qat-q2k:Q2_K # Run inference directly in the terminal: llama cli -hf AMR5210/qwen3-0.6b-qat-q2k:Q2_K
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf AMR5210/qwen3-0.6b-qat-q2k:Q2_K # Run inference directly in the terminal: llama cli -hf AMR5210/qwen3-0.6b-qat-q2k:Q2_K
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf AMR5210/qwen3-0.6b-qat-q2k:Q2_K # Run inference directly in the terminal: ./llama-cli -hf AMR5210/qwen3-0.6b-qat-q2k:Q2_K
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf AMR5210/qwen3-0.6b-qat-q2k:Q2_K # Run inference directly in the terminal: ./build/bin/llama-cli -hf AMR5210/qwen3-0.6b-qat-q2k:Q2_K
Use Docker
docker model run hf.co/AMR5210/qwen3-0.6b-qat-q2k:Q2_K
- LM Studio
- Jan
- vLLM
How to use AMR5210/qwen3-0.6b-qat-q2k with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "AMR5210/qwen3-0.6b-qat-q2k" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "AMR5210/qwen3-0.6b-qat-q2k", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/AMR5210/qwen3-0.6b-qat-q2k:Q2_K
- Ollama
How to use AMR5210/qwen3-0.6b-qat-q2k with Ollama:
ollama run hf.co/AMR5210/qwen3-0.6b-qat-q2k:Q2_K
- Unsloth Desktop
- Pi
How to use AMR5210/qwen3-0.6b-qat-q2k with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf AMR5210/qwen3-0.6b-qat-q2k:Q2_K
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "AMR5210/qwen3-0.6b-qat-q2k:Q2_K" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use AMR5210/qwen3-0.6b-qat-q2k with Docker Model Runner:
docker model run hf.co/AMR5210/qwen3-0.6b-qat-q2k:Q2_K
- Lemonade
How to use AMR5210/qwen3-0.6b-qat-q2k with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull AMR5210/qwen3-0.6b-qat-q2k:Q2_K
Run and chat with the model
lemonade run user.qwen3-0.6b-qat-q2k-Q2_K
List all available models
lemonade list
- Hermes Agent
How to use AMR5210/qwen3-0.6b-qat-q2k with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf AMR5210/qwen3-0.6b-qat-q2k:Q2_K
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default AMR5210/qwen3-0.6b-qat-q2k:Q2_K
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use AMR5210/qwen3-0.6b-qat-q2k with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf AMR5210/qwen3-0.6b-qat-q2k:Q2_K
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "AMR5210/qwen3-0.6b-qat-q2k:Q2_K" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Qwen3-0.6B, 2-bit QAT (Q2_K)
Quantization-aware training of Qwen/Qwen3-0.6B
down to 2 bits, exported to GGUF Q2_K for CPU inference with llama.cpp.
The point of this repo is a controlled comparison: the same base model and the same final on-device format, with the quantization method as the only variable. Naive post-training quantization to 2 bits collapses this model (WikiText-2 perplexity 21.37 → 220.93). QAT recovers it.
Measured on an iPhone 17 Pro Max (A19 Pro, CPU backend, 4 threads), the 2-bit model runs at 67.3 generated tokens/sec in 2.2 GB of RAM.
Training code, benchmark harness, and full result records: AMR5210/arm-create-mobileai.
Featured model
The model this card describes is the fineweb-blend variant — the best of the
recipes tried here on both evaluation sets.
| File | fineweb-blend/qwen3-0.6b-qat-fineweb-blend-q2_k.gguf |
| Size | 495,193,952 bytes (472.3 MiB) |
| SHA-256 | a861b8924a2b1881720d38123ec32aae7f99ac05732d819644a6584f3cc38fef |
| Tensors | 187 Q2_K / 9 F16 skip-layers / 114 copied |
Download
Local scripts and the iOS harness in the companion repo always read one fixed path,
models/qwen3-0.6b-qat-q2_k.gguf, regardless of which variant is current. The
download therefore renames as it saves:
curl -L -o models/qwen3-0.6b-qat-q2_k.gguf \
https://huggingface.co/AMR5210/qwen3-0.6b-qat-q2k/resolve/main/fineweb-blend/qwen3-0.6b-qat-fineweb-blend-q2_k.gguf
Verify what landed on disk before using it. A filename is not an identity — an
earlier round of measurements in this project was invalidated by a same-named
export, which is why the companion repo ships verify_model_signatures.py:
shasum -a 256 models/qwen3-0.6b-qat-q2_k.gguf
# expected: a861b8924a2b1881720d38123ec32aae7f99ac05732d819644a6584f3cc38fef
No authentication is required. The repo is public and ungated.
The
qwen3-0.6b-qat-q2_k.ggufat the root of this repo is a superseded export (510,971,712 bytes,ec36e450…, 12 skip-layers) kept for history. It is not the model described here and matches no recorded result. Use thefineweb-blend/path above.
Model details
Base model. Qwen3-0.6B. 28 layers, tied token embeddings, Apache-2.0.
Training data. A three-way blend, 44,000 examples at --max-examples 44000:
- 44,000 Alpaca instruction examples
- 4,889 WikiText-2 train examples (
--wikitext-frac 0.05) - 43,705 FineWeb web-text examples (
--fineweb-frac 0.5)
The WikiText-2 share was deliberately kept light — a prior recipe used 0.5 and scored worse on both eval sets. No test split was used in training.
QAT method.
- Affine (asymmetric) fake quantization with a straight-through estimator, applied to 187 linear layers in the forward pass
- CAKLD distillation against the fp16 teacher (
--distill-loss cakld --distill-weight 0.5), the objective from BitDistiller (arXiv:2402.10631) - PTQ-quality initialization (
--init-mode ptq_q2k): shadow weights start from real dequantizedQ2_Kvalues rather than the fp16 weights, so training begins on the grid it will be deployed on - Group size 16 (
--group-size 16), matchingQ2_K's 16-element sub-blocks - 9 outlier-heavy layers held at full precision:
layers.{8,16,21,27}.self_attn.k_proj,layers.{26,27}.mlp.{gate,up}_proj,layers.2.mlp.down_proj - 48,828 steps (~600M tokens), AdamW, batch 24, lr 2.5e-5, forward bit-width annealed 4→2 over the first 150 steps
- Trained on one AMD Instinct MI300X. No non-finite loss or gradient events.
Export. A custom pure-numpy Q2_K encoder
(qat/gguf_q2k.py)
writes the GGUF directly. llama-quantize re-derives its own Q2_K scales and
minima, which would mean the deployed weights are not the ones training optimized.
The encoder is validated against gguf-py's own dequantizer.
Evaluation
llama.cpp llama-perplexity, full corpus, n_ctx=512, non-overlapping chunks.
| Model | WikiText-2 | C4 |
|---|---|---|
| fp16 base, unadapted | 21.37 | 29.46 |
PTQ Q2_K (naive 2-bit) |
220.93 | 279.30 |
This model (QAT Q2_K) |
18.49 ± 0.15 | 32.25 ± 0.28 |
Both 2-bit rows use F16 tied embeddings, so neither is credited an embedding-precision advantage the other lacks.
On-device (iPhone 17 Pro Max, CPU backend, 4 threads):
| Variant | Disk | Peak RAM | Prompt tok/s | Gen tok/s | WikiText-2 |
|---|---|---|---|---|---|
| fp16 | 1509.3 MB | 4220.6 MB | 819.23 | 47.34 | 21.37 |
| PTQ 2-bit | 479.8 MB | 2149.3 MB | 686.39 | 64.27 | 220.91 |
| QAT 2-bit | 495.2 MB | 2181.0 MB | 758.27 | 67.30 | 18.46 |
The on-device perplexity figures reproduce desktop llama-perplexity to within
0.011%.
Limitations
WikiText-2 18.49 is below the fp16 base's 21.37. This is not evidence that 2-bit QAT beats fp16. The training blend contains WikiText-2 train data, so the comparison rewards domain adaptation. The off-domain C4 column is the honest read, and there the unadapted fp16 base is still ahead (29.46 vs 32.25).
No fp16 ceiling run. Any ratio against the fp16 base compares QAT to a model that never saw the training blend, conflating domain adaptation with the cost of quantization. The fair ceiling — fp16 fine-tuned on the identical blend for the same step count — has not been run. Until it exists, those ratios are raw data, not support for a comparison in either direction.
KleidiAI does not accelerate this model. Arm's KleidiAI microkernels are
compiled in, but supports_op covers Q4_0, Q8_0, F32 and one F16 path — not
Q2_K. The 2-bit weight matmuls run on stock ggml CPU kernels. The throughput
figures above are unaccelerated.
Instruction-following was not measurably improved and is not reported. Across three scoring methods every variant sat at or below the 25% chance line on a 4-way multiple-choice slice; under forced choice this model selected "A" for all 100 questions. Perplexity is the metric this work stands on.
The training blend includes Alpaca, which is CC BY-NC 4.0 — non-commercial.
Whether that restriction reaches model weights trained on the data is unsettled.
Stanford's own Alpaca release took the position that it does: its usage notice states
the dataset allows only non-commercial use, and that models trained on it should not
be used outside research purposes. One part of their reasoning does not carry over —
their base model was LLaMA, itself under a non-commercial license, whereas this model's
base is Qwen3-0.6B under Apache-2.0. The data question remains open regardless, and
the apache-2.0 label on this repo, inherited from Qwen3-0.6B, does not resolve it.
Anyone considering commercial use should evaluate this independently or consult
counsel. Full license inventory:
THIRD_PARTY_NOTICES.md.
hf/ subfolders are fp16, not 2-bit. They hold materialized fake-quant values
— 2-bit-representable numbers stored at full width — and give no size or speed
benefit. They exist for further training or re-export. Use the .gguf for inference.
Other variants
Supplementary, for reproducibility. Each is a different recipe on the same base model; they are not separately documented to the depth of the featured model.
| Folder | Recipe | WikiText-2 |
|---|---|---|
fineweb-blend/ |
Featured. PTQ init + light WikiText-2 + FineWeb | 18.49 |
ptq-init-blend/ |
PTQ init, WikiText-2-heavy blend, no FineWeb | 26.95 |
seq-roundtrip/ |
ParetoQ SEQ quantizer, roundtrip init | 26.59 |
cakld/ |
CAKLD, earlier recipe | — |
alpaca-only/ |
Alpaca only, no corpus blend | — |
fp16-blend/ |
fp16 fine-tune, ceiling control for ptq-init-blend |
— |
Negative results are recorded with the same detail as positive ones in the
companion repo's
results/.
License
Apache-2.0, matching the base model (Qwen/Qwen3-0.6B).
- Downloads last month
- 166
2-bit
16-bit