Text Generation
MLX
Safetensors
gemma4
4-bit precision
4bit
apple-silicon
chat
conversational
edge-ai
function-calling
gemma
gemma-4
instruct
local-llm
m1
m2
m3
m4
mac
mac-mini
mac-studio
macbook-air
macbook-pro
macos
metal
mixture-of-experts
mlx-lm
mmlu-verified
Mixture of Experts
multilingual
no-cloud
offline
on-device
outlier
outlier-app
private
private-ai
quantized
reasoning
thinking
tool-use
Eval Results (legacy)
Instructions to use Outlier-Ai/Outlier-Quick-26B-MLX-4bit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use Outlier-Ai/Outlier-Quick-26B-MLX-4bit with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("Outlier-Ai/Outlier-Quick-26B-MLX-4bit") prompt = "Write a story about Einstein" messages = [{"role": "user", "content": prompt}] prompt = tokenizer.apply_chat_template( messages, add_generation_prompt=True ) text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use Outlier-Ai/Outlier-Quick-26B-MLX-4bit with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "Outlier-Ai/Outlier-Quick-26B-MLX-4bit"
Configure the model in Pi
# Install Pi: npm install -g @mariozechner/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "Outlier-Ai/Outlier-Quick-26B-MLX-4bit" } ] } } }Run Pi
# Start Pi in your project directory: pi
- MLX LM
How to use Outlier-Ai/Outlier-Quick-26B-MLX-4bit with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Interactive chat REPL mlx_lm.chat --model "Outlier-Ai/Outlier-Quick-26B-MLX-4bit"
Run an OpenAI-compatible server
# Install MLX LM uv tool install mlx-lm # Start the server mlx_lm.server --model "Outlier-Ai/Outlier-Quick-26B-MLX-4bit" # Calling the OpenAI-compatible server with curl curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Outlier-Ai/Outlier-Quick-26B-MLX-4bit", "messages": [ {"role": "user", "content": "Hello"} ] }' - Hermes Agent
How to use Outlier-Ai/Outlier-Quick-26B-MLX-4bit with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "Outlier-Ai/Outlier-Quick-26B-MLX-4bit"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default Outlier-Ai/Outlier-Quick-26B-MLX-4bit
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use Outlier-Ai/Outlier-Quick-26B-MLX-4bit with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "Outlier-Ai/Outlier-Quick-26B-MLX-4bit"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "Outlier-Ai/Outlier-Quick-26B-MLX-4bit" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
File size: 3,732 Bytes
cbfad44 dcd8777 24e224b cbfad44 dcd8777 cbfad44 24e224b ec1106a dcd8777 24e224b dcd8777 cbfad44 ebc6928 cbfad44 ebc6928 cbfad44 ebc6928 cbfad44 ebc6928 cbfad44 ebc6928 cbfad44 ebc6928 dcd8777 ebc6928 dcd8777 ebc6928 dcd8777 ebc6928 cbfad44 dcd8777 ebc6928 dcd8777 cbfad44 dcd8777 ebc6928 dcd8777 ebc6928 dcd8777 ebc6928 dcd8777 ebc6928 dcd8777 ebc6928 cbfad44 ebc6928 cbfad44 ebc6928 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 | ---
language:
- en
- zh
- fr
- es
- pt
- de
- it
- ru
- ja
- ko
- ar
- vi
- th
- nl
- pl
license: apache-2.0
library_name: mlx
base_model: google/gemma-4-26b-a4b-it
tags:
- 4-bit
- 4bit
- apple-silicon
- chat
- conversational
- edge-ai
- function-calling
- gemma
- gemma-4
- gemma4
- instruct
- local-llm
- m1
- m2
- m3
- m4
- mac
- mac-mini
- mac-studio
- macbook-air
- macbook-pro
- macos
- metal
- mixture-of-experts
- mlx
- mlx-lm
- mmlu-verified
- moe
- multilingual
- no-cloud
- offline
- on-device
- outlier
- outlier-app
- private
- private-ai
- quantized
- reasoning
- safetensors
- text-generation
- thinking
- tool-use
pipeline_tag: text-generation
model-index:
- name: Outlier-Ai/Outlier-Quick-26B-MLX-4bit
results:
- task:
type: text-generation
name: Text Generation
dataset:
name: MMLU (stratified n=300)
type: cais/mmlu
config: all
split: test
metrics:
- type: acc
name: accuracy
value: 0.7933
verified: false
- task:
type: text-generation
name: Text Generation
dataset:
name: HumanEval
type: openai_humaneval
split: test
metrics:
- type: pass@1
name: pass@1
value: 0.128
verified: false
---
> **Part of the [Outlier](https://outlier.host/?utm_source=hf&utm_medium=modelcard&utm_campaign=outlier_quick_26b_mlx_4bit) shipping lineup.** Outlier is a free macOS app that runs this model locally, with one click. Apple Silicon only.
# Outlier Quick 26B-A4B (MLX 4-bit)
Sparse MoE tier (26B params, ~4B active per token). Sits between Lite and Core in latency, with stronger thinking-mode reasoning. Optimized for general chat and reasoning, not for code generation.
## Try it in Outlier
The simplest way to use this model is through the Outlier app — open the tier picker, select **Outlier Quick**, click download, and chat. No setup, no Python, no MLX install, no token quotas.
➡ **[Download Outlier — outlier.host](https://outlier.host/?utm_source=hf&utm_medium=modelcard&utm_campaign=outlier_quick_26b_mlx_4bit)**
A screenshot of the tier picker is at [outlier.host/screenshots/tier-picker.png](https://outlier.host/screenshots/tier-picker.png?utm_source=hf&utm_medium=modelcard&utm_campaign=outlier_quick_26b_mlx_4bit).
## Load this directly (power users)
If you want the raw MLX-4bit weights without the app:
```bash
pip install mlx-lm
python -m mlx_lm.generate \
--model Outlier-Ai/Outlier-Quick-26B-MLX-4bit \
--prompt "Write a quicksort in Python." \
--max-tokens 512
```
```python
from mlx_lm import load, generate
model, tokenizer = load("Outlier-Ai/Outlier-Quick-26B-MLX-4bit")
print(generate(model, tokenizer, prompt="Hello", max_tokens=256))
```
## Verified benchmarks
For σ-qualified MMLU, HumanEval, and Mac inference-speed numbers — with full provenance (source file, command, n, stderr, date) — see **[outlier.host/benchmarks](https://outlier.host/benchmarks?utm_source=hf&utm_medium=modelcard&utm_campaign=outlier_quick_26b_mlx_4bit)**.
## Other Outlier shipping tiers
- [Outlier Nano 4B (entry tier, ~3 GB)](https://huggingface.co/Outlier-Ai/Outlier-Nano-4B-MLX-4bit)
- [Outlier Lite 9B (balanced, ~6 GB)](https://huggingface.co/Outlier-Ai/Outlier-Lite-9B-MLX-4bit)
- [Outlier Core 27B (default, ~16 GB)](https://huggingface.co/Outlier-Ai/Outlier-Core-27B-MLX-4bit)
- [Outlier Code 27B (code-tuned, ~16 GB)](https://huggingface.co/Outlier-Ai/Outlier-Code-27B-MLX-4bit)
- [Outlier Vision 35B-A3B (multimodal, ~20 GB)](https://huggingface.co/Outlier-Ai/Outlier-Vision-35B-A3B-MLX-4bit)
## License
Apache 2.0 (inherits from upstream base model). Conversion artifact only — the underlying weights are governed by the base model's license.
|