Text Generation
MLX
Safetensors
qwen3_5
16gb-mac
4-bit precision
4bit
apple-silicon
balanced
chat
conversational
edge-ai
everyday
function-calling
instruct
lite
local-llm
m1
m2
m3
m4
mac
mac-mini
mac-studio
macbook-air
macbook-pro
macos
metal
mlx-lm
mmlu-verified
no-cloud
offline
on-device
outlier
outlier-app
private
private-ai
quantized
qwen
qwen3.5
reasoning
thinking
tool-use
Eval Results (legacy)
Instructions to use Outlier-Ai/Outlier-Lite-9B-MLX-4bit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use Outlier-Ai/Outlier-Lite-9B-MLX-4bit with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("Outlier-Ai/Outlier-Lite-9B-MLX-4bit") prompt = "Write a story about Einstein" messages = [{"role": "user", "content": prompt}] prompt = tokenizer.apply_chat_template( messages, add_generation_prompt=True ) text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use Outlier-Ai/Outlier-Lite-9B-MLX-4bit with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "Outlier-Ai/Outlier-Lite-9B-MLX-4bit"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "Outlier-Ai/Outlier-Lite-9B-MLX-4bit" } ] } } }Run Pi
# Start Pi in your project directory: pi
- MLX LM
How to use Outlier-Ai/Outlier-Lite-9B-MLX-4bit with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Interactive chat REPL mlx_lm.chat --model "Outlier-Ai/Outlier-Lite-9B-MLX-4bit"
Run an OpenAI-compatible server
# Install MLX LM uv tool install mlx-lm # Start the server mlx_lm.server --model "Outlier-Ai/Outlier-Lite-9B-MLX-4bit" # Calling the OpenAI-compatible server with curl curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Outlier-Ai/Outlier-Lite-9B-MLX-4bit", "messages": [ {"role": "user", "content": "Hello"} ] }' - Hermes Agent
How to use Outlier-Ai/Outlier-Lite-9B-MLX-4bit with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "Outlier-Ai/Outlier-Lite-9B-MLX-4bit"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default Outlier-Ai/Outlier-Lite-9B-MLX-4bit
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use Outlier-Ai/Outlier-Lite-9B-MLX-4bit with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "Outlier-Ai/Outlier-Lite-9B-MLX-4bit"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "Outlier-Ai/Outlier-Lite-9B-MLX-4bit" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
docs(card): SEO + cross-link refresh (HF-DISCOVERABILITY-001)
Browse files
README.md
CHANGED
|
@@ -94,72 +94,50 @@ model-index:
|
|
| 94 |
value: 0.6951
|
| 95 |
verified: false
|
| 96 |
---
|
|
|
|
| 97 |
|
| 98 |
-
# Outlier Lite 9B (MLX-
|
| 99 |
|
| 100 |
-
|
| 101 |
|
| 102 |
-
|
| 103 |
|
| 104 |
-
|
| 105 |
-
|
| 106 |
-
## At a glance
|
| 107 |
-
|
| 108 |
-
| Property | Value |
|
| 109 |
-
|---|---|
|
| 110 |
-
| Architecture | Qwen3.5, hybrid linear:full attention 3:1 |
|
| 111 |
-
| Parameters | 9B |
|
| 112 |
-
| Quantization | MLX 4-bit |
|
| 113 |
-
| Disk size | **5.04 GB** (vision tower stripped) |
|
| 114 |
-
| Min RAM | **12 GB** unified memory |
|
| 115 |
-
| Speed | **53.4 tok/s** (M1 Ultra 64 GB) |
|
| 116 |
-
| Context (default) | 32K (native: 256K) |
|
| 117 |
-
| Thinking mode | ✅ |
|
| 118 |
-
| License | Apache 2.0 |
|
| 119 |
-
|
| 120 |
-
---
|
| 121 |
-
|
| 122 |
-
## Verified benchmarks
|
| 123 |
|
| 124 |
-
|
| 125 |
-
|---|---|---|---|---|
|
| 126 |
-
| MMLU (5-shot) | **0.7846** ± 0.0035 | 14,042 (full test set) | 0.003498 | 2026-04-30 |
|
| 127 |
-
| HumanEval pass@1 | **0.6951** ± 0.0361 | 164 | 0.036058 | 2026-04-30 |
|
| 128 |
|
| 129 |
-
|
| 130 |
|
| 131 |
-
|
| 132 |
|
| 133 |
-
|
| 134 |
|
| 135 |
```bash
|
| 136 |
pip install mlx-lm
|
| 137 |
-
mlx_lm.generate
|
| 138 |
-
--
|
| 139 |
-
--
|
|
|
|
| 140 |
```
|
| 141 |
|
| 142 |
```python
|
| 143 |
-
from mlx_lm import load,
|
| 144 |
-
|
| 145 |
model, tokenizer = load("Outlier-Ai/Outlier-Lite-9B-MLX-4bit")
|
| 146 |
-
|
| 147 |
-
print(chunk.text, end="", flush=True)
|
| 148 |
```
|
| 149 |
|
| 150 |
-
|
|
|
|
|
|
|
| 151 |
|
| 152 |
-
## Outlier
|
| 153 |
|
| 154 |
-
|
| 155 |
-
|
| 156 |
-
|
| 157 |
-
|
| 158 |
-
|
| 159 |
-
| [Core](https://hf.co/Outlier-Ai/Outlier-Core-27B-MLX-4bit) | 27B | 20.7 tok/s | 24 GB |
|
| 160 |
-
| [Code](https://hf.co/Outlier-Ai/Outlier-Code-27B-MLX-4bit) | 27B | 20.7 tok/s | 24 GB |
|
| 161 |
-
| [Vision](https://hf.co/Outlier-Ai/Outlier-Vision-35B-A3B-MLX-4bit) | 35B MoE | ~61 tok/s | 24 GB |
|
| 162 |
|
| 163 |
## License
|
| 164 |
|
| 165 |
-
Apache 2.0 (
|
|
|
|
| 94 |
value: 0.6951
|
| 95 |
verified: false
|
| 96 |
---
|
| 97 |
+
> **Part of the [Outlier](https://outlier.host/?utm_source=hf&utm_medium=modelcard&utm_campaign=outlier_lite_9b_mlx_4bit) shipping lineup.** Outlier is a free macOS app that runs this model locally, with one click. Apple Silicon only.
|
| 98 |
|
| 99 |
+
# Outlier Lite 9B (MLX 4-bit)
|
| 100 |
|
| 101 |
+
The mid-tier shipping model. 9B dense, text-only, MLX 4-bit. Recommended for 12 GB+ Macs that want stronger reasoning than Nano without paying the latency cost of a 27B-class model.
|
| 102 |
|
| 103 |
+
## Try it in Outlier
|
| 104 |
|
| 105 |
+
The simplest way to use this model is through the Outlier app — open the tier picker, select **Outlier Lite**, click download, and chat. No setup, no Python, no MLX install, no token quotas.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 106 |
|
| 107 |
+
➡ **[Download Outlier — outlier.host](https://outlier.host/?utm_source=hf&utm_medium=modelcard&utm_campaign=outlier_lite_9b_mlx_4bit)**
|
|
|
|
|
|
|
|
|
|
| 108 |
|
| 109 |
+
A screenshot of the tier picker is at [outlier.host/screenshots/tier-picker.png](https://outlier.host/screenshots/tier-picker.png?utm_source=hf&utm_medium=modelcard&utm_campaign=outlier_lite_9b_mlx_4bit).
|
| 110 |
|
| 111 |
+
## Load this directly (power users)
|
| 112 |
|
| 113 |
+
If you want the raw MLX-4bit weights without the app:
|
| 114 |
|
| 115 |
```bash
|
| 116 |
pip install mlx-lm
|
| 117 |
+
python -m mlx_lm.generate \
|
| 118 |
+
--model Outlier-Ai/Outlier-Lite-9B-MLX-4bit \
|
| 119 |
+
--prompt "Write a quicksort in Python." \
|
| 120 |
+
--max-tokens 512
|
| 121 |
```
|
| 122 |
|
| 123 |
```python
|
| 124 |
+
from mlx_lm import load, generate
|
|
|
|
| 125 |
model, tokenizer = load("Outlier-Ai/Outlier-Lite-9B-MLX-4bit")
|
| 126 |
+
print(generate(model, tokenizer, prompt="Hello", max_tokens=256))
|
|
|
|
| 127 |
```
|
| 128 |
|
| 129 |
+
## Verified benchmarks
|
| 130 |
+
|
| 131 |
+
For σ-qualified MMLU, HumanEval, and Mac inference-speed numbers — with full provenance (source file, command, n, stderr, date) — see **[outlier.host/benchmarks](https://outlier.host/benchmarks?utm_source=hf&utm_medium=modelcard&utm_campaign=outlier_lite_9b_mlx_4bit)**.
|
| 132 |
|
| 133 |
+
## Other Outlier shipping tiers
|
| 134 |
|
| 135 |
+
- [Outlier Nano 4B (entry tier, ~3 GB)](https://huggingface.co/Outlier-Ai/Outlier-Nano-4B-MLX-4bit)
|
| 136 |
+
- [Outlier Quick 26B-A4B MoE (~16 GB)](https://huggingface.co/Outlier-Ai/Outlier-Quick-26B-MLX-4bit)
|
| 137 |
+
- [Outlier Core 27B (default, ~16 GB)](https://huggingface.co/Outlier-Ai/Outlier-Core-27B-MLX-4bit)
|
| 138 |
+
- [Outlier Code 27B (code-tuned, ~16 GB)](https://huggingface.co/Outlier-Ai/Outlier-Code-27B-MLX-4bit)
|
| 139 |
+
- [Outlier Vision 35B-A3B (multimodal, ~20 GB)](https://huggingface.co/Outlier-Ai/Outlier-Vision-35B-A3B-MLX-4bit)
|
|
|
|
|
|
|
|
|
|
| 140 |
|
| 141 |
## License
|
| 142 |
|
| 143 |
+
Apache 2.0 (inherits from upstream base model). Conversion artifact only — the underlying weights are governed by the base model's license.
|