ur-dad-matt commited on
Commit
519040b
·
verified ·
1 Parent(s): 01dbe7b

docs(card): SEO + cross-link refresh (HF-DISCOVERABILITY-001)

Browse files
Files changed (1) hide show
  1. README.md +25 -47
README.md CHANGED
@@ -94,72 +94,50 @@ model-index:
94
  value: 0.6951
95
  verified: false
96
  ---
 
97
 
98
- # Outlier Lite 9B (MLX-4bit)
99
 
100
- > **53.4 tok/s on M1 Ultra. MMLU 0.7846 (+2.22 σ). Balanced everyday AI for 12+ GB Macs.**
101
 
102
- **Outlier Lite** is a text-only 4-bit MLX build of [Qwen3.5-9B](https://huggingface.co/Qwen/Qwen3.5-9B) for Apple Silicon. The balanced tier in the [Outlier](https://outlier.host) lineup — fast, capable, fits on MacBook Pro 16 GB.
103
 
104
- ---
105
-
106
- ## At a glance
107
-
108
- | Property | Value |
109
- |---|---|
110
- | Architecture | Qwen3.5, hybrid linear:full attention 3:1 |
111
- | Parameters | 9B |
112
- | Quantization | MLX 4-bit |
113
- | Disk size | **5.04 GB** (vision tower stripped) |
114
- | Min RAM | **12 GB** unified memory |
115
- | Speed | **53.4 tok/s** (M1 Ultra 64 GB) |
116
- | Context (default) | 32K (native: 256K) |
117
- | Thinking mode | ✅ |
118
- | License | Apache 2.0 |
119
-
120
- ---
121
-
122
- ## Verified benchmarks
123
 
124
- | Task | Metric | n | stderr | Date |
125
- |---|---|---|---|---|
126
- | MMLU (5-shot) | **0.7846** ± 0.0035 | 14,042 (full test set) | 0.003498 | 2026-04-30 |
127
- | HumanEval pass@1 | **0.6951** ± 0.0361 | 164 | 0.036058 | 2026-04-30 |
128
 
129
- MMLU: **+1.10 pp / 2.22 σ above base** — clears the strict 2σ bar. One of the most accurately benchmarked 9B models on HuggingFace.
130
 
131
- ---
132
 
133
- ## Quick start
134
 
135
  ```bash
136
  pip install mlx-lm
137
- mlx_lm.generate --model Outlier-Ai/Outlier-Lite-9B-MLX-4bit \
138
- --prompt "Summarize the key ideas in transformers architecture." \
139
- --max-tokens 256
 
140
  ```
141
 
142
  ```python
143
- from mlx_lm import load, stream_generate
144
-
145
  model, tokenizer = load("Outlier-Ai/Outlier-Lite-9B-MLX-4bit")
146
- for chunk in stream_generate(model, tokenizer, "What is Apple Silicon?", max_tokens=128):
147
- print(chunk.text, end="", flush=True)
148
  ```
149
 
150
- ---
 
 
151
 
152
- ## Outlier lineup
153
 
154
- | Tier | Params | Speed | Min RAM |
155
- |---|---|---|---|
156
- | [Nano](https://hf.co/Outlier-Ai/Outlier-Nano-4B-MLX-4bit) | 4B | 71.7 tok/s | 6 GB |
157
- | **Lite** you are here | 9B | **53.4 tok/s** | 12 GB |
158
- | [Quick](https://hf.co/Outlier-Ai/Outlier-Quick-26B-MLX-4bit) | 26B MoE | 14.6 tok/s | 16 GB |
159
- | [Core](https://hf.co/Outlier-Ai/Outlier-Core-27B-MLX-4bit) | 27B | 20.7 tok/s | 24 GB |
160
- | [Code](https://hf.co/Outlier-Ai/Outlier-Code-27B-MLX-4bit) | 27B | 20.7 tok/s | 24 GB |
161
- | [Vision](https://hf.co/Outlier-Ai/Outlier-Vision-35B-A3B-MLX-4bit) | 35B MoE | ~61 tok/s | 24 GB |
162
 
163
  ## License
164
 
165
- Apache 2.0 (inherited from Qwen3.5-9B).
 
94
  value: 0.6951
95
  verified: false
96
  ---
97
+ > **Part of the [Outlier](https://outlier.host/?utm_source=hf&utm_medium=modelcard&utm_campaign=outlier_lite_9b_mlx_4bit) shipping lineup.** Outlier is a free macOS app that runs this model locally, with one click. Apple Silicon only.
98
 
99
+ # Outlier Lite 9B (MLX 4-bit)
100
 
101
+ The mid-tier shipping model. 9B dense, text-only, MLX 4-bit. Recommended for 12 GB+ Macs that want stronger reasoning than Nano without paying the latency cost of a 27B-class model.
102
 
103
+ ## Try it in Outlier
104
 
105
+ The simplest way to use this model is through the Outlier app — open the tier picker, select **Outlier Lite**, click download, and chat. No setup, no Python, no MLX install, no token quotas.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
106
 
107
+ **[Download Outlier outlier.host](https://outlier.host/?utm_source=hf&utm_medium=modelcard&utm_campaign=outlier_lite_9b_mlx_4bit)**
 
 
 
108
 
109
+ A screenshot of the tier picker is at [outlier.host/screenshots/tier-picker.png](https://outlier.host/screenshots/tier-picker.png?utm_source=hf&utm_medium=modelcard&utm_campaign=outlier_lite_9b_mlx_4bit).
110
 
111
+ ## Load this directly (power users)
112
 
113
+ If you want the raw MLX-4bit weights without the app:
114
 
115
  ```bash
116
  pip install mlx-lm
117
+ python -m mlx_lm.generate \
118
+ --model Outlier-Ai/Outlier-Lite-9B-MLX-4bit \
119
+ --prompt "Write a quicksort in Python." \
120
+ --max-tokens 512
121
  ```
122
 
123
  ```python
124
+ from mlx_lm import load, generate
 
125
  model, tokenizer = load("Outlier-Ai/Outlier-Lite-9B-MLX-4bit")
126
+ print(generate(model, tokenizer, prompt="Hello", max_tokens=256))
 
127
  ```
128
 
129
+ ## Verified benchmarks
130
+
131
+ For σ-qualified MMLU, HumanEval, and Mac inference-speed numbers — with full provenance (source file, command, n, stderr, date) — see **[outlier.host/benchmarks](https://outlier.host/benchmarks?utm_source=hf&utm_medium=modelcard&utm_campaign=outlier_lite_9b_mlx_4bit)**.
132
 
133
+ ## Other Outlier shipping tiers
134
 
135
+ - [Outlier Nano 4B (entry tier, ~3 GB)](https://huggingface.co/Outlier-Ai/Outlier-Nano-4B-MLX-4bit)
136
+ - [Outlier Quick 26B-A4B MoE (~16 GB)](https://huggingface.co/Outlier-Ai/Outlier-Quick-26B-MLX-4bit)
137
+ - [Outlier Core 27B (default, ~16 GB)](https://huggingface.co/Outlier-Ai/Outlier-Core-27B-MLX-4bit)
138
+ - [Outlier Code 27B (code-tuned, ~16 GB)](https://huggingface.co/Outlier-Ai/Outlier-Code-27B-MLX-4bit)
139
+ - [Outlier Vision 35B-A3B (multimodal, ~20 GB)](https://huggingface.co/Outlier-Ai/Outlier-Vision-35B-A3B-MLX-4bit)
 
 
 
140
 
141
  ## License
142
 
143
+ Apache 2.0 (inherits from upstream base model). Conversion artifact only — the underlying weights are governed by the base model's license.