Text Generation
GGUF
llama-cpp
llama.cpp
ollama
lm-studio
jan
quantized
4bit
4-bit precision
5-bit
8-bit precision
local-llm
on-device
cpu
edge-ai
offline
outlier
outlier-app
qwen2.5
qwen
conversational
chat
instruct
windows
linux
mac
function-calling
Instructions to use Outlier-Ai/Outlier-Max-32B-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use Outlier-Ai/Outlier-Max-32B-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf Outlier-Ai/Outlier-Max-32B-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf Outlier-Ai/Outlier-Max-32B-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf Outlier-Ai/Outlier-Max-32B-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf Outlier-Ai/Outlier-Max-32B-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf Outlier-Ai/Outlier-Max-32B-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf Outlier-Ai/Outlier-Max-32B-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf Outlier-Ai/Outlier-Max-32B-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf Outlier-Ai/Outlier-Max-32B-GGUF:Q4_K_M
Use Docker
docker model run hf.co/Outlier-Ai/Outlier-Max-32B-GGUF:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use Outlier-Ai/Outlier-Max-32B-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Outlier-Ai/Outlier-Max-32B-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Outlier-Ai/Outlier-Max-32B-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Outlier-Ai/Outlier-Max-32B-GGUF:Q4_K_M
- Ollama
How to use Outlier-Ai/Outlier-Max-32B-GGUF with Ollama:
ollama run hf.co/Outlier-Ai/Outlier-Max-32B-GGUF:Q4_K_M
- Unsloth Desktop
- Pi
How to use Outlier-Ai/Outlier-Max-32B-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Outlier-Ai/Outlier-Max-32B-GGUF:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "Outlier-Ai/Outlier-Max-32B-GGUF:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use Outlier-Ai/Outlier-Max-32B-GGUF with Docker Model Runner:
docker model run hf.co/Outlier-Ai/Outlier-Max-32B-GGUF:Q4_K_M
- Lemonade
How to use Outlier-Ai/Outlier-Max-32B-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull Outlier-Ai/Outlier-Max-32B-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.Outlier-Max-32B-GGUF-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use Outlier-Ai/Outlier-Max-32B-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Outlier-Ai/Outlier-Max-32B-GGUF:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default Outlier-Ai/Outlier-Max-32B-GGUF:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use Outlier-Ai/Outlier-Max-32B-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Outlier-Ai/Outlier-Max-32B-GGUF:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "Outlier-Ai/Outlier-Max-32B-GGUF:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
docs(card): SEO + cross-link refresh (HF-DISCOVERABILITY-001)
Browse files
README.md
CHANGED
|
@@ -67,99 +67,28 @@ widget:
|
|
| 67 |
content: Summarize the differences between 4-bit and 8-bit quantization in two
|
| 68 |
short paragraphs.
|
| 69 |
---
|
| 70 |
-
>
|
| 71 |
-
>
|
| 72 |
-
> This is a **Max (deferred to v1.9+ — see Plus tier on Outlier-Ai org) GGUF** build from a pre-v1.8 era. The current Outlier app (v1.8+) ships **MLX 4-bit** versions optimized for Apple Silicon — much faster than GGUF on Mac.
|
| 73 |
-
>
|
| 74 |
-
> | Want | Use |
|
| 75 |
-
> |---|---|
|
| 76 |
-
> | **Run on macOS** (any M-series Mac) | [Outlier-Ai (MLX 4-bit lineup)](https://huggingface.co/Outlier-Ai) |
|
| 77 |
-
> | **Run on Windows / Linux** | This GGUF still works with [llama.cpp](https://github.com/ggerganov/llama.cpp), [Ollama](https://ollama.ai), [LM Studio](https://lmstudio.ai), [Jan](https://jan.ai) |
|
| 78 |
-
>
|
| 79 |
-
> [📥 Download the Outlier desktop app](https://outlier.host)
|
| 80 |
|
| 81 |
-
--
|
| 82 |
-
|
| 83 |
-
|
| 84 |
-
# Outlier Max 32B (GGUF)
|
| 85 |
-
|
| 86 |
-
Cross-platform build for llama.cpp, Ollama, LM Studio, and Jan. Runs on macOS, Windows, and
|
| 87 |
-
Linux. Multiple quant levels included so you can pick by RAM budget.
|
| 88 |
-
|
| 89 |
-
## Quick facts
|
| 90 |
-
|
| 91 |
-
- **Formats included:** Q4_K_M, Q5_K_M, Q8_0
|
| 92 |
-
- **Frozen base:** [Qwen2.5-32B-Instruct](https://huggingface.co/Qwen/Qwen2.5-32B-Instruct)
|
| 93 |
-
- **License:** Apache 2.0
|
| 94 |
-
- **Runtime:** llama.cpp, Ollama, LM Studio, Jan, any GGUF consumer
|
| 95 |
-
|
| 96 |
-
## Quickstart (Ollama)
|
| 97 |
-
|
| 98 |
-
```bash
|
| 99 |
-
ollama run hf.co/Outlier-Ai/Outlier-Max-32B-GGUF:Q4_K_M
|
| 100 |
-
```
|
| 101 |
-
|
| 102 |
-
## Quickstart (llama.cpp)
|
| 103 |
-
|
| 104 |
-
```bash
|
| 105 |
-
./llama-cli \
|
| 106 |
-
-m Outlier-Ai/Outlier-Max-32B-GGUF/model-Q4_K_M.gguf \
|
| 107 |
-
-p "Explain 4-bit quantization in one paragraph." \
|
| 108 |
-
-n 256
|
| 109 |
-
```
|
| 110 |
-
|
| 111 |
-
## Quant picker
|
| 112 |
-
|
| 113 |
-
| Quant | Recommended for |
|
| 114 |
-
|---|---|
|
| 115 |
-
| Q4_K_M | most users, laptops with 8 GB plus of free RAM |
|
| 116 |
-
| Q5_K_M | better quality trade, 12 GB plus free RAM |
|
| 117 |
-
| Q8_0 | highest quality, 24 GB plus free RAM |
|
| 118 |
-
|
| 119 |
-
## Related
|
| 120 |
-
|
| 121 |
-
- **Outlier organization:** [Outlier-Ai](https://huggingface.co/Outlier-Ai)
|
| 122 |
-
- **Consumer Edition collection:** [Outlier-Ai/outlier-consumer-edition-69e2fb4a0df119ea1747275e](https://huggingface.co/collections/Outlier-Ai/outlier-consumer-edition-69e2fb4a0df119ea1747275e)
|
| 123 |
-
- **Research collection:** [Outlier-Ai/outlier-research-69e2fb3a71984614b3c7a279](https://huggingface.co/collections/Outlier-Ai/outlier-research-69e2fb3a71984614b3c7a279)
|
| 124 |
-
- **Server V3.2 collection:** [Outlier-Ai/outlier-server-v32-69e2fb4b71984614b3c7a4a3](https://huggingface.co/collections/Outlier-Ai/outlier-server-v32-69e2fb4b71984614b3c7a4a3)
|
| 125 |
-
- **Desktop app:** [https://outlier.host](https://outlier.host)
|
| 126 |
-
- **Founders lifetime ($200, 500-seat cap):** [https://buy.polar.sh/polar_cl_mJfYZsEpEMDcYrgxzvTdnahSeSQNq1UYLqV0l08CUhW](https://buy.polar.sh/polar_cl_mJfYZsEpEMDcYrgxzvTdnahSeSQNq1UYLqV0l08CUhW)
|
| 127 |
-
- **Discord:** [discord.gg/Hapennmdn9](https://discord.gg/Hapennmdn9)
|
| 128 |
-
|
| 129 |
-
## What is Outlier?
|
| 130 |
-
|
| 131 |
-
Outlier is a Mac-native, offline by default AI platform. One desktop app, a curated model library,
|
| 132 |
-
Bring-Your-Own-Model support, projects with codebase indexing, a 9-tool coding agent, an
|
| 133 |
-
artifacts panel, SQLite memory, and an OpenAI-compatible local API - all Apache 2.0,
|
| 134 |
-
all local, no cloud round-trip, no per-token billing.
|
| 135 |
-
|
| 136 |
-
- **Desktop app:** [https://outlier.host](https://outlier.host)
|
| 137 |
-
- **Founders lifetime (500-seat cap, $200 one-time):** [support Outlier](https://buy.polar.sh/polar_cl_mJfYZsEpEMDcYrgxzvTdnahSeSQNq1UYLqV0l08CUhW)
|
| 138 |
-
- **Discord:** [discord.gg/Hapennmdn9](https://discord.gg/Hapennmdn9)
|
| 139 |
-
- **Org:** [huggingface.co/Outlier-Ai](https://huggingface.co/Outlier-Ai)
|
| 140 |
|
| 141 |
-
|
| 142 |
-
$1,200 in compute spend. Mission: a local AI model that codes as well as the cloud flagships,
|
| 143 |
-
funded by users who want it - not investors who need exits.
|
| 144 |
|
|
|
|
| 145 |
|
| 146 |
-
|
| 147 |
|
| 148 |
-
-
|
| 149 |
-
|
| 150 |
-
-
|
| 151 |
-
|
| 152 |
-
-
|
| 153 |
|
| 154 |
-
|
| 155 |
|
| 156 |
-
|
| 157 |
-
Capability credit for the base belongs to upstream.
|
| 158 |
|
| 159 |
-
|
| 160 |
-
- Qwen 2.5 release: [arxiv.org/abs/2412.15115](https://arxiv.org/abs/2412.15115)
|
| 161 |
|
| 162 |
-
## License
|
| 163 |
|
| 164 |
-
|
| 165 |
-
contributed by Outlier-Ai and released under the same terms.
|
|
|
|
| 67 |
content: Summarize the differences between 4-bit and 8-bit quantization in two
|
| 68 |
short paragraphs.
|
| 69 |
---
|
| 70 |
+
> **Superseded.** This repo is a research artifact from an earlier Outlier lineage and is no longer the recommended download. The current shipping tier is at **[outlier.host](https://outlier.host/?utm_source=hf&utm_medium=modelcard&utm_campaign=outlier_max_32b_gguf)**.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 71 |
|
| 72 |
+
# Outlier-Max-32B (GGUF)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 73 |
|
| 74 |
+
This repo predates the v1.8 Outlier lineup. It is preserved here for reproducibility and historical reference, not as a production recommendation.
|
|
|
|
|
|
|
| 75 |
|
| 76 |
+
## What replaced it
|
| 77 |
|
| 78 |
+
Current shipping tiers (see Outlier app v1.8+):
|
| 79 |
|
| 80 |
+
- [Outlier Nano 4B — current entry tier](https://huggingface.co/Outlier-Ai/Outlier-Nano-4B-MLX-4bit)
|
| 81 |
+
- [Outlier Core 27B — current default tier](https://huggingface.co/Outlier-Ai/Outlier-Core-27B-MLX-4bit)
|
| 82 |
+
- [Outlier Vision 35B-A3B — current multimodal tier](https://huggingface.co/Outlier-Ai/Outlier-Vision-35B-A3B-MLX-4bit)
|
| 83 |
+
- [DeepSeek-R1-Distill-Qwen-7B — popular reasoning model](https://huggingface.co/Outlier-Ai/DeepSeek-R1-Distill-Qwen-7B-MLX-4bit)
|
| 84 |
+
- [Qwen3-Coder-30B-A3B — popular coding model](https://huggingface.co/Outlier-Ai/Qwen3-Coder-30B-A3B-Instruct-MLX-4bit)
|
| 85 |
|
| 86 |
+
For the latest verified benchmarks and downloads, visit **[outlier.host](https://outlier.host/?utm_source=hf&utm_medium=modelcard&utm_campaign=outlier_max_32b_gguf)**.
|
| 87 |
|
| 88 |
+
## Original notes
|
|
|
|
| 89 |
|
| 90 |
+
This was a research / preview artifact. It may contain experimental adapters, overlays, or quantization variants that did not graduate into the shipping product. Treat any technical claims in earlier revisions of this card as provisional.
|
|
|
|
| 91 |
|
| 92 |
+
## License
|
| 93 |
|
| 94 |
+
See YAML frontmatter above. Original license terms preserved.
|
|
|