⚡ Each donation = another big MoE quantized

I host 30+ free APEX MoE quantizations as independent research. My only local hardware is an NVIDIA DGX Spark (122 GB unified memory), enough for ~30-50B-class MoEs, but bigger ones (200B+) require rented compute on H100/H200/Blackwell, typically $20-100 per quant.
If APEX quants are useful to you, your support directly funds those bigger runs.

🎉 Patreon (Monthly)  |  ☕ Buy Me a Coffee  |  ⭐ GitHub Sponsors

Ornith-1.5-35B-A3B APEX GGUF

APEX quantizations of ornith-ai/Ornith-1.5-35B-A3B.

Brought to you by the LocalAI team | APEX Project

These are the standard quants. For versions that bundle the MTP draft head for speculative decoding, see Ornith-1.5-35B-A3B-APEX-MTP-GGUF.

Files

File Size For
Ornith-1.5-35B-A3B-APEX-Quality.gguf 22.82 GB highest quality
Ornith-1.5-35B-A3B-APEX-Balanced.gguf 25.27 GB general purpose
Ornith-1.5-35B-A3B-APEX-Compact.gguf 16.54 GB consumer GPUs
Ornith-1.5-35B-A3B-APEX-I-Mini.gguf 13.47 GB smallest, imatrix only
mmproj.gguf 0.90 GB vision projector, pair with any of the above

I- files use an importance matrix built from diverse calibration data (chat, code, reasoning, tool-calling, agentic traces, Wikipedia). Quality, Balanced and Compact also ship without it.

The model

Ornith-1.5-35B-A3B is a 36 B parameter Mixture-of-Experts model with 256 routed experts and 8 active per token, plus a shared expert. It has 40 layers with hybrid attention, interleaving three linear-attention layers per full-attention layer, and a vision tower.

How APEX quantizes it

Routed experts are 89.6% of the weights here but only 8 of 256 fire for any given token, so they tolerate lower precision than the parts every token passes through. APEX classifies each tensor by role and applies a layer-wise precision gradient: the first and last layers keep higher precision, middle layers compress harder, and the always-active shared expert is kept high.

Attention is only 3.6% of the weights on this model (2.8% linear, 0.8% full), so it is not where the size is and is not treated as a lever.

Usage

# text
llama-cli -m Ornith-1.5-35B-A3B-APEX-Balanced.gguf -p "Your prompt" -ngl 99

# vision
llama-mtmd-cli -m Ornith-1.5-35B-A3B-APEX-Balanced.gguf --mmproj mmproj.gguf -ngl 99

Needs a recent llama.cpp with qwen3_5_moe support.

Notes

Sizes and quantization recipes are published in the APEX repository. No throughput benchmarks were run on these files.

Downloads last month
15,677
GGUF
Model size
35B params
Architecture
qwen35moe
Hardware compatibility
Log In to add your hardware

We're not able to determine the quantization variants.

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for mudler/Ornith-1.5-35B-A3B-APEX-GGUF

Quantized
(132)
this model

Collection including mudler/Ornith-1.5-35B-A3B-APEX-GGUF