MiniCPM5-2B-MLX — Pollard

Pollard shrank this model for Apple Silicon: 5.04 GB (f16) → 1.71 GB66% smaller, 3.0× down.

The smallest rung here; larger, higher-fidelity rungs are listed below.

Pollard builds of openbmb/MiniCPM5-2B made with Pollard Weights — a ladder of measured-allocation quants (bits placed by per-layer sensitivity, not a uniform crush).

Model details

Parameter count ~2.5B
Architecture llama
Input support text
imatrix no
Perplexity measured yes — table below

Which file should I choose?

Every rung is the same weights, sized to a different RAM budget by the measured allocation. Pick the largest one that fits your machine with room for context:

  • ~5 GB RAM / VRAMq8/model.safetensors (2.67 GB).
  • ~4 GB RAM / VRAMmodel.safetensors (2.04 GB).
  • ~4 GB RAM / VRAMq4/model.safetensors (1.71 GB).

Available files

file PPL size Mean KLD notes
q4/model.safetensors 1.71 GB q4/model.safetensors
model.safetensors 2.04 GB model.safetensors
q8/model.safetensors 2.67 GB q8/model.safetensors

Available rungs

rung bpw size notes path
q8 8.50 2.5 GB near-lossless q8/
mix (recommended) 6.48 1.9 GB measured 4/8 mixed-precision repo root
q4 5.43 1.6 GB smallest q4/

PPL / Mean-KLD benchmarking pending — sizes and allocation are final.

Download a specific file

pip install -U "huggingface_hub[cli]"
hf download PollardWeights/MiniCPM5-2B-Pollard-MLX \
  --include "q4/model.safetensors" --local-dir ./

How to run

mlx_lm.generate --model PollardWeights/MiniCPM5-2B-Pollard-MLX --prompt "Hello"

Errata

  • Measured allocation places bits by per-layer sensitivity under a size budget.
  • Single machine; replication invited.

Credits & license

Built with Pollard Weights — frontier models, small hardware, no compromise.

Downloads last month
687
Safetensors
Model size
3B params
Tensor type
U32
·
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for PollardWeights/MiniCPM5-2B-Pollard-MLX

Quantized
(69)
this model