kernelpool/GLM-5.3-5bit-UVMAX

Mixed-precision (UVMAX) quantization of zai-org/GLM-5.3-BF16, converted with mlx-lm from the bf16 release.

What is UVMAX?

UVMAX assigns bit widths per tensor class instead of quantizing uniformly.

tensor class precision parameters size share
Expert FFN gate/up 4-bit, group 128 483B 239.1 GiB 59.1%
Expert FFN down 5-bit, group 128 242B 147.7 GiB 36.5%
Attention + DSA indexer + dense MLP 8-bit, group 64 13.7B 13.6 GiB 3.4%
Shared experts 8-bit, group 64 2.8B 2.8 GiB 0.7%
Embeddings, LM head 4-bit, group 64 1.9B 1.0 GiB 0.2%
Routers bf16 0.1B 0.2 GiB <0.1%
Norms bf16 <0.1%
total 4.67 bits/weight 743B 404 GiB

Quality

Teacher-forced against the bf16 release on identical tokens, 48 windows of 1025 tokens. KLD is KL(bf16 ‖ quant) over the full output distribution and is corpus-specific.

UVMAX (4.67 bits/weight)

corpus ppl bf16 ppl UVMAX ratio mean KLD median KLD top-1 agreement
Linux kernel C 1.434 1.459 1.02× 0.040 0.0002 96.8%
XNU kernel C 2.435 2.472 1.02× 0.056 0.0027 93.4%
JavaScriptCore C++ 1.757 1.791 1.02× 0.055 0.0007 95.1%
English prose 2.687 2.750 1.02× 0.063 0.0076 92.4%
all 2.015 2.053 1.02× 0.053 0.0012 94.4%

Uniform 4-bit (4.50 bits/weight)

corpus ppl bf16 ppl 4-bit ratio mean KLD median KLD top-1 agreement
Linux kernel C 1.434 1.492 1.04× 0.076 0.0004 95.3%
XNU kernel C 2.435 2.522 1.04× 0.109 0.0063 91.0%
JavaScriptCore C++ 1.757 1.841 1.05× 0.102 0.0014 93.4%
English prose 2.687 2.909 1.08× 0.143 0.0209 88.6%
all 2.015 2.119 1.05× 0.107 0.0028 92.1%

Use with mlx

Requires a recent mlx-lm with GLM-5.3 (glm_moe_dsa) support.

The model fits on a single 512 GB machine:

mlx_lm.server --model kernelpool/GLM-5.3-5bit-UVMAX

Sampling follows the base model: temperature 1.0, top-p 0.95.

Downloads last month
668
Safetensors
Model size
743B params
Tensor type
U32
·
BF16
·
F32
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for kernelpool/GLM-5.3-5bit-UVMAX

Quantized
(16)
this model