AutoRound & ASHQ1-Remix Suite
Version 2.3.1 β Activation-aware GGUF quantization whose every ratio, floor, and cap traces to a measured experiment. Plain-BF16-native first; AutoRound lineage supported with explicit saturation bounds. Full seven-tier ladder validated across six model families.
π Highlights
ββββββββββββββββββββββββββ
β Safetensors (Raw/BF16) β
ββββββββββββββ¬ββββββββββββ
β 00_SAFETENSORS-to-AutoRound-BF16-GGUF.py (optional, int4 sources)
βΌ
ββββββββββββββββββββββββββ
β BF16 GGUF + provenance β
ββββββββββββββ¬ββββββββββββ
β 01_create-calibration-dataset-and-imatrix.py (lineage-aware)
βΌ
βββββββββββββββββββββββββββββββββββββββββββββββββββ
β 02_BF16-GGUF-to-ASHQ1.py β
β ASHQ1-Remix optimizer (knapsack + floors) β
β + ASHQ1-mmproj.py (vision tower, imatrix-free) |
ββββββββββββββ¬βββββββββββββββββββββββββββββββββββββ
βΌ
ββββββββββββββββββββββββββββββββ
β Tiers: Pico 24 Β· Nano 27 Β· β
| Mini 30 Β· Compact 33 Β· β
| Quality 36 Β· Precision 42 Β· β
| Fidelity 48 (%) β
ββββββββββββββ¬ββββββββββββββββββ
βΌ
ββββββββββββββββββββββββββββββββββββββββββββββββ
β 03_perplexity_test.py + 90_attribution- β
β probe.py PPL/KLD/RMS/top-p Β· probes β
ββββββββββββββββββββββββββββββββββββββββββββββββ
- Measurement-backed: the Calibration Ledger L0βL8 documents eight laws
plus the 2.3.0 refinements β L6 completed into a three-regime
U-curve, L7 refined into an addition-only lever law β each rooted in
knockout probes and stock-twin duels archived in
attribution-results.csv. - Seven catalogue ratios on the arithmetic 24/27/30/33/36/42/48 ladder: Precision-42 and the duel-rehabilitated Fidelity-48 join the monotone core across all supported families; Pico-24 remains an opt-in service tier.
- Cross-family proven: qwen35 hybrids (4B & 9B), vision/OCR models (OvisOCR2), dense architectures (SmolLM3 3B tied, llama 1B), and the mixer-dominant LFM2.5-2.6B β including a mapped sub-Nano validity domain.
- Module-complete: vision towers and standalone speculative drafts (DSpark) quantize through dedicated imatrix-free engines β the trunk grid never touches a module it cannot rank, and drafts score on acceptance rate, not corpus proxies.
- Honest failures: floors never lie β impossible targets warn instead of silently degrading safety classes; pipeline gates compare RATIOS against a measured reference, never absolute thresholds that would condemn a healthy build for merely being off-domain (Charter Β§5).
- Stock-defended: every tier is ranked against its size-matched stock twin (L6): allocation owns the sub-wall band, uniform+imatrix the mid-band, and near-lossless archival defers to stock Q8_0's measured dual crown (fidelity + throughput).
π Tiers & Selection Policy
| Tier | Plain ratio | Role |
|---|---|---|
| Pico | 24% | Service tier, opt-in ASHQ1_INCLUDE_PICO=1 β tight-VRAM serving; validity β₯ |
| Nano | 27% | Edge cases / maximum compression (blocked by default on int4 lineage) |
| Mini | 30% | Minimum for β₯9B serving |
| Compact | 33% | Minimum for 3β4B; balanced deployment |
| Quality | 36% | Minimum for ~1B; near-lossless general deployment (flat builder on routed families) |
| Precision | 42% | Flat Q6_K + embd lever β default quasi-lossless top tier across all families (KLD 0.0021β0.0132) |
| Fidelity | 48% | Archival flat Q6_K + embd lever β disabled by default, opt-in via ASHQ1_INCLUDE_FIDELITY=1 |
Minimum-tier policy (measured, not folklore β see Charter Β§7). Smaller
models hit the constructibility wall sooner: a 1B below 33% stops
differentiating adjacent tiers, while a 9B tolerates 30% comfortably and
rides Pico at upper-usable. Validity is architectural, not merely scalar:
dispersed-projector hybrids absorb the Pico step, mixer-dominant ones
(LFM2.5 family) bottom out at Mini (law L8). AutoRound-int4 lineage:
practical ceiling is Compact; Nano AND Pico are excluded unless
ASHQ1_INCLUDE_NANO=1 (shared double-quantization guard).
π Release Benchmarks (wiki.test.raw, symmetric FA-auto reference)
Ornith-1.5-9B (hybrid, plain lineage)
| Tier | Size | PPL | KLD | RMS Ξp | top-p |
|---|---|---|---|---|---|
| Fidelity-48pc | 9464 MiB | 9.5239 | 0.0081 | 2.43% | 97.6% |
| Precision-42pc | 8414 MiB | 9.4347 | 0.0132 | 3.04% | 96.6% |
| Quality-36pc | 6330 MiB | 9.3692 | 0.0366 | 5.00% | 93.3% |
| Compact-33pc π₯ Second Choice | 5803 MiB | 9.6043 | 0.0517 | 5.91% | 91.6% |
| Mini-30pc β Recommended | 5385 MiB | 9.8564 | 0.0649 | 6.67% | 90.3% |
| Nano-27pc | 4750 MiB | 10.1061 | 0.0907 | 7.90% | 87.8% |
| Pico-24pc | 4389 MiB | 10.1078 | 0.1309 | 9.63% | 85.0% |
Qwen3.8-4B-Distill (hybrid, plain lineage)
| Tier | Size | PPL | KLD | RMS Ξp | top-p |
|---|---|---|---|---|---|
| Fidelity-48pc | 3866 MiB | 9.0744 | 0.0014 | 1.01% | 98.1% |
| Precision-42pc | 3509 MiB | 9.0818 | 0.0021 | 1.20% | 97.6% |
| Quality-36pc | 2902 MiB | 9.1124 | 0.0095 | 2.56% | 95.1% |
| Compact-33pc π₯ Second Choice | 2661 MiB | 9.1963 | 0.0160 | 3.31% | 93.7% |
| Mini-30pc β Recommended | 2420 MiB | 9.2973 | 0.0258 | 4.26% | 92.1% |
| Nano-27pc | 2245 MiB | 9.5957 | 0.0546 | 6.61% | 88.7% |
| Pico-24pc | 2168 MiB | 9.6579 | 0.0636 | 6.95% | 87.9% |
TwIL-LM3 (dense SmolLM3-arch, tied token_embd, 128k vocab, L7-fixed)
| Tier | Size | PPL | KLD | RMS Ξp | top-p |
|---|---|---|---|---|---|
| Fidelity-48pc | 2826 MiB | 9.4677 | 0.0024 | 1.26% | 97.0% |
| Precision-42pc | 2474 MiB | 9.4744 | 0.0036 | 1.53% | 96.4% |
| Quality-36pc β Recommended | 2121 MiB | 9.5785 | 0.0163 | 3.35% | 92.4% |
| Compact-33pc π₯ Second Choice | 1945 MiB | 9.6292 | 0.0235 | 3.94% | 91.2% |
| Mini-30pc | 1769 MiB | 9.6890 | 0.0304 | 4.42% | 90.2% |
| Nano-27pc | 1593 MiB | 9.9084 | 0.0499 | 5.75% | 88.1% |
| Pico-24pc | 1445 MiB | 10.3800 | 0.0959 | 7.95% | 84.3% |
LFM2.5-2.6B (shortconv-mixer hybrid, tied readout β 2.3.0 ladder)
| Tier | Size | PPL | KLD | RMS Ξp | top-p |
|---|---|---|---|---|---|
| Fidelity-48pc | 2481 MiB | 56.0022 | 0.0072 | 1.87% | 95.8% |
| Precision-42pc | 2179 MiB | 55.2953 | 0.0108 | 2.32% | 94.9% |
| Quality-36pc β Recommended | 1850 MiB | 55.5619 | 0.0354 | 4.12% | 91.0% |
| Compact-33pc π₯ Second Choice | 1708 MiB | 54.6921β | 0.0762 | 6.05% | 87.0% |
| Mini-30pc | 1554 MiB | 56.5119 | 0.1199 | 7.53% | 83.9% |
| Nano-27pc | 1399 MiB | 61.1460 | 0.2131 | 10.06% | 78.9% |
| Pico-24pc | 1276 MiB | 56.7726β | 0.3707 | 13.08% | 72.7% |
Stock twins (same protocol): Q8_0 2742 MiB Β· 0.0030 Β· 97.3% β Q6_K 2119 Β· 0.0143 Β· 94.1% β Q5_K_M 1850 Β· 0.0467 Β· 89.8% β Q4_K_M 1597 Β· 0.1552 Β· 81.7%. Quality-36 rides the flat builder (mid-band law); Precision and Fidelity deliver consistent quasi-lossless fidelity across all measured architectures.
ReaderLM-v2 (dense HTML/markdown extractor, 1.5B, 28 layers)
| Tier | Size | PPL | KLD | RMS Ξp | top-p |
|---|---|---|---|---|---|
| Fidelity-48pc | 1422 MiB | 17.0871 | 0.0050 | 1.65% | 96.3% |
| Precision-42pc π₯ Second Choice | 1268 MiB | 17.0608 | 0.0076 | 2.03% | 95.5% |
| Quality-36pc β Recommended | 1068 MiB | 17.0610 | 0.0291 | 4.01% | 91.2% |
| Compact-33pc | 980 MiB | 17.5504 | 0.1211 | 7.74% | 87.1% |
| Mini-30pc | 891 MiB | 17.2444 | 0.1145 | 7.74% | 85.4% |
| Nano-27pc | 803 MiB | 17.3237 | 0.1513 | 9.04% | 81.4% |
| Pico-24pc | 763 MiB | 18.0892 | 0.2087 | 10.87% | 77.9% |
MiniCPM5-1B (dense llama-arch)
| Tier | Size | PPL | KLD | RMS Ξp | top-p |
|---|---|---|---|---|---|
| Fidelity-48pc | 997 MiB | 26.3758 | 0.0056 | 1.63% | 95.3% |
| Precision-42pc π₯ Second Choice | 943 MiB | 26.4353 | 0.0072 | 1.84% | 94.5% |
| Quality-36pc β Recommended | 749 MiB | 27.2456 | 0.0555 | 5.16% | 86.3% |
| Compact-33pc | 687 MiB | 28.1717 | 0.0997 | 6.69% | 81.9% |
| Mini-30pc | 663 MiB | 28.5485 | 0.1153 | 7.16% | 80.4% |
| Nano-27pc | 563 MiB | 28.8736 | 0.1382 | 7.98% | 78.4% |
| Pico-24pc | 503 MiB | 37.5561 | 0.3854 | 14.27% | 66.3% |
PaddleOCR-VL-1.6 (vision-language OCR decoder, ERNIE-4.5-0.3B)
| Tier | Size | PPL | KLD | RMS Ξp | top-p |
|---|---|---|---|---|---|
| Fidelity-48pc | 537 MiB | 365.8789 | 0.0366 | 2.99% | 91.1% |
| Precision-42pc β Recommended | 484 MiB | 362.9912 | 0.0337 | 2.57% | 91.5% |
| Quality-36pc | 430 MiB | 400.7598 | 0.1747 | 5.72% | 80.7% |
| Compact-33pc | 404 MiB | 351.6550 | 0.2320 | 6.82% | 77.3% |
| Mini-30pc | 378 MiB | 333.6259 | 0.2712 | 7.33% | 74.7% |
| Nano-27pc | 351 MiB | 385.6355 | 0.2612 | 7.73% | 72.5% |
| Pico-24pc | 324 MiB | 509.1721 | 0.5632 | 11.01% | 63.5% |
OvisOCR2 (vision-aligned text backbone, ~0.3B)
| Tier | Size | PPL | KLD | RMS Ξp | top-p |
|---|---|---|---|---|---|
| Fidelity-48pc | 704 MiB | 31.0673 | 0.0021 | 0.98% | 97.4% |
| Precision-42pc π₯ Second Choice | 668 MiB | 31.0809 | 0.0026 | 1.06% | 97.1% |
| Quality-36pc β Recommended | 531 MiB | 32.0240 | 0.0155 | 2.68% | 93.0% |
| Compact-33pc | 487 MiB | 32.7223 | 0.0306 | 3.90% | 90.3% |
| Mini-30pc | 474 MiB | 32.8232 | 0.0352 | 4.17% | 89.8% |
| Nano-27pc | 457 MiB | 33.1136 | 0.0648 | 5.95% | 86.7% |
| Pico-24pc | 442 MiB | 34.7822 | 0.0775 | 6.33% | 85.2% |
π Sub-Nano Compendium (full-corpus protocol, v2.2.0)
Orderings replicate the tables above; absolutes shift a few percent with full 559β580-chunk passes (never mix protocols in one column β Charter Β§5).
| Family | Arch | NanoβPico KLD step | Pico verdict |
|---|---|---|---|
| Qwen3.8-4B | GDN hybrid | +16.5% | β 2168 MiB Β· 0.0636 Β· top-p 87.9% (effective ~27%β ) |
| OvisOCR2 | vision backbone | +19.6% | β 442 MiB Β· 0.0775 Β· top-p 85.2% |
| Ornith-1.5-9B | GDN hybrid | +44% | β 4389 MiB Β· 0.1309 Β· top-p 85.0% |
| LFM2.5-2.6B | shortconv mixer | +73% | β 0.3707 Β· 72.7% β floor is Mini (L8) |
| TwIL-LM3 | dense tied | +92% | β οΈ 1445 MiB Β· 0.0959 Β· 84.3% β priced, allowed |
| MiniCPM5-1B | dense untied | +179% | β 0.3854 Β· 66.3% β collapse; ends at Nano |
β no-mtp trunk + carved heads raise effective coverage to ~27%.
Highlights of the campaign behind these rows: the cheapest marginal ever
measured here is the readout rung embd Q6_KβQ8_0 (Pico+, TwIL
0.0959β0.0914 / 85.0% for +60 MiB); a size-matched engineered stock twin
TIES Pico-24 (ΞKLD 0.0018 < noise) β allocation is inert at constant
codebook support; and on LFM, raising only the 8 attention layers
recovers 0.371β0.213 while shielding the whole shortconv stack recovers
just 0.030 β law L8: cut locality beats cut depth. Full per-tier
tables and KO records: CHARTER.md Β§6b + attribution-results.csv.
π οΈ Suite Components
| Script | Purpose |
|---|---|
00_SAFETENSORS-to-AutoRound-BF16-GGUF.py |
Optional int4 reconditioning + provenance sidecars + standalone draft-module builds (shard-index audit, parent-tokenizer borrowing). |
00b_BF16-GGUF-MTP-extract.py |
Split the MTP draft head off the trunk. |
01_create-calibration-dataset-and-imatrix.py |
Multi-domain corpus + lineage-aware imatrix; GPU autotune behind a relative split-integrity gate (--cpu-only bypasses offload on any architecture). |
01b_BF16-GGUF-modules-fusion.py |
Reattach quantized modules (MTP/mmproj). |
02_BF16-GGUF-to-ASHQ1.py |
Batch tier orchestration (auto-runs mmproj + draft modules; --module dspark builds one standalone). |
03_perplexity_test.py |
KL reference base + PPL/KLD/RMS/top-p/per-chunk sweep. |
ASHQ1.py |
Core optimizer (single target CLI + tier runner). |
ASHQ1-mmproj.py |
Vision-tower tiers, deliberately imatrix-free. |
ASHQ1-dspark.py |
Speculative-draft tiers, imatrix-free, acceptance-rate domain. |
90_attribution-probe.py |
Knockout attribution harness (methodology instrument). |
β‘ Quick Start
pip install gguf numpy huggingface_hub safetensors # + auto-round/torchvision for step 00
python 00_SAFETENSORS-to-AutoRound-BF16-GGUF.py # create BF16 gguf files from safetensors project, create AutoRound version with `--autoround`
python 01_create-calibration-dataset-and-imatrix.py # imatrix.gguf (plain lineage)
python 02_BF16-GGUF-to-ASHQ1.py # batch tiers + modules (ASHQ1_INCLUDE_PICO=1 or ASHQ1_INCLUDE_FIDELITY=1 for opt-in tiers)
python 02_BF16-GGUF-to-ASHQ1.py --module dspark # standalone draft module only
python 03_perplexity_test.py # evaluate
python 90_attribution-probe.py --list # optional: inspect probes
Windows/NTFS: compact /c /exe:xpress8k <file> shrinks KLD logits ~85% and
BF16 GGUFs ~18% without touching results.
π Citation & Credits
- ASHQ1 (Autonomous Selective Hybrid Quantization) by wepiqx β priority-queue knapsack formulation, tied-group activation hashing, MSE scheduling.
- Empero AI (Qwen3.8-27B-Ridge) β GDN state preservation (
ssm_alpha/ssm_beta@ Q8_0) and native MTP draft heads. - Intel AutoRound β sign-gradient low-bit optimization with Hessian compensation.
- llama.cpp by Georgi Gerganov & ggml contributors β GGUF/GGML runtime and tools.
- Calibration recipes inspired by Bartowski; multi-imatrix max-combination per community practice (cHunter789 KV-cache recipe referenced by the orchestrator).
License: apache-2.0