AutoRound & ASHQ1-Remix Suite

Version 2.3.1 β€” Activation-aware GGUF quantization whose every ratio, floor, and cap traces to a measured experiment. Plain-BF16-native first; AutoRound lineage supported with explicit saturation bounds. Full seven-tier ladder validated across six model families.


🌟 Highlights

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ Safetensors (Raw/BF16) β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
             β”‚ 00_SAFETENSORS-to-AutoRound-BF16-GGUF.py (optional, int4 sources)
             β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ BF16 GGUF + provenance β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
             β”‚ 01_create-calibration-dataset-and-imatrix.py (lineage-aware)
             β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ 02_BF16-GGUF-to-ASHQ1.py                        β”‚
β”‚  ASHQ1-Remix optimizer (knapsack + floors)      β”‚
β”‚  + ASHQ1-mmproj.py (vision tower, imatrix-free) |
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
             β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ Tiers: Pico 24 Β· Nano 27 Β·   β”‚
|  Mini 30 Β· Compact 33 Β·      β”‚
|  Quality 36 Β· Precision 42 Β· β”‚
|  Fidelity 48 (%)             β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
             β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ 03_perplexity_test.py  +  90_attribution-    β”‚
β”‚ probe.py          PPL/KLD/RMS/top-p Β· probes β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
  • Measurement-backed: the Calibration Ledger L0–L8 documents eight laws plus the 2.3.0 refinements β€” L6 completed into a three-regime U-curve, L7 refined into an addition-only lever law β€” each rooted in knockout probes and stock-twin duels archived in attribution-results.csv.
  • Seven catalogue ratios on the arithmetic 24/27/30/33/36/42/48 ladder: Precision-42 and the duel-rehabilitated Fidelity-48 join the monotone core across all supported families; Pico-24 remains an opt-in service tier.
  • Cross-family proven: qwen35 hybrids (4B & 9B), vision/OCR models (OvisOCR2), dense architectures (SmolLM3 3B tied, llama 1B), and the mixer-dominant LFM2.5-2.6B β€” including a mapped sub-Nano validity domain.
  • Module-complete: vision towers and standalone speculative drafts (DSpark) quantize through dedicated imatrix-free engines β€” the trunk grid never touches a module it cannot rank, and drafts score on acceptance rate, not corpus proxies.
  • Honest failures: floors never lie β€” impossible targets warn instead of silently degrading safety classes; pipeline gates compare RATIOS against a measured reference, never absolute thresholds that would condemn a healthy build for merely being off-domain (Charter Β§5).
  • Stock-defended: every tier is ranked against its size-matched stock twin (L6): allocation owns the sub-wall band, uniform+imatrix the mid-band, and near-lossless archival defers to stock Q8_0's measured dual crown (fidelity + throughput).

πŸ“Š Tiers & Selection Policy

Tier Plain ratio Role
Pico 24% Service tier, opt-in ASHQ1_INCLUDE_PICO=1 β€” tight-VRAM serving; validity β‰₯4B dispersed hybrids, β‰₯3B dense tied (see Charter Β§7)
Nano 27% Edge cases / maximum compression (blocked by default on int4 lineage)
Mini 30% Minimum for β‰₯9B serving
Compact 33% Minimum for 3–4B; balanced deployment
Quality 36% Minimum for ~1B; near-lossless general deployment (flat builder on routed families)
Precision 42% Flat Q6_K + embd lever β€” default quasi-lossless top tier across all families (KLD 0.0021–0.0132)
Fidelity 48% Archival flat Q6_K + embd lever β€” disabled by default, opt-in via ASHQ1_INCLUDE_FIDELITY=1

Minimum-tier policy (measured, not folklore β€” see Charter Β§7). Smaller models hit the constructibility wall sooner: a 1B below 33% stops differentiating adjacent tiers, while a 9B tolerates 30% comfortably and rides Pico at upper-usable. Validity is architectural, not merely scalar: dispersed-projector hybrids absorb the Pico step, mixer-dominant ones (LFM2.5 family) bottom out at Mini (law L8). AutoRound-int4 lineage: practical ceiling is Compact; Nano AND Pico are excluded unless ASHQ1_INCLUDE_NANO=1 (shared double-quantization guard).


πŸ“ˆ Release Benchmarks (wiki.test.raw, symmetric FA-auto reference)

Ornith-1.5-9B (hybrid, plain lineage)

Tier Size PPL KLD RMS Ξ”p top-p
Fidelity-48pc 9464 MiB 9.5239 0.0081 2.43% 97.6%
Precision-42pc 8414 MiB 9.4347 0.0132 3.04% 96.6%
Quality-36pc 6330 MiB 9.3692 0.0366 5.00% 93.3%
Compact-33pc πŸ₯ˆ Second Choice 5803 MiB 9.6043 0.0517 5.91% 91.6%
Mini-30pc ⭐ Recommended 5385 MiB 9.8564 0.0649 6.67% 90.3%
Nano-27pc 4750 MiB 10.1061 0.0907 7.90% 87.8%
Pico-24pc 4389 MiB 10.1078 0.1309 9.63% 85.0%

Qwen3.8-4B-Distill (hybrid, plain lineage)

Tier Size PPL KLD RMS Ξ”p top-p
Fidelity-48pc 3866 MiB 9.0744 0.0014 1.01% 98.1%
Precision-42pc 3509 MiB 9.0818 0.0021 1.20% 97.6%
Quality-36pc 2902 MiB 9.1124 0.0095 2.56% 95.1%
Compact-33pc πŸ₯ˆ Second Choice 2661 MiB 9.1963 0.0160 3.31% 93.7%
Mini-30pc ⭐ Recommended 2420 MiB 9.2973 0.0258 4.26% 92.1%
Nano-27pc 2245 MiB 9.5957 0.0546 6.61% 88.7%
Pico-24pc 2168 MiB 9.6579 0.0636 6.95% 87.9%

TwIL-LM3 (dense SmolLM3-arch, tied token_embd, 128k vocab, L7-fixed)

Tier Size PPL KLD RMS Ξ”p top-p
Fidelity-48pc 2826 MiB 9.4677 0.0024 1.26% 97.0%
Precision-42pc 2474 MiB 9.4744 0.0036 1.53% 96.4%
Quality-36pc ⭐ Recommended 2121 MiB 9.5785 0.0163 3.35% 92.4%
Compact-33pc πŸ₯ˆ Second Choice 1945 MiB 9.6292 0.0235 3.94% 91.2%
Mini-30pc 1769 MiB 9.6890 0.0304 4.42% 90.2%
Nano-27pc 1593 MiB 9.9084 0.0499 5.75% 88.1%
Pico-24pc 1445 MiB 10.3800 0.0959 7.95% 84.3%

LFM2.5-2.6B (shortconv-mixer hybrid, tied readout β€” 2.3.0 ladder)

Tier Size PPL KLD RMS Ξ”p top-p
Fidelity-48pc 2481 MiB 56.0022 0.0072 1.87% 95.8%
Precision-42pc 2179 MiB 55.2953 0.0108 2.32% 94.9%
Quality-36pc ⭐ Recommended 1850 MiB 55.5619 0.0354 4.12% 91.0%
Compact-33pc πŸ₯ˆ Second Choice 1708 MiB 54.6921† 0.0762 6.05% 87.0%
Mini-30pc 1554 MiB 56.5119 0.1199 7.53% 83.9%
Nano-27pc 1399 MiB 61.1460 0.2131 10.06% 78.9%
Pico-24pc 1276 MiB 56.7726† 0.3707 13.08% 72.7%

Stock twins (same protocol): Q8_0 2742 MiB Β· 0.0030 Β· 97.3% β€” Q6_K 2119 Β· 0.0143 Β· 94.1% β€” Q5_K_M 1850 Β· 0.0467 Β· 89.8% β€” Q4_K_M 1597 Β· 0.1552 Β· 81.7%. Quality-36 rides the flat builder (mid-band law); Precision and Fidelity deliver consistent quasi-lossless fidelity across all measured architectures.

ReaderLM-v2 (dense HTML/markdown extractor, 1.5B, 28 layers)

Tier Size PPL KLD RMS Ξ”p top-p
Fidelity-48pc 1422 MiB 17.0871 0.0050 1.65% 96.3%
Precision-42pc πŸ₯ˆ Second Choice 1268 MiB 17.0608 0.0076 2.03% 95.5%
Quality-36pc ⭐ Recommended 1068 MiB 17.0610 0.0291 4.01% 91.2%
Compact-33pc 980 MiB 17.5504 0.1211 7.74% 87.1%
Mini-30pc 891 MiB 17.2444 0.1145 7.74% 85.4%
Nano-27pc 803 MiB 17.3237 0.1513 9.04% 81.4%
Pico-24pc 763 MiB 18.0892 0.2087 10.87% 77.9%

MiniCPM5-1B (dense llama-arch)

Tier Size PPL KLD RMS Ξ”p top-p
Fidelity-48pc 997 MiB 26.3758 0.0056 1.63% 95.3%
Precision-42pc πŸ₯ˆ Second Choice 943 MiB 26.4353 0.0072 1.84% 94.5%
Quality-36pc ⭐ Recommended 749 MiB 27.2456 0.0555 5.16% 86.3%
Compact-33pc 687 MiB 28.1717 0.0997 6.69% 81.9%
Mini-30pc 663 MiB 28.5485 0.1153 7.16% 80.4%
Nano-27pc 563 MiB 28.8736 0.1382 7.98% 78.4%
Pico-24pc 503 MiB 37.5561 0.3854 14.27% 66.3%

PaddleOCR-VL-1.6 (vision-language OCR decoder, ERNIE-4.5-0.3B)

Tier Size PPL KLD RMS Ξ”p top-p
Fidelity-48pc 537 MiB 365.8789 0.0366 2.99% 91.1%
Precision-42pc ⭐ Recommended 484 MiB 362.9912 0.0337 2.57% 91.5%
Quality-36pc 430 MiB 400.7598 0.1747 5.72% 80.7%
Compact-33pc 404 MiB 351.6550 0.2320 6.82% 77.3%
Mini-30pc 378 MiB 333.6259 0.2712 7.33% 74.7%
Nano-27pc 351 MiB 385.6355 0.2612 7.73% 72.5%
Pico-24pc 324 MiB 509.1721 0.5632 11.01% 63.5%

OvisOCR2 (vision-aligned text backbone, ~0.3B)

Tier Size PPL KLD RMS Ξ”p top-p
Fidelity-48pc 704 MiB 31.0673 0.0021 0.98% 97.4%
Precision-42pc πŸ₯ˆ Second Choice 668 MiB 31.0809 0.0026 1.06% 97.1%
Quality-36pc ⭐ Recommended 531 MiB 32.0240 0.0155 2.68% 93.0%
Compact-33pc 487 MiB 32.7223 0.0306 3.90% 90.3%
Mini-30pc 474 MiB 32.8232 0.0352 4.17% 89.8%
Nano-27pc 457 MiB 33.1136 0.0648 5.95% 86.7%
Pico-24pc 442 MiB 34.7822 0.0775 6.33% 85.2%

πŸ“‰ Sub-Nano Compendium (full-corpus protocol, v2.2.0)

Orderings replicate the tables above; absolutes shift a few percent with full 559–580-chunk passes (never mix protocols in one column β€” Charter Β§5).

Family Arch Nano→Pico KLD step Pico verdict
Qwen3.8-4B GDN hybrid +16.5% βœ… 2168 MiB Β· 0.0636 Β· top-p 87.9% (effective ~27%†)
OvisOCR2 vision backbone +19.6% βœ… 442 MiB Β· 0.0775 Β· top-p 85.2%
Ornith-1.5-9B GDN hybrid +44% βœ… 4389 MiB Β· 0.1309 Β· top-p 85.0%
LFM2.5-2.6B shortconv mixer +73% ❌ 0.3707 Β· 72.7% β€” floor is Mini (L8)
TwIL-LM3 dense tied +92% ⚠️ 1445 MiB Β· 0.0959 Β· 84.3% β€” priced, allowed
MiniCPM5-1B dense untied +179% ❌ 0.3854 Β· 66.3% β€” collapse; ends at Nano

† no-mtp trunk + carved heads raise effective coverage to ~27%.

Highlights of the campaign behind these rows: the cheapest marginal ever measured here is the readout rung embd Q6_K→Q8_0 (Pico+, TwIL 0.0959→0.0914 / 85.0% for +60 MiB); a size-matched engineered stock twin TIES Pico-24 (ΔKLD 0.0018 < noise) — allocation is inert at constant codebook support; and on LFM, raising only the 8 attention layers recovers 0.371→0.213 while shielding the whole shortconv stack recovers just 0.030 — law L8: cut locality beats cut depth. Full per-tier tables and KO records: CHARTER.md §6b + attribution-results.csv.


πŸ› οΈ Suite Components

Script Purpose
00_SAFETENSORS-to-AutoRound-BF16-GGUF.py Optional int4 reconditioning + provenance sidecars + standalone draft-module builds (shard-index audit, parent-tokenizer borrowing).
00b_BF16-GGUF-MTP-extract.py Split the MTP draft head off the trunk.
01_create-calibration-dataset-and-imatrix.py Multi-domain corpus + lineage-aware imatrix; GPU autotune behind a relative split-integrity gate (--cpu-only bypasses offload on any architecture).
01b_BF16-GGUF-modules-fusion.py Reattach quantized modules (MTP/mmproj).
02_BF16-GGUF-to-ASHQ1.py Batch tier orchestration (auto-runs mmproj + draft modules; --module dspark builds one standalone).
03_perplexity_test.py KL reference base + PPL/KLD/RMS/top-p/per-chunk sweep.
ASHQ1.py Core optimizer (single target CLI + tier runner).
ASHQ1-mmproj.py Vision-tower tiers, deliberately imatrix-free.
ASHQ1-dspark.py Speculative-draft tiers, imatrix-free, acceptance-rate domain.
90_attribution-probe.py Knockout attribution harness (methodology instrument).

⚑ Quick Start

pip install gguf numpy huggingface_hub safetensors         # + auto-round/torchvision for step 00
python 00_SAFETENSORS-to-AutoRound-BF16-GGUF.py            # create BF16 gguf files from safetensors project, create AutoRound version with `--autoround`
python 01_create-calibration-dataset-and-imatrix.py        # imatrix.gguf (plain lineage)
python 02_BF16-GGUF-to-ASHQ1.py                            # batch tiers + modules (ASHQ1_INCLUDE_PICO=1 or ASHQ1_INCLUDE_FIDELITY=1 for opt-in tiers)
python 02_BF16-GGUF-to-ASHQ1.py --module dspark            # standalone draft module only
python 03_perplexity_test.py                               # evaluate
python 90_attribution-probe.py --list                      # optional: inspect probes

Windows/NTFS: compact /c /exe:xpress8k <file> shrinks KLD logits ~85% and BF16 GGUFs ~18% without touching results.


πŸ“œ Citation & Credits

  • ASHQ1 (Autonomous Selective Hybrid Quantization) by wepiqx β€” priority-queue knapsack formulation, tied-group activation hashing, MSE scheduling.
  • Empero AI (Qwen3.8-27B-Ridge) β€” GDN state preservation (ssm_alpha/ssm_beta @ Q8_0) and native MTP draft heads.
  • Intel AutoRound β€” sign-gradient low-bit optimization with Hessian compensation.
  • llama.cpp by Georgi Gerganov & ggml contributors β€” GGUF/GGML runtime and tools.
  • Calibration recipes inspired by Bartowski; multi-imatrix max-combination per community practice (cHunter789 KV-cache recipe referenced by the orchestrator).

License: apache-2.0

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support