Upload folder using huggingface_hub
Browse files- Qwen3.8-27B-MTP-MoQ-vs-Unsloth-KLD.csv +18 -0
- Qwen3.8-27B-MTP-MoQ-vs-Unsloth-KLD.html +0 -0
- README.md +117 -0
- mean_kld.png +0 -0
- p999_kld.png +0 -0
- ppl.png +0 -0
Qwen3.8-27B-MTP-MoQ-vs-Unsloth-KLD.csv
ADDED
|
@@ -0,0 +1,18 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
series,label,file_bytes,size_gb,size_gib,file_bpw,ppl,ppl_se,base_ppl,base_ppl_se,mean_kld,mean_kld_se,max_kld,p999_kld,p99_kld,p95_kld,median_kld,rms_delta_p_pct,same_top_pct,elapsed_sec,tensor_count,parameter_count,tensor_types
|
| 2 |
+
Our MoQ,MoQ-3.2,10571830880,10.57183088,9.845784754,3.09562543,7.408463,0.047353,6.950282,0.044932,0.111637,0.000682,14.055522,3.348125,1.084846,0.376798,0.053885,9.634,85.679,139.8635048,866,27320697856,"IQ3_XXS 51.8% (192), IQ2_S 17.0% (52), IQ3_S 13.6% (107), Q2_K 9.3% (25), IQ4_XS 8.3% (34)"
|
| 3 |
+
Our MoQ,MoQ-3.6,11891684960,11.89168496,11.07499465,3.482102843,7.22108,0.045761,6.950282,0.044932,0.072244,0.000504,14.646406,2.351531,0.71737,0.238477,0.034486,7.682,88.722,136.9859063,866,27320697856,"IQ3_S 43.5% (195), IQ3_XXS 36.8% (116), IQ4_XS 7.8% (65), Q5_K 7.0% (27), Q2_K 4.7% (1), Q6_K 0.1% (6)"
|
| 4 |
+
Our MoQ,MoQ-3.8,12715042400,12.7150424,11.84180602,3.72319696,7.122706,0.045942,6.950282,0.044932,0.04819,0.000382,14.569729,1.624749,0.437946,0.161417,0.022377,6.077,90.109,134.4476103,866,27320697856,"IQ3_S 60.9% (227), IQ4_XS 29.2% (142), Q3_K 4.7% (1), IQ3_XXS 3.7% (32), Q6_K 0.7% (4), Q5_K 0.7% (2), Q8_0 0.2% (2), BF16 0.1% (96)"
|
| 5 |
+
Our MoQ,MoQ-4.1,14136825440,14.13682544,13.16594467,4.139521037,7.068801,0.045795,6.950282,0.044932,0.03146,0.000297,16.211906,1.056327,0.285069,0.102749,0.014744,4.915,91.872,135.2048469,866,27320697856,"IQ4_XS 59.9% (254), IQ3_S 26.1% (80), Q4_K 6.5% (20), IQ4_NL 3.7% (16), Q5_K 2.8% (34), Q6_K 0.7% (4), Q8_0 0.2% (2), BF16 0.1% (96)"
|
| 6 |
+
Our MoQ,MoQ-4.3,14889711200,14.8897112,13.86712417,4.359979757,7.045628,0.045636,6.950282,0.044932,0.0245,0.000233,12.517167,0.805452,0.216612,0.077038,0.012293,4.332,92.648,135.7567087,866,27320697856,"IQ4_XS 67.6% (258), Q4_K 30.4% (128), Q6_K 0.7% (4), Q5_K 0.7% (2), IQ3_XXS 0.3% (16), Q8_0 0.2% (2), Q2_K 0.1% (96)"
|
| 7 |
+
Our MoQ,MoQ-4.6,15129368160,15.12936816,14.09032211,4.430155698,7.014436,0.045464,6.950282,0.044932,0.020171,0.000218,12.15505,0.758995,0.200817,0.065882,0.008432,3.936,93.809,134.8057683,866,27320697856,"IQ4_XS 86.6% (352), Q5_K 11.8% (20), Q4_K 0.7% (128), Q6_K 0.7% (4), Q8_0 0.2% (2)"
|
| 8 |
+
Our MoQ,MoQ-4.8,16037615200,16.0376152,14.93619308,4.696107042,7.008832,0.045411,6.950282,0.044932,0.017282,0.00021,12.204391,0.650152,0.169743,0.056426,0.007149,3.584,94.339,137.5041961,866,27320697856,"IQ4_XS 65.7% (288), Q5_K 32.7% (84), Q6_K 0.7% (4), Q4_K 0.7% (80), Q8_0 0.2% (2)"
|
| 9 |
+
Our MoQ,MoQ-4.9,16411252320,16.41125232,15.28416976,4.805514824,7.021034,0.04555,6.950282,0.044932,0.016342,0.0002,17.640537,0.651529,0.162814,0.053539,0.006743,3.53,94.423,141.1051134,866,27320697856,"IQ4_XS 56.8% (256), Q5_K 41.9% (132), Q6_K 0.7% (4), Q4_K 0.4% (112), Q8_0 0.2% (2)"
|
| 10 |
+
Our MoQ,MoQ-5.1,17339078240,17.33907824,16.14827499,5.077199223,7.020549,0.045583,6.950282,0.044932,0.012857,0.000154,8.997525,0.487299,0.123491,0.041153,0.005481,3.134,94.968,143.431954,866,27320697856,"Q5_K 62.8% (196), IQ4_XS 35.6% (176), Q6_K 0.7% (4), Q4_K 0.6% (32), Q8_0 0.2% (2), BF16 0.1% (96)"
|
| 11 |
+
Unsloth,UD-IQ2_XXS,9010048064,9.010048064,8.39126116,2.638306858,7.957195,0.052513,6.950282,0.044932,0.163689,0.000991,15.348619,4.714475,1.739405,0.555111,0.077528,11.822,82.389,152.8016501,866,27320697856,"IQ2_S 64.5% (208), IQ3_XXS 9.5% (80), IQ2_XS 9.2% (48), IQ2_XXS 5.5% (48), Q3_K 4.8% (2), Q2_K 4.7% (1), IQ4_XS 1.7% (23), IQ1_M 0.1% (96)"
|
| 12 |
+
Unsloth,UD-IQ2_M,10319907904,10.3199079,9.611163199,3.021857775,7.507406,0.049109,6.950282,0.044932,0.103564,0.000712,17.452929,3.590751,1.076759,0.33183,0.049609,9.332,85.496,143.0807505,866,27320697856,"IQ3_XXS 56.5% (224), IQ2_S 20.9% (64), IQ3_S 11.7% (99), Q2_K 4.7% (1), Q3_K 4.7% (1), IQ4_XS 1.5% (21), IQ1_M 0.1% (96)"
|
| 13 |
+
Unsloth,UD-Q2_K_XL,10676423744,10.67642374,9.943194449,3.126252133,7.403969,0.048319,6.950282,0.044932,0.091579,0.000642,15.962362,3.107871,0.925859,0.28854,0.045107,8.708,86.16,139.3403591,866,27320697856,"IQ3_XXS 77.4% (288), IQ3_S 11.7% (99), Q2_K 4.7% (1), Q3_K 4.7% (1), IQ4_XS 1.5% (21), IQ1_M 0.1% (96)"
|
| 14 |
+
Unsloth,UD-IQ3_XXS,11913559104,11.9135591,11.09536654,3.488507993,7.24624,0.046971,6.950282,0.044932,0.054517,0.000471,14.718587,2.25904,0.627471,0.185082,0.020561,6.766,90.359,140.511719,866,27320697856,"IQ3_S 53.6% (195), IQ3_XXS 26.4% (112), IQ4_XS 10.7% (149), Q2_K 4.7% (1), Q5_K 4.7% (1)"
|
| 15 |
+
Unsloth,UD-Q3_K_XL,13441059904,13.4410599,12.51796252,3.935788163,7.111811,0.045892,6.950282,0.044932,0.031355,0.000313,15.984297,1.24832,0.353498,0.103949,0.012276,5.129,92.396,139.6360067,866,27320697856,"IQ4_XS 47.9% (357), IQ3_S 42.4% (130), Q5_K 5.0% (18), Q3_K 4.7% (1)"
|
| 16 |
+
Unsloth,UD-Q4_K_XL,17923394624,17.92339462,16.69246203,5.248297747,6.976017,0.045138,6.950282,0.044932,0.00864,0.000167,16.768026,0.42288,0.090719,0.027377,0.003062,2.614,96.073,145.8000825,866,27320697856,"Q5_K 68.9% (325), IQ4_XS 21.2% (65), Q6_K 5.2% (19), Q4_K 4.7% (97)"
|
| 17 |
+
Unsloth,UD-Q5_K_XL,20218178624,20.21817862,18.82964617,5.920252471,6.9673,0.045101,6.950282,0.044932,0.004526,0.000169,13.261263,0.200576,0.040594,0.013022,0.001629,1.849,97.153,155.024853,866,27320697856,"Q5_K 61.8% (275), Q6_K 37.6% (165), Q8_0 0.5% (18)"
|
| 18 |
+
Unsloth,UD-Q6_K_XL,25924152384,25.92415238,24.14374834,7.591065944,6.953912,0.044986,6.950282,0.044932,0.001331,5.10E-05,3.6362,0.0628,0.011536,0.003536,0.000445,1.107,98.493,156.2973565,866,27320697856,"Q8_0 52.8% (279), Q6_K 47.1% (131), Q5_K 0.1% (96)"
|
Qwen3.8-27B-MTP-MoQ-vs-Unsloth-KLD.html
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|
README.md
CHANGED
|
@@ -1,3 +1,120 @@
|
|
| 1 |
---
|
| 2 |
license: apache-2.0
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 3 |
---
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
---
|
| 2 |
license: apache-2.0
|
| 3 |
+
base_model:
|
| 4 |
+
- Qwen/Qwen3.8-27B
|
| 5 |
+
base_model_relation: quantized
|
| 6 |
+
tags:
|
| 7 |
+
- gguf
|
| 8 |
+
- llama.cpp
|
| 9 |
+
- moq
|
| 10 |
+
- mtp
|
| 11 |
+
- quantized
|
| 12 |
---
|
| 13 |
+
|
| 14 |
+
# Qwen3.8-27B-MoQ-GGUF
|
| 15 |
+
|
| 16 |
+
This repository contains imatrix-aware, layer-aware Mixture-of-Quantization (MoQ) GGUF files for the original [Qwen/Qwen3.8-27B](https://huggingface.co/Qwen/Qwen3.8-27B) model. These files preserve the model's MTP tensors and are quantizations of the base Qwen model. In terms of overall quality, MoQ is comparable to Unsloth's Dynamic Quantization 3.0; however, our quantization series offers a more granular range of options (4.0 bpw – 4.75 bpw) than the Unsloth team's offerings, while maintaining a slight performance edge over Dynamic Quantization 3.0 within this range.
|
| 17 |
+
|
| 18 |
+
The complete series was evaluated locally alongside the Unsloth Dynamic GGUF series under identical conditions. Lower is better in all three charts.
|
| 19 |
+
|
| 20 |
+

|
| 21 |
+
|
| 22 |
+

|
| 23 |
+
|
| 24 |
+

|
| 25 |
+
|
| 26 |
+
The [interactive comparison report](Qwen3.8-27B-MTP-MoQ-vs-Unsloth-KLD.html) supports pan, zoom, view reset, and per-series visibility controls. The [companion CSV](Qwen3.8-27B-MTP-MoQ-vs-Unsloth-KLD.csv) contains all measured values and tensor-composition summaries.
|
| 27 |
+
|
| 28 |
+
## Comparison Summary
|
| 29 |
+
|
| 30 |
+
All 17 measured GGUF files contain 866 tensors and 27,320,697,856 parameters. They were evaluated against the same BF16 reference logits on WikiText-2 with a fixed context length of 512.
|
| 31 |
+
|
| 32 |
+
At the exact size of each of the 9 MoQ files, linearly interpolating the Unsloth curve favors this MoQ series on:
|
| 33 |
+
|
| 34 |
+
- PPL: 8 of 9 points
|
| 35 |
+
- p999 KLD: 7 of 9 points
|
| 36 |
+
- Mean KLD: 2 of 9 points
|
| 37 |
+
|
| 38 |
+
The results show the intended tradeoff clearly. The layer-aware MoQ recipes are especially effective on PPL and tail divergence, while the Unsloth Dynamic recipes remain very strong on Mean KLD, particularly at low and middle file sizes. The full table is included below so that users can choose based on the metric that matters for their workload.
|
| 39 |
+
|
| 40 |
+
## Full Quality Results
|
| 41 |
+
|
| 42 |
+
`Actual BPW` is computed from the complete GGUF file size, including metadata and alignment. GB is decimal. PPL and all KLD values are lower-is-better; same top-p is higher-is-better.
|
| 43 |
+
|
| 44 |
+
| Series | Recipe | Actual BPW | GB | PPL | Mean KLD | p999 KLD | p99 KLD | RMS delta-p | Same top-p |
|
| 45 |
+
|---|---|---:|---:|---:|---:|---:|---:|---:|---:|
|
| 46 |
+
| Jianqiao1 MoQ | 3.2 | 3.0956 | 10.572 | 7.408463 | 0.111637 | 3.348125 | 1.084846 | 9.634% | 85.679% |
|
| 47 |
+
| Jianqiao1 MoQ | 3.6 | 3.4821 | 11.892 | 7.221080 | 0.072244 | 2.351531 | 0.717370 | 7.682% | 88.722% |
|
| 48 |
+
| Jianqiao1 MoQ | 3.8 | 3.7232 | 12.715 | 7.122706 | 0.048190 | 1.624749 | 0.437946 | 6.077% | 90.109% |
|
| 49 |
+
| Jianqiao1 MoQ | 4.1 | 4.1395 | 14.137 | 7.068801 | 0.031460 | 1.056327 | 0.285069 | 4.915% | 91.872% |
|
| 50 |
+
| Jianqiao1 MoQ | 4.3 | 4.3600 | 14.890 | 7.045628 | 0.024500 | 0.805452 | 0.216612 | 4.332% | 92.648% |
|
| 51 |
+
| Jianqiao1 MoQ | 4.6 | 4.4302 | 15.129 | 7.014436 | 0.020171 | 0.758995 | 0.200817 | 3.936% | 93.809% |
|
| 52 |
+
| Jianqiao1 MoQ | 4.8 | 4.6961 | 16.038 | 7.008832 | 0.017282 | 0.650152 | 0.169743 | 3.584% | 94.339% |
|
| 53 |
+
| Jianqiao1 MoQ | 4.9 | 4.8055 | 16.411 | 7.021034 | 0.016342 | 0.651529 | 0.162814 | 3.530% | 94.423% |
|
| 54 |
+
| Jianqiao1 MoQ | 5.1 | 5.0772 | 17.339 | 7.020549 | 0.012857 | 0.487299 | 0.123491 | 3.134% | 94.968% |
|
| 55 |
+
| Unsloth | UD-IQ2_XXS | 2.6383 | 9.010 | 7.957195 | 0.163689 | 4.714475 | 1.739405 | 11.822% | 82.389% |
|
| 56 |
+
| Unsloth | UD-IQ2_M | 3.0219 | 10.320 | 7.507406 | 0.103564 | 3.590751 | 1.076759 | 9.332% | 85.496% |
|
| 57 |
+
| Unsloth | UD-Q2_K_XL | 3.1263 | 10.676 | 7.403969 | 0.091579 | 3.107871 | 0.925859 | 8.708% | 86.160% |
|
| 58 |
+
| Unsloth | UD-IQ3_XXS | 3.4885 | 11.914 | 7.246240 | 0.054517 | 2.259040 | 0.627471 | 6.766% | 90.359% |
|
| 59 |
+
| Unsloth | UD-Q3_K_XL | 3.9358 | 13.441 | 7.111811 | 0.031355 | 1.248320 | 0.353498 | 5.129% | 92.396% |
|
| 60 |
+
| Unsloth | UD-Q4_K_XL | 5.2483 | 17.923 | 6.976017 | 0.008640 | 0.422880 | 0.090719 | 2.614% | 96.073% |
|
| 61 |
+
| Unsloth | UD-Q5_K_XL | 5.9203 | 20.218 | 6.967300 | 0.004526 | 0.200576 | 0.040594 | 1.849% | 97.153% |
|
| 62 |
+
| Unsloth | UD-Q6_K_XL | 7.5911 | 25.924 | 6.953912 | 0.001331 | 0.062800 | 0.011536 | 1.107% | 98.493% |
|
| 63 |
+
|
| 64 |
+
The Unsloth rows are local comparison measurements only. Unsloth GGUF files are not redistributed in this repository.
|
| 65 |
+
|
| 66 |
+
## Available Models
|
| 67 |
+
|
| 68 |
+
The recipe number is a series label. `Actual BPW` below is calculated from the complete file and is the value to use for exact size comparisons.
|
| 69 |
+
|
| 70 |
+
| File | Actual BPW | Size GB | Size GiB |
|
| 71 |
+
|---|---:|---:|---:|
|
| 72 |
+
| `Qwen3.8-27B-MTP-MoQ-3.2.gguf` | 3.0956 | 10.572 | 9.846 |
|
| 73 |
+
| `Qwen3.8-27B-MTP-MoQ-3.6.gguf` | 3.4821 | 11.892 | 11.075 |
|
| 74 |
+
| `Qwen3.8-27B-MTP-MoQ-3.8.gguf` | 3.7232 | 12.715 | 11.842 |
|
| 75 |
+
| `Qwen3.8-27B-MTP-MoQ-4.1.gguf` | 4.1395 | 14.137 | 13.166 |
|
| 76 |
+
| `Qwen3.8-27B-MTP-MoQ-4.3.gguf` | 4.3600 | 14.890 | 13.867 |
|
| 77 |
+
| `Qwen3.8-27B-MTP-MoQ-4.6.gguf` | 4.4302 | 15.129 | 14.090 |
|
| 78 |
+
| `Qwen3.8-27B-MTP-MoQ-4.8.gguf` | 4.6961 | 16.038 | 14.936 |
|
| 79 |
+
| `Qwen3.8-27B-MTP-MoQ-4.9.gguf` | 4.8055 | 16.411 | 15.284 |
|
| 80 |
+
| `Qwen3.8-27B-MTP-MoQ-5.1.gguf` | 5.0772 | 17.339 | 16.148 |
|
| 81 |
+
|
| 82 |
+
The repository also includes the locally generated calibration imatrix and a BF16 multimodal projector.
|
| 83 |
+
|
| 84 |
+
## Quantization Approach
|
| 85 |
+
|
| 86 |
+
This series combines the advantages of the MoQ (Mixed-precision Quantization) scheme at the layer level with tensor-specific strategies derived from Unsloth Dynamic 3.0's low-bit GGUF implementations. Since Qwen3.8 and Qwen3.6 share a matching tensor-level architecture, we were able to transfer and fine-tune layer-specific quantization strategies without being forced to apply a uniform quantization type across all tensor families.
|
| 87 |
+
|
| 88 |
+
The specific process is as follows:
|
| 89 |
+
|
| 90 |
+
1. Generate a calibration imatrix specifically for Qwen3.8 using a calibration corpus.
|
| 91 |
+
2. Transfer the layer-wise MoQ allocation scheme from Qwen3.6 to the corresponding tensors in Qwen3.8.
|
| 92 |
+
3. Compare each tensor family and layer against the original weights.
|
| 93 |
+
4. Retain MoQ layer-wise precision adjustments (whether increasing or decreasing precision) if they improve perplexity (PPL) or KLD tail performance.
|
| 94 |
+
5. While maintaining the layer-wise structure, optimize average KLD performance for the low-BPW (bits per weight) schemes of versions 3.2 and 3.6 by referencing Unsloth's low-bit quantization strategies.
|
| 95 |
+
|
| 96 |
+
This is not a single, global quantization preset; different precision settings can be applied to specific tensor families and layers based on their measured or inferred sensitivity.
|
| 97 |
+
|
| 98 |
+
## MTP Precision
|
| 99 |
+
|
| 100 |
+
Compared to the MTP layer in Qwen 3.6, the MTP layer in Qwen 3.8 employs lower-bit quantization; this is because our tests revealed that the prediction performance of Qwen 3.8's MTP layer—even at Q8 precision—falls short of that of Qwen 3.6, rendering the maintenance of high precision for the MTP layer less meaningful.
|
| 101 |
+
|
| 102 |
+
## Evaluation Conditions
|
| 103 |
+
|
| 104 |
+
The comparison used llama.cpp build 10276, commit `6ea215d17`, with CUDA on an NVIDIA GeForce RTX 5090. WikiText-2 `wiki.test.raw` was evaluated against BF16 logits from the original Qwen3.8-27B GGUF. Context length was fixed at 512; logical batch was 2048, ubatch was 8192, GPU layers were selected automatically with fit enabled, op offload and flash attention were enabled, and 16 CPU threads were used. GPU evaluations were run strictly one at a time.
|
| 105 |
+
|
| 106 |
+
The BF16 reference PPL reported by the evaluator was 6.950282.
|
| 107 |
+
|
| 108 |
+
## Usage
|
| 109 |
+
|
| 110 |
+
Use a recent llama.cpp build with Qwen3.8 and MTP support. The files can be used for ordinary generation or MTP speculative decoding. Choose a file according to available memory and the quality curves above; the nominal recipe label is useful for navigating the series, while the actual file BPW and GB columns provide precise memory-planning values.
|
| 111 |
+
|
| 112 |
+
## License and Acknowledgements
|
| 113 |
+
|
| 114 |
+
Released under the Apache License 2.0, following the base model license metadata.
|
| 115 |
+
|
| 116 |
+
Thanks to:
|
| 117 |
+
|
| 118 |
+
- the Qwen team for Qwen3.8-27B and its MTP architecture;
|
| 119 |
+
- the llama.cpp project and contributors for GGUF quantization, MTP support, and evaluation tooling;
|
| 120 |
+
- the Unsloth team for the Dynamic GGUF series.
|
mean_kld.png
ADDED
|
p999_kld.png
ADDED
|
ppl.png
ADDED
|