rwkv7-g1h-2.9b i1 GGUF

Model-specific imatrix quants generated from rwkv7-g1h-2.9b-20260710-ctx10240-BF16.gguf. The BF16 master remains in the linked static-GGUF repository and is not duplicated here.

Calibration

  • Dataset: lemon07r/pile-calibration-v5
  • Dataset revision: ea863bb930b9959dd78095c165aa1376d14e698b
  • llama.cpp revision: c92e806d1c81091c9035edce99c35374da1b465e
  • Context per chunk: 1024 tokens
  • Chunks: 512
  • Approximate evaluated tokens: 524288
  • Output-weight collection: disabled, following llama.cpp's default recommendation
  • Raw JSONL SHA-256: 54d1f2bd6a80cc75a72e7f8a62e23438e23f06287aab3855c5a7826127fae7be
  • Deterministically curated corpus SHA-256: 8c7de8adad55f0be5e00f394b7c7efca7c6d13fdd4a2f476bc08c0491fe98027

The source dataset is diverse and duplicate-free, but contains a few book-length outliers. Before calibration, corrupt/severely repetitive records are removed, long records are capped using four separated excerpts, records are deterministically shuffled, and rare scripts are lightly interleaved into the early calibration window.

Files

Quant PPL BF16 retained
i1-Q6_K 6.8764 99.62%
i1-Q5_1 6.952 98.54%
i1-Q5_K_M 6.9247 98.93%
i1-Q5_K_S 6.9551 98.49%
i1-Q5_0 7.179 95.42%
i1-Q4_1 23.2667 29.44%
i1-Q4_K_M 6.9945 97.94%
i1-Q4_K_S 16.5402 41.42%
i1-Q4_0 231.5514 2.96%
i1-IQ4_NL 8.8656 77.27%
i1-IQ4_XS 9.0385 75.79%
i1-Q3_K_L 7.2029 95.10%
i1-Q3_K_M 7.231 94.74%
i1-IQ3_S 554.0051 1.24%
i1-Q3_K_S 8229.9054 0.0832%
i1-Q2_K 12573.2758 0.0545%
i1-IQ3_XXS 475.9699 1.44%
i1-IQ2_M 1073.0608 0.638%
i1-IQ2_S 1242.044 0.552%
i1-IQ2_XS 42504.01 0.0161%
i1-IQ2_XXS 49975.9054 0.0137%
i1-IQ1_M 22106.1854 0.0310%
i1-IQ1_S 23790.2745 0.0288%

RWKV-aware mixed quantizations

The Q3_K_M, Q3_K_L, Q4_K_M, and Q5_K_M files use custom RWKV-aware recipes with explicit tensor assignments. Higher precision is used for the token embeddings and selected value, time-mix output, and channel-mix tensors where it is expected to preserve the most quality.

Earlier automated files with these names were removed after verification showed that llama.cpp's generic mixed recipes did not recognize RWKV's time_mix_* and channel_mix_* tensor roles. Because of that, the automated M and L variants had collapsed to the same effective layouts as the retained S variants.

These replacement files have genuinely different tensor layouts, providing additional size and quality choices between the existing S variants and the larger quantizations.

The files in this i1 repository use the same RWKV-aware tensor layouts together with this model's existing importance matrix.

Downloads last month
539
GGUF
Model size
3B params
Architecture
rwkv7
Hardware compatibility
Log In to add your hardware

1-bit

2-bit

3-bit

4-bit

5-bit

6-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for RemySkye/rwkv7-g1h-2.9b-i1-GGUF

Base model

BlinkDL/rwkv7-g1
Quantized
(28)
this model

Collection including RemySkye/rwkv7-g1h-2.9b-i1-GGUF