PY-AI-Dev commited on
Commit
4448b82
·
verified ·
1 Parent(s): 9a4d323

Add iMatrix GGUF quantizations for Grug-12B

Browse files
.gitattributes CHANGED
@@ -33,3 +33,10 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
 
 
 
 
 
 
 
 
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
36
+ Grug-12B-IQ2_M.gguf filter=lfs diff=lfs merge=lfs -text
37
+ Grug-12B-IQ3_M.gguf filter=lfs diff=lfs merge=lfs -text
38
+ Grug-12B-IQ4_XS.gguf filter=lfs diff=lfs merge=lfs -text
39
+ Grug-12B-Q4_K_M.gguf filter=lfs diff=lfs merge=lfs -text
40
+ Grug-12B-Q5_K_M.gguf filter=lfs diff=lfs merge=lfs -text
41
+ Grug-12B-Q6_K.gguf filter=lfs diff=lfs merge=lfs -text
42
+ Grug-12B-Q8_0.gguf filter=lfs diff=lfs merge=lfs -text
Grug-12B-IQ2_M.gguf ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:987bc908bfacabf52508c5d81d5399ed872fc9a265797642016360f7340e35f2
3
+ size 4371807360
Grug-12B-IQ3_M.gguf ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:991a536ddad31bd751ddfda70f0cc33283120618c8933b906b89d8de9bb2a666
3
+ size 5733993600
Grug-12B-IQ4_XS.gguf ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:b69771cd3a04a52c14ad50db9d26209f52f021a20f8f3a114f2580603bc428e9
3
+ size 6635256960
Grug-12B-Q4_K_M.gguf ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:7140956fc7f38b5b60977da99dba7468306ad988d4c24ea963ff6ee30df10f49
3
+ size 7381384320
Grug-12B-Q5_K_M.gguf ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:dfef2724918149063d13dfb33b2fe5ea24af26a528743b0be82f9e6e394858b9
3
+ size 8547269760
Grug-12B-Q6_K.gguf ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:eccfafcd94a0f30c9dc9b3f3f3b91bb0819a591689bdec7d617e1a4e99b1f899
3
+ size 9786023040
Grug-12B-Q8_0.gguf ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:f162f65d291c597e13018673100698b6346c7dc486bacfd29d7e7c2a004cf541
3
+ size 12669648000
README.md ADDED
@@ -0,0 +1,65 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: other
3
+ base_model: kai-os/Grug-12B
4
+ pipeline_tag: text-generation
5
+ tags:
6
+ - gguf
7
+ - local-llm
8
+ - llama.cpp
9
+ - lm-studio
10
+ - quantized
11
+ - imatrix
12
+ - sub-4-bit
13
+ - gemma4_unified
14
+ - gemma-4
15
+ ---
16
+
17
+ # Grug-12B — iMatrix GGUF
18
+
19
+ GGUF quantizations of [kai-os/Grug-12B](https://huggingface.co/kai-os/Grug-12B), published by [Liodon AI](https://huggingface.co/liodon-ai).
20
+
21
+ ## Quick Start
22
+
23
+ **llama.cpp**
24
+ ```bash
25
+ llama-cli -hf liodon-ai/Grug-12B-imatrix-GGUF:Q4_K_M
26
+ ```
27
+
28
+ **Ollama**
29
+ ```bash
30
+ ollama run hf.co/liodon-ai/Grug-12B-imatrix-GGUF:Q4_K_M
31
+ ```
32
+
33
+ **LM Studio / Jan** — search `liodon-ai/Grug-12B-imatrix-GGUF` and pick your quant.
34
+
35
+ ## Quants
36
+
37
+ | Quant | Size | VRAM est. | Notes |
38
+ |-------|------|-----------|-------|
39
+ | `IQ2_M` | 4.37 GB | ~5 GB | 2-bit, iMatrix — smallest usable |
40
+ | `IQ3_M` | 5.73 GB | ~7 GB | 3-bit, iMatrix — great quality/size tradeoff |
41
+ | `IQ4_XS` | 6.64 GB | ~8 GB | 4-bit extra-small, iMatrix |
42
+ | `Q4_K_M` | 7.38 GB | ~8 GB | 4-bit, iMatrix-calibrated (recommended) |
43
+ | `Q5_K_M` | 8.55 GB | ~10 GB | 5-bit, iMatrix-calibrated |
44
+ | `Q6_K` | 9.79 GB | ~11 GB | 6-bit, iMatrix-calibrated, near-lossless |
45
+ | `Q8_0` | 12.67 GB | ~15 GB | 8-bit, essentially lossless |
46
+
47
+
48
+ ## What is iMatrix?
49
+
50
+ Standard quantization treats all weights equally. iMatrix runs 128 calibration chunks through
51
+ the full-precision model to find which weights matter most, then allocates more precision where
52
+ it counts. At Q2/Q3/Q4 this means noticeably better coherence and instruction-following —
53
+ **same file size, better output**.
54
+
55
+ Calibration: 2M tokens of [WikiText-103](https://huggingface.co/datasets/wikitext).
56
+
57
+ > Also see plain (non-iMatrix) quants: `liodon-ai/Grug-12B-GGUF`
58
+
59
+ ## Source
60
+
61
+ - **Model**: [kai-os/Grug-12B](https://huggingface.co/kai-os/Grug-12B)
62
+ - **License**: other
63
+
64
+ ---
65
+ *Quantized by [Liodon AI](https://huggingface.co/liodon-ai)*