ngquocvinh commited on
Commit
b38c8b3
Β·
verified Β·
1 Parent(s): a321375

Remove raw validation logs from public package

Browse files
Files changed (24) hide show
  1. reproducibility/validation/convert-dry-run.log +0 -1297
  2. reproducibility/validation/convert.log +0 -1292
  3. reproducibility/validation/imatrix-combine.log +0 -6
  4. reproducibility/validation/imatrix-k2-corpus-gpu1.log +0 -10
  5. reproducibility/validation/imatrix-wikitext-gpu1.log +0 -19
  6. reproducibility/validation/quantize-logs/IQ1_M.log +0 -318
  7. reproducibility/validation/quantize-logs/IQ2_XS.log +0 -318
  8. reproducibility/validation/quantize-logs/Q1_0.log +0 -318
  9. reproducibility/validation/quantize-logs/Q2_K.log +0 -318
  10. reproducibility/validation/quantize-logs/Q3_K_M.log +0 -318
  11. reproducibility/validation/quantize-logs/Q4_K_M.log +0 -318
  12. reproducibility/validation/quantize-logs/Q5_K_M.log +0 -318
  13. reproducibility/validation/quantize-logs/Q6_K.log +0 -318
  14. reproducibility/validation/quantize-logs/Q8_0.log +0 -309
  15. reproducibility/validation/smoke-test-gpu1.tsv +0 -10
  16. reproducibility/validation/smoke-tests-gpu1/K2-Horizon-0.9B-IQ1_M.log +0 -56
  17. reproducibility/validation/smoke-tests-gpu1/K2-Horizon-0.9B-IQ2_XS.log +0 -33
  18. reproducibility/validation/smoke-tests-gpu1/K2-Horizon-0.9B-Q1_0.log +0 -33
  19. reproducibility/validation/smoke-tests-gpu1/K2-Horizon-0.9B-Q2_K.log +0 -33
  20. reproducibility/validation/smoke-tests-gpu1/K2-Horizon-0.9B-Q3_K_M.log +0 -33
  21. reproducibility/validation/smoke-tests-gpu1/K2-Horizon-0.9B-Q4_K_M.log +0 -33
  22. reproducibility/validation/smoke-tests-gpu1/K2-Horizon-0.9B-Q5_K_M.log +0 -33
  23. reproducibility/validation/smoke-tests-gpu1/K2-Horizon-0.9B-Q6_K.log +0 -33
  24. reproducibility/validation/smoke-tests-gpu1/K2-Horizon-0.9B-Q8_0.log +0 -33
reproducibility/validation/convert-dry-run.log DELETED
@@ -1,1297 +0,0 @@
1
- INFO:hf-to-gguf:Loading model: source-hf-k2-horizon-0.9b
2
- WARNING:hf-to-gguf:Failed to load model config from source-hf-k2-horizon-0.9b: The repository source-hf-k2-horizon-0.9b contains custom code which must be executed to correctly load the model. You can inspect the repository content at /root/workspace/HF/source-hf-k2-horizon-0.9b .
3
- You can inspect the repository content at https://hf.co/source-hf-k2-horizon-0.9b.
4
- Please pass the argument `trust_remote_code=True` to allow custom code to be run.
5
- WARNING:hf-to-gguf:Trying to load config.json instead
6
- INFO:hf-to-gguf:Model architecture: K2HorizonForCausalLM
7
- WARNING:hf-to-gguf:Failed to load model config from source-hf-k2-horizon-0.9b: The repository source-hf-k2-horizon-0.9b contains custom code which must be executed to correctly load the model. You can inspect the repository content at /root/workspace/HF/source-hf-k2-horizon-0.9b .
8
- You can inspect the repository content at https://hf.co/source-hf-k2-horizon-0.9b.
9
- Please pass the argument `trust_remote_code=True` to allow custom code to be run.
10
- WARNING:hf-to-gguf:Trying to load config.json instead
11
- INFO:hf-to-gguf:gguf: loading model weight map from 'model.safetensors.index.json'
12
- INFO:hf-to-gguf:gguf: indexing model part 'model-00000-of-00001.safetensors'
13
- INFO:gguf.gguf_writer:gguf: This GGUF file is for Little Endian only
14
- INFO:hf-to-gguf:Exporting model...
15
- INFO:hf-to-gguf:output.weight, torch.bfloat16 --> BF16, shape = {1536, 64256}
16
- INFO:hf-to-gguf:token_embd.weight, torch.bfloat16 --> BF16, shape = {1536, 64256}
17
- INFO:hf-to-gguf:blk.0.attn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
18
- INFO:hf-to-gguf:blk.0.ffn_down.weight, torch.bfloat16 --> BF16, shape = {5120, 1536}
19
- INFO:hf-to-gguf:blk.0.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
20
- INFO:hf-to-gguf:blk.0.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
21
- INFO:hf-to-gguf:blk.0.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
22
- INFO:hf-to-gguf:blk.0.attn_k.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
23
- INFO:hf-to-gguf:blk.0.attn_output.weight, torch.bfloat16 --> BF16, shape = {2048, 1536}
24
- INFO:hf-to-gguf:blk.0.attn_q.weight, torch.bfloat16 --> BF16, shape = {1536, 2048}
25
- INFO:hf-to-gguf:blk.0.attn_v.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
26
- INFO:hf-to-gguf:blk.1.attn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
27
- INFO:hf-to-gguf:blk.1.ffn_down.weight, torch.bfloat16 --> BF16, shape = {5120, 1536}
28
- INFO:hf-to-gguf:blk.1.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
29
- INFO:hf-to-gguf:blk.1.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
30
- INFO:hf-to-gguf:blk.1.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
31
- INFO:hf-to-gguf:blk.1.attn_k.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
32
- INFO:hf-to-gguf:blk.1.attn_output.weight, torch.bfloat16 --> BF16, shape = {2048, 1536}
33
- INFO:hf-to-gguf:blk.1.attn_q.weight, torch.bfloat16 --> BF16, shape = {1536, 2048}
34
- INFO:hf-to-gguf:blk.1.attn_v.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
35
- INFO:hf-to-gguf:blk.10.attn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
36
- INFO:hf-to-gguf:blk.10.ffn_down.weight, torch.bfloat16 --> BF16, shape = {5120, 1536}
37
- INFO:hf-to-gguf:blk.10.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
38
- INFO:hf-to-gguf:blk.10.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
39
- INFO:hf-to-gguf:blk.10.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
40
- INFO:hf-to-gguf:blk.10.attn_k.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
41
- INFO:hf-to-gguf:blk.10.attn_output.weight, torch.bfloat16 --> BF16, shape = {2048, 1536}
42
- INFO:hf-to-gguf:blk.10.attn_q.weight, torch.bfloat16 --> BF16, shape = {1536, 2048}
43
- INFO:hf-to-gguf:blk.10.attn_v.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
44
- INFO:hf-to-gguf:blk.11.attn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
45
- INFO:hf-to-gguf:blk.11.ffn_down.weight, torch.bfloat16 --> BF16, shape = {5120, 1536}
46
- INFO:hf-to-gguf:blk.11.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
47
- INFO:hf-to-gguf:blk.11.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
48
- INFO:hf-to-gguf:blk.11.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
49
- INFO:hf-to-gguf:blk.11.attn_k.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
50
- INFO:hf-to-gguf:blk.11.attn_output.weight, torch.bfloat16 --> BF16, shape = {2048, 1536}
51
- INFO:hf-to-gguf:blk.11.attn_q.weight, torch.bfloat16 --> BF16, shape = {1536, 2048}
52
- INFO:hf-to-gguf:blk.11.attn_v.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
53
- INFO:hf-to-gguf:blk.12.attn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
54
- INFO:hf-to-gguf:blk.12.ffn_down.weight, torch.bfloat16 --> BF16, shape = {5120, 1536}
55
- INFO:hf-to-gguf:blk.12.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
56
- INFO:hf-to-gguf:blk.12.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
57
- INFO:hf-to-gguf:blk.12.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
58
- INFO:hf-to-gguf:blk.12.attn_k.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
59
- INFO:hf-to-gguf:blk.12.attn_output.weight, torch.bfloat16 --> BF16, shape = {2048, 1536}
60
- INFO:hf-to-gguf:blk.12.attn_q.weight, torch.bfloat16 --> BF16, shape = {1536, 2048}
61
- INFO:hf-to-gguf:blk.12.attn_v.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
62
- INFO:hf-to-gguf:blk.13.attn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
63
- INFO:hf-to-gguf:blk.13.ffn_down.weight, torch.bfloat16 --> BF16, shape = {5120, 1536}
64
- INFO:hf-to-gguf:blk.13.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
65
- INFO:hf-to-gguf:blk.13.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
66
- INFO:hf-to-gguf:blk.13.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
67
- INFO:hf-to-gguf:blk.13.attn_k.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
68
- INFO:hf-to-gguf:blk.13.attn_output.weight, torch.bfloat16 --> BF16, shape = {2048, 1536}
69
- INFO:hf-to-gguf:blk.13.attn_q.weight, torch.bfloat16 --> BF16, shape = {1536, 2048}
70
- INFO:hf-to-gguf:blk.13.attn_v.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
71
- INFO:hf-to-gguf:blk.14.attn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
72
- INFO:hf-to-gguf:blk.14.ffn_down.weight, torch.bfloat16 --> BF16, shape = {5120, 1536}
73
- INFO:hf-to-gguf:blk.14.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
74
- INFO:hf-to-gguf:blk.14.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
75
- INFO:hf-to-gguf:blk.14.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
76
- INFO:hf-to-gguf:blk.14.attn_k.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
77
- INFO:hf-to-gguf:blk.14.attn_output.weight, torch.bfloat16 --> BF16, shape = {2048, 1536}
78
- INFO:hf-to-gguf:blk.14.attn_q.weight, torch.bfloat16 --> BF16, shape = {1536, 2048}
79
- INFO:hf-to-gguf:blk.14.attn_v.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
80
- INFO:hf-to-gguf:blk.15.attn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
81
- INFO:hf-to-gguf:blk.15.ffn_down.weight, torch.bfloat16 --> BF16, shape = {5120, 1536}
82
- INFO:hf-to-gguf:blk.15.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
83
- INFO:hf-to-gguf:blk.15.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
84
- INFO:hf-to-gguf:blk.15.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
85
- INFO:hf-to-gguf:blk.15.attn_k.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
86
- INFO:hf-to-gguf:blk.15.attn_output.weight, torch.bfloat16 --> BF16, shape = {2048, 1536}
87
- INFO:hf-to-gguf:blk.15.attn_q.weight, torch.bfloat16 --> BF16, shape = {1536, 2048}
88
- INFO:hf-to-gguf:blk.15.attn_v.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
89
- INFO:hf-to-gguf:blk.16.attn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
90
- INFO:hf-to-gguf:blk.16.ffn_down.weight, torch.bfloat16 --> BF16, shape = {5120, 1536}
91
- INFO:hf-to-gguf:blk.16.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
92
- INFO:hf-to-gguf:blk.16.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
93
- INFO:hf-to-gguf:blk.16.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
94
- INFO:hf-to-gguf:blk.16.attn_k.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
95
- INFO:hf-to-gguf:blk.16.attn_output.weight, torch.bfloat16 --> BF16, shape = {2048, 1536}
96
- INFO:hf-to-gguf:blk.16.attn_q.weight, torch.bfloat16 --> BF16, shape = {1536, 2048}
97
- INFO:hf-to-gguf:blk.16.attn_v.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
98
- INFO:hf-to-gguf:blk.17.attn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
99
- INFO:hf-to-gguf:blk.17.ffn_down.weight, torch.bfloat16 --> BF16, shape = {5120, 1536}
100
- INFO:hf-to-gguf:blk.17.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
101
- INFO:hf-to-gguf:blk.17.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
102
- INFO:hf-to-gguf:blk.17.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
103
- INFO:hf-to-gguf:blk.17.attn_k.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
104
- INFO:hf-to-gguf:blk.17.attn_output.weight, torch.bfloat16 --> BF16, shape = {2048, 1536}
105
- INFO:hf-to-gguf:blk.17.attn_q.weight, torch.bfloat16 --> BF16, shape = {1536, 2048}
106
- INFO:hf-to-gguf:blk.17.attn_v.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
107
- INFO:hf-to-gguf:blk.18.attn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
108
- INFO:hf-to-gguf:blk.18.ffn_down.weight, torch.bfloat16 --> BF16, shape = {5120, 1536}
109
- INFO:hf-to-gguf:blk.18.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
110
- INFO:hf-to-gguf:blk.18.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
111
- INFO:hf-to-gguf:blk.18.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
112
- INFO:hf-to-gguf:blk.18.attn_k.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
113
- INFO:hf-to-gguf:blk.18.attn_output.weight, torch.bfloat16 --> BF16, shape = {2048, 1536}
114
- INFO:hf-to-gguf:blk.18.attn_q.weight, torch.bfloat16 --> BF16, shape = {1536, 2048}
115
- INFO:hf-to-gguf:blk.18.attn_v.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
116
- INFO:hf-to-gguf:blk.19.attn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
117
- INFO:hf-to-gguf:blk.19.ffn_down.weight, torch.bfloat16 --> BF16, shape = {5120, 1536}
118
- INFO:hf-to-gguf:blk.19.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
119
- INFO:hf-to-gguf:blk.19.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
120
- INFO:hf-to-gguf:blk.19.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
121
- INFO:hf-to-gguf:blk.19.attn_k.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
122
- INFO:hf-to-gguf:blk.19.attn_output.weight, torch.bfloat16 --> BF16, shape = {2048, 1536}
123
- INFO:hf-to-gguf:blk.19.attn_q.weight, torch.bfloat16 --> BF16, shape = {1536, 2048}
124
- INFO:hf-to-gguf:blk.19.attn_v.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
125
- INFO:hf-to-gguf:blk.2.attn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
126
- INFO:hf-to-gguf:blk.2.ffn_down.weight, torch.bfloat16 --> BF16, shape = {5120, 1536}
127
- INFO:hf-to-gguf:blk.2.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
128
- INFO:hf-to-gguf:blk.2.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
129
- INFO:hf-to-gguf:blk.2.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
130
- INFO:hf-to-gguf:blk.2.attn_k.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
131
- INFO:hf-to-gguf:blk.2.attn_output.weight, torch.bfloat16 --> BF16, shape = {2048, 1536}
132
- INFO:hf-to-gguf:blk.2.attn_q.weight, torch.bfloat16 --> BF16, shape = {1536, 2048}
133
- INFO:hf-to-gguf:blk.2.attn_v.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
134
- INFO:hf-to-gguf:blk.20.attn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
135
- INFO:hf-to-gguf:blk.20.ffn_down.weight, torch.bfloat16 --> BF16, shape = {5120, 1536}
136
- INFO:hf-to-gguf:blk.20.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
137
- INFO:hf-to-gguf:blk.20.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
138
- INFO:hf-to-gguf:blk.20.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
139
- INFO:hf-to-gguf:blk.20.attn_k.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
140
- INFO:hf-to-gguf:blk.20.attn_output.weight, torch.bfloat16 --> BF16, shape = {2048, 1536}
141
- INFO:hf-to-gguf:blk.20.attn_q.weight, torch.bfloat16 --> BF16, shape = {1536, 2048}
142
- INFO:hf-to-gguf:blk.20.attn_v.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
143
- INFO:hf-to-gguf:blk.21.attn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
144
- INFO:hf-to-gguf:blk.21.ffn_down.weight, torch.bfloat16 --> BF16, shape = {5120, 1536}
145
- INFO:hf-to-gguf:blk.21.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
146
- INFO:hf-to-gguf:blk.21.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
147
- INFO:hf-to-gguf:blk.21.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
148
- INFO:hf-to-gguf:blk.21.attn_k.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
149
- INFO:hf-to-gguf:blk.21.attn_output.weight, torch.bfloat16 --> BF16, shape = {2048, 1536}
150
- INFO:hf-to-gguf:blk.21.attn_q.weight, torch.bfloat16 --> BF16, shape = {1536, 2048}
151
- INFO:hf-to-gguf:blk.21.attn_v.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
152
- INFO:hf-to-gguf:blk.22.attn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
153
- INFO:hf-to-gguf:blk.22.ffn_down.weight, torch.bfloat16 --> BF16, shape = {5120, 1536}
154
- INFO:hf-to-gguf:blk.22.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
155
- INFO:hf-to-gguf:blk.22.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
156
- INFO:hf-to-gguf:blk.22.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
157
- INFO:hf-to-gguf:blk.22.attn_k.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
158
- INFO:hf-to-gguf:blk.22.attn_output.weight, torch.bfloat16 --> BF16, shape = {2048, 1536}
159
- INFO:hf-to-gguf:blk.22.attn_q.weight, torch.bfloat16 --> BF16, shape = {1536, 2048}
160
- INFO:hf-to-gguf:blk.22.attn_v.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
161
- INFO:hf-to-gguf:blk.23.attn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
162
- INFO:hf-to-gguf:blk.23.ffn_down.weight, torch.bfloat16 --> BF16, shape = {5120, 1536}
163
- INFO:hf-to-gguf:blk.23.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
164
- INFO:hf-to-gguf:blk.23.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
165
- INFO:hf-to-gguf:blk.23.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
166
- INFO:hf-to-gguf:blk.23.attn_k.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
167
- INFO:hf-to-gguf:blk.23.attn_output.weight, torch.bfloat16 --> BF16, shape = {2048, 1536}
168
- INFO:hf-to-gguf:blk.23.attn_q.weight, torch.bfloat16 --> BF16, shape = {1536, 2048}
169
- INFO:hf-to-gguf:blk.23.attn_v.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
170
- INFO:hf-to-gguf:blk.24.attn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
171
- INFO:hf-to-gguf:blk.24.ffn_down.weight, torch.bfloat16 --> BF16, shape = {5120, 1536}
172
- INFO:hf-to-gguf:blk.24.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
173
- INFO:hf-to-gguf:blk.24.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
174
- INFO:hf-to-gguf:blk.24.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
175
- INFO:hf-to-gguf:blk.24.attn_k.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
176
- INFO:hf-to-gguf:blk.24.attn_output.weight, torch.bfloat16 --> BF16, shape = {2048, 1536}
177
- INFO:hf-to-gguf:blk.24.attn_q.weight, torch.bfloat16 --> BF16, shape = {1536, 2048}
178
- INFO:hf-to-gguf:blk.24.attn_v.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
179
- INFO:hf-to-gguf:blk.25.attn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
180
- INFO:hf-to-gguf:blk.25.ffn_down.weight, torch.bfloat16 --> BF16, shape = {5120, 1536}
181
- INFO:hf-to-gguf:blk.25.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
182
- INFO:hf-to-gguf:blk.25.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
183
- INFO:hf-to-gguf:blk.25.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
184
- INFO:hf-to-gguf:blk.25.attn_k.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
185
- INFO:hf-to-gguf:blk.25.attn_output.weight, torch.bfloat16 --> BF16, shape = {2048, 1536}
186
- INFO:hf-to-gguf:blk.25.attn_q.weight, torch.bfloat16 --> BF16, shape = {1536, 2048}
187
- INFO:hf-to-gguf:blk.25.attn_v.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
188
- INFO:hf-to-gguf:blk.26.attn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
189
- INFO:hf-to-gguf:blk.26.ffn_down.weight, torch.bfloat16 --> BF16, shape = {5120, 1536}
190
- INFO:hf-to-gguf:blk.26.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
191
- INFO:hf-to-gguf:blk.26.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
192
- INFO:hf-to-gguf:blk.26.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
193
- INFO:hf-to-gguf:blk.26.attn_k.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
194
- INFO:hf-to-gguf:blk.26.attn_output.weight, torch.bfloat16 --> BF16, shape = {2048, 1536}
195
- INFO:hf-to-gguf:blk.26.attn_q.weight, torch.bfloat16 --> BF16, shape = {1536, 2048}
196
- INFO:hf-to-gguf:blk.26.attn_v.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
197
- INFO:hf-to-gguf:blk.27.attn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
198
- INFO:hf-to-gguf:blk.27.ffn_down.weight, torch.bfloat16 --> BF16, shape = {5120, 1536}
199
- INFO:hf-to-gguf:blk.27.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
200
- INFO:hf-to-gguf:blk.27.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
201
- INFO:hf-to-gguf:blk.27.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
202
- INFO:hf-to-gguf:blk.27.attn_k.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
203
- INFO:hf-to-gguf:blk.27.attn_output.weight, torch.bfloat16 --> BF16, shape = {2048, 1536}
204
- INFO:hf-to-gguf:blk.27.attn_q.weight, torch.bfloat16 --> BF16, shape = {1536, 2048}
205
- INFO:hf-to-gguf:blk.27.attn_v.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
206
- INFO:hf-to-gguf:blk.3.attn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
207
- INFO:hf-to-gguf:blk.3.ffn_down.weight, torch.bfloat16 --> BF16, shape = {5120, 1536}
208
- INFO:hf-to-gguf:blk.3.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
209
- INFO:hf-to-gguf:blk.3.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
210
- INFO:hf-to-gguf:blk.3.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
211
- INFO:hf-to-gguf:blk.3.attn_k.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
212
- INFO:hf-to-gguf:blk.3.attn_output.weight, torch.bfloat16 --> BF16, shape = {2048, 1536}
213
- INFO:hf-to-gguf:blk.3.attn_q.weight, torch.bfloat16 --> BF16, shape = {1536, 2048}
214
- INFO:hf-to-gguf:blk.3.attn_v.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
215
- INFO:hf-to-gguf:blk.4.attn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
216
- INFO:hf-to-gguf:blk.4.ffn_down.weight, torch.bfloat16 --> BF16, shape = {5120, 1536}
217
- INFO:hf-to-gguf:blk.4.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
218
- INFO:hf-to-gguf:blk.4.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
219
- INFO:hf-to-gguf:blk.4.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
220
- INFO:hf-to-gguf:blk.4.attn_k.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
221
- INFO:hf-to-gguf:blk.4.attn_output.weight, torch.bfloat16 --> BF16, shape = {2048, 1536}
222
- INFO:hf-to-gguf:blk.4.attn_q.weight, torch.bfloat16 --> BF16, shape = {1536, 2048}
223
- INFO:hf-to-gguf:blk.4.attn_v.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
224
- INFO:hf-to-gguf:blk.5.attn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
225
- INFO:hf-to-gguf:blk.5.ffn_down.weight, torch.bfloat16 --> BF16, shape = {5120, 1536}
226
- INFO:hf-to-gguf:blk.5.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
227
- INFO:hf-to-gguf:blk.5.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
228
- INFO:hf-to-gguf:blk.5.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
229
- INFO:hf-to-gguf:blk.5.attn_k.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
230
- INFO:hf-to-gguf:blk.5.attn_output.weight, torch.bfloat16 --> BF16, shape = {2048, 1536}
231
- INFO:hf-to-gguf:blk.5.attn_q.weight, torch.bfloat16 --> BF16, shape = {1536, 2048}
232
- INFO:hf-to-gguf:blk.5.attn_v.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
233
- INFO:hf-to-gguf:blk.6.attn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
234
- INFO:hf-to-gguf:blk.6.ffn_down.weight, torch.bfloat16 --> BF16, shape = {5120, 1536}
235
- INFO:hf-to-gguf:blk.6.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
236
- INFO:hf-to-gguf:blk.6.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
237
- INFO:hf-to-gguf:blk.6.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
238
- INFO:hf-to-gguf:blk.6.attn_k.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
239
- INFO:hf-to-gguf:blk.6.attn_output.weight, torch.bfloat16 --> BF16, shape = {2048, 1536}
240
- INFO:hf-to-gguf:blk.6.attn_q.weight, torch.bfloat16 --> BF16, shape = {1536, 2048}
241
- INFO:hf-to-gguf:blk.6.attn_v.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
242
- INFO:hf-to-gguf:blk.7.attn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
243
- INFO:hf-to-gguf:blk.7.ffn_down.weight, torch.bfloat16 --> BF16, shape = {5120, 1536}
244
- INFO:hf-to-gguf:blk.7.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
245
- INFO:hf-to-gguf:blk.7.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
246
- INFO:hf-to-gguf:blk.7.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
247
- INFO:hf-to-gguf:blk.7.attn_k.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
248
- INFO:hf-to-gguf:blk.7.attn_output.weight, torch.bfloat16 --> BF16, shape = {2048, 1536}
249
- INFO:hf-to-gguf:blk.7.attn_q.weight, torch.bfloat16 --> BF16, shape = {1536, 2048}
250
- INFO:hf-to-gguf:blk.7.attn_v.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
251
- INFO:hf-to-gguf:blk.8.attn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
252
- INFO:hf-to-gguf:blk.8.ffn_down.weight, torch.bfloat16 --> BF16, shape = {5120, 1536}
253
- INFO:hf-to-gguf:blk.8.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
254
- INFO:hf-to-gguf:blk.8.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
255
- INFO:hf-to-gguf:blk.8.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
256
- INFO:hf-to-gguf:blk.8.attn_k.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
257
- INFO:hf-to-gguf:blk.8.attn_output.weight, torch.bfloat16 --> BF16, shape = {2048, 1536}
258
- INFO:hf-to-gguf:blk.8.attn_q.weight, torch.bfloat16 --> BF16, shape = {1536, 2048}
259
- INFO:hf-to-gguf:blk.8.attn_v.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
260
- INFO:hf-to-gguf:blk.9.attn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
261
- INFO:hf-to-gguf:blk.9.ffn_down.weight, torch.bfloat16 --> BF16, shape = {5120, 1536}
262
- INFO:hf-to-gguf:blk.9.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
263
- INFO:hf-to-gguf:blk.9.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
264
- INFO:hf-to-gguf:blk.9.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
265
- INFO:hf-to-gguf:blk.9.attn_k.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
266
- INFO:hf-to-gguf:blk.9.attn_output.weight, torch.bfloat16 --> BF16, shape = {2048, 1536}
267
- INFO:hf-to-gguf:blk.9.attn_q.weight, torch.bfloat16 --> BF16, shape = {1536, 2048}
268
- INFO:hf-to-gguf:blk.9.attn_v.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
269
- INFO:hf-to-gguf:output_norm.weight, torch.bfloat16 --> F32, shape = {1536}
270
- INFO:hf-to-gguf:Set meta model
271
- INFO:hf-to-gguf:Set model parameters
272
- INFO:hf-to-gguf:gguf: context length = 131072
273
- INFO:hf-to-gguf:gguf: embedding length = 1536
274
- INFO:hf-to-gguf:gguf: feed forward length = 5120
275
- INFO:hf-to-gguf:gguf: head count = 32
276
- INFO:hf-to-gguf:gguf: key-value head count = 8
277
- INFO:hf-to-gguf:gguf: rope scaling type = YARN
278
- INFO:hf-to-gguf:gguf: rope theta = 1000000.0
279
- INFO:hf-to-gguf:gguf: rms norm epsilon = 1e-06
280
- INFO:hf-to-gguf:gguf: expert count = 0
281
- INFO:hf-to-gguf:gguf: experts used count = 0
282
- INFO:hf-to-gguf:gguf: file type = 32
283
- INFO:hf-to-gguf:Set model quantization version
284
- INFO:hf-to-gguf:Set model tokenizer
285
- DEBUG:hf-to-gguf:chktok: [200, 5834, 40416, 222, 32719, 47979, 7928, 5828, 15421, 5544, 49219, 4643, 250, 224, 383, 18099, 10, 30533, 116, 16441, 61457, 106, 16995, 383, 25548, 4753, 952, 16767, 291, 3708, 8102, 709, 10, 29816, 229, 10005, 101, 249, 4643, 101, 249, 222, 20, 222, 837, 222, 4669, 222, 4669, 20, 222, 4669, 837, 222, 4669, 4669, 222, 4669, 4669, 20, 222, 4669, 4669, 837, 222, 20, 15, 20, 222, 20, 456, 20, 222, 20, 955, 20, 222, 29178, 224, 29178, 116, 29178, 243, 159, 255, 235, 29178, 239, 159, 255, 226, 29178, 246, 29178, 117, 29178, 255, 159, 255, 225, 29178, 255, 29178, 97, 29178, 116, 29178, 229, 32182, 225, 3278, 3897, 9729, 2421, 31158, 15026, 6198, 9970, 18, 6809, 19291, 24894, 392, 1544, 43212, 3199, 2919, 420, 2154, 13544, 16497, 707, 1753, 4130, 51639, 28024, 37225, 6277, 6277, 14741, 4185, 4185, 19855, 19245, 3344, 21484, 6441, 372, 2947, 1292, 473, 85, 1120, 592, 552, 1322, 13, 473, 1540, 403, 3256, 32, 473, 46, 667, 3256, 372, 2750, 1468, 497, 13, 473, 37, 403, 1293, 1262, 16943, 32, 1215, 8, 23892, 265, 61312, 45]
286
- DEBUG:hf-to-gguf:chkhsh: 1f9825a388f700a6b591722f17d470cbbcf10973ece35d2fd14239a14110ae1a
287
- DEBUG:hf-to-gguf:tokenizer.ggml.pre: 'k2-horizon'
288
- DEBUG:hf-to-gguf:chkhsh: 1f9825a388f700a6b591722f17d470cbbcf10973ece35d2fd14239a14110ae1a
289
- INFO:gguf.vocab:Adding 63742 merge(s).
290
- INFO:gguf.vocab:Setting special token type bos to 0
291
- INFO:gguf.vocab:Setting special token type eos to 1
292
- INFO:gguf.vocab:Setting special token type pad to 64255
293
- INFO:gguf.vocab:Setting chat_template to {%- if tool_presentation is defined -%}
294
- {{- raise_exception("Unsupported argument: tool_presentation. Use tool_presentation_format with one of: json, xml, markdown.") -}}
295
- {%- endif -%}
296
- {%- if tool_calling_format is defined -%}
297
- {{- raise_exception("Unsupported argument: tool_calling_format. Use tool_call_format with one of: json, xml, xml_typed.") -}}
298
- {%- endif -%}
299
- {%- if tool_format is defined -%}
300
- {{- raise_exception("Unsupported argument: tool_format. Use tool_call_format with one of: json, xml, xml_typed.") -}}
301
- {%- endif -%}
302
- {%- set tool_presentation_fmt = tool_presentation_format | default('markdown') -%}
303
- {%- set tool_call_fmt = tool_call_format | default('xml') -%}
304
- {%- if tool_presentation_fmt != 'json' and tool_presentation_fmt != 'xml' and tool_presentation_fmt != 'markdown' -%}
305
- {{- raise_exception("Unsupported tool_presentation_format: '" ~ tool_presentation_fmt ~ "'. Supported formats: json, xml, markdown.") -}}
306
- {%- endif -%}
307
- {%- if tool_call_fmt != 'json' and tool_call_fmt != 'xml' and tool_call_fmt != 'xml_typed' -%}
308
- {{- raise_exception("Unsupported tool_call_format: '" ~ tool_call_fmt ~ "'. Supported formats: json, xml, xml_typed.") -}}
309
- {%- endif -%}
310
-
311
- {#- Renderability state, computed during validate_tools (single walk, no extra -#}
312
- {#- traversal at render time): ok = working flag for the tool being validated; -#}
313
- {#- bad = pipe-delimited indices of tools that must render as verbatim JSON. -#}
314
- {%- set RB = namespace(ok=true, bad='|') -%}
315
-
316
- {%- macro value_contains_mapping(v) -%}
317
- {%- if v is mapping -%}
318
- true
319
- {%- elif v is sequence and v is not string -%}
320
- {%- set f = namespace(x='false') -%}
321
- {%- for c in v -%}{%- if value_contains_mapping(c) == 'true' -%}{%- set f.x = 'true' -%}{%- endif -%}{%- endfor -%}
322
- {{- f.x -}}
323
- {%- else -%}
324
- false
325
- {%- endif -%}
326
- {%- endmacro -%}
327
-
328
- {#- $ref inlining state: defs = local $defs of the tool being rendered; seen = -#}
329
- {#- pipe-delimited names already expanded for this tool (each def inlines at most -#}
330
- {#- once; later references render by def name; cycles terminate immediately). -#}
331
- {#- $ref-sibling annotations (description/default/...) merge OVER the def at -#}
332
- {#- the inline site, so use-site annotations win and are never dropped. -#}
333
- {%- set REFS = namespace(defs={}, seen='|') -%}
334
-
335
- {%- macro render_compact_type_name(type_name, spec) -%}
336
- {%- if type_name == "array" -%}
337
- array[{%- if 'items' in spec -%}{{ render_compact_type(spec['items']) }}{%- else -%}any{%- endif -%}]
338
- {%- elif type_name -%}
339
- {{- type_name -}}
340
- {%- else -%}
341
- any
342
- {%- endif -%}
343
- {%- endmacro -%}
344
-
345
- {%- macro render_compact_type(spec) -%}
346
- {%- if spec is not mapping -%}
347
- any
348
- {%- elif spec.type is defined and spec.type is sequence and spec.type is not string and spec.type | length > 0 -%}
349
- {%- for type_name in spec.type -%}{{ render_compact_type_name(type_name, spec) }}{%- if not loop.last -%}|{%- endif -%}{%- endfor -%}
350
- {%- elif spec.type is defined and spec.type is sequence and spec.type is not string -%}
351
- any
352
- {%- elif spec.type -%}
353
- {{- render_compact_type_name(spec.type, spec) -}}
354
- {%- elif spec['$ref'] is string -%}
355
- {{- spec['$ref'].split('/') | last -}}
356
- {%- elif spec.oneOf -%}
357
- oneOf[{%- for variant in spec.oneOf -%}{{ render_compact_type(variant) }}{%- if not loop.last -%}|{%- endif -%}{%- endfor -%}]
358
- {%- elif spec.anyOf -%}
359
- anyOf[{%- for variant in spec.anyOf -%}{{ render_compact_type(variant) }}{%- if not loop.last -%}|{%- endif -%}{%- endfor -%}]
360
- {%- elif spec.properties -%}
361
- object
362
- {%- elif 'items' in spec -%}
363
- array[{{ render_compact_type(spec['items']) }}]
364
- {%- else -%}
365
- any
366
- {%- endif -%}
367
- {%- endmacro -%}
368
-
369
- {%- macro render_markdown_type_name(type_name, spec) -%}
370
- {%- if type_name == "array" -%}
371
- array of {% if 'items' in spec %}{{ render_markdown_type(spec['items']) }}{% else %}any{% endif %}
372
- {%- elif type_name -%}
373
- {{- type_name -}}
374
- {%- else -%}
375
- any
376
- {%- endif -%}
377
- {%- endmacro -%}
378
-
379
- {%- macro render_markdown_type(spec) -%}
380
- {%- if spec is sameas true -%}
381
- True
382
- {%- elif spec is sameas false -%}
383
- False
384
- {%- elif spec is not mapping -%}
385
- any
386
- {%- elif spec.type is defined and spec.type is sequence and spec.type is not string and spec.type | length > 0 -%}
387
- {%- for type_name in spec.type -%}{{ render_markdown_type_name(type_name, spec) }}{% if not loop.last %} or {% endif %}{%- endfor -%}
388
- {%- elif spec.type is defined and spec.type is sequence and spec.type is not string -%}
389
- any
390
- {%- elif spec.type -%}
391
- {{- render_markdown_type_name(spec.type, spec) -}}
392
- {%- elif spec['$ref'] is string -%}
393
- {{- spec['$ref'].split('/') | last -}}
394
- {%- elif spec.oneOf -%}
395
- oneOf[{%- for variant in spec.oneOf -%}{{ render_markdown_type(variant) }}{% if not loop.last %} or {% endif %}{%- endfor -%}]
396
- {%- elif spec.anyOf -%}
397
- anyOf[{%- for variant in spec.anyOf -%}{{ render_markdown_type(variant) }}{% if not loop.last %} or {% endif %}{%- endfor -%}]
398
- {%- elif spec.properties -%}
399
- object
400
- {%- elif 'items' in spec -%}
401
- array of {{ render_markdown_type(spec['items']) }}
402
- {%- else -%}
403
- any
404
- {%- endif -%}
405
- {%- endmacro -%}
406
-
407
- {%- macro render_xml_text(value) -%}
408
- {{- value.split() | join(" ") -}}
409
- {%- endmacro -%}
410
-
411
- {%- macro render_python_string(value) -%}
412
- '{{- value.split() | join(" ") | replace("\\", "\\\\") | replace("'", "\\'") -}}'
413
- {%- endmacro -%}
414
-
415
- {%- macro render_python_repr(value) -%}
416
- {%- if value is string -%}
417
- {{ render_python_string(value) }}
418
- {%- elif value is sameas true -%}
419
- True
420
- {%- elif value is sameas false -%}
421
- False
422
- {%- elif value is none -%}
423
- None
424
- {%- elif value is mapping -%}
425
- {{- "{" -}}
426
- {%- for key, child in value | items -%}
427
- {{ render_python_repr(key) }}: {{ render_python_repr(child) }}{%- if not loop.last -%}, {% endif -%}
428
- {%- endfor -%}
429
- {{- "}" -}}
430
- {%- elif value is sequence -%}
431
- {{- "[" -}}
432
- {%- for child in value -%}
433
- {{ render_python_repr(child) }}{%- if not loop.last -%}, {% endif -%}
434
- {%- endfor -%}
435
- {{- "]" -}}
436
- {%- else -%}
437
- {{- value -}}
438
- {%- endif -%}
439
- {%- endmacro -%}
440
-
441
- {%- macro render_xml_value(value) -%}
442
- {%- if value is string -%}{{ render_xml_text(value) }}{%- else -%}{{ render_python_repr(value) }}{%- endif -%}
443
- {%- endmacro -%}
444
-
445
- {%- macro render_xml_enum_value(value) -%}
446
- {%- if value is string -%}"{{- value | replace("\\", "\\\\") | replace("\"", "\\\"") -}}"{%- else -%}"{{- render_python_repr(value) | replace("\\", "\\\\") | replace("\"", "\\\"") -}}"{%- endif -%}
447
- {%- endmacro -%}
448
-
449
- {%- macro render_xml_enum(values) -%}
450
- {%- for value in values -%}{{ render_xml_enum_value(value) }}{%- if not loop.last -%}|{%- endif -%}{%- endfor -%}
451
- {%- endmacro -%}
452
-
453
- {%- macro render_xml_default_attr(value) -%}
454
- {{- " default=" }}{%- if value is string -%}"{{- value | replace("\\", "\\\\") | replace("\"", "\\\"") -}}"{%- else -%}{{ render_xml_value(value) }}{%- endif -%}
455
- {%- endmacro -%}
456
-
457
- {%- macro render_xml_attr(name, value) -%}
458
- {{- " " + name + "=" }}{%- if value == "" -%}""{%- else -%}{{ render_xml_value(value) }}{%- endif -%}
459
- {%- endmacro -%}
460
-
461
- {%- macro validate_schema(spec, path, lenient=false, classify=true, in_variant=false) -%}
462
- {%- if spec is mapping -%}
463
- {%- if not lenient -%}
464
- {%- if spec.required is defined -%}
465
- {%- if spec.required is string or spec.required is not sequence -%}
466
- {{- raise_exception("Schema '" + path + "' has 'required' but it is not a list.") -}}
467
- {%- endif -%}
468
- {%- if spec.required | length > 0 and not spec.properties and not in_variant -%}
469
- {{- raise_exception("Schema '" + path + "' has required fields but no properties object to define them.") -}}
470
- {%- endif -%}
471
- {%- if spec.properties -%}
472
- {%- for required_name in spec.required -%}
473
- {%- if required_name not in spec.properties -%}
474
- {{- raise_exception("Schema '" + path + "' marks '" + required_name + "' as required, but that property is not defined in properties.") -}}
475
- {%- endif -%}
476
- {%- endfor -%}
477
- {%- endif -%}
478
- {%- endif -%}
479
- {%- endif -%}
480
- {#- renderability classification, piggybacking on this walk (no raises here): -#}
481
- {#- constructs the pretty renderer does not fully handle flip RB.ok so the -#}
482
- {#- tool falls back to verbatim JSON. Skipped entirely for json presentation. -#}
483
- {%- if classify -%}
484
- {%- for key, value in spec | items -%}
485
- {%- if key == '$ref' -%}
486
- {%- if value is not string -%}{%- set RB.ok = false -%}
487
- {%- elif not (value.startswith('#/$defs/') or value.startswith('#/definitions/')) -%}{%- set RB.ok = false -%}{%- endif -%}
488
- {%- elif key == '$defs' or key == 'definitions' -%}
489
- {%- if value is mapping -%}
490
- {%- for dk, dv in value | items -%}
491
- {{- validate_schema(dv, path + ".$defs." + dk, true) -}}
492
- {%- endfor -%}
493
- {%- else -%}{%- set RB.ok = false -%}{%- endif -%}
494
- {%- elif key == 'type' -%}
495
- {%- if value is mapping -%}{%- set RB.ok = false -%}{%- endif -%}
496
- {%- elif key == 'enum' -%}
497
- {%- if value is string or value is mapping or value is not sequence -%}{%- set RB.ok = false -%}{%- endif -%}
498
- {%- elif key == 'items' -%}
499
- {#- any items shape renders: mapping structurally, others via repr detail -#}
500
- {%- elif key == 'oneOf' or key == 'anyOf' -%}
501
- {%- if value is mapping or value is string or value is not sequence -%}{%- set RB.ok = false -%}{%- endif -%}
502
- {%- elif key == 'required' -%}
503
- {%- if value and not spec.properties -%}{%- set RB.ok = false -%}{%- endif -%}
504
- {%- elif ('|' ~ key ~ '|') in '|description|default|title|examples|properties|patternProperties|additionalProperties|returns|' -%}
505
- {%- elif value is mapping -%}
506
- {%- for uk, uv in value | items -%}
507
- {%- if value_contains_mapping(uv) == 'true' -%}{%- set RB.ok = false -%}{%- endif -%}
508
- {%- endfor -%}
509
- {%- elif value is sequence and value is not string -%}
510
- {%- if value_contains_mapping(value) == 'true' -%}{%- set RB.ok = false -%}{%- endif -%}
511
- {%- endif -%}
512
- {%- endfor -%}
513
- {%- endif -%}
514
- {%- if spec.properties -%}
515
- {%- for child_name, child_spec in spec.properties | items -%}
516
- {{- validate_schema(child_spec, path + "." + child_name, lenient, classify) -}}
517
- {%- endfor -%}
518
- {%- endif -%}
519
- {%- if 'items' in spec -%}{{- validate_schema(spec['items'], path + "[]", lenient, classify) -}}{%- endif -%}
520
- {%- if spec.oneOf -%}
521
- {%- for variant in spec.oneOf -%}{{- validate_schema(variant, path + ".oneOf[" + (loop.index0 | string) + "]", lenient, classify, true) -}}{%- endfor -%}
522
- {%- endif -%}
523
- {%- if spec.anyOf -%}
524
- {%- for variant in spec.anyOf -%}{{- validate_schema(variant, path + ".anyOf[" + (loop.index0 | string) + "]", lenient, classify, true) -}}{%- endfor -%}
525
- {%- endif -%}
526
- {%- if spec.additionalProperties is mapping -%}{{- validate_schema(spec.additionalProperties, path + ".additionalProperties", lenient, classify) -}}{%- endif -%}
527
- {%- if spec.patternProperties is mapping -%}
528
- {%- for pattern, pattern_spec in spec.patternProperties | items -%}
529
- {{- validate_schema(pattern_spec, path + ".patternProperties[" + pattern + "]", lenient, classify) -}}
530
- {%- endfor -%}
531
- {%- endif -%}
532
- {%- if spec.returns is mapping -%}{{- validate_schema(spec.returns, path + ".returns", lenient, classify) -}}{%- endif -%}
533
- {%- endif -%}
534
- {%- endmacro -%}
535
-
536
- {%- macro validate_tools(tools_list, classify=true) -%}
537
- {%- set RB.bad = '|' -%}
538
- {%- for tool in tools_list -%}
539
- {%- set fn = tool.function if tool.function is defined else tool -%}
540
- {%- set RB.ok = true -%}
541
- {%- if fn.parameters is defined and fn.parameters is string -%}
542
- {{- raise_exception("tool.function.parameters must be a dict, not a JSON string. Parse it before passing to the template.") -}}
543
- {%- endif -%}
544
- {%- if fn.parameters is not defined or fn.parameters is none -%}
545
- {%- if fn.arguments is defined -%}
546
- {{- raise_exception("Tool '" + fn.name + "' has 'arguments' instead of 'parameters'. Rename 'arguments' to 'parameters'.") -}}
547
- {%- else -%}
548
- {{- raise_exception("Tool '" + fn.name + "' is missing required 'parameters' field. Each tool must have a 'parameters' dict with 'type', 'properties', and 'required' keys.") -}}
549
- {%- endif -%}
550
- {%- endif -%}
551
- {{- validate_schema(fn.parameters, "tool." + fn.name + ".parameters", false, classify) -}}
552
- {%- if classify -%}
553
- {%- if fn.parameters is mapping -%}
554
- {#- unknown container-valued keys at the parameters ROOT are never rendered -#}
555
- {#- by the pretty path (root extras are dropped) -> verbatim fallback. -#}
556
- {%- for rk, rv in fn.parameters | items -%}
557
- {%- if rk not in ['type', 'description', 'enum', 'default', 'properties', 'required', 'optional', 'title', 'items', 'oneOf', 'anyOf', 'additionalProperties', 'patternProperties', 'returns', 'examples', '$defs', 'definitions', '$ref'] -%}
558
- {%- if rv is mapping or (rv is sequence and rv is not string) -%}{%- set RB.ok = false -%}{%- endif -%}
559
- {%- endif -%}
560
- {%- endfor -%}
561
- {%- else -%}
562
- {%- set RB.ok = false -%}
563
- {%- endif -%}
564
- {%- endif -%}
565
- {%- if fn.returns is mapping -%}{{- validate_schema(fn.returns, "tool." + fn.name + ".returns", false, classify) -}}{%- endif -%}
566
- {%- if classify and fn.returns is not defined and fn.response is mapping -%}{{- validate_schema(fn.response, "tool." + fn.name + ".response", true) -}}{%- endif -%}
567
- {#- unknown container-valued keys at the FUNCTION level are never rendered -> fallback. -#}
568
- {%- if classify -%}
569
- {%- for fk, fv in fn | items -%}
570
- {%- if fk not in ['name', 'description', 'parameters', 'returns', 'response', 'type', 'function'] -%}
571
- {%- if fv is mapping or (fv is sequence and fv is not string) -%}{%- set RB.ok = false -%}{%- endif -%}
572
- {%- endif -%}
573
- {%- endfor -%}
574
- {%- endif -%}
575
- {%- if not RB.ok -%}{%- set RB.bad = RB.bad ~ loop.index0 ~ '|' -%}{%- endif -%}
576
- {%- endfor -%}
577
- {%- endmacro -%}
578
-
579
- {%- macro render_tools_json(tools_list) -%}
580
- {{- "<ifm|tools>" }}
581
- {%- for tool in tools_list %}
582
- {{- "\n" }}
583
- {{- tool | tojson }}
584
- {%- endfor %}
585
- {{- "\n</ifm|tools>" }}
586
- {%- endmacro -%}
587
-
588
- {%- macro render_xml_schema_attrs(spec, include_value_attrs) -%}
589
- {%- if spec is mapping -%}
590
- {%- set structural_keys = ["type", "description", "enum", "default", "properties", "required", "items", "oneOf", "anyOf", "additionalProperties", "patternProperties", "returns"] -%}
591
- {%- if include_value_attrs and spec.enum -%}{{- " enum=" }}{{ render_xml_enum(spec.enum) }}{%- endif -%}
592
- {%- if include_value_attrs and spec.default is defined -%}{{ render_xml_default_attr(spec.default) }}{%- endif -%}
593
- {%- if spec.additionalProperties is defined and spec.additionalProperties is not mapping -%}{{ render_xml_attr("additionalProperties", spec.additionalProperties) }}{%- endif -%}
594
- {%- if spec.patternProperties is defined and spec.patternProperties is not mapping -%}{{ render_xml_attr("patternProperties", spec.patternProperties) }}{%- endif -%}
595
- {%- for key, value in spec | items -%}
596
- {%- if key not in structural_keys -%}
597
- {{ render_xml_attr(key, value) }}
598
- {%- endif -%}
599
- {%- endfor -%}
600
- {%- endif -%}
601
- {%- endmacro -%}
602
-
603
- {%- macro xml_schema_has_children(spec, include_properties, include_description) -%}
604
- {%- if spec is not mapping -%}
605
- false
606
- {%- elif (include_description and spec.description is defined) or (include_properties and spec.properties) or 'items' in spec or spec.oneOf or spec.anyOf or spec.additionalProperties is mapping or spec.patternProperties is mapping or spec.returns is defined -%}
607
- true
608
- {%- else -%}
609
- false
610
- {%- endif -%}
611
- {%- endmacro -%}
612
-
613
- {%- macro render_xml_schema_node(tag, spec, include_properties) -%}
614
- {%- if spec is mapping and spec['$ref'] is string -%}
615
- {%- set _r = spec['$ref'] -%}
616
- {%- set _k = _r[8:] if _r.startswith('#/$defs/') else (_r[14:] if _r.startswith('#/definitions/') else none) -%}
617
- {%- if _k is not none and ('|' + _k + '|') not in REFS.seen and REFS.defs[_k] is mapping -%}
618
- {%- set spec = dict((REFS.defs[_k] | items | list) + (spec | items | rejectattr('0', 'equalto', '$ref') | list)) -%}
619
- {%- set REFS.seen = REFS.seen + _k + '|' -%}
620
- {%- if spec['$ref'] is string -%}
621
- {%- set _r2 = spec['$ref'] -%}
622
- {%- set _k2 = _r2[8:] if _r2.startswith('#/$defs/') else (_r2[14:] if _r2.startswith('#/definitions/') else none) -%}
623
- {%- if _k2 is not none and REFS.defs[_k2] is mapping -%}
624
- {%- set spec = dict((REFS.defs[_k2] | items | list) + (spec | items | rejectattr('0', 'equalto', '$ref') | list)) -%}
625
- {%- set REFS.seen = REFS.seen + _k2 + '|' -%}
626
- {%- endif -%}
627
- {%- endif -%}
628
- {%- endif -%}
629
- {%- endif -%}
630
- {%- if spec is mapping -%}
631
- {{- "<" + tag + " type=" + render_compact_type(spec) }}{{ render_xml_schema_attrs(spec, true) }}
632
- {%- if xml_schema_has_children(spec, include_properties, true) == 'true' -%}
633
- {{- ">" }}{{ render_xml_schema_children(spec, include_properties, true) }}{{- "</" + tag + ">" }}
634
- {%- else -%}
635
- {{- "/>" }}
636
- {%- endif -%}
637
- {%- else -%}
638
- {{- "<" + tag + ">" }}{{ render_xml_value(spec) }}{{- "</" + tag + ">" }}
639
- {%- endif -%}
640
- {%- endmacro -%}
641
-
642
- {%- macro render_xml_pattern_property(pattern, spec) -%}
643
- {%- if spec is mapping and spec['$ref'] is string -%}
644
- {%- set _r = spec['$ref'] -%}
645
- {%- set _k = _r[8:] if _r.startswith('#/$defs/') else (_r[14:] if _r.startswith('#/definitions/') else none) -%}
646
- {%- if _k is not none and ('|' + _k + '|') not in REFS.seen and REFS.defs[_k] is mapping -%}
647
- {%- set spec = dict((REFS.defs[_k] | items | list) + (spec | items | rejectattr('0', 'equalto', '$ref') | list)) -%}
648
- {%- set REFS.seen = REFS.seen + _k + '|' -%}
649
- {%- if spec['$ref'] is string -%}
650
- {%- set _r2 = spec['$ref'] -%}
651
- {%- set _k2 = _r2[8:] if _r2.startswith('#/$defs/') else (_r2[14:] if _r2.startswith('#/definitions/') else none) -%}
652
- {%- if _k2 is not none and REFS.defs[_k2] is mapping -%}
653
- {%- set spec = dict((REFS.defs[_k2] | items | list) + (spec | items | rejectattr('0', 'equalto', '$ref') | list)) -%}
654
- {%- set REFS.seen = REFS.seen + _k2 + '|' -%}
655
- {%- endif -%}
656
- {%- endif -%}
657
- {%- endif -%}
658
- {%- endif -%}
659
- {%- if spec is mapping -%}
660
- {{- "<patternProperty" }}{{ render_xml_attr("pattern", pattern) }}{{- " type=" + render_compact_type(spec) }}{{ render_xml_schema_attrs(spec, true) }}
661
- {%- if xml_schema_has_children(spec, true, true) == 'true' -%}
662
- {{- ">" }}{{ render_xml_schema_children(spec, true, true) }}{{- "</patternProperty>" }}
663
- {%- else -%}
664
- {{- "/>" }}
665
- {%- endif -%}
666
- {%- else -%}
667
- {{- "<patternProperty" }}{{ render_xml_attr("pattern", pattern) }}{{- ">" }}{{ render_xml_value(spec) }}{{- "</patternProperty>" }}
668
- {%- endif -%}
669
- {%- endmacro -%}
670
-
671
- {%- macro render_xml_schema_children(spec, include_properties, include_description) -%}
672
- {%- if include_description and spec.description is defined -%}{{- "<description>" }}{{ spec.description }}{{- "</description>" }}{%- endif -%}
673
- {%- if include_properties and spec.properties -%}
674
- {%- for child_name, child_spec in spec.properties | items -%}
675
- {{- render_xml_param(child_name, child_spec, spec.required or []) }}
676
- {%- endfor -%}
677
- {%- endif -%}
678
- {%- if 'items' in spec -%}{{ render_xml_schema_node("items", spec['items'], true) }}{%- endif -%}
679
- {%- if spec.oneOf -%}
680
- {{- "<oneOf>" }}
681
- {%- for variant in spec.oneOf -%}{{ render_xml_schema_node("variant", variant, true) }}{%- endfor -%}
682
- {{- "</oneOf>" }}
683
- {%- endif -%}
684
- {%- if spec.anyOf -%}
685
- {{- "<anyOf>" }}
686
- {%- for variant in spec.anyOf -%}{{ render_xml_schema_node("variant", variant, true) }}{%- endfor -%}
687
- {{- "</anyOf>" }}
688
- {%- endif -%}
689
- {%- if spec.additionalProperties is mapping -%}{{ render_xml_schema_node("additionalProperties", spec.additionalProperties, true) }}{%- endif -%}
690
- {%- if spec.patternProperties is mapping -%}
691
- {{- "<patternProperties>" }}
692
- {%- for pattern, pattern_spec in spec.patternProperties | items -%}{{ render_xml_pattern_property(pattern, pattern_spec) }}{%- endfor -%}
693
- {{- "</patternProperties>" }}
694
- {%- elif spec.patternProperties is defined -%}<patternProperties>{{ render_xml_value(spec.patternProperties) }}</patternProperties>{%- endif -%}
695
- {%- if spec.returns is mapping -%}{{ render_xml_schema_node("returns", spec.returns, true) }}{%- elif spec.returns is defined -%}<returns>{{ render_xml_value(spec.returns) }}</returns>{%- endif -%}
696
- {%- endmacro -%}
697
-
698
- {%- macro render_xml_param(name, spec, required_list) -%}
699
- {%- if spec is mapping and spec['$ref'] is string -%}
700
- {%- set _r = spec['$ref'] -%}
701
- {%- set _k = _r[8:] if _r.startswith('#/$defs/') else (_r[14:] if _r.startswith('#/definitions/') else none) -%}
702
- {%- if _k is not none and ('|' + _k + '|') not in REFS.seen and REFS.defs[_k] is mapping -%}
703
- {%- set spec = dict((REFS.defs[_k] | items | list) + (spec | items | rejectattr('0', 'equalto', '$ref') | list)) -%}
704
- {%- set REFS.seen = REFS.seen + _k + '|' -%}
705
- {%- if spec['$ref'] is string -%}
706
- {%- set _r2 = spec['$ref'] -%}
707
- {%- set _k2 = _r2[8:] if _r2.startswith('#/$defs/') else (_r2[14:] if _r2.startswith('#/definitions/') else none) -%}
708
- {%- if _k2 is not none and REFS.defs[_k2] is mapping -%}
709
- {%- set spec = dict((REFS.defs[_k2] | items | list) + (spec | items | rejectattr('0', 'equalto', '$ref') | list)) -%}
710
- {%- set REFS.seen = REFS.seen + _k2 + '|' -%}
711
- {%- endif -%}
712
- {%- endif -%}
713
- {%- endif -%}
714
- {%- endif -%}
715
- {{- "<param name=" + name + " type=" + render_compact_type(spec) }}
716
- {%- if name in (required_list or []) -%}{{- " required=true" }}{%- endif -%}
717
- {%- if spec.enum -%}{{- " enum=" }}{{ render_xml_enum(spec.enum) }}{%- endif -%}
718
- {%- if spec.default is defined -%}{{ render_xml_default_attr(spec.default) }}{%- endif -%}
719
- {{- render_xml_schema_attrs(spec, false) }}
720
- {%- if spec.description or xml_schema_has_children(spec, true, false) == 'true' -%}
721
- {{- ">" }}
722
- {%- if spec.description -%}{{ spec.description }}{%- endif -%}
723
- {{- render_xml_schema_children(spec, true, false) }}
724
- {{- "</param>" }}
725
- {%- else -%}
726
- {{- "/>" }}
727
- {%- endif -%}
728
- {%- endmacro -%}
729
-
730
- {%- macro render_tools_xml(tools_list) -%}
731
- {{- "<ifm|tools>" }}
732
- {%- for tool in tools_list -%}
733
- {%- set fn = tool.function if tool.function is defined else tool -%}
734
- {%- set REFS.defs = fn.parameters['$defs'] if (fn.parameters is mapping and fn.parameters['$defs'] is mapping) else (fn.parameters['definitions'] if (fn.parameters is mapping and fn.parameters['definitions'] is mapping) else {}) -%}
735
- {%- set REFS.seen = '|' -%}
736
- {%- set fnp = namespace(p=fn.parameters) -%}
737
- {%- if fnp.p is mapping and fnp.p['$ref'] is string -%}
738
- {%- set _r = fnp.p['$ref'] -%}
739
- {%- set _k = _r[8:] if _r.startswith('#/$defs/') else (_r[14:] if _r.startswith('#/definitions/') else none) -%}
740
- {%- if _k is not none and REFS.defs[_k] is mapping -%}
741
- {%- set fnp.p = dict((REFS.defs[_k] | items | list) + (fnp.p | items | rejectattr('0', 'equalto', '$ref') | list)) -%}
742
- {%- set REFS.seen = REFS.seen + _k + '|' -%}
743
- {%- endif -%}
744
- {%- endif -%}
745
- {{- "\n<function name=" + fn.name + ">" }}
746
- {%- if fn.description -%}
747
- {{- "<description>" }}{{ fn.description }}{{- "</description>" }}
748
- {%- endif -%}
749
- {{- "<parameters>" }}
750
- {%- if fnp.p and fnp.p.properties -%}
751
- {%- for pname, pspec in fnp.p.properties | items -%}
752
- {{- render_xml_param(pname, pspec, fnp.p.required or []) }}
753
- {%- endfor -%}
754
- {%- elif fnp.p is mapping and (fnp.p.oneOf or fnp.p.anyOf or 'items' in fnp.p) -%}
755
- {{- render_xml_schema_children(fnp.p, true, false) }}
756
- {%- endif -%}
757
- {{- "</parameters>" }}
758
- {%- set fn_ret = fn.returns if fn.returns is defined else fn.response -%}
759
- {%- if fn_ret is mapping -%}{{ render_xml_schema_node("returns", fn_ret, true) }}{%- elif fn_ret is defined -%}<returns>{{ render_xml_value(fn_ret) }}</returns>{%- endif -%}
760
- {{- "</function>" }}
761
- {%- endfor -%}
762
- {{- "\n</ifm|tools>" }}
763
- {%- endmacro -%}
764
-
765
- {%- macro render_markdown_literal(value) -%}
766
- {%- if value is string and value == "" -%}""
767
- {%- elif value is string -%}`{{ value | replace("\n", "\\n") }}`
768
- {%- else -%}`{{ render_python_repr(value) }}`
769
- {%- endif -%}
770
- {%- endmacro -%}
771
-
772
- {%- macro render_allowed_values(values) -%}
773
- {%- for value in values -%}{{ render_markdown_literal(value) }}{% if not loop.last %}, {% endif %}{%- endfor -%}
774
- {%- endmacro -%}
775
-
776
- {%- macro render_markdown_value(value) -%}
777
- {%- if value is string and value == "" -%}""{%- elif value is string -%}{{ value }}{%- else -%}{{ render_python_repr(value) }}{%- endif -%}
778
- {%- endmacro -%}
779
-
780
- {%- macro render_markdown_detail(indent, label, value) -%}
781
- {{- "\n" + indent + " - " + label + ": " }}{{ render_markdown_value(value) }}
782
- {%- endmacro -%}
783
-
784
- {%- macro render_markdown_metadata_detail(label, value) -%}
785
- {{- "\n- " + label + ": " }}{{ render_markdown_value(value) }}
786
- {%- endmacro -%}
787
-
788
- {%- macro render_markdown_schema_annotations(spec, indent, include_value_details) -%}
789
- {%- if include_value_details and spec.description is defined -%}{{ render_markdown_detail(indent, "Description", spec.description | replace("\n", "\n" + indent + " ")) }}{%- endif -%}
790
- {%- if include_value_details and spec.enum is defined -%}{{- "\n" + indent + " - Allowed values: " }}{{ render_allowed_values(spec.enum) }}{%- endif -%}
791
- {%- if include_value_details and spec.default is defined -%}{{- "\n" + indent + " - Default: " }}{{ render_markdown_literal(spec.default) }}{%- endif -%}
792
- {%- if spec.additionalProperties is defined -%}
793
- {%- if spec.additionalProperties is mapping -%}
794
- {{- "\n" + indent + " - Additional properties *(" + render_markdown_type(spec.additionalProperties) + ")*" }}
795
- {{- render_markdown_schema_details(spec.additionalProperties, indent + " ", true) }}
796
- {%- else -%}
797
- {{ render_markdown_detail(indent, "Additional properties", spec.additionalProperties) }}
798
- {%- endif -%}
799
- {%- endif -%}
800
- {%- endmacro -%}
801
-
802
- {%- macro render_markdown_metadata_annotations(spec) -%}
803
- {%- if spec.description is defined -%}{{ render_markdown_metadata_detail("Description", spec.description | replace("\n", "\n ")) }}{%- endif -%}
804
- {%- if spec.enum is defined -%}{{- "\n- Allowed values: " }}{{ render_allowed_values(spec.enum) }}{%- endif -%}
805
- {%- if spec.default is defined -%}{{- "\n- Default: " }}{{ render_markdown_literal(spec.default) }}{%- endif -%}
806
- {%- if spec.additionalProperties is defined -%}
807
- {%- if spec.additionalProperties is mapping -%}
808
- {{- "\n- Additional properties *(" + render_markdown_type(spec.additionalProperties) + ")*" }}
809
- {{- render_markdown_schema_details(spec.additionalProperties, "", true) }}
810
- {%- else -%}
811
- {{ render_markdown_metadata_detail("Additional properties", spec.additionalProperties) }}
812
- {%- endif -%}
813
- {%- endif -%}
814
- {%- endmacro -%}
815
-
816
- {%- macro render_markdown_schema_extras(spec, indent) -%}
817
- {%- set rendered_keys = ["type", "description", "enum", "default", "properties", "required", "items", "oneOf", "anyOf", "additionalProperties", "patternProperties", "returns"] -%}
818
- {%- for key, value in spec | items -%}
819
- {%- if key not in rendered_keys -%}
820
- {{- "\n" + indent + " - " + key + ": " }}{{ render_markdown_value(value) }}
821
- {%- endif -%}
822
- {%- endfor -%}
823
- {%- endmacro -%}
824
-
825
- {%- macro render_markdown_metadata_extras(spec) -%}
826
- {%- set rendered_keys = ["type", "description", "enum", "default", "properties", "required", "items", "oneOf", "anyOf", "additionalProperties", "patternProperties", "returns"] -%}
827
- {%- for key, value in spec | items -%}
828
- {%- if key not in rendered_keys -%}
829
- {{- "\n- " + key + ": " }}{{ render_markdown_value(value) }}
830
- {%- endif -%}
831
- {%- endfor -%}
832
- {%- endmacro -%}
833
-
834
- {%- macro markdown_schema_has_extra(spec) -%}
835
- {%- set rendered_keys = ["type", "description", "enum", "default", "properties", "required", "items", "oneOf", "anyOf", "additionalProperties", "patternProperties", "returns"] -%}
836
- {%- set found = namespace(value='false') -%}
837
- {%- for key, value in spec | items -%}
838
- {%- if key not in rendered_keys -%}{%- set found.value = 'true' -%}{%- endif -%}
839
- {%- endfor -%}
840
- {{- found.value -}}
841
- {%- endmacro -%}
842
-
843
- {%- macro markdown_parameter_schema_has_details(spec) -%}
844
- {%- if spec.description is defined or spec.enum is defined or spec.default is defined or spec.additionalProperties is defined or spec.patternProperties is defined or 'items' in spec or spec.oneOf or spec.anyOf or spec.returns is defined or markdown_schema_has_extra(spec) == 'true' -%}
845
- true
846
- {%- else -%}
847
- false
848
- {%- endif -%}
849
- {%- endmacro -%}
850
-
851
- {%- macro render_markdown_schema_structure(spec, indent, include_properties) -%}
852
- {%- if include_properties and spec.properties -%}
853
- {%- for child_name, child_spec in spec.properties | items -%}
854
- {{- render_markdown_param(child_name, child_spec, spec.required or [], indent + " ") }}
855
- {%- endfor -%}
856
- {%- endif -%}
857
- {%- if 'items' in spec and spec['items'] is mapping -%}
858
- {{- "\n" + indent + " - Items *(" + render_markdown_type(spec['items']) + ")*" }}
859
- {{- render_markdown_schema_details(spec['items'], indent + " ", true) }}
860
- {%- elif 'items' in spec -%}
861
- {{ render_markdown_detail(indent, "Items", spec['items']) }}
862
- {%- endif -%}
863
- {%- if spec.oneOf -%}
864
- {{- "\n" + indent + " - oneOf:" }}
865
- {%- for variant in spec.oneOf -%}
866
- {{- "\n" + indent + " - Variant " }}{{ loop.index }}{{- " *(" + render_markdown_type(variant) + ")*" }}
867
- {{- render_markdown_schema_details(variant, indent + " ", true) }}
868
- {%- endfor -%}
869
- {%- endif -%}
870
- {%- if spec.anyOf -%}
871
- {{- "\n" + indent + " - anyOf:" }}
872
- {%- for variant in spec.anyOf -%}
873
- {{- "\n" + indent + " - Variant " }}{{ loop.index }}{{- " *(" + render_markdown_type(variant) + ")*" }}
874
- {{- render_markdown_schema_details(variant, indent + " ", true) }}
875
- {%- endfor -%}
876
- {%- endif -%}
877
- {%- if spec.patternProperties is mapping -%}
878
- {{- "\n" + indent + " - Pattern properties:" }}
879
- {%- for pattern, pattern_spec in spec.patternProperties | items -%}
880
- {%- if pattern_spec is mapping -%}
881
- {{- "\n" + indent + " - `" + pattern + "` *(" + render_markdown_type(pattern_spec) + ")*" }}
882
- {{- render_markdown_schema_details(pattern_spec, indent + " ", true) }}
883
- {%- else -%}
884
- {{- "\n" + indent + " - `" + pattern + "`: " }}{{ render_markdown_value(pattern_spec) }}
885
- {%- endif -%}
886
- {%- endfor -%}
887
- {%- elif spec.patternProperties is defined -%}
888
- {{ render_markdown_detail(indent, "Pattern properties", spec.patternProperties) }}
889
- {%- endif -%}
890
- {%- if spec.returns is mapping -%}
891
- {{- "\n" + indent + " - Returns *(" + render_markdown_type(spec.returns) + ")*" }}
892
- {{- render_markdown_schema_details(spec.returns, indent + " ", true) }}
893
- {%- elif spec.returns is defined -%}
894
- {{ render_markdown_detail(indent, "Returns", spec.returns) }}
895
- {%- endif -%}
896
- {%- endmacro -%}
897
-
898
- {%- macro render_markdown_schema_details(spec, indent, include_value_details) -%}
899
- {%- if spec is mapping and spec['$ref'] is string -%}
900
- {%- set _r = spec['$ref'] -%}
901
- {%- set _k = _r[8:] if _r.startswith('#/$defs/') else (_r[14:] if _r.startswith('#/definitions/') else none) -%}
902
- {%- if _k is not none and ('|' + _k + '|') not in REFS.seen and REFS.defs[_k] is mapping -%}
903
- {%- set spec = dict((REFS.defs[_k] | items | list) + (spec | items | rejectattr('0', 'equalto', '$ref') | list)) -%}
904
- {%- set REFS.seen = REFS.seen + _k + '|' -%}
905
- {%- if spec['$ref'] is string -%}
906
- {%- set _r2 = spec['$ref'] -%}
907
- {%- set _k2 = _r2[8:] if _r2.startswith('#/$defs/') else (_r2[14:] if _r2.startswith('#/definitions/') else none) -%}
908
- {%- if _k2 is not none and REFS.defs[_k2] is mapping -%}
909
- {%- set spec = dict((REFS.defs[_k2] | items | list) + (spec | items | rejectattr('0', 'equalto', '$ref') | list)) -%}
910
- {%- set REFS.seen = REFS.seen + _k2 + '|' -%}
911
- {%- endif -%}
912
- {%- endif -%}
913
- {%- endif -%}
914
- {%- endif -%}
915
- {%- if spec is mapping -%}
916
- {{- render_markdown_schema_annotations(spec, indent, include_value_details) }}
917
- {{- render_markdown_schema_structure(spec, indent, true) }}
918
- {{- render_markdown_schema_extras(spec, indent) }}
919
- {%- elif spec is not sameas true and spec is not sameas false -%}
920
- {{- "\n" + indent + " - Value: " }}{{ render_markdown_literal(spec) }}
921
- {%- endif -%}
922
- {%- endmacro -%}
923
-
924
- {%- macro render_markdown_parameter_schema(spec) -%}
925
- {%- if spec is mapping and spec['$ref'] is string -%}
926
- {%- set _r = spec['$ref'] -%}
927
- {%- set _k = _r[8:] if _r.startswith('#/$defs/') else (_r[14:] if _r.startswith('#/definitions/') else none) -%}
928
- {%- if _k is not none and ('|' + _k + '|') not in REFS.seen and REFS.defs[_k] is mapping -%}
929
- {%- set spec = dict((REFS.defs[_k] | items | list) + (spec | items | rejectattr('0', 'equalto', '$ref') | list)) -%}
930
- {%- set REFS.seen = REFS.seen + _k + '|' -%}
931
- {%- if spec['$ref'] is string -%}
932
- {%- set _r2 = spec['$ref'] -%}
933
- {%- set _k2 = _r2[8:] if _r2.startswith('#/$defs/') else (_r2[14:] if _r2.startswith('#/definitions/') else none) -%}
934
- {%- if _k2 is not none and REFS.defs[_k2] is mapping -%}
935
- {%- set spec = dict((REFS.defs[_k2] | items | list) + (spec | items | rejectattr('0', 'equalto', '$ref') | list)) -%}
936
- {%- set REFS.seen = REFS.seen + _k2 + '|' -%}
937
- {%- endif -%}
938
- {%- endif -%}
939
- {%- endif -%}
940
- {%- endif -%}
941
- {%- if spec is mapping -%}
942
- {{- render_markdown_metadata_annotations(spec) }}
943
- {%- if 'items' in spec and spec['items'] is mapping -%}
944
- {{- "\n- Items *(" + render_markdown_type(spec['items']) + ")*" }}
945
- {{- render_markdown_schema_details(spec['items'], "", true) }}
946
- {%- elif 'items' in spec -%}
947
- {{ render_markdown_metadata_detail("Items", spec['items']) }}
948
- {%- endif -%}
949
- {%- if spec.oneOf -%}
950
- {{- "\n- oneOf:" }}
951
- {%- for variant in spec.oneOf -%}
952
- {{- "\n - Variant " }}{{ loop.index }}{{- " *(" + render_markdown_type(variant) + ")*" }}
953
- {{- render_markdown_schema_details(variant, " ", true) }}
954
- {%- endfor -%}
955
- {%- endif -%}
956
- {%- if spec.anyOf -%}
957
- {{- "\n- anyOf:" }}
958
- {%- for variant in spec.anyOf -%}
959
- {{- "\n - Variant " }}{{ loop.index }}{{- " *(" + render_markdown_type(variant) + ")*" }}
960
- {{- render_markdown_schema_details(variant, " ", true) }}
961
- {%- endfor -%}
962
- {%- endif -%}
963
- {%- if spec.patternProperties is mapping -%}
964
- {{- "\n- Pattern properties:" }}
965
- {%- for pattern, pattern_spec in spec.patternProperties | items -%}
966
- {%- if pattern_spec is mapping -%}
967
- {{- "\n - `" + pattern + "` *(" + render_markdown_type(pattern_spec) + ")*" }}
968
- {{- render_markdown_schema_details(pattern_spec, " ", true) }}
969
- {%- else -%}
970
- {{- "\n - `" + pattern + "`: " }}{{ render_markdown_value(pattern_spec) }}
971
- {%- endif -%}
972
- {%- endfor -%}
973
- {%- elif spec.patternProperties is defined -%}
974
- {{ render_markdown_metadata_detail("Pattern properties", spec.patternProperties) }}
975
- {%- endif -%}
976
- {%- if spec.returns is mapping -%}
977
- {{- "\n- Returns *(" + render_markdown_type(spec.returns) + ")*" }}
978
- {{- render_markdown_schema_details(spec.returns, "", true) }}
979
- {%- elif spec.returns is defined -%}
980
- {{ render_markdown_metadata_detail("Returns", spec.returns) }}
981
- {%- endif -%}
982
- {{- render_markdown_metadata_extras(spec) }}
983
- {%- endif -%}
984
- {%- endmacro -%}
985
-
986
- {%- macro render_markdown_param(name, spec, required_list, indent) -%}
987
- {%- if spec is mapping and spec['$ref'] is string -%}
988
- {%- set _r = spec['$ref'] -%}
989
- {%- set _k = _r[8:] if _r.startswith('#/$defs/') else (_r[14:] if _r.startswith('#/definitions/') else none) -%}
990
- {%- if _k is not none and ('|' + _k + '|') not in REFS.seen and REFS.defs[_k] is mapping -%}
991
- {%- set spec = dict((REFS.defs[_k] | items | list) + (spec | items | rejectattr('0', 'equalto', '$ref') | list)) -%}
992
- {%- set REFS.seen = REFS.seen + _k + '|' -%}
993
- {%- if spec['$ref'] is string -%}
994
- {%- set _r2 = spec['$ref'] -%}
995
- {%- set _k2 = _r2[8:] if _r2.startswith('#/$defs/') else (_r2[14:] if _r2.startswith('#/definitions/') else none) -%}
996
- {%- if _k2 is not none and REFS.defs[_k2] is mapping -%}
997
- {%- set spec = dict((REFS.defs[_k2] | items | list) + (spec | items | rejectattr('0', 'equalto', '$ref') | list)) -%}
998
- {%- set REFS.seen = REFS.seen + _k2 + '|' -%}
999
- {%- endif -%}
1000
- {%- endif -%}
1001
- {%- endif -%}
1002
- {%- endif -%}
1003
- {{- "\n" + indent + "- `" + name + "` *(" + render_markdown_type(spec) }}
1004
- {%- if name in (required_list or []) -%}{{- ", required" }}{%- endif -%}
1005
- {{- ")*" }}
1006
- {%- if spec.description -%}{{- " - " + spec.description | replace("\n", "\n" + indent + " ") }}{%- endif -%}
1007
- {%- if spec.enum -%}
1008
- {{- "\n" + indent + " - Allowed values: " }}{{ render_allowed_values(spec.enum) }}
1009
- {%- endif -%}
1010
- {%- if spec.default is defined -%}
1011
- {{- "\n" + indent + " - Default: " }}{{ render_markdown_literal(spec.default) }}
1012
- {%- endif -%}
1013
- {{- render_markdown_schema_details(spec, indent, false) }}
1014
- {%- endmacro -%}
1015
-
1016
- {%- macro render_tools_markdown(tools_list) -%}
1017
- {{- "<ifm|tools>" }}
1018
- {%- for tool in tools_list -%}
1019
- {%- set fn = tool.function if tool.function is defined else tool -%}
1020
- {%- set REFS.defs = fn.parameters['$defs'] if (fn.parameters is mapping and fn.parameters['$defs'] is mapping) else (fn.parameters['definitions'] if (fn.parameters is mapping and fn.parameters['definitions'] is mapping) else {}) -%}
1021
- {%- set REFS.seen = '|' -%}
1022
- {%- set fnp = namespace(p=fn.parameters) -%}
1023
- {%- if fnp.p is mapping and fnp.p['$ref'] is string -%}
1024
- {%- set _r = fnp.p['$ref'] -%}
1025
- {%- set _k = _r[8:] if _r.startswith('#/$defs/') else (_r[14:] if _r.startswith('#/definitions/') else none) -%}
1026
- {%- if _k is not none and REFS.defs[_k] is mapping -%}
1027
- {%- set fnp.p = dict((REFS.defs[_k] | items | list) + (fnp.p | items | rejectattr('0', 'equalto', '$ref') | list)) -%}
1028
- {%- set REFS.seen = REFS.seen + _k + '|' -%}
1029
- {%- endif -%}
1030
- {%- endif -%}
1031
- {{- "\n## " + fn.name }}
1032
- {%- if fn.description -%}
1033
- {{- "\n" + fn.description }}
1034
- {%- endif -%}
1035
- {{- "\n\n**Parameters**" }}
1036
- {%- if fnp.p and fnp.p.properties -%}
1037
- {%- for pname, pspec in fnp.p.properties | items -%}
1038
- {{- render_markdown_param(pname, pspec, fnp.p.required or [], "") }}
1039
- {%- endfor -%}
1040
- {%- elif fnp.p is mapping and (fnp.p.oneOf or fnp.p.anyOf or 'items' in fnp.p) -%}
1041
- {{- render_markdown_parameter_schema(fnp.p) }}
1042
- {%- else -%}
1043
- {{- "\n- None" }}
1044
- {%- endif -%}
1045
- {%- set fn_ret = fn.returns if fn.returns is defined else fn.response -%}
1046
- {%- if fn_ret is mapping -%}
1047
- {{- "\n\n**Returns**" }}
1048
- {{- "\n- Return *(" + render_markdown_type(fn_ret) + ")*" }}
1049
- {{- render_markdown_schema_details(fn_ret, "", true) }}
1050
- {%- elif fn_ret is defined -%}
1051
- {{- "\n\n**Returns**\n- " }}{{ render_markdown_value(fn_ret) }}
1052
- {%- endif -%}
1053
- {%- if not loop.last -%}{{- "\n" }}{%- endif -%}
1054
- {%- endfor -%}
1055
- {{- "\n</ifm|tools>" }}
1056
- {%- endmacro -%}
1057
-
1058
- {%- macro render_tool_presentation(tools_list, fmt) -%}
1059
- {%- if fmt == 'json' -%}
1060
- {{- render_tools_json(tools_list) }}
1061
- {%- elif RB.bad != '|' -%}
1062
- {#- some tool uses constructs the pretty renderers cannot represent (verdicts -#}
1063
- {#- computed during validate_tools): render the WHOLE toolset exactly as the -#}
1064
- {#- json presentation would, so the block stays uniform and model-familiar. -#}
1065
- {{- render_tools_json(tools_list) }}
1066
- {%- elif fmt == 'xml' -%}
1067
- {{- render_tools_xml(tools_list) }}
1068
- {%- elif fmt == 'markdown' -%}
1069
- {{- render_tools_markdown(tools_list) }}
1070
- {%- else -%}
1071
- {{- raise_exception("Unsupported tool_presentation_format: '" + fmt + "'. Supported formats: json, xml, markdown.") }}
1072
- {%- endif -%}
1073
- {%- endmacro -%}
1074
-
1075
- {%- macro render_call_instructions(fmt) -%}
1076
- {%- if fmt == 'json' -%}
1077
- {{- "Wrap all tool calls in a single <ifm|tool_calls></ifm|tool_calls> block. For each call, emit one JSON object with the function name and arguments on the same line inside <ifm|tool_call></ifm|tool_call> tags:\n\n<ifm|tool_calls>\n<ifm|tool_call>{\"name\": <function-name>, \"arguments\": <args-json-object>}</ifm|tool_call>\n</ifm|tool_calls>" }}
1078
- {%- elif fmt == 'xml' -%}
1079
- {{- "Wrap all tool calls in a single <ifm|tool_calls></ifm|tool_calls> block. For each call, write the function name at the start of <ifm|tool_call>, followed by paired <ifm|arg_key> and <ifm|arg_value> tags for each argument:\n\n<ifm|tool_calls>\n<ifm|tool_call>$FUNCTION_NAME\n<ifm|arg_key>$PARAMETER_NAME</ifm|arg_key>\n<ifm|arg_value>$PARAMETER_VALUE</ifm|arg_value>\n...\n</ifm|tool_call>\n</ifm|tool_calls>\n\nString and scalar parameters should be written as plain text. Array and object parameters should be written as JSON literals." }}
1080
- {%- elif fmt == 'xml_typed' -%}
1081
- {{- "Wrap all tool calls in a single <ifm|tool_calls></ifm|tool_calls> block. For each call, write the function name at the start of <ifm|tool_call>, followed by <ifm|arg_key>, <ifm|arg_type>, and <ifm|arg_value> tags for each argument:\n\n<ifm|tool_calls>\n<ifm|tool_call>$FUNCTION_NAME\n<ifm|arg_key>$PARAMETER_NAME</ifm|arg_key>\n<ifm|arg_type>$ARGUMENT_TYPE</ifm|arg_type>\n<ifm|arg_value>$PARAMETER_VALUE</ifm|arg_value>\n...\n</ifm|tool_call>\n</ifm|tool_calls>\n\nUse the parameter type shown in the tool definition. If that type contains anyOf or oneOf, use the actual argument value type instead. String and scalar parameters should be written as plain text. Array and object parameters should be written as JSON literals." }}
1082
- {%- else -%}
1083
- {{- raise_exception("Unsupported tool_call_format: '" + fmt + "'. Supported formats: json, xml, xml_typed.") }}
1084
- {%- endif -%}
1085
- {%- endmacro -%}
1086
-
1087
- {%- macro render_system_with_tools(tools_list, system_content, presentation_fmt, call_fmt) -%}
1088
- {{- "<|ifm|im_start|>system\n# Tools\nYou may call one or more tools to assist with the user query.\n\nAvailable tools are:\n\n" }}
1089
- {{- render_tool_presentation(tools_list, presentation_fmt) }}
1090
- {{- "\n\nWhen calling tools, you MUST follow the tool-call format below:\n\n" }}
1091
- {{- render_call_instructions(call_fmt) }}
1092
- {%- if system_content -%}
1093
- {{- "\n\n" + system_content }}
1094
- {%- endif -%}
1095
- {{- "<|ifm|im_end|>" }}
1096
- {%- endmacro -%}
1097
-
1098
- {%- macro render_argument_value(value) -%}
1099
- {%- if value is string -%}{{- value -}}{%- else -%}{{- value | tojson -}}{%- endif -%}
1100
- {%- endmacro -%}
1101
-
1102
- {%- macro render_value_type(value) -%}
1103
- {%- if value is none -%}null
1104
- {%- elif value is boolean -%}boolean
1105
- {%- elif value is integer -%}integer
1106
- {%- elif value is number -%}number
1107
- {%- elif value is string -%}string
1108
- {%- elif value is mapping -%}object
1109
- {%- elif value is sequence -%}array
1110
- {%- else -%}any
1111
- {%- endif -%}
1112
- {%- endmacro -%}
1113
-
1114
- {%- macro schema_has_combinator(spec) -%}
1115
- {%- if spec.oneOf or spec.anyOf -%}
1116
- true
1117
- {%- elif spec.type is defined and spec.type is sequence and spec.type is not string and spec.type | length > 1 -%}
1118
- true
1119
- {%- elif spec.type == "array" and 'items' in spec -%}
1120
- {{- schema_has_combinator(spec['items']) -}}
1121
- {%- elif spec.properties -%}
1122
- {%- set found = namespace(value='false') -%}
1123
- {%- for child_name, child_spec in spec.properties | items -%}
1124
- {%- if schema_has_combinator(child_spec) == 'true' -%}
1125
- {%- set found.value = 'true' -%}
1126
- {%- endif -%}
1127
- {%- endfor -%}
1128
- {{- found.value -}}
1129
- {%- else -%}
1130
- false
1131
- {%- endif -%}
1132
- {%- endmacro -%}
1133
-
1134
- {%- macro render_arg_type(tools_list, tool_name, arg_name, value) -%}
1135
- {%- set found = namespace(type='any') -%}
1136
- {%- for tool in tools_list -%}
1137
- {%- set fn = tool.function if tool.function is defined else tool -%}
1138
- {%- if fn.name == tool_name and fn.parameters and fn.parameters.properties and arg_name in fn.parameters.properties -%}
1139
- {%- set spec = fn.parameters.properties[arg_name] -%}
1140
- {%- if spec is mapping and spec['$ref'] is string -%}
1141
- {%- set _r = spec['$ref'] -%}
1142
- {%- set _k = _r[8:] if _r.startswith('#/$defs/') else (_r[14:] if _r.startswith('#/definitions/') else none) -%}
1143
- {%- set _d = fn.parameters['$defs'] if fn.parameters['$defs'] is mapping else fn.parameters['definitions'] -%}
1144
- {%- set spec = dict((_d[_k] | items | list) + (spec | items | rejectattr('0', 'equalto', '$ref') | list)) if (_k is not none and _d is mapping and _d[_k] is mapping) else spec -%}
1145
- {%- endif -%}
1146
- {%- if schema_has_combinator(spec) == 'true' -%}
1147
- {%- set found.type = render_value_type(value) -%}
1148
- {%- else -%}
1149
- {%- set found.type = render_compact_type(spec) -%}
1150
- {%- endif -%}
1151
- {%- endif -%}
1152
- {%- endfor -%}
1153
- {{- found.type -}}
1154
- {%- endmacro -%}
1155
-
1156
- {%- macro render_tool_calls_block(tool_calls, fmt, tools_list) -%}
1157
- {{- "<ifm|tool_calls>" }}
1158
- {%- for raw_tool_call in tool_calls -%}
1159
- {%- set tool_call = raw_tool_call.function if raw_tool_call.function else raw_tool_call -%}
1160
- {%- if tool_call.arguments is string -%}
1161
- {{- raise_exception("tool_call.arguments must be a dict, not a JSON string. Parse it before passing to the template.") -}}
1162
- {%- endif -%}
1163
- {%- if fmt == 'json' -%}
1164
- {{- "\n<ifm|tool_call>{\"name\": \"" + tool_call.name + "\", \"arguments\": " }}{{ tool_call.arguments | tojson }}{{- "}</ifm|tool_call>" }}
1165
- {%- elif fmt == 'xml' or fmt == 'xml_typed' -%}
1166
- {{- "\n<ifm|tool_call>" + tool_call.name + "\n" }}
1167
- {%- for key, value in tool_call.arguments | items -%}
1168
- {{- "<ifm|arg_key>" + key + "</ifm|arg_key>\n" }}
1169
- {%- if fmt == 'xml_typed' -%}
1170
- {{- "<ifm|arg_type>" + render_arg_type(tools_list, tool_call.name, key, value) + "</ifm|arg_type>\n" }}
1171
- {%- endif -%}
1172
- {{- "<ifm|arg_value>" }}{{ render_argument_value(value) }}{{- "</ifm|arg_value>\n" }}
1173
- {%- endfor -%}
1174
- {{- "</ifm|tool_call>" }}
1175
- {%- else -%}
1176
- {{- raise_exception("Unsupported tool_call_format: '" + fmt + "'. Supported formats: json, xml, xml_typed.") -}}
1177
- {%- endif -%}
1178
- {%- endfor -%}
1179
- {{- "\n</ifm|tool_calls>" }}
1180
- {%- endmacro -%}
1181
-
1182
- {%- macro render_tool_response_messages(raw_content) -%}
1183
- {%- if raw_content is string -%}
1184
- {{- '<|ifm|im_start|>tool\n' + raw_content + '<|ifm|im_end|>' }}
1185
- {%- elif raw_content is sequence and raw_content is not string and raw_content is not mapping -%}
1186
- {%- if raw_content | length == 0 -%}
1187
- {{- raise_exception("tool message content list must not be empty.") -}}
1188
- {%- endif -%}
1189
- {{- '<|ifm|im_start|>tool\n' -}}
1190
- {%- for item in raw_content -%}
1191
- {%- if not loop.first -%}{{- '\n' -}}{%- endif -%}
1192
- {%- if item is string -%}
1193
- {{- item -}}
1194
- {%- elif item is mapping and item.text is string -%}
1195
- {{- item.text -}}
1196
- {%- else -%}
1197
- {{- (item | tojson) -}}
1198
- {%- endif -%}
1199
- {%- endfor -%}
1200
- {{- '<|ifm|im_end|>' -}}
1201
- {%- else -%}
1202
- {{- '<|ifm|im_start|>tool\n' }}{{ raw_content | tojson }}{{- '<|ifm|im_end|>' }}
1203
- {%- endif -%}
1204
- {%- endmacro -%}
1205
-
1206
- {%- set available_tools = tools if tools else [] -%}
1207
- {%- if (not available_tools) and messages[0].role == 'system' and messages[0].get('tools') -%}
1208
- {%- set available_tools = messages[0]['tools'] -%}
1209
- {%- endif -%}
1210
- {%- if available_tools -%}
1211
- {{- validate_tools(available_tools, tool_presentation_fmt != 'json') }}
1212
- {%- set system_content = '' -%}
1213
- {%- if messages[0].role == 'system' and messages[0].content -%}
1214
- {%- set system_content = messages[0].content -%}
1215
- {%- endif -%}
1216
- {{- render_system_with_tools(available_tools, system_content, tool_presentation_fmt, tool_call_fmt) }}
1217
- {%- else -%}
1218
- {%- if messages[0].role == 'system' -%}
1219
- {{- '<|ifm|im_start|>system\n' + messages[0].content + '<|ifm|im_end|>' }}
1220
- {%- endif -%}
1221
- {%- endif -%}
1222
-
1223
- {%- for message in messages -%}
1224
- {%- if message.content is string -%}
1225
- {%- set content = message.content -%}
1226
- {%- else -%}
1227
- {%- set content = '' -%}
1228
- {%- endif -%}
1229
- {%- if (message.role == "user") or (message.role == "system" and not loop.first) -%}
1230
- {{- '<|ifm|im_start|>' + message.role + '\n' + content + '<|ifm|im_end|>' }}
1231
- {%- elif message.role == "assistant" -%}
1232
- {%- set thinking_content = '' -%}
1233
- {%- set think_tag = '' -%}
1234
- {%- if message.think is defined and message.think is string -%}
1235
- {%- set thinking_content = message.think -%}
1236
- {%- set think_tag = 'ifm|think' -%}
1237
- {%- elif message.think_fast is defined and message.think_fast is string -%}
1238
- {%- set thinking_content = message.think_fast -%}
1239
- {%- set think_tag = 'ifm|think_fast' -%}
1240
- {%- elif message.think_faster is defined and message.think_faster is string -%}
1241
- {%- set thinking_content = message.think_faster -%}
1242
- {%- set think_tag = 'ifm|think_faster' -%}
1243
- {%- else -%}
1244
- {%- if '</ifm|think>' in content -%}
1245
- {%- set thinking_content = content.split('</ifm|think>')[0].rstrip('\n').split('<ifm|think>')[-1].lstrip('\n') -%}
1246
- {%- set content = content.split('</ifm|think>')[-1].lstrip('\n') -%}
1247
- {%- set think_tag = 'ifm|think' -%}
1248
- {%- elif '</ifm|think_fast>' in content -%}
1249
- {%- set thinking_content = content.split('</ifm|think_fast>')[0].rstrip('\n').split('<ifm|think_fast>')[-1].lstrip('\n') -%}
1250
- {%- set content = content.split('</ifm|think_fast>')[-1].lstrip('\n') -%}
1251
- {%- set think_tag = 'ifm|think_fast' -%}
1252
- {%- elif '</ifm|think_faster>' in content -%}
1253
- {%- set thinking_content = content.split('</ifm|think_faster>')[0].rstrip('\n').split('<ifm|think_faster>')[-1].lstrip('\n') -%}
1254
- {%- set content = content.split('</ifm|think_faster>')[-1].lstrip('\n') -%}
1255
- {%- set think_tag = 'ifm|think_faster' -%}
1256
- {%- endif -%}
1257
- {%- endif -%}
1258
- {{- '<|ifm|im_start|>' + message.role }}
1259
- {% generation %}
1260
- {%- if think_tag -%}
1261
- {%- if thinking_content -%}
1262
- {{- '<' + think_tag + '>\n' + thinking_content + '\n</' + think_tag + '>\n' + content.lstrip('\n') }}
1263
- {%- else -%}
1264
- {{- '<' + think_tag + '>\n</' + think_tag + '>\n' + content.lstrip('\n') }}
1265
- {%- endif -%}
1266
- {%- else -%}
1267
- {{- content }}
1268
- {%- endif -%}
1269
- {%- if message.tool_calls -%}
1270
- {%- if content -%}
1271
- {{- '\n' }}
1272
- {%- endif -%}
1273
- {{- render_tool_calls_block(message.tool_calls, tool_call_fmt, available_tools) }}
1274
- {%- endif -%}
1275
- {{- '<|ifm|im_end|>' -}}
1276
- {%- endgeneration -%}
1277
- {%- elif message.role == "tool" -%}
1278
- {{- render_tool_response_messages(message.content) }}
1279
- {%- endif -%}
1280
- {%- endfor -%}
1281
- {%- if add_generation_prompt -%}
1282
- {%- set effort = reasoning_effort | default('high') -%}
1283
- {%- if effort == 'high' -%}
1284
- {{- '<|ifm|im_start|>assistant\n<ifm|think>\n' }}
1285
- {%- elif effort == 'medium' -%}
1286
- {{- '<|ifm|im_start|>assistant\n<ifm|think_fast>\n' }}
1287
- {%- elif effort == 'low' -%}
1288
- {{- '<|ifm|im_start|>assistant\n<ifm|think_faster>\n' }}
1289
- {%- else -%}
1290
- {{- '<|ifm|im_start|>assistant\n<ifm|think_fast>\n' }}
1291
- {%- endif -%}
1292
- {%- endif -%}
1293
-
1294
- INFO:gguf.gguf_writer:Writing the following files:
1295
- INFO:gguf.gguf_writer:source-official-bf16-k2-horizon-0.9b/K2-Horizon-0.9B-BF16.gguf: n_tensors = 255, total_size = 2.2G
1296
- INFO:gguf.gguf_writer:Dry run, not writing files
1297
- source-official-bf16-k2-horizon-0.9b/K2-Horizon-0.9B-BF16.gguf
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
reproducibility/validation/convert.log DELETED
@@ -1,1292 +0,0 @@
1
- INFO:hf-to-gguf:Loading model: source-hf-k2-horizon-0.9b
2
- WARNING:hf-to-gguf:Failed to load model config from source-hf-k2-horizon-0.9b: The repository source-hf-k2-horizon-0.9b contains custom code which must be executed to correctly load the model. You can inspect the repository content at /root/workspace/HF/source-hf-k2-horizon-0.9b .
3
- You can inspect the repository content at https://hf.co/source-hf-k2-horizon-0.9b.
4
- Please pass the argument `trust_remote_code=True` to allow custom code to be run.
5
- WARNING:hf-to-gguf:Trying to load config.json instead
6
- INFO:hf-to-gguf:Model architecture: K2HorizonForCausalLM
7
- WARNING:hf-to-gguf:Failed to load model config from source-hf-k2-horizon-0.9b: The repository source-hf-k2-horizon-0.9b contains custom code which must be executed to correctly load the model. You can inspect the repository content at /root/workspace/HF/source-hf-k2-horizon-0.9b .
8
- You can inspect the repository content at https://hf.co/source-hf-k2-horizon-0.9b.
9
- Please pass the argument `trust_remote_code=True` to allow custom code to be run.
10
- WARNING:hf-to-gguf:Trying to load config.json instead
11
- INFO:hf-to-gguf:gguf: loading model weight map from 'model.safetensors.index.json'
12
- INFO:hf-to-gguf:gguf: indexing model part 'model-00000-of-00001.safetensors'
13
- INFO:gguf.gguf_writer:gguf: This GGUF file is for Little Endian only
14
- INFO:hf-to-gguf:Exporting model...
15
- INFO:hf-to-gguf:output.weight, torch.bfloat16 --> BF16, shape = {1536, 64256}
16
- INFO:hf-to-gguf:token_embd.weight, torch.bfloat16 --> BF16, shape = {1536, 64256}
17
- INFO:hf-to-gguf:blk.0.attn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
18
- INFO:hf-to-gguf:blk.0.ffn_down.weight, torch.bfloat16 --> BF16, shape = {5120, 1536}
19
- INFO:hf-to-gguf:blk.0.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
20
- INFO:hf-to-gguf:blk.0.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
21
- INFO:hf-to-gguf:blk.0.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
22
- INFO:hf-to-gguf:blk.0.attn_k.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
23
- INFO:hf-to-gguf:blk.0.attn_output.weight, torch.bfloat16 --> BF16, shape = {2048, 1536}
24
- INFO:hf-to-gguf:blk.0.attn_q.weight, torch.bfloat16 --> BF16, shape = {1536, 2048}
25
- INFO:hf-to-gguf:blk.0.attn_v.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
26
- INFO:hf-to-gguf:blk.1.attn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
27
- INFO:hf-to-gguf:blk.1.ffn_down.weight, torch.bfloat16 --> BF16, shape = {5120, 1536}
28
- INFO:hf-to-gguf:blk.1.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
29
- INFO:hf-to-gguf:blk.1.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
30
- INFO:hf-to-gguf:blk.1.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
31
- INFO:hf-to-gguf:blk.1.attn_k.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
32
- INFO:hf-to-gguf:blk.1.attn_output.weight, torch.bfloat16 --> BF16, shape = {2048, 1536}
33
- INFO:hf-to-gguf:blk.1.attn_q.weight, torch.bfloat16 --> BF16, shape = {1536, 2048}
34
- INFO:hf-to-gguf:blk.1.attn_v.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
35
- INFO:hf-to-gguf:blk.10.attn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
36
- INFO:hf-to-gguf:blk.10.ffn_down.weight, torch.bfloat16 --> BF16, shape = {5120, 1536}
37
- INFO:hf-to-gguf:blk.10.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
38
- INFO:hf-to-gguf:blk.10.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
39
- INFO:hf-to-gguf:blk.10.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
40
- INFO:hf-to-gguf:blk.10.attn_k.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
41
- INFO:hf-to-gguf:blk.10.attn_output.weight, torch.bfloat16 --> BF16, shape = {2048, 1536}
42
- INFO:hf-to-gguf:blk.10.attn_q.weight, torch.bfloat16 --> BF16, shape = {1536, 2048}
43
- INFO:hf-to-gguf:blk.10.attn_v.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
44
- INFO:hf-to-gguf:blk.11.attn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
45
- INFO:hf-to-gguf:blk.11.ffn_down.weight, torch.bfloat16 --> BF16, shape = {5120, 1536}
46
- INFO:hf-to-gguf:blk.11.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
47
- INFO:hf-to-gguf:blk.11.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
48
- INFO:hf-to-gguf:blk.11.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
49
- INFO:hf-to-gguf:blk.11.attn_k.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
50
- INFO:hf-to-gguf:blk.11.attn_output.weight, torch.bfloat16 --> BF16, shape = {2048, 1536}
51
- INFO:hf-to-gguf:blk.11.attn_q.weight, torch.bfloat16 --> BF16, shape = {1536, 2048}
52
- INFO:hf-to-gguf:blk.11.attn_v.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
53
- INFO:hf-to-gguf:blk.12.attn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
54
- INFO:hf-to-gguf:blk.12.ffn_down.weight, torch.bfloat16 --> BF16, shape = {5120, 1536}
55
- INFO:hf-to-gguf:blk.12.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
56
- INFO:hf-to-gguf:blk.12.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
57
- INFO:hf-to-gguf:blk.12.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
58
- INFO:hf-to-gguf:blk.12.attn_k.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
59
- INFO:hf-to-gguf:blk.12.attn_output.weight, torch.bfloat16 --> BF16, shape = {2048, 1536}
60
- INFO:hf-to-gguf:blk.12.attn_q.weight, torch.bfloat16 --> BF16, shape = {1536, 2048}
61
- INFO:hf-to-gguf:blk.12.attn_v.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
62
- INFO:hf-to-gguf:blk.13.attn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
63
- INFO:hf-to-gguf:blk.13.ffn_down.weight, torch.bfloat16 --> BF16, shape = {5120, 1536}
64
- INFO:hf-to-gguf:blk.13.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
65
- INFO:hf-to-gguf:blk.13.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
66
- INFO:hf-to-gguf:blk.13.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
67
- INFO:hf-to-gguf:blk.13.attn_k.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
68
- INFO:hf-to-gguf:blk.13.attn_output.weight, torch.bfloat16 --> BF16, shape = {2048, 1536}
69
- INFO:hf-to-gguf:blk.13.attn_q.weight, torch.bfloat16 --> BF16, shape = {1536, 2048}
70
- INFO:hf-to-gguf:blk.13.attn_v.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
71
- INFO:hf-to-gguf:blk.14.attn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
72
- INFO:hf-to-gguf:blk.14.ffn_down.weight, torch.bfloat16 --> BF16, shape = {5120, 1536}
73
- INFO:hf-to-gguf:blk.14.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
74
- INFO:hf-to-gguf:blk.14.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
75
- INFO:hf-to-gguf:blk.14.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
76
- INFO:hf-to-gguf:blk.14.attn_k.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
77
- INFO:hf-to-gguf:blk.14.attn_output.weight, torch.bfloat16 --> BF16, shape = {2048, 1536}
78
- INFO:hf-to-gguf:blk.14.attn_q.weight, torch.bfloat16 --> BF16, shape = {1536, 2048}
79
- INFO:hf-to-gguf:blk.14.attn_v.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
80
- INFO:hf-to-gguf:blk.15.attn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
81
- INFO:hf-to-gguf:blk.15.ffn_down.weight, torch.bfloat16 --> BF16, shape = {5120, 1536}
82
- INFO:hf-to-gguf:blk.15.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
83
- INFO:hf-to-gguf:blk.15.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
84
- INFO:hf-to-gguf:blk.15.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
85
- INFO:hf-to-gguf:blk.15.attn_k.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
86
- INFO:hf-to-gguf:blk.15.attn_output.weight, torch.bfloat16 --> BF16, shape = {2048, 1536}
87
- INFO:hf-to-gguf:blk.15.attn_q.weight, torch.bfloat16 --> BF16, shape = {1536, 2048}
88
- INFO:hf-to-gguf:blk.15.attn_v.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
89
- INFO:hf-to-gguf:blk.16.attn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
90
- INFO:hf-to-gguf:blk.16.ffn_down.weight, torch.bfloat16 --> BF16, shape = {5120, 1536}
91
- INFO:hf-to-gguf:blk.16.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
92
- INFO:hf-to-gguf:blk.16.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
93
- INFO:hf-to-gguf:blk.16.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
94
- INFO:hf-to-gguf:blk.16.attn_k.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
95
- INFO:hf-to-gguf:blk.16.attn_output.weight, torch.bfloat16 --> BF16, shape = {2048, 1536}
96
- INFO:hf-to-gguf:blk.16.attn_q.weight, torch.bfloat16 --> BF16, shape = {1536, 2048}
97
- INFO:hf-to-gguf:blk.16.attn_v.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
98
- INFO:hf-to-gguf:blk.17.attn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
99
- INFO:hf-to-gguf:blk.17.ffn_down.weight, torch.bfloat16 --> BF16, shape = {5120, 1536}
100
- INFO:hf-to-gguf:blk.17.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
101
- INFO:hf-to-gguf:blk.17.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
102
- INFO:hf-to-gguf:blk.17.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
103
- INFO:hf-to-gguf:blk.17.attn_k.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
104
- INFO:hf-to-gguf:blk.17.attn_output.weight, torch.bfloat16 --> BF16, shape = {2048, 1536}
105
- INFO:hf-to-gguf:blk.17.attn_q.weight, torch.bfloat16 --> BF16, shape = {1536, 2048}
106
- INFO:hf-to-gguf:blk.17.attn_v.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
107
- INFO:hf-to-gguf:blk.18.attn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
108
- INFO:hf-to-gguf:blk.18.ffn_down.weight, torch.bfloat16 --> BF16, shape = {5120, 1536}
109
- INFO:hf-to-gguf:blk.18.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
110
- INFO:hf-to-gguf:blk.18.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
111
- INFO:hf-to-gguf:blk.18.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
112
- INFO:hf-to-gguf:blk.18.attn_k.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
113
- INFO:hf-to-gguf:blk.18.attn_output.weight, torch.bfloat16 --> BF16, shape = {2048, 1536}
114
- INFO:hf-to-gguf:blk.18.attn_q.weight, torch.bfloat16 --> BF16, shape = {1536, 2048}
115
- INFO:hf-to-gguf:blk.18.attn_v.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
116
- INFO:hf-to-gguf:blk.19.attn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
117
- INFO:hf-to-gguf:blk.19.ffn_down.weight, torch.bfloat16 --> BF16, shape = {5120, 1536}
118
- INFO:hf-to-gguf:blk.19.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
119
- INFO:hf-to-gguf:blk.19.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
120
- INFO:hf-to-gguf:blk.19.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
121
- INFO:hf-to-gguf:blk.19.attn_k.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
122
- INFO:hf-to-gguf:blk.19.attn_output.weight, torch.bfloat16 --> BF16, shape = {2048, 1536}
123
- INFO:hf-to-gguf:blk.19.attn_q.weight, torch.bfloat16 --> BF16, shape = {1536, 2048}
124
- INFO:hf-to-gguf:blk.19.attn_v.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
125
- INFO:hf-to-gguf:blk.2.attn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
126
- INFO:hf-to-gguf:blk.2.ffn_down.weight, torch.bfloat16 --> BF16, shape = {5120, 1536}
127
- INFO:hf-to-gguf:blk.2.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
128
- INFO:hf-to-gguf:blk.2.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
129
- INFO:hf-to-gguf:blk.2.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
130
- INFO:hf-to-gguf:blk.2.attn_k.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
131
- INFO:hf-to-gguf:blk.2.attn_output.weight, torch.bfloat16 --> BF16, shape = {2048, 1536}
132
- INFO:hf-to-gguf:blk.2.attn_q.weight, torch.bfloat16 --> BF16, shape = {1536, 2048}
133
- INFO:hf-to-gguf:blk.2.attn_v.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
134
- INFO:hf-to-gguf:blk.20.attn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
135
- INFO:hf-to-gguf:blk.20.ffn_down.weight, torch.bfloat16 --> BF16, shape = {5120, 1536}
136
- INFO:hf-to-gguf:blk.20.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
137
- INFO:hf-to-gguf:blk.20.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
138
- INFO:hf-to-gguf:blk.20.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
139
- INFO:hf-to-gguf:blk.20.attn_k.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
140
- INFO:hf-to-gguf:blk.20.attn_output.weight, torch.bfloat16 --> BF16, shape = {2048, 1536}
141
- INFO:hf-to-gguf:blk.20.attn_q.weight, torch.bfloat16 --> BF16, shape = {1536, 2048}
142
- INFO:hf-to-gguf:blk.20.attn_v.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
143
- INFO:hf-to-gguf:blk.21.attn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
144
- INFO:hf-to-gguf:blk.21.ffn_down.weight, torch.bfloat16 --> BF16, shape = {5120, 1536}
145
- INFO:hf-to-gguf:blk.21.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
146
- INFO:hf-to-gguf:blk.21.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
147
- INFO:hf-to-gguf:blk.21.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
148
- INFO:hf-to-gguf:blk.21.attn_k.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
149
- INFO:hf-to-gguf:blk.21.attn_output.weight, torch.bfloat16 --> BF16, shape = {2048, 1536}
150
- INFO:hf-to-gguf:blk.21.attn_q.weight, torch.bfloat16 --> BF16, shape = {1536, 2048}
151
- INFO:hf-to-gguf:blk.21.attn_v.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
152
- INFO:hf-to-gguf:blk.22.attn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
153
- INFO:hf-to-gguf:blk.22.ffn_down.weight, torch.bfloat16 --> BF16, shape = {5120, 1536}
154
- INFO:hf-to-gguf:blk.22.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
155
- INFO:hf-to-gguf:blk.22.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
156
- INFO:hf-to-gguf:blk.22.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
157
- INFO:hf-to-gguf:blk.22.attn_k.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
158
- INFO:hf-to-gguf:blk.22.attn_output.weight, torch.bfloat16 --> BF16, shape = {2048, 1536}
159
- INFO:hf-to-gguf:blk.22.attn_q.weight, torch.bfloat16 --> BF16, shape = {1536, 2048}
160
- INFO:hf-to-gguf:blk.22.attn_v.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
161
- INFO:hf-to-gguf:blk.23.attn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
162
- INFO:hf-to-gguf:blk.23.ffn_down.weight, torch.bfloat16 --> BF16, shape = {5120, 1536}
163
- INFO:hf-to-gguf:blk.23.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
164
- INFO:hf-to-gguf:blk.23.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
165
- INFO:hf-to-gguf:blk.23.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
166
- INFO:hf-to-gguf:blk.23.attn_k.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
167
- INFO:hf-to-gguf:blk.23.attn_output.weight, torch.bfloat16 --> BF16, shape = {2048, 1536}
168
- INFO:hf-to-gguf:blk.23.attn_q.weight, torch.bfloat16 --> BF16, shape = {1536, 2048}
169
- INFO:hf-to-gguf:blk.23.attn_v.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
170
- INFO:hf-to-gguf:blk.24.attn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
171
- INFO:hf-to-gguf:blk.24.ffn_down.weight, torch.bfloat16 --> BF16, shape = {5120, 1536}
172
- INFO:hf-to-gguf:blk.24.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
173
- INFO:hf-to-gguf:blk.24.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
174
- INFO:hf-to-gguf:blk.24.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
175
- INFO:hf-to-gguf:blk.24.attn_k.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
176
- INFO:hf-to-gguf:blk.24.attn_output.weight, torch.bfloat16 --> BF16, shape = {2048, 1536}
177
- INFO:hf-to-gguf:blk.24.attn_q.weight, torch.bfloat16 --> BF16, shape = {1536, 2048}
178
- INFO:hf-to-gguf:blk.24.attn_v.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
179
- INFO:hf-to-gguf:blk.25.attn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
180
- INFO:hf-to-gguf:blk.25.ffn_down.weight, torch.bfloat16 --> BF16, shape = {5120, 1536}
181
- INFO:hf-to-gguf:blk.25.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
182
- INFO:hf-to-gguf:blk.25.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
183
- INFO:hf-to-gguf:blk.25.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
184
- INFO:hf-to-gguf:blk.25.attn_k.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
185
- INFO:hf-to-gguf:blk.25.attn_output.weight, torch.bfloat16 --> BF16, shape = {2048, 1536}
186
- INFO:hf-to-gguf:blk.25.attn_q.weight, torch.bfloat16 --> BF16, shape = {1536, 2048}
187
- INFO:hf-to-gguf:blk.25.attn_v.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
188
- INFO:hf-to-gguf:blk.26.attn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
189
- INFO:hf-to-gguf:blk.26.ffn_down.weight, torch.bfloat16 --> BF16, shape = {5120, 1536}
190
- INFO:hf-to-gguf:blk.26.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
191
- INFO:hf-to-gguf:blk.26.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
192
- INFO:hf-to-gguf:blk.26.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
193
- INFO:hf-to-gguf:blk.26.attn_k.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
194
- INFO:hf-to-gguf:blk.26.attn_output.weight, torch.bfloat16 --> BF16, shape = {2048, 1536}
195
- INFO:hf-to-gguf:blk.26.attn_q.weight, torch.bfloat16 --> BF16, shape = {1536, 2048}
196
- INFO:hf-to-gguf:blk.26.attn_v.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
197
- INFO:hf-to-gguf:blk.27.attn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
198
- INFO:hf-to-gguf:blk.27.ffn_down.weight, torch.bfloat16 --> BF16, shape = {5120, 1536}
199
- INFO:hf-to-gguf:blk.27.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
200
- INFO:hf-to-gguf:blk.27.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
201
- INFO:hf-to-gguf:blk.27.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
202
- INFO:hf-to-gguf:blk.27.attn_k.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
203
- INFO:hf-to-gguf:blk.27.attn_output.weight, torch.bfloat16 --> BF16, shape = {2048, 1536}
204
- INFO:hf-to-gguf:blk.27.attn_q.weight, torch.bfloat16 --> BF16, shape = {1536, 2048}
205
- INFO:hf-to-gguf:blk.27.attn_v.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
206
- INFO:hf-to-gguf:blk.3.attn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
207
- INFO:hf-to-gguf:blk.3.ffn_down.weight, torch.bfloat16 --> BF16, shape = {5120, 1536}
208
- INFO:hf-to-gguf:blk.3.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
209
- INFO:hf-to-gguf:blk.3.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
210
- INFO:hf-to-gguf:blk.3.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
211
- INFO:hf-to-gguf:blk.3.attn_k.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
212
- INFO:hf-to-gguf:blk.3.attn_output.weight, torch.bfloat16 --> BF16, shape = {2048, 1536}
213
- INFO:hf-to-gguf:blk.3.attn_q.weight, torch.bfloat16 --> BF16, shape = {1536, 2048}
214
- INFO:hf-to-gguf:blk.3.attn_v.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
215
- INFO:hf-to-gguf:blk.4.attn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
216
- INFO:hf-to-gguf:blk.4.ffn_down.weight, torch.bfloat16 --> BF16, shape = {5120, 1536}
217
- INFO:hf-to-gguf:blk.4.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
218
- INFO:hf-to-gguf:blk.4.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
219
- INFO:hf-to-gguf:blk.4.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
220
- INFO:hf-to-gguf:blk.4.attn_k.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
221
- INFO:hf-to-gguf:blk.4.attn_output.weight, torch.bfloat16 --> BF16, shape = {2048, 1536}
222
- INFO:hf-to-gguf:blk.4.attn_q.weight, torch.bfloat16 --> BF16, shape = {1536, 2048}
223
- INFO:hf-to-gguf:blk.4.attn_v.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
224
- INFO:hf-to-gguf:blk.5.attn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
225
- INFO:hf-to-gguf:blk.5.ffn_down.weight, torch.bfloat16 --> BF16, shape = {5120, 1536}
226
- INFO:hf-to-gguf:blk.5.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
227
- INFO:hf-to-gguf:blk.5.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
228
- INFO:hf-to-gguf:blk.5.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
229
- INFO:hf-to-gguf:blk.5.attn_k.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
230
- INFO:hf-to-gguf:blk.5.attn_output.weight, torch.bfloat16 --> BF16, shape = {2048, 1536}
231
- INFO:hf-to-gguf:blk.5.attn_q.weight, torch.bfloat16 --> BF16, shape = {1536, 2048}
232
- INFO:hf-to-gguf:blk.5.attn_v.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
233
- INFO:hf-to-gguf:blk.6.attn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
234
- INFO:hf-to-gguf:blk.6.ffn_down.weight, torch.bfloat16 --> BF16, shape = {5120, 1536}
235
- INFO:hf-to-gguf:blk.6.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
236
- INFO:hf-to-gguf:blk.6.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
237
- INFO:hf-to-gguf:blk.6.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
238
- INFO:hf-to-gguf:blk.6.attn_k.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
239
- INFO:hf-to-gguf:blk.6.attn_output.weight, torch.bfloat16 --> BF16, shape = {2048, 1536}
240
- INFO:hf-to-gguf:blk.6.attn_q.weight, torch.bfloat16 --> BF16, shape = {1536, 2048}
241
- INFO:hf-to-gguf:blk.6.attn_v.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
242
- INFO:hf-to-gguf:blk.7.attn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
243
- INFO:hf-to-gguf:blk.7.ffn_down.weight, torch.bfloat16 --> BF16, shape = {5120, 1536}
244
- INFO:hf-to-gguf:blk.7.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
245
- INFO:hf-to-gguf:blk.7.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
246
- INFO:hf-to-gguf:blk.7.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
247
- INFO:hf-to-gguf:blk.7.attn_k.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
248
- INFO:hf-to-gguf:blk.7.attn_output.weight, torch.bfloat16 --> BF16, shape = {2048, 1536}
249
- INFO:hf-to-gguf:blk.7.attn_q.weight, torch.bfloat16 --> BF16, shape = {1536, 2048}
250
- INFO:hf-to-gguf:blk.7.attn_v.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
251
- INFO:hf-to-gguf:blk.8.attn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
252
- INFO:hf-to-gguf:blk.8.ffn_down.weight, torch.bfloat16 --> BF16, shape = {5120, 1536}
253
- INFO:hf-to-gguf:blk.8.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
254
- INFO:hf-to-gguf:blk.8.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
255
- INFO:hf-to-gguf:blk.8.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
256
- INFO:hf-to-gguf:blk.8.attn_k.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
257
- INFO:hf-to-gguf:blk.8.attn_output.weight, torch.bfloat16 --> BF16, shape = {2048, 1536}
258
- INFO:hf-to-gguf:blk.8.attn_q.weight, torch.bfloat16 --> BF16, shape = {1536, 2048}
259
- INFO:hf-to-gguf:blk.8.attn_v.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
260
- INFO:hf-to-gguf:blk.9.attn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
261
- INFO:hf-to-gguf:blk.9.ffn_down.weight, torch.bfloat16 --> BF16, shape = {5120, 1536}
262
- INFO:hf-to-gguf:blk.9.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
263
- INFO:hf-to-gguf:blk.9.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1536, 5120}
264
- INFO:hf-to-gguf:blk.9.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1536}
265
- INFO:hf-to-gguf:blk.9.attn_k.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
266
- INFO:hf-to-gguf:blk.9.attn_output.weight, torch.bfloat16 --> BF16, shape = {2048, 1536}
267
- INFO:hf-to-gguf:blk.9.attn_q.weight, torch.bfloat16 --> BF16, shape = {1536, 2048}
268
- INFO:hf-to-gguf:blk.9.attn_v.weight, torch.bfloat16 --> BF16, shape = {1536, 512}
269
- INFO:hf-to-gguf:output_norm.weight, torch.bfloat16 --> F32, shape = {1536}
270
- INFO:hf-to-gguf:Set meta model
271
- INFO:hf-to-gguf:Set model parameters
272
- INFO:hf-to-gguf:gguf: context length = 131072
273
- INFO:hf-to-gguf:gguf: embedding length = 1536
274
- INFO:hf-to-gguf:gguf: feed forward length = 5120
275
- INFO:hf-to-gguf:gguf: head count = 32
276
- INFO:hf-to-gguf:gguf: key-value head count = 8
277
- INFO:hf-to-gguf:gguf: rope scaling type = YARN
278
- INFO:hf-to-gguf:gguf: rope theta = 1000000.0
279
- INFO:hf-to-gguf:gguf: rms norm epsilon = 1e-06
280
- INFO:hf-to-gguf:gguf: expert count = 0
281
- INFO:hf-to-gguf:gguf: experts used count = 0
282
- INFO:hf-to-gguf:gguf: file type = 32
283
- INFO:hf-to-gguf:Set model quantization version
284
- INFO:hf-to-gguf:Set model tokenizer
285
- INFO:gguf.vocab:Adding 63742 merge(s).
286
- INFO:gguf.vocab:Setting special token type bos to 0
287
- INFO:gguf.vocab:Setting special token type eos to 1
288
- INFO:gguf.vocab:Setting special token type pad to 64255
289
- INFO:gguf.vocab:Setting chat_template to {%- if tool_presentation is defined -%}
290
- {{- raise_exception("Unsupported argument: tool_presentation. Use tool_presentation_format with one of: json, xml, markdown.") -}}
291
- {%- endif -%}
292
- {%- if tool_calling_format is defined -%}
293
- {{- raise_exception("Unsupported argument: tool_calling_format. Use tool_call_format with one of: json, xml, xml_typed.") -}}
294
- {%- endif -%}
295
- {%- if tool_format is defined -%}
296
- {{- raise_exception("Unsupported argument: tool_format. Use tool_call_format with one of: json, xml, xml_typed.") -}}
297
- {%- endif -%}
298
- {%- set tool_presentation_fmt = tool_presentation_format | default('markdown') -%}
299
- {%- set tool_call_fmt = tool_call_format | default('xml') -%}
300
- {%- if tool_presentation_fmt != 'json' and tool_presentation_fmt != 'xml' and tool_presentation_fmt != 'markdown' -%}
301
- {{- raise_exception("Unsupported tool_presentation_format: '" ~ tool_presentation_fmt ~ "'. Supported formats: json, xml, markdown.") -}}
302
- {%- endif -%}
303
- {%- if tool_call_fmt != 'json' and tool_call_fmt != 'xml' and tool_call_fmt != 'xml_typed' -%}
304
- {{- raise_exception("Unsupported tool_call_format: '" ~ tool_call_fmt ~ "'. Supported formats: json, xml, xml_typed.") -}}
305
- {%- endif -%}
306
-
307
- {#- Renderability state, computed during validate_tools (single walk, no extra -#}
308
- {#- traversal at render time): ok = working flag for the tool being validated; -#}
309
- {#- bad = pipe-delimited indices of tools that must render as verbatim JSON. -#}
310
- {%- set RB = namespace(ok=true, bad='|') -%}
311
-
312
- {%- macro value_contains_mapping(v) -%}
313
- {%- if v is mapping -%}
314
- true
315
- {%- elif v is sequence and v is not string -%}
316
- {%- set f = namespace(x='false') -%}
317
- {%- for c in v -%}{%- if value_contains_mapping(c) == 'true' -%}{%- set f.x = 'true' -%}{%- endif -%}{%- endfor -%}
318
- {{- f.x -}}
319
- {%- else -%}
320
- false
321
- {%- endif -%}
322
- {%- endmacro -%}
323
-
324
- {#- $ref inlining state: defs = local $defs of the tool being rendered; seen = -#}
325
- {#- pipe-delimited names already expanded for this tool (each def inlines at most -#}
326
- {#- once; later references render by def name; cycles terminate immediately). -#}
327
- {#- $ref-sibling annotations (description/default/...) merge OVER the def at -#}
328
- {#- the inline site, so use-site annotations win and are never dropped. -#}
329
- {%- set REFS = namespace(defs={}, seen='|') -%}
330
-
331
- {%- macro render_compact_type_name(type_name, spec) -%}
332
- {%- if type_name == "array" -%}
333
- array[{%- if 'items' in spec -%}{{ render_compact_type(spec['items']) }}{%- else -%}any{%- endif -%}]
334
- {%- elif type_name -%}
335
- {{- type_name -}}
336
- {%- else -%}
337
- any
338
- {%- endif -%}
339
- {%- endmacro -%}
340
-
341
- {%- macro render_compact_type(spec) -%}
342
- {%- if spec is not mapping -%}
343
- any
344
- {%- elif spec.type is defined and spec.type is sequence and spec.type is not string and spec.type | length > 0 -%}
345
- {%- for type_name in spec.type -%}{{ render_compact_type_name(type_name, spec) }}{%- if not loop.last -%}|{%- endif -%}{%- endfor -%}
346
- {%- elif spec.type is defined and spec.type is sequence and spec.type is not string -%}
347
- any
348
- {%- elif spec.type -%}
349
- {{- render_compact_type_name(spec.type, spec) -}}
350
- {%- elif spec['$ref'] is string -%}
351
- {{- spec['$ref'].split('/') | last -}}
352
- {%- elif spec.oneOf -%}
353
- oneOf[{%- for variant in spec.oneOf -%}{{ render_compact_type(variant) }}{%- if not loop.last -%}|{%- endif -%}{%- endfor -%}]
354
- {%- elif spec.anyOf -%}
355
- anyOf[{%- for variant in spec.anyOf -%}{{ render_compact_type(variant) }}{%- if not loop.last -%}|{%- endif -%}{%- endfor -%}]
356
- {%- elif spec.properties -%}
357
- object
358
- {%- elif 'items' in spec -%}
359
- array[{{ render_compact_type(spec['items']) }}]
360
- {%- else -%}
361
- any
362
- {%- endif -%}
363
- {%- endmacro -%}
364
-
365
- {%- macro render_markdown_type_name(type_name, spec) -%}
366
- {%- if type_name == "array" -%}
367
- array of {% if 'items' in spec %}{{ render_markdown_type(spec['items']) }}{% else %}any{% endif %}
368
- {%- elif type_name -%}
369
- {{- type_name -}}
370
- {%- else -%}
371
- any
372
- {%- endif -%}
373
- {%- endmacro -%}
374
-
375
- {%- macro render_markdown_type(spec) -%}
376
- {%- if spec is sameas true -%}
377
- True
378
- {%- elif spec is sameas false -%}
379
- False
380
- {%- elif spec is not mapping -%}
381
- any
382
- {%- elif spec.type is defined and spec.type is sequence and spec.type is not string and spec.type | length > 0 -%}
383
- {%- for type_name in spec.type -%}{{ render_markdown_type_name(type_name, spec) }}{% if not loop.last %} or {% endif %}{%- endfor -%}
384
- {%- elif spec.type is defined and spec.type is sequence and spec.type is not string -%}
385
- any
386
- {%- elif spec.type -%}
387
- {{- render_markdown_type_name(spec.type, spec) -}}
388
- {%- elif spec['$ref'] is string -%}
389
- {{- spec['$ref'].split('/') | last -}}
390
- {%- elif spec.oneOf -%}
391
- oneOf[{%- for variant in spec.oneOf -%}{{ render_markdown_type(variant) }}{% if not loop.last %} or {% endif %}{%- endfor -%}]
392
- {%- elif spec.anyOf -%}
393
- anyOf[{%- for variant in spec.anyOf -%}{{ render_markdown_type(variant) }}{% if not loop.last %} or {% endif %}{%- endfor -%}]
394
- {%- elif spec.properties -%}
395
- object
396
- {%- elif 'items' in spec -%}
397
- array of {{ render_markdown_type(spec['items']) }}
398
- {%- else -%}
399
- any
400
- {%- endif -%}
401
- {%- endmacro -%}
402
-
403
- {%- macro render_xml_text(value) -%}
404
- {{- value.split() | join(" ") -}}
405
- {%- endmacro -%}
406
-
407
- {%- macro render_python_string(value) -%}
408
- '{{- value.split() | join(" ") | replace("\\", "\\\\") | replace("'", "\\'") -}}'
409
- {%- endmacro -%}
410
-
411
- {%- macro render_python_repr(value) -%}
412
- {%- if value is string -%}
413
- {{ render_python_string(value) }}
414
- {%- elif value is sameas true -%}
415
- True
416
- {%- elif value is sameas false -%}
417
- False
418
- {%- elif value is none -%}
419
- None
420
- {%- elif value is mapping -%}
421
- {{- "{" -}}
422
- {%- for key, child in value | items -%}
423
- {{ render_python_repr(key) }}: {{ render_python_repr(child) }}{%- if not loop.last -%}, {% endif -%}
424
- {%- endfor -%}
425
- {{- "}" -}}
426
- {%- elif value is sequence -%}
427
- {{- "[" -}}
428
- {%- for child in value -%}
429
- {{ render_python_repr(child) }}{%- if not loop.last -%}, {% endif -%}
430
- {%- endfor -%}
431
- {{- "]" -}}
432
- {%- else -%}
433
- {{- value -}}
434
- {%- endif -%}
435
- {%- endmacro -%}
436
-
437
- {%- macro render_xml_value(value) -%}
438
- {%- if value is string -%}{{ render_xml_text(value) }}{%- else -%}{{ render_python_repr(value) }}{%- endif -%}
439
- {%- endmacro -%}
440
-
441
- {%- macro render_xml_enum_value(value) -%}
442
- {%- if value is string -%}"{{- value | replace("\\", "\\\\") | replace("\"", "\\\"") -}}"{%- else -%}"{{- render_python_repr(value) | replace("\\", "\\\\") | replace("\"", "\\\"") -}}"{%- endif -%}
443
- {%- endmacro -%}
444
-
445
- {%- macro render_xml_enum(values) -%}
446
- {%- for value in values -%}{{ render_xml_enum_value(value) }}{%- if not loop.last -%}|{%- endif -%}{%- endfor -%}
447
- {%- endmacro -%}
448
-
449
- {%- macro render_xml_default_attr(value) -%}
450
- {{- " default=" }}{%- if value is string -%}"{{- value | replace("\\", "\\\\") | replace("\"", "\\\"") -}}"{%- else -%}{{ render_xml_value(value) }}{%- endif -%}
451
- {%- endmacro -%}
452
-
453
- {%- macro render_xml_attr(name, value) -%}
454
- {{- " " + name + "=" }}{%- if value == "" -%}""{%- else -%}{{ render_xml_value(value) }}{%- endif -%}
455
- {%- endmacro -%}
456
-
457
- {%- macro validate_schema(spec, path, lenient=false, classify=true, in_variant=false) -%}
458
- {%- if spec is mapping -%}
459
- {%- if not lenient -%}
460
- {%- if spec.required is defined -%}
461
- {%- if spec.required is string or spec.required is not sequence -%}
462
- {{- raise_exception("Schema '" + path + "' has 'required' but it is not a list.") -}}
463
- {%- endif -%}
464
- {%- if spec.required | length > 0 and not spec.properties and not in_variant -%}
465
- {{- raise_exception("Schema '" + path + "' has required fields but no properties object to define them.") -}}
466
- {%- endif -%}
467
- {%- if spec.properties -%}
468
- {%- for required_name in spec.required -%}
469
- {%- if required_name not in spec.properties -%}
470
- {{- raise_exception("Schema '" + path + "' marks '" + required_name + "' as required, but that property is not defined in properties.") -}}
471
- {%- endif -%}
472
- {%- endfor -%}
473
- {%- endif -%}
474
- {%- endif -%}
475
- {%- endif -%}
476
- {#- renderability classification, piggybacking on this walk (no raises here): -#}
477
- {#- constructs the pretty renderer does not fully handle flip RB.ok so the -#}
478
- {#- tool falls back to verbatim JSON. Skipped entirely for json presentation. -#}
479
- {%- if classify -%}
480
- {%- for key, value in spec | items -%}
481
- {%- if key == '$ref' -%}
482
- {%- if value is not string -%}{%- set RB.ok = false -%}
483
- {%- elif not (value.startswith('#/$defs/') or value.startswith('#/definitions/')) -%}{%- set RB.ok = false -%}{%- endif -%}
484
- {%- elif key == '$defs' or key == 'definitions' -%}
485
- {%- if value is mapping -%}
486
- {%- for dk, dv in value | items -%}
487
- {{- validate_schema(dv, path + ".$defs." + dk, true) -}}
488
- {%- endfor -%}
489
- {%- else -%}{%- set RB.ok = false -%}{%- endif -%}
490
- {%- elif key == 'type' -%}
491
- {%- if value is mapping -%}{%- set RB.ok = false -%}{%- endif -%}
492
- {%- elif key == 'enum' -%}
493
- {%- if value is string or value is mapping or value is not sequence -%}{%- set RB.ok = false -%}{%- endif -%}
494
- {%- elif key == 'items' -%}
495
- {#- any items shape renders: mapping structurally, others via repr detail -#}
496
- {%- elif key == 'oneOf' or key == 'anyOf' -%}
497
- {%- if value is mapping or value is string or value is not sequence -%}{%- set RB.ok = false -%}{%- endif -%}
498
- {%- elif key == 'required' -%}
499
- {%- if value and not spec.properties -%}{%- set RB.ok = false -%}{%- endif -%}
500
- {%- elif ('|' ~ key ~ '|') in '|description|default|title|examples|properties|patternProperties|additionalProperties|returns|' -%}
501
- {%- elif value is mapping -%}
502
- {%- for uk, uv in value | items -%}
503
- {%- if value_contains_mapping(uv) == 'true' -%}{%- set RB.ok = false -%}{%- endif -%}
504
- {%- endfor -%}
505
- {%- elif value is sequence and value is not string -%}
506
- {%- if value_contains_mapping(value) == 'true' -%}{%- set RB.ok = false -%}{%- endif -%}
507
- {%- endif -%}
508
- {%- endfor -%}
509
- {%- endif -%}
510
- {%- if spec.properties -%}
511
- {%- for child_name, child_spec in spec.properties | items -%}
512
- {{- validate_schema(child_spec, path + "." + child_name, lenient, classify) -}}
513
- {%- endfor -%}
514
- {%- endif -%}
515
- {%- if 'items' in spec -%}{{- validate_schema(spec['items'], path + "[]", lenient, classify) -}}{%- endif -%}
516
- {%- if spec.oneOf -%}
517
- {%- for variant in spec.oneOf -%}{{- validate_schema(variant, path + ".oneOf[" + (loop.index0 | string) + "]", lenient, classify, true) -}}{%- endfor -%}
518
- {%- endif -%}
519
- {%- if spec.anyOf -%}
520
- {%- for variant in spec.anyOf -%}{{- validate_schema(variant, path + ".anyOf[" + (loop.index0 | string) + "]", lenient, classify, true) -}}{%- endfor -%}
521
- {%- endif -%}
522
- {%- if spec.additionalProperties is mapping -%}{{- validate_schema(spec.additionalProperties, path + ".additionalProperties", lenient, classify) -}}{%- endif -%}
523
- {%- if spec.patternProperties is mapping -%}
524
- {%- for pattern, pattern_spec in spec.patternProperties | items -%}
525
- {{- validate_schema(pattern_spec, path + ".patternProperties[" + pattern + "]", lenient, classify) -}}
526
- {%- endfor -%}
527
- {%- endif -%}
528
- {%- if spec.returns is mapping -%}{{- validate_schema(spec.returns, path + ".returns", lenient, classify) -}}{%- endif -%}
529
- {%- endif -%}
530
- {%- endmacro -%}
531
-
532
- {%- macro validate_tools(tools_list, classify=true) -%}
533
- {%- set RB.bad = '|' -%}
534
- {%- for tool in tools_list -%}
535
- {%- set fn = tool.function if tool.function is defined else tool -%}
536
- {%- set RB.ok = true -%}
537
- {%- if fn.parameters is defined and fn.parameters is string -%}
538
- {{- raise_exception("tool.function.parameters must be a dict, not a JSON string. Parse it before passing to the template.") -}}
539
- {%- endif -%}
540
- {%- if fn.parameters is not defined or fn.parameters is none -%}
541
- {%- if fn.arguments is defined -%}
542
- {{- raise_exception("Tool '" + fn.name + "' has 'arguments' instead of 'parameters'. Rename 'arguments' to 'parameters'.") -}}
543
- {%- else -%}
544
- {{- raise_exception("Tool '" + fn.name + "' is missing required 'parameters' field. Each tool must have a 'parameters' dict with 'type', 'properties', and 'required' keys.") -}}
545
- {%- endif -%}
546
- {%- endif -%}
547
- {{- validate_schema(fn.parameters, "tool." + fn.name + ".parameters", false, classify) -}}
548
- {%- if classify -%}
549
- {%- if fn.parameters is mapping -%}
550
- {#- unknown container-valued keys at the parameters ROOT are never rendered -#}
551
- {#- by the pretty path (root extras are dropped) -> verbatim fallback. -#}
552
- {%- for rk, rv in fn.parameters | items -%}
553
- {%- if rk not in ['type', 'description', 'enum', 'default', 'properties', 'required', 'optional', 'title', 'items', 'oneOf', 'anyOf', 'additionalProperties', 'patternProperties', 'returns', 'examples', '$defs', 'definitions', '$ref'] -%}
554
- {%- if rv is mapping or (rv is sequence and rv is not string) -%}{%- set RB.ok = false -%}{%- endif -%}
555
- {%- endif -%}
556
- {%- endfor -%}
557
- {%- else -%}
558
- {%- set RB.ok = false -%}
559
- {%- endif -%}
560
- {%- endif -%}
561
- {%- if fn.returns is mapping -%}{{- validate_schema(fn.returns, "tool." + fn.name + ".returns", false, classify) -}}{%- endif -%}
562
- {%- if classify and fn.returns is not defined and fn.response is mapping -%}{{- validate_schema(fn.response, "tool." + fn.name + ".response", true) -}}{%- endif -%}
563
- {#- unknown container-valued keys at the FUNCTION level are never rendered -> fallback. -#}
564
- {%- if classify -%}
565
- {%- for fk, fv in fn | items -%}
566
- {%- if fk not in ['name', 'description', 'parameters', 'returns', 'response', 'type', 'function'] -%}
567
- {%- if fv is mapping or (fv is sequence and fv is not string) -%}{%- set RB.ok = false -%}{%- endif -%}
568
- {%- endif -%}
569
- {%- endfor -%}
570
- {%- endif -%}
571
- {%- if not RB.ok -%}{%- set RB.bad = RB.bad ~ loop.index0 ~ '|' -%}{%- endif -%}
572
- {%- endfor -%}
573
- {%- endmacro -%}
574
-
575
- {%- macro render_tools_json(tools_list) -%}
576
- {{- "<ifm|tools>" }}
577
- {%- for tool in tools_list %}
578
- {{- "\n" }}
579
- {{- tool | tojson }}
580
- {%- endfor %}
581
- {{- "\n</ifm|tools>" }}
582
- {%- endmacro -%}
583
-
584
- {%- macro render_xml_schema_attrs(spec, include_value_attrs) -%}
585
- {%- if spec is mapping -%}
586
- {%- set structural_keys = ["type", "description", "enum", "default", "properties", "required", "items", "oneOf", "anyOf", "additionalProperties", "patternProperties", "returns"] -%}
587
- {%- if include_value_attrs and spec.enum -%}{{- " enum=" }}{{ render_xml_enum(spec.enum) }}{%- endif -%}
588
- {%- if include_value_attrs and spec.default is defined -%}{{ render_xml_default_attr(spec.default) }}{%- endif -%}
589
- {%- if spec.additionalProperties is defined and spec.additionalProperties is not mapping -%}{{ render_xml_attr("additionalProperties", spec.additionalProperties) }}{%- endif -%}
590
- {%- if spec.patternProperties is defined and spec.patternProperties is not mapping -%}{{ render_xml_attr("patternProperties", spec.patternProperties) }}{%- endif -%}
591
- {%- for key, value in spec | items -%}
592
- {%- if key not in structural_keys -%}
593
- {{ render_xml_attr(key, value) }}
594
- {%- endif -%}
595
- {%- endfor -%}
596
- {%- endif -%}
597
- {%- endmacro -%}
598
-
599
- {%- macro xml_schema_has_children(spec, include_properties, include_description) -%}
600
- {%- if spec is not mapping -%}
601
- false
602
- {%- elif (include_description and spec.description is defined) or (include_properties and spec.properties) or 'items' in spec or spec.oneOf or spec.anyOf or spec.additionalProperties is mapping or spec.patternProperties is mapping or spec.returns is defined -%}
603
- true
604
- {%- else -%}
605
- false
606
- {%- endif -%}
607
- {%- endmacro -%}
608
-
609
- {%- macro render_xml_schema_node(tag, spec, include_properties) -%}
610
- {%- if spec is mapping and spec['$ref'] is string -%}
611
- {%- set _r = spec['$ref'] -%}
612
- {%- set _k = _r[8:] if _r.startswith('#/$defs/') else (_r[14:] if _r.startswith('#/definitions/') else none) -%}
613
- {%- if _k is not none and ('|' + _k + '|') not in REFS.seen and REFS.defs[_k] is mapping -%}
614
- {%- set spec = dict((REFS.defs[_k] | items | list) + (spec | items | rejectattr('0', 'equalto', '$ref') | list)) -%}
615
- {%- set REFS.seen = REFS.seen + _k + '|' -%}
616
- {%- if spec['$ref'] is string -%}
617
- {%- set _r2 = spec['$ref'] -%}
618
- {%- set _k2 = _r2[8:] if _r2.startswith('#/$defs/') else (_r2[14:] if _r2.startswith('#/definitions/') else none) -%}
619
- {%- if _k2 is not none and REFS.defs[_k2] is mapping -%}
620
- {%- set spec = dict((REFS.defs[_k2] | items | list) + (spec | items | rejectattr('0', 'equalto', '$ref') | list)) -%}
621
- {%- set REFS.seen = REFS.seen + _k2 + '|' -%}
622
- {%- endif -%}
623
- {%- endif -%}
624
- {%- endif -%}
625
- {%- endif -%}
626
- {%- if spec is mapping -%}
627
- {{- "<" + tag + " type=" + render_compact_type(spec) }}{{ render_xml_schema_attrs(spec, true) }}
628
- {%- if xml_schema_has_children(spec, include_properties, true) == 'true' -%}
629
- {{- ">" }}{{ render_xml_schema_children(spec, include_properties, true) }}{{- "</" + tag + ">" }}
630
- {%- else -%}
631
- {{- "/>" }}
632
- {%- endif -%}
633
- {%- else -%}
634
- {{- "<" + tag + ">" }}{{ render_xml_value(spec) }}{{- "</" + tag + ">" }}
635
- {%- endif -%}
636
- {%- endmacro -%}
637
-
638
- {%- macro render_xml_pattern_property(pattern, spec) -%}
639
- {%- if spec is mapping and spec['$ref'] is string -%}
640
- {%- set _r = spec['$ref'] -%}
641
- {%- set _k = _r[8:] if _r.startswith('#/$defs/') else (_r[14:] if _r.startswith('#/definitions/') else none) -%}
642
- {%- if _k is not none and ('|' + _k + '|') not in REFS.seen and REFS.defs[_k] is mapping -%}
643
- {%- set spec = dict((REFS.defs[_k] | items | list) + (spec | items | rejectattr('0', 'equalto', '$ref') | list)) -%}
644
- {%- set REFS.seen = REFS.seen + _k + '|' -%}
645
- {%- if spec['$ref'] is string -%}
646
- {%- set _r2 = spec['$ref'] -%}
647
- {%- set _k2 = _r2[8:] if _r2.startswith('#/$defs/') else (_r2[14:] if _r2.startswith('#/definitions/') else none) -%}
648
- {%- if _k2 is not none and REFS.defs[_k2] is mapping -%}
649
- {%- set spec = dict((REFS.defs[_k2] | items | list) + (spec | items | rejectattr('0', 'equalto', '$ref') | list)) -%}
650
- {%- set REFS.seen = REFS.seen + _k2 + '|' -%}
651
- {%- endif -%}
652
- {%- endif -%}
653
- {%- endif -%}
654
- {%- endif -%}
655
- {%- if spec is mapping -%}
656
- {{- "<patternProperty" }}{{ render_xml_attr("pattern", pattern) }}{{- " type=" + render_compact_type(spec) }}{{ render_xml_schema_attrs(spec, true) }}
657
- {%- if xml_schema_has_children(spec, true, true) == 'true' -%}
658
- {{- ">" }}{{ render_xml_schema_children(spec, true, true) }}{{- "</patternProperty>" }}
659
- {%- else -%}
660
- {{- "/>" }}
661
- {%- endif -%}
662
- {%- else -%}
663
- {{- "<patternProperty" }}{{ render_xml_attr("pattern", pattern) }}{{- ">" }}{{ render_xml_value(spec) }}{{- "</patternProperty>" }}
664
- {%- endif -%}
665
- {%- endmacro -%}
666
-
667
- {%- macro render_xml_schema_children(spec, include_properties, include_description) -%}
668
- {%- if include_description and spec.description is defined -%}{{- "<description>" }}{{ spec.description }}{{- "</description>" }}{%- endif -%}
669
- {%- if include_properties and spec.properties -%}
670
- {%- for child_name, child_spec in spec.properties | items -%}
671
- {{- render_xml_param(child_name, child_spec, spec.required or []) }}
672
- {%- endfor -%}
673
- {%- endif -%}
674
- {%- if 'items' in spec -%}{{ render_xml_schema_node("items", spec['items'], true) }}{%- endif -%}
675
- {%- if spec.oneOf -%}
676
- {{- "<oneOf>" }}
677
- {%- for variant in spec.oneOf -%}{{ render_xml_schema_node("variant", variant, true) }}{%- endfor -%}
678
- {{- "</oneOf>" }}
679
- {%- endif -%}
680
- {%- if spec.anyOf -%}
681
- {{- "<anyOf>" }}
682
- {%- for variant in spec.anyOf -%}{{ render_xml_schema_node("variant", variant, true) }}{%- endfor -%}
683
- {{- "</anyOf>" }}
684
- {%- endif -%}
685
- {%- if spec.additionalProperties is mapping -%}{{ render_xml_schema_node("additionalProperties", spec.additionalProperties, true) }}{%- endif -%}
686
- {%- if spec.patternProperties is mapping -%}
687
- {{- "<patternProperties>" }}
688
- {%- for pattern, pattern_spec in spec.patternProperties | items -%}{{ render_xml_pattern_property(pattern, pattern_spec) }}{%- endfor -%}
689
- {{- "</patternProperties>" }}
690
- {%- elif spec.patternProperties is defined -%}<patternProperties>{{ render_xml_value(spec.patternProperties) }}</patternProperties>{%- endif -%}
691
- {%- if spec.returns is mapping -%}{{ render_xml_schema_node("returns", spec.returns, true) }}{%- elif spec.returns is defined -%}<returns>{{ render_xml_value(spec.returns) }}</returns>{%- endif -%}
692
- {%- endmacro -%}
693
-
694
- {%- macro render_xml_param(name, spec, required_list) -%}
695
- {%- if spec is mapping and spec['$ref'] is string -%}
696
- {%- set _r = spec['$ref'] -%}
697
- {%- set _k = _r[8:] if _r.startswith('#/$defs/') else (_r[14:] if _r.startswith('#/definitions/') else none) -%}
698
- {%- if _k is not none and ('|' + _k + '|') not in REFS.seen and REFS.defs[_k] is mapping -%}
699
- {%- set spec = dict((REFS.defs[_k] | items | list) + (spec | items | rejectattr('0', 'equalto', '$ref') | list)) -%}
700
- {%- set REFS.seen = REFS.seen + _k + '|' -%}
701
- {%- if spec['$ref'] is string -%}
702
- {%- set _r2 = spec['$ref'] -%}
703
- {%- set _k2 = _r2[8:] if _r2.startswith('#/$defs/') else (_r2[14:] if _r2.startswith('#/definitions/') else none) -%}
704
- {%- if _k2 is not none and REFS.defs[_k2] is mapping -%}
705
- {%- set spec = dict((REFS.defs[_k2] | items | list) + (spec | items | rejectattr('0', 'equalto', '$ref') | list)) -%}
706
- {%- set REFS.seen = REFS.seen + _k2 + '|' -%}
707
- {%- endif -%}
708
- {%- endif -%}
709
- {%- endif -%}
710
- {%- endif -%}
711
- {{- "<param name=" + name + " type=" + render_compact_type(spec) }}
712
- {%- if name in (required_list or []) -%}{{- " required=true" }}{%- endif -%}
713
- {%- if spec.enum -%}{{- " enum=" }}{{ render_xml_enum(spec.enum) }}{%- endif -%}
714
- {%- if spec.default is defined -%}{{ render_xml_default_attr(spec.default) }}{%- endif -%}
715
- {{- render_xml_schema_attrs(spec, false) }}
716
- {%- if spec.description or xml_schema_has_children(spec, true, false) == 'true' -%}
717
- {{- ">" }}
718
- {%- if spec.description -%}{{ spec.description }}{%- endif -%}
719
- {{- render_xml_schema_children(spec, true, false) }}
720
- {{- "</param>" }}
721
- {%- else -%}
722
- {{- "/>" }}
723
- {%- endif -%}
724
- {%- endmacro -%}
725
-
726
- {%- macro render_tools_xml(tools_list) -%}
727
- {{- "<ifm|tools>" }}
728
- {%- for tool in tools_list -%}
729
- {%- set fn = tool.function if tool.function is defined else tool -%}
730
- {%- set REFS.defs = fn.parameters['$defs'] if (fn.parameters is mapping and fn.parameters['$defs'] is mapping) else (fn.parameters['definitions'] if (fn.parameters is mapping and fn.parameters['definitions'] is mapping) else {}) -%}
731
- {%- set REFS.seen = '|' -%}
732
- {%- set fnp = namespace(p=fn.parameters) -%}
733
- {%- if fnp.p is mapping and fnp.p['$ref'] is string -%}
734
- {%- set _r = fnp.p['$ref'] -%}
735
- {%- set _k = _r[8:] if _r.startswith('#/$defs/') else (_r[14:] if _r.startswith('#/definitions/') else none) -%}
736
- {%- if _k is not none and REFS.defs[_k] is mapping -%}
737
- {%- set fnp.p = dict((REFS.defs[_k] | items | list) + (fnp.p | items | rejectattr('0', 'equalto', '$ref') | list)) -%}
738
- {%- set REFS.seen = REFS.seen + _k + '|' -%}
739
- {%- endif -%}
740
- {%- endif -%}
741
- {{- "\n<function name=" + fn.name + ">" }}
742
- {%- if fn.description -%}
743
- {{- "<description>" }}{{ fn.description }}{{- "</description>" }}
744
- {%- endif -%}
745
- {{- "<parameters>" }}
746
- {%- if fnp.p and fnp.p.properties -%}
747
- {%- for pname, pspec in fnp.p.properties | items -%}
748
- {{- render_xml_param(pname, pspec, fnp.p.required or []) }}
749
- {%- endfor -%}
750
- {%- elif fnp.p is mapping and (fnp.p.oneOf or fnp.p.anyOf or 'items' in fnp.p) -%}
751
- {{- render_xml_schema_children(fnp.p, true, false) }}
752
- {%- endif -%}
753
- {{- "</parameters>" }}
754
- {%- set fn_ret = fn.returns if fn.returns is defined else fn.response -%}
755
- {%- if fn_ret is mapping -%}{{ render_xml_schema_node("returns", fn_ret, true) }}{%- elif fn_ret is defined -%}<returns>{{ render_xml_value(fn_ret) }}</returns>{%- endif -%}
756
- {{- "</function>" }}
757
- {%- endfor -%}
758
- {{- "\n</ifm|tools>" }}
759
- {%- endmacro -%}
760
-
761
- {%- macro render_markdown_literal(value) -%}
762
- {%- if value is string and value == "" -%}""
763
- {%- elif value is string -%}`{{ value | replace("\n", "\\n") }}`
764
- {%- else -%}`{{ render_python_repr(value) }}`
765
- {%- endif -%}
766
- {%- endmacro -%}
767
-
768
- {%- macro render_allowed_values(values) -%}
769
- {%- for value in values -%}{{ render_markdown_literal(value) }}{% if not loop.last %}, {% endif %}{%- endfor -%}
770
- {%- endmacro -%}
771
-
772
- {%- macro render_markdown_value(value) -%}
773
- {%- if value is string and value == "" -%}""{%- elif value is string -%}{{ value }}{%- else -%}{{ render_python_repr(value) }}{%- endif -%}
774
- {%- endmacro -%}
775
-
776
- {%- macro render_markdown_detail(indent, label, value) -%}
777
- {{- "\n" + indent + " - " + label + ": " }}{{ render_markdown_value(value) }}
778
- {%- endmacro -%}
779
-
780
- {%- macro render_markdown_metadata_detail(label, value) -%}
781
- {{- "\n- " + label + ": " }}{{ render_markdown_value(value) }}
782
- {%- endmacro -%}
783
-
784
- {%- macro render_markdown_schema_annotations(spec, indent, include_value_details) -%}
785
- {%- if include_value_details and spec.description is defined -%}{{ render_markdown_detail(indent, "Description", spec.description | replace("\n", "\n" + indent + " ")) }}{%- endif -%}
786
- {%- if include_value_details and spec.enum is defined -%}{{- "\n" + indent + " - Allowed values: " }}{{ render_allowed_values(spec.enum) }}{%- endif -%}
787
- {%- if include_value_details and spec.default is defined -%}{{- "\n" + indent + " - Default: " }}{{ render_markdown_literal(spec.default) }}{%- endif -%}
788
- {%- if spec.additionalProperties is defined -%}
789
- {%- if spec.additionalProperties is mapping -%}
790
- {{- "\n" + indent + " - Additional properties *(" + render_markdown_type(spec.additionalProperties) + ")*" }}
791
- {{- render_markdown_schema_details(spec.additionalProperties, indent + " ", true) }}
792
- {%- else -%}
793
- {{ render_markdown_detail(indent, "Additional properties", spec.additionalProperties) }}
794
- {%- endif -%}
795
- {%- endif -%}
796
- {%- endmacro -%}
797
-
798
- {%- macro render_markdown_metadata_annotations(spec) -%}
799
- {%- if spec.description is defined -%}{{ render_markdown_metadata_detail("Description", spec.description | replace("\n", "\n ")) }}{%- endif -%}
800
- {%- if spec.enum is defined -%}{{- "\n- Allowed values: " }}{{ render_allowed_values(spec.enum) }}{%- endif -%}
801
- {%- if spec.default is defined -%}{{- "\n- Default: " }}{{ render_markdown_literal(spec.default) }}{%- endif -%}
802
- {%- if spec.additionalProperties is defined -%}
803
- {%- if spec.additionalProperties is mapping -%}
804
- {{- "\n- Additional properties *(" + render_markdown_type(spec.additionalProperties) + ")*" }}
805
- {{- render_markdown_schema_details(spec.additionalProperties, "", true) }}
806
- {%- else -%}
807
- {{ render_markdown_metadata_detail("Additional properties", spec.additionalProperties) }}
808
- {%- endif -%}
809
- {%- endif -%}
810
- {%- endmacro -%}
811
-
812
- {%- macro render_markdown_schema_extras(spec, indent) -%}
813
- {%- set rendered_keys = ["type", "description", "enum", "default", "properties", "required", "items", "oneOf", "anyOf", "additionalProperties", "patternProperties", "returns"] -%}
814
- {%- for key, value in spec | items -%}
815
- {%- if key not in rendered_keys -%}
816
- {{- "\n" + indent + " - " + key + ": " }}{{ render_markdown_value(value) }}
817
- {%- endif -%}
818
- {%- endfor -%}
819
- {%- endmacro -%}
820
-
821
- {%- macro render_markdown_metadata_extras(spec) -%}
822
- {%- set rendered_keys = ["type", "description", "enum", "default", "properties", "required", "items", "oneOf", "anyOf", "additionalProperties", "patternProperties", "returns"] -%}
823
- {%- for key, value in spec | items -%}
824
- {%- if key not in rendered_keys -%}
825
- {{- "\n- " + key + ": " }}{{ render_markdown_value(value) }}
826
- {%- endif -%}
827
- {%- endfor -%}
828
- {%- endmacro -%}
829
-
830
- {%- macro markdown_schema_has_extra(spec) -%}
831
- {%- set rendered_keys = ["type", "description", "enum", "default", "properties", "required", "items", "oneOf", "anyOf", "additionalProperties", "patternProperties", "returns"] -%}
832
- {%- set found = namespace(value='false') -%}
833
- {%- for key, value in spec | items -%}
834
- {%- if key not in rendered_keys -%}{%- set found.value = 'true' -%}{%- endif -%}
835
- {%- endfor -%}
836
- {{- found.value -}}
837
- {%- endmacro -%}
838
-
839
- {%- macro markdown_parameter_schema_has_details(spec) -%}
840
- {%- if spec.description is defined or spec.enum is defined or spec.default is defined or spec.additionalProperties is defined or spec.patternProperties is defined or 'items' in spec or spec.oneOf or spec.anyOf or spec.returns is defined or markdown_schema_has_extra(spec) == 'true' -%}
841
- true
842
- {%- else -%}
843
- false
844
- {%- endif -%}
845
- {%- endmacro -%}
846
-
847
- {%- macro render_markdown_schema_structure(spec, indent, include_properties) -%}
848
- {%- if include_properties and spec.properties -%}
849
- {%- for child_name, child_spec in spec.properties | items -%}
850
- {{- render_markdown_param(child_name, child_spec, spec.required or [], indent + " ") }}
851
- {%- endfor -%}
852
- {%- endif -%}
853
- {%- if 'items' in spec and spec['items'] is mapping -%}
854
- {{- "\n" + indent + " - Items *(" + render_markdown_type(spec['items']) + ")*" }}
855
- {{- render_markdown_schema_details(spec['items'], indent + " ", true) }}
856
- {%- elif 'items' in spec -%}
857
- {{ render_markdown_detail(indent, "Items", spec['items']) }}
858
- {%- endif -%}
859
- {%- if spec.oneOf -%}
860
- {{- "\n" + indent + " - oneOf:" }}
861
- {%- for variant in spec.oneOf -%}
862
- {{- "\n" + indent + " - Variant " }}{{ loop.index }}{{- " *(" + render_markdown_type(variant) + ")*" }}
863
- {{- render_markdown_schema_details(variant, indent + " ", true) }}
864
- {%- endfor -%}
865
- {%- endif -%}
866
- {%- if spec.anyOf -%}
867
- {{- "\n" + indent + " - anyOf:" }}
868
- {%- for variant in spec.anyOf -%}
869
- {{- "\n" + indent + " - Variant " }}{{ loop.index }}{{- " *(" + render_markdown_type(variant) + ")*" }}
870
- {{- render_markdown_schema_details(variant, indent + " ", true) }}
871
- {%- endfor -%}
872
- {%- endif -%}
873
- {%- if spec.patternProperties is mapping -%}
874
- {{- "\n" + indent + " - Pattern properties:" }}
875
- {%- for pattern, pattern_spec in spec.patternProperties | items -%}
876
- {%- if pattern_spec is mapping -%}
877
- {{- "\n" + indent + " - `" + pattern + "` *(" + render_markdown_type(pattern_spec) + ")*" }}
878
- {{- render_markdown_schema_details(pattern_spec, indent + " ", true) }}
879
- {%- else -%}
880
- {{- "\n" + indent + " - `" + pattern + "`: " }}{{ render_markdown_value(pattern_spec) }}
881
- {%- endif -%}
882
- {%- endfor -%}
883
- {%- elif spec.patternProperties is defined -%}
884
- {{ render_markdown_detail(indent, "Pattern properties", spec.patternProperties) }}
885
- {%- endif -%}
886
- {%- if spec.returns is mapping -%}
887
- {{- "\n" + indent + " - Returns *(" + render_markdown_type(spec.returns) + ")*" }}
888
- {{- render_markdown_schema_details(spec.returns, indent + " ", true) }}
889
- {%- elif spec.returns is defined -%}
890
- {{ render_markdown_detail(indent, "Returns", spec.returns) }}
891
- {%- endif -%}
892
- {%- endmacro -%}
893
-
894
- {%- macro render_markdown_schema_details(spec, indent, include_value_details) -%}
895
- {%- if spec is mapping and spec['$ref'] is string -%}
896
- {%- set _r = spec['$ref'] -%}
897
- {%- set _k = _r[8:] if _r.startswith('#/$defs/') else (_r[14:] if _r.startswith('#/definitions/') else none) -%}
898
- {%- if _k is not none and ('|' + _k + '|') not in REFS.seen and REFS.defs[_k] is mapping -%}
899
- {%- set spec = dict((REFS.defs[_k] | items | list) + (spec | items | rejectattr('0', 'equalto', '$ref') | list)) -%}
900
- {%- set REFS.seen = REFS.seen + _k + '|' -%}
901
- {%- if spec['$ref'] is string -%}
902
- {%- set _r2 = spec['$ref'] -%}
903
- {%- set _k2 = _r2[8:] if _r2.startswith('#/$defs/') else (_r2[14:] if _r2.startswith('#/definitions/') else none) -%}
904
- {%- if _k2 is not none and REFS.defs[_k2] is mapping -%}
905
- {%- set spec = dict((REFS.defs[_k2] | items | list) + (spec | items | rejectattr('0', 'equalto', '$ref') | list)) -%}
906
- {%- set REFS.seen = REFS.seen + _k2 + '|' -%}
907
- {%- endif -%}
908
- {%- endif -%}
909
- {%- endif -%}
910
- {%- endif -%}
911
- {%- if spec is mapping -%}
912
- {{- render_markdown_schema_annotations(spec, indent, include_value_details) }}
913
- {{- render_markdown_schema_structure(spec, indent, true) }}
914
- {{- render_markdown_schema_extras(spec, indent) }}
915
- {%- elif spec is not sameas true and spec is not sameas false -%}
916
- {{- "\n" + indent + " - Value: " }}{{ render_markdown_literal(spec) }}
917
- {%- endif -%}
918
- {%- endmacro -%}
919
-
920
- {%- macro render_markdown_parameter_schema(spec) -%}
921
- {%- if spec is mapping and spec['$ref'] is string -%}
922
- {%- set _r = spec['$ref'] -%}
923
- {%- set _k = _r[8:] if _r.startswith('#/$defs/') else (_r[14:] if _r.startswith('#/definitions/') else none) -%}
924
- {%- if _k is not none and ('|' + _k + '|') not in REFS.seen and REFS.defs[_k] is mapping -%}
925
- {%- set spec = dict((REFS.defs[_k] | items | list) + (spec | items | rejectattr('0', 'equalto', '$ref') | list)) -%}
926
- {%- set REFS.seen = REFS.seen + _k + '|' -%}
927
- {%- if spec['$ref'] is string -%}
928
- {%- set _r2 = spec['$ref'] -%}
929
- {%- set _k2 = _r2[8:] if _r2.startswith('#/$defs/') else (_r2[14:] if _r2.startswith('#/definitions/') else none) -%}
930
- {%- if _k2 is not none and REFS.defs[_k2] is mapping -%}
931
- {%- set spec = dict((REFS.defs[_k2] | items | list) + (spec | items | rejectattr('0', 'equalto', '$ref') | list)) -%}
932
- {%- set REFS.seen = REFS.seen + _k2 + '|' -%}
933
- {%- endif -%}
934
- {%- endif -%}
935
- {%- endif -%}
936
- {%- endif -%}
937
- {%- if spec is mapping -%}
938
- {{- render_markdown_metadata_annotations(spec) }}
939
- {%- if 'items' in spec and spec['items'] is mapping -%}
940
- {{- "\n- Items *(" + render_markdown_type(spec['items']) + ")*" }}
941
- {{- render_markdown_schema_details(spec['items'], "", true) }}
942
- {%- elif 'items' in spec -%}
943
- {{ render_markdown_metadata_detail("Items", spec['items']) }}
944
- {%- endif -%}
945
- {%- if spec.oneOf -%}
946
- {{- "\n- oneOf:" }}
947
- {%- for variant in spec.oneOf -%}
948
- {{- "\n - Variant " }}{{ loop.index }}{{- " *(" + render_markdown_type(variant) + ")*" }}
949
- {{- render_markdown_schema_details(variant, " ", true) }}
950
- {%- endfor -%}
951
- {%- endif -%}
952
- {%- if spec.anyOf -%}
953
- {{- "\n- anyOf:" }}
954
- {%- for variant in spec.anyOf -%}
955
- {{- "\n - Variant " }}{{ loop.index }}{{- " *(" + render_markdown_type(variant) + ")*" }}
956
- {{- render_markdown_schema_details(variant, " ", true) }}
957
- {%- endfor -%}
958
- {%- endif -%}
959
- {%- if spec.patternProperties is mapping -%}
960
- {{- "\n- Pattern properties:" }}
961
- {%- for pattern, pattern_spec in spec.patternProperties | items -%}
962
- {%- if pattern_spec is mapping -%}
963
- {{- "\n - `" + pattern + "` *(" + render_markdown_type(pattern_spec) + ")*" }}
964
- {{- render_markdown_schema_details(pattern_spec, " ", true) }}
965
- {%- else -%}
966
- {{- "\n - `" + pattern + "`: " }}{{ render_markdown_value(pattern_spec) }}
967
- {%- endif -%}
968
- {%- endfor -%}
969
- {%- elif spec.patternProperties is defined -%}
970
- {{ render_markdown_metadata_detail("Pattern properties", spec.patternProperties) }}
971
- {%- endif -%}
972
- {%- if spec.returns is mapping -%}
973
- {{- "\n- Returns *(" + render_markdown_type(spec.returns) + ")*" }}
974
- {{- render_markdown_schema_details(spec.returns, "", true) }}
975
- {%- elif spec.returns is defined -%}
976
- {{ render_markdown_metadata_detail("Returns", spec.returns) }}
977
- {%- endif -%}
978
- {{- render_markdown_metadata_extras(spec) }}
979
- {%- endif -%}
980
- {%- endmacro -%}
981
-
982
- {%- macro render_markdown_param(name, spec, required_list, indent) -%}
983
- {%- if spec is mapping and spec['$ref'] is string -%}
984
- {%- set _r = spec['$ref'] -%}
985
- {%- set _k = _r[8:] if _r.startswith('#/$defs/') else (_r[14:] if _r.startswith('#/definitions/') else none) -%}
986
- {%- if _k is not none and ('|' + _k + '|') not in REFS.seen and REFS.defs[_k] is mapping -%}
987
- {%- set spec = dict((REFS.defs[_k] | items | list) + (spec | items | rejectattr('0', 'equalto', '$ref') | list)) -%}
988
- {%- set REFS.seen = REFS.seen + _k + '|' -%}
989
- {%- if spec['$ref'] is string -%}
990
- {%- set _r2 = spec['$ref'] -%}
991
- {%- set _k2 = _r2[8:] if _r2.startswith('#/$defs/') else (_r2[14:] if _r2.startswith('#/definitions/') else none) -%}
992
- {%- if _k2 is not none and REFS.defs[_k2] is mapping -%}
993
- {%- set spec = dict((REFS.defs[_k2] | items | list) + (spec | items | rejectattr('0', 'equalto', '$ref') | list)) -%}
994
- {%- set REFS.seen = REFS.seen + _k2 + '|' -%}
995
- {%- endif -%}
996
- {%- endif -%}
997
- {%- endif -%}
998
- {%- endif -%}
999
- {{- "\n" + indent + "- `" + name + "` *(" + render_markdown_type(spec) }}
1000
- {%- if name in (required_list or []) -%}{{- ", required" }}{%- endif -%}
1001
- {{- ")*" }}
1002
- {%- if spec.description -%}{{- " - " + spec.description | replace("\n", "\n" + indent + " ") }}{%- endif -%}
1003
- {%- if spec.enum -%}
1004
- {{- "\n" + indent + " - Allowed values: " }}{{ render_allowed_values(spec.enum) }}
1005
- {%- endif -%}
1006
- {%- if spec.default is defined -%}
1007
- {{- "\n" + indent + " - Default: " }}{{ render_markdown_literal(spec.default) }}
1008
- {%- endif -%}
1009
- {{- render_markdown_schema_details(spec, indent, false) }}
1010
- {%- endmacro -%}
1011
-
1012
- {%- macro render_tools_markdown(tools_list) -%}
1013
- {{- "<ifm|tools>" }}
1014
- {%- for tool in tools_list -%}
1015
- {%- set fn = tool.function if tool.function is defined else tool -%}
1016
- {%- set REFS.defs = fn.parameters['$defs'] if (fn.parameters is mapping and fn.parameters['$defs'] is mapping) else (fn.parameters['definitions'] if (fn.parameters is mapping and fn.parameters['definitions'] is mapping) else {}) -%}
1017
- {%- set REFS.seen = '|' -%}
1018
- {%- set fnp = namespace(p=fn.parameters) -%}
1019
- {%- if fnp.p is mapping and fnp.p['$ref'] is string -%}
1020
- {%- set _r = fnp.p['$ref'] -%}
1021
- {%- set _k = _r[8:] if _r.startswith('#/$defs/') else (_r[14:] if _r.startswith('#/definitions/') else none) -%}
1022
- {%- if _k is not none and REFS.defs[_k] is mapping -%}
1023
- {%- set fnp.p = dict((REFS.defs[_k] | items | list) + (fnp.p | items | rejectattr('0', 'equalto', '$ref') | list)) -%}
1024
- {%- set REFS.seen = REFS.seen + _k + '|' -%}
1025
- {%- endif -%}
1026
- {%- endif -%}
1027
- {{- "\n## " + fn.name }}
1028
- {%- if fn.description -%}
1029
- {{- "\n" + fn.description }}
1030
- {%- endif -%}
1031
- {{- "\n\n**Parameters**" }}
1032
- {%- if fnp.p and fnp.p.properties -%}
1033
- {%- for pname, pspec in fnp.p.properties | items -%}
1034
- {{- render_markdown_param(pname, pspec, fnp.p.required or [], "") }}
1035
- {%- endfor -%}
1036
- {%- elif fnp.p is mapping and (fnp.p.oneOf or fnp.p.anyOf or 'items' in fnp.p) -%}
1037
- {{- render_markdown_parameter_schema(fnp.p) }}
1038
- {%- else -%}
1039
- {{- "\n- None" }}
1040
- {%- endif -%}
1041
- {%- set fn_ret = fn.returns if fn.returns is defined else fn.response -%}
1042
- {%- if fn_ret is mapping -%}
1043
- {{- "\n\n**Returns**" }}
1044
- {{- "\n- Return *(" + render_markdown_type(fn_ret) + ")*" }}
1045
- {{- render_markdown_schema_details(fn_ret, "", true) }}
1046
- {%- elif fn_ret is defined -%}
1047
- {{- "\n\n**Returns**\n- " }}{{ render_markdown_value(fn_ret) }}
1048
- {%- endif -%}
1049
- {%- if not loop.last -%}{{- "\n" }}{%- endif -%}
1050
- {%- endfor -%}
1051
- {{- "\n</ifm|tools>" }}
1052
- {%- endmacro -%}
1053
-
1054
- {%- macro render_tool_presentation(tools_list, fmt) -%}
1055
- {%- if fmt == 'json' -%}
1056
- {{- render_tools_json(tools_list) }}
1057
- {%- elif RB.bad != '|' -%}
1058
- {#- some tool uses constructs the pretty renderers cannot represent (verdicts -#}
1059
- {#- computed during validate_tools): render the WHOLE toolset exactly as the -#}
1060
- {#- json presentation would, so the block stays uniform and model-familiar. -#}
1061
- {{- render_tools_json(tools_list) }}
1062
- {%- elif fmt == 'xml' -%}
1063
- {{- render_tools_xml(tools_list) }}
1064
- {%- elif fmt == 'markdown' -%}
1065
- {{- render_tools_markdown(tools_list) }}
1066
- {%- else -%}
1067
- {{- raise_exception("Unsupported tool_presentation_format: '" + fmt + "'. Supported formats: json, xml, markdown.") }}
1068
- {%- endif -%}
1069
- {%- endmacro -%}
1070
-
1071
- {%- macro render_call_instructions(fmt) -%}
1072
- {%- if fmt == 'json' -%}
1073
- {{- "Wrap all tool calls in a single <ifm|tool_calls></ifm|tool_calls> block. For each call, emit one JSON object with the function name and arguments on the same line inside <ifm|tool_call></ifm|tool_call> tags:\n\n<ifm|tool_calls>\n<ifm|tool_call>{\"name\": <function-name>, \"arguments\": <args-json-object>}</ifm|tool_call>\n</ifm|tool_calls>" }}
1074
- {%- elif fmt == 'xml' -%}
1075
- {{- "Wrap all tool calls in a single <ifm|tool_calls></ifm|tool_calls> block. For each call, write the function name at the start of <ifm|tool_call>, followed by paired <ifm|arg_key> and <ifm|arg_value> tags for each argument:\n\n<ifm|tool_calls>\n<ifm|tool_call>$FUNCTION_NAME\n<ifm|arg_key>$PARAMETER_NAME</ifm|arg_key>\n<ifm|arg_value>$PARAMETER_VALUE</ifm|arg_value>\n...\n</ifm|tool_call>\n</ifm|tool_calls>\n\nString and scalar parameters should be written as plain text. Array and object parameters should be written as JSON literals." }}
1076
- {%- elif fmt == 'xml_typed' -%}
1077
- {{- "Wrap all tool calls in a single <ifm|tool_calls></ifm|tool_calls> block. For each call, write the function name at the start of <ifm|tool_call>, followed by <ifm|arg_key>, <ifm|arg_type>, and <ifm|arg_value> tags for each argument:\n\n<ifm|tool_calls>\n<ifm|tool_call>$FUNCTION_NAME\n<ifm|arg_key>$PARAMETER_NAME</ifm|arg_key>\n<ifm|arg_type>$ARGUMENT_TYPE</ifm|arg_type>\n<ifm|arg_value>$PARAMETER_VALUE</ifm|arg_value>\n...\n</ifm|tool_call>\n</ifm|tool_calls>\n\nUse the parameter type shown in the tool definition. If that type contains anyOf or oneOf, use the actual argument value type instead. String and scalar parameters should be written as plain text. Array and object parameters should be written as JSON literals." }}
1078
- {%- else -%}
1079
- {{- raise_exception("Unsupported tool_call_format: '" + fmt + "'. Supported formats: json, xml, xml_typed.") }}
1080
- {%- endif -%}
1081
- {%- endmacro -%}
1082
-
1083
- {%- macro render_system_with_tools(tools_list, system_content, presentation_fmt, call_fmt) -%}
1084
- {{- "<|ifm|im_start|>system\n# Tools\nYou may call one or more tools to assist with the user query.\n\nAvailable tools are:\n\n" }}
1085
- {{- render_tool_presentation(tools_list, presentation_fmt) }}
1086
- {{- "\n\nWhen calling tools, you MUST follow the tool-call format below:\n\n" }}
1087
- {{- render_call_instructions(call_fmt) }}
1088
- {%- if system_content -%}
1089
- {{- "\n\n" + system_content }}
1090
- {%- endif -%}
1091
- {{- "<|ifm|im_end|>" }}
1092
- {%- endmacro -%}
1093
-
1094
- {%- macro render_argument_value(value) -%}
1095
- {%- if value is string -%}{{- value -}}{%- else -%}{{- value | tojson -}}{%- endif -%}
1096
- {%- endmacro -%}
1097
-
1098
- {%- macro render_value_type(value) -%}
1099
- {%- if value is none -%}null
1100
- {%- elif value is boolean -%}boolean
1101
- {%- elif value is integer -%}integer
1102
- {%- elif value is number -%}number
1103
- {%- elif value is string -%}string
1104
- {%- elif value is mapping -%}object
1105
- {%- elif value is sequence -%}array
1106
- {%- else -%}any
1107
- {%- endif -%}
1108
- {%- endmacro -%}
1109
-
1110
- {%- macro schema_has_combinator(spec) -%}
1111
- {%- if spec.oneOf or spec.anyOf -%}
1112
- true
1113
- {%- elif spec.type is defined and spec.type is sequence and spec.type is not string and spec.type | length > 1 -%}
1114
- true
1115
- {%- elif spec.type == "array" and 'items' in spec -%}
1116
- {{- schema_has_combinator(spec['items']) -}}
1117
- {%- elif spec.properties -%}
1118
- {%- set found = namespace(value='false') -%}
1119
- {%- for child_name, child_spec in spec.properties | items -%}
1120
- {%- if schema_has_combinator(child_spec) == 'true' -%}
1121
- {%- set found.value = 'true' -%}
1122
- {%- endif -%}
1123
- {%- endfor -%}
1124
- {{- found.value -}}
1125
- {%- else -%}
1126
- false
1127
- {%- endif -%}
1128
- {%- endmacro -%}
1129
-
1130
- {%- macro render_arg_type(tools_list, tool_name, arg_name, value) -%}
1131
- {%- set found = namespace(type='any') -%}
1132
- {%- for tool in tools_list -%}
1133
- {%- set fn = tool.function if tool.function is defined else tool -%}
1134
- {%- if fn.name == tool_name and fn.parameters and fn.parameters.properties and arg_name in fn.parameters.properties -%}
1135
- {%- set spec = fn.parameters.properties[arg_name] -%}
1136
- {%- if spec is mapping and spec['$ref'] is string -%}
1137
- {%- set _r = spec['$ref'] -%}
1138
- {%- set _k = _r[8:] if _r.startswith('#/$defs/') else (_r[14:] if _r.startswith('#/definitions/') else none) -%}
1139
- {%- set _d = fn.parameters['$defs'] if fn.parameters['$defs'] is mapping else fn.parameters['definitions'] -%}
1140
- {%- set spec = dict((_d[_k] | items | list) + (spec | items | rejectattr('0', 'equalto', '$ref') | list)) if (_k is not none and _d is mapping and _d[_k] is mapping) else spec -%}
1141
- {%- endif -%}
1142
- {%- if schema_has_combinator(spec) == 'true' -%}
1143
- {%- set found.type = render_value_type(value) -%}
1144
- {%- else -%}
1145
- {%- set found.type = render_compact_type(spec) -%}
1146
- {%- endif -%}
1147
- {%- endif -%}
1148
- {%- endfor -%}
1149
- {{- found.type -}}
1150
- {%- endmacro -%}
1151
-
1152
- {%- macro render_tool_calls_block(tool_calls, fmt, tools_list) -%}
1153
- {{- "<ifm|tool_calls>" }}
1154
- {%- for raw_tool_call in tool_calls -%}
1155
- {%- set tool_call = raw_tool_call.function if raw_tool_call.function else raw_tool_call -%}
1156
- {%- if tool_call.arguments is string -%}
1157
- {{- raise_exception("tool_call.arguments must be a dict, not a JSON string. Parse it before passing to the template.") -}}
1158
- {%- endif -%}
1159
- {%- if fmt == 'json' -%}
1160
- {{- "\n<ifm|tool_call>{\"name\": \"" + tool_call.name + "\", \"arguments\": " }}{{ tool_call.arguments | tojson }}{{- "}</ifm|tool_call>" }}
1161
- {%- elif fmt == 'xml' or fmt == 'xml_typed' -%}
1162
- {{- "\n<ifm|tool_call>" + tool_call.name + "\n" }}
1163
- {%- for key, value in tool_call.arguments | items -%}
1164
- {{- "<ifm|arg_key>" + key + "</ifm|arg_key>\n" }}
1165
- {%- if fmt == 'xml_typed' -%}
1166
- {{- "<ifm|arg_type>" + render_arg_type(tools_list, tool_call.name, key, value) + "</ifm|arg_type>\n" }}
1167
- {%- endif -%}
1168
- {{- "<ifm|arg_value>" }}{{ render_argument_value(value) }}{{- "</ifm|arg_value>\n" }}
1169
- {%- endfor -%}
1170
- {{- "</ifm|tool_call>" }}
1171
- {%- else -%}
1172
- {{- raise_exception("Unsupported tool_call_format: '" + fmt + "'. Supported formats: json, xml, xml_typed.") -}}
1173
- {%- endif -%}
1174
- {%- endfor -%}
1175
- {{- "\n</ifm|tool_calls>" }}
1176
- {%- endmacro -%}
1177
-
1178
- {%- macro render_tool_response_messages(raw_content) -%}
1179
- {%- if raw_content is string -%}
1180
- {{- '<|ifm|im_start|>tool\n' + raw_content + '<|ifm|im_end|>' }}
1181
- {%- elif raw_content is sequence and raw_content is not string and raw_content is not mapping -%}
1182
- {%- if raw_content | length == 0 -%}
1183
- {{- raise_exception("tool message content list must not be empty.") -}}
1184
- {%- endif -%}
1185
- {{- '<|ifm|im_start|>tool\n' -}}
1186
- {%- for item in raw_content -%}
1187
- {%- if not loop.first -%}{{- '\n' -}}{%- endif -%}
1188
- {%- if item is string -%}
1189
- {{- item -}}
1190
- {%- elif item is mapping and item.text is string -%}
1191
- {{- item.text -}}
1192
- {%- else -%}
1193
- {{- (item | tojson) -}}
1194
- {%- endif -%}
1195
- {%- endfor -%}
1196
- {{- '<|ifm|im_end|>' -}}
1197
- {%- else -%}
1198
- {{- '<|ifm|im_start|>tool\n' }}{{ raw_content | tojson }}{{- '<|ifm|im_end|>' }}
1199
- {%- endif -%}
1200
- {%- endmacro -%}
1201
-
1202
- {%- set available_tools = tools if tools else [] -%}
1203
- {%- if (not available_tools) and messages[0].role == 'system' and messages[0].get('tools') -%}
1204
- {%- set available_tools = messages[0]['tools'] -%}
1205
- {%- endif -%}
1206
- {%- if available_tools -%}
1207
- {{- validate_tools(available_tools, tool_presentation_fmt != 'json') }}
1208
- {%- set system_content = '' -%}
1209
- {%- if messages[0].role == 'system' and messages[0].content -%}
1210
- {%- set system_content = messages[0].content -%}
1211
- {%- endif -%}
1212
- {{- render_system_with_tools(available_tools, system_content, tool_presentation_fmt, tool_call_fmt) }}
1213
- {%- else -%}
1214
- {%- if messages[0].role == 'system' -%}
1215
- {{- '<|ifm|im_start|>system\n' + messages[0].content + '<|ifm|im_end|>' }}
1216
- {%- endif -%}
1217
- {%- endif -%}
1218
-
1219
- {%- for message in messages -%}
1220
- {%- if message.content is string -%}
1221
- {%- set content = message.content -%}
1222
- {%- else -%}
1223
- {%- set content = '' -%}
1224
- {%- endif -%}
1225
- {%- if (message.role == "user") or (message.role == "system" and not loop.first) -%}
1226
- {{- '<|ifm|im_start|>' + message.role + '\n' + content + '<|ifm|im_end|>' }}
1227
- {%- elif message.role == "assistant" -%}
1228
- {%- set thinking_content = '' -%}
1229
- {%- set think_tag = '' -%}
1230
- {%- if message.think is defined and message.think is string -%}
1231
- {%- set thinking_content = message.think -%}
1232
- {%- set think_tag = 'ifm|think' -%}
1233
- {%- elif message.think_fast is defined and message.think_fast is string -%}
1234
- {%- set thinking_content = message.think_fast -%}
1235
- {%- set think_tag = 'ifm|think_fast' -%}
1236
- {%- elif message.think_faster is defined and message.think_faster is string -%}
1237
- {%- set thinking_content = message.think_faster -%}
1238
- {%- set think_tag = 'ifm|think_faster' -%}
1239
- {%- else -%}
1240
- {%- if '</ifm|think>' in content -%}
1241
- {%- set thinking_content = content.split('</ifm|think>')[0].rstrip('\n').split('<ifm|think>')[-1].lstrip('\n') -%}
1242
- {%- set content = content.split('</ifm|think>')[-1].lstrip('\n') -%}
1243
- {%- set think_tag = 'ifm|think' -%}
1244
- {%- elif '</ifm|think_fast>' in content -%}
1245
- {%- set thinking_content = content.split('</ifm|think_fast>')[0].rstrip('\n').split('<ifm|think_fast>')[-1].lstrip('\n') -%}
1246
- {%- set content = content.split('</ifm|think_fast>')[-1].lstrip('\n') -%}
1247
- {%- set think_tag = 'ifm|think_fast' -%}
1248
- {%- elif '</ifm|think_faster>' in content -%}
1249
- {%- set thinking_content = content.split('</ifm|think_faster>')[0].rstrip('\n').split('<ifm|think_faster>')[-1].lstrip('\n') -%}
1250
- {%- set content = content.split('</ifm|think_faster>')[-1].lstrip('\n') -%}
1251
- {%- set think_tag = 'ifm|think_faster' -%}
1252
- {%- endif -%}
1253
- {%- endif -%}
1254
- {{- '<|ifm|im_start|>' + message.role }}
1255
- {% generation %}
1256
- {%- if think_tag -%}
1257
- {%- if thinking_content -%}
1258
- {{- '<' + think_tag + '>\n' + thinking_content + '\n</' + think_tag + '>\n' + content.lstrip('\n') }}
1259
- {%- else -%}
1260
- {{- '<' + think_tag + '>\n</' + think_tag + '>\n' + content.lstrip('\n') }}
1261
- {%- endif -%}
1262
- {%- else -%}
1263
- {{- content }}
1264
- {%- endif -%}
1265
- {%- if message.tool_calls -%}
1266
- {%- if content -%}
1267
- {{- '\n' }}
1268
- {%- endif -%}
1269
- {{- render_tool_calls_block(message.tool_calls, tool_call_fmt, available_tools) }}
1270
- {%- endif -%}
1271
- {{- '<|ifm|im_end|>' -}}
1272
- {%- endgeneration -%}
1273
- {%- elif message.role == "tool" -%}
1274
- {{- render_tool_response_messages(message.content) }}
1275
- {%- endif -%}
1276
- {%- endfor -%}
1277
- {%- if add_generation_prompt -%}
1278
- {%- set effort = reasoning_effort | default('high') -%}
1279
- {%- if effort == 'high' -%}
1280
- {{- '<|ifm|im_start|>assistant\n<ifm|think>\n' }}
1281
- {%- elif effort == 'medium' -%}
1282
- {{- '<|ifm|im_start|>assistant\n<ifm|think_fast>\n' }}
1283
- {%- elif effort == 'low' -%}
1284
- {{- '<|ifm|im_start|>assistant\n<ifm|think_faster>\n' }}
1285
- {%- else -%}
1286
- {{- '<|ifm|im_start|>assistant\n<ifm|think_fast>\n' }}
1287
- {%- endif -%}
1288
- {%- endif -%}
1289
-
1290
- INFO:gguf.gguf_writer:Writing the following files:
1291
- INFO:gguf.gguf_writer:source-official-bf16-k2-horizon-0.9b/K2-Horizon-0.9B-BF16.gguf: n_tensors = 255, total_size = 2.2G
1292
- INFO:hf-to-gguf:Model successfully exported to source-official-bf16-k2-horizon-0.9b/K2-Horizon-0.9B-BF16.gguf
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
reproducibility/validation/imatrix-combine.log DELETED
@@ -1,6 +0,0 @@
1
- 0.00.047.117 W DEPRECATED: argument '--in-file' specified multiple times, use comma-separated values instead (only last value will be used)
2
- 0.00.051.191 I main : loading imatrix from 'calibration/k2-horizon-0.9b/k2_horizon_wikitext.imatrix.gguf'
3
- 0.00.055.600 I main : loading imatrix from 'calibration/k2-horizon-0.9b/k2_horizon_en_zh_code_tool.imatrix.gguf'
4
- 0.00.057.918 I No prompt provided; combining precomputed matrices only.
5
- 0.00.057.922 I main : saving combined imatrix to 'calibration/k2-horizon-0.9b/k2_horizon_combined.imatrix.gguf'
6
-
 
 
 
 
 
 
 
reproducibility/validation/imatrix-k2-corpus-gpu1.log DELETED
@@ -1,10 +0,0 @@
1
- 0.01.388.894 I cmn init: llama threadpool init, n_threads = 128
2
- 0.01.391.290 I
3
- 0.01.391.499 I system_info: n_threads = 128 (n_threads_batch = 128) / 256 | CUDA : ARCHS = 860 | USE_GRAPHS = 1 | CPU : SSE3 = 1 | SSSE3 = 1 | AVX = 1 | AVX2 = 1 | F16C = 1 | FMA = 1 | BMI2 = 1 | LLAMAFILE = 1 | OPENMP = 1 | REPACK = 1 |
4
- 0.01.391.506 I compute_imatrix: tokenizing the input ..
5
- 0.01.403.919 I compute_imatrix: tokenization took 12.411 ms
6
- 0.01.403.943 I compute_imatrix: computing over 3 chunks, n_ctx=512, batch_size=512, n_seq=1
7
- 0.02.116.253 I compute_imatrix: 0.71 seconds per pass - ETA 0.03 minutes
8
-
9
-
10
-
 
 
 
 
 
 
 
 
 
 
 
reproducibility/validation/imatrix-wikitext-gpu1.log DELETED
@@ -1,19 +0,0 @@
1
- 0.01.330.346 I cmn init: llama threadpool init, n_threads = 128
2
- 0.01.331.937 I
3
- 0.01.332.143 I system_info: n_threads = 128 (n_threads_batch = 128) / 256 | CUDA : ARCHS = 860 | USE_GRAPHS = 1 | CPU : SSE3 = 1 | SSSE3 = 1 | AVX = 1 | AVX2 = 1 | F16C = 1 | FMA = 1 | BMI2 = 1 | LLAMAFILE = 1 | OPENMP = 1 | REPACK = 1 |
4
- 0.01.332.151 I compute_imatrix: tokenizing the input ..
5
- 0.08.581.630 I compute_imatrix: tokenization took 7249.44 ms
6
- 0.08.581.709 I compute_imatrix: computing over 96 chunks, n_ctx=512, batch_size=512, n_seq=1
7
- 0.09.047.081 I compute_imatrix: 0.47 seconds per pass - ETA 0.73 minutes
8
-
9
-
10
-
11
-
12
-
13
-
14
-
15
-
16
-
17
-
18
-
19
-
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
reproducibility/validation/quantize-logs/IQ1_M.log DELETED
@@ -1,318 +0,0 @@
1
- ggml_cuda_init: failed to initialize CUDA: no CUDA-capable device is detected
2
- version: 0.3.0-dev (build 10671, commit 35999d101)
3
- built with GNU 13.3.0 for Linux x86_64
4
- llama_quantize: quantizing '/root/workspace/HF/source-official-bf16-k2-horizon-0.9b/K2-Horizon-0.9B-BF16.gguf' to '/root/workspace/HF/quantized-k2-horizon-0.9b/K2-Horizon-0.9B-IQ1_M.gguf' as IQ1_M using 128 threads
5
- llama_model_loader: loaded meta data with 41 key-value pairs and 255 tensors from /root/workspace/HF/source-official-bf16-k2-horizon-0.9b/K2-Horizon-0.9B-BF16.gguf (version GGUF V3 (latest))
6
- llama_model_loader: Dumping metadata keys/values. Note: KV overrides do not apply in this output.
7
- llama_model_loader: - kv 0: general.architecture str = k2-horizon
8
- llama_model_loader: - kv 1: general.type str = model
9
- llama_model_loader: - kv 2: general.name str = K2-Horizon-0.9B
10
- llama_model_loader: - kv 3: general.basename str = source-hf-k2-horizon
11
- llama_model_loader: - kv 4: general.size_label str = 0.9B
12
- llama_model_loader: - kv 5: general.license str = apache-2.0
13
- llama_model_loader: - kv 6: general.license.name str = internal-only
14
- llama_model_loader: - kv 7: general.license.link str = LICENSE
15
- llama_model_loader: - kv 8: general.tags arr[str,7] = ["k2-horizon", "0.9b", "dense", "reas...
16
- llama_model_loader: - kv 9: general.languages arr[str,2] = ["en", "zh"]
17
- llama_model_loader: - kv 10: k2-horizon.block_count u32 = 28
18
- llama_model_loader: - kv 11: k2-horizon.context_length u32 = 131072
19
- llama_model_loader: - kv 12: k2-horizon.embedding_length u32 = 1536
20
- llama_model_loader: - kv 13: k2-horizon.feed_forward_length u32 = 5120
21
- llama_model_loader: - kv 14: k2-horizon.attention.head_count u32 = 32
22
- llama_model_loader: - kv 15: k2-horizon.attention.head_count_kv u32 = 8
23
- llama_model_loader: - kv 16: k2-horizon.rope.scaling.type str = yarn
24
- llama_model_loader: - kv 17: k2-horizon.rope.scaling.factor f32 = 16.000000
25
- llama_model_loader: - kv 18: k2-horizon.rope.scaling.original_context_length u32 = 8192
26
- llama_model_loader: - kv 19: k2-horizon.rope.scaling.yarn_attn_factor f32 = 1.277259
27
- llama_model_loader: - kv 20: k2-horizon.rope.scaling.yarn_beta_fast f32 = 128.000000
28
- llama_model_loader: - kv 21: k2-horizon.rope.scaling.yarn_beta_slow f32 = 4.000000
29
- llama_model_loader: - kv 22: k2-horizon.rope.freq_base f32 = 1000000.000000
30
- llama_model_loader: - kv 23: k2-horizon.attention.layer_norm_rms_epsilon f32 = 0.000001
31
- llama_model_loader: - kv 24: k2-horizon.expert_count u32 = 0
32
- llama_model_loader: - kv 25: k2-horizon.expert_used_count u32 = 0
33
- llama_model_loader: - kv 26: k2-horizon.attention.key_length u32 = 64
34
- llama_model_loader: - kv 27: k2-horizon.attention.value_length u32 = 64
35
- llama_model_loader: - kv 28: general.file_type u32 = 32
36
- llama_model_loader: - kv 29: k2-horizon.attention.group_norm_groups u32 = 1
37
- llama_model_loader: - kv 30: k2-horizon.rope.dimension_count u32 = 64
38
- llama_model_loader: - kv 31: general.quantization_version u32 = 2
39
- llama_model_loader: - kv 32: tokenizer.ggml.model str = gpt2
40
- llama_model_loader: - kv 33: tokenizer.ggml.pre str = k2-horizon
41
- llama_model_loader: - kv 34: tokenizer.ggml.tokens arr[str,64256] = ["<|begin_of_text|>", "<|endoftext|>"...
42
- llama_model_loader: - kv 35: tokenizer.ggml.token_type arr[i32,64256] = [3, 3, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, ...
43
- llama_model_loader: - kv 36: tokenizer.ggml.merges arr[str,63742] = ["Δ  Δ ", "Ø Β§", "Γ™ Δ¦", "Γ  Β€", ...
44
- llama_model_loader: - kv 37: tokenizer.ggml.bos_token_id u32 = 0
45
- llama_model_loader: - kv 38: tokenizer.ggml.eos_token_id u32 = 1
46
- llama_model_loader: - kv 39: tokenizer.ggml.padding_token_id u32 = 64255
47
- llama_model_loader: - kv 40: tokenizer.chat_template str = {{- bos_token }}\n{%- if tool_presenta...
48
- llama_model_loader: - type f32: 57 tensors
49
- llama_model_loader: - type bf16: 198 tensors
50
-
51
- llama_model_quantize_impl: have importance matrix data with 196 entries
52
- [ 1/ 255] output.weight - [ 1536, 64256, 1, 1], type = bf16,
53
- ====== llama_model_quantize_impl: did not find weights for output.weight
54
- converting to q5_K .. load_imatrix: imatrix datasets=['calibration/wikitext-2-raw/wiki.train.raw', 'calibration/k2-horizon-0.9b/k2_horizon_en_zh_code_tool.txt']
55
- load_imatrix: loaded 196 importance matrix entries from /root/workspace/HF/calibration/k2-horizon-0.9b/k2_horizon_combined.imatrix.gguf computed on 99 chunks
56
- prepare_imatrix: have 196 importance matrix entries
57
- size = 188.25 MiB -> 64.71 MiB
58
- [ 2/ 255] output_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
59
- [ 3/ 255] token_embd.weight - [ 1536, 64256, 1, 1], type = bf16,
60
- ====== llama_model_quantize_impl: did not find weights for token_embd.weight
61
- converting to q2_K .. size = 188.25 MiB -> 30.88 MiB
62
- [ 4/ 255] blk.0.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to iq1_m .. size = 1.50 MiB -> 0.16 MiB
63
- [ 5/ 255] blk.0.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
64
- [ 6/ 255] blk.0.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to iq2_xxs .. size = 6.00 MiB -> 0.77 MiB
65
- [ 7/ 255] blk.0.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to iq1_m .. size = 6.00 MiB -> 0.66 MiB
66
- [ 8/ 255] blk.0.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
67
- [ 9/ 255] blk.0.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q2_K .. size = 15.00 MiB -> 2.46 MiB
68
- [ 10/ 255] blk.0.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 15.00 MiB -> 1.64 MiB
69
- [ 11/ 255] blk.0.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
70
- [ 12/ 255] blk.0.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 15.00 MiB -> 1.64 MiB
71
- [ 13/ 255] blk.1.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to iq1_m .. size = 1.50 MiB -> 0.16 MiB
72
- [ 14/ 255] blk.1.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
73
- [ 15/ 255] blk.1.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to iq2_xxs .. size = 6.00 MiB -> 0.77 MiB
74
- [ 16/ 255] blk.1.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to iq1_m .. size = 6.00 MiB -> 0.66 MiB
75
- [ 17/ 255] blk.1.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
76
- [ 18/ 255] blk.1.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q2_K .. size = 15.00 MiB -> 2.46 MiB
77
- [ 19/ 255] blk.1.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 15.00 MiB -> 1.64 MiB
78
- [ 20/ 255] blk.1.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
79
- [ 21/ 255] blk.1.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 15.00 MiB -> 1.64 MiB
80
- [ 22/ 255] blk.2.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to iq1_m .. size = 1.50 MiB -> 0.16 MiB
81
- [ 23/ 255] blk.2.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
82
- [ 24/ 255] blk.2.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to iq2_xxs .. size = 6.00 MiB -> 0.77 MiB
83
- [ 25/ 255] blk.2.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to iq1_m .. size = 6.00 MiB -> 0.66 MiB
84
- [ 26/ 255] blk.2.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
85
- [ 27/ 255] blk.2.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q2_K .. size = 15.00 MiB -> 2.46 MiB
86
- [ 28/ 255] blk.2.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 15.00 MiB -> 1.64 MiB
87
- [ 29/ 255] blk.2.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
88
- [ 30/ 255] blk.2.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 15.00 MiB -> 1.64 MiB
89
- [ 31/ 255] blk.3.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to iq1_m .. size = 1.50 MiB -> 0.16 MiB
90
- [ 32/ 255] blk.3.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
91
- [ 33/ 255] blk.3.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to iq2_xxs .. size = 6.00 MiB -> 0.77 MiB
92
- [ 34/ 255] blk.3.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to iq1_m .. size = 6.00 MiB -> 0.66 MiB
93
- [ 35/ 255] blk.3.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
94
- [ 36/ 255] blk.3.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to iq1_m .. size = 15.00 MiB -> 1.64 MiB
95
- [ 37/ 255] blk.3.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 15.00 MiB -> 1.64 MiB
96
- [ 38/ 255] blk.3.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
97
- [ 39/ 255] blk.3.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 15.00 MiB -> 1.64 MiB
98
- [ 40/ 255] blk.4.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to iq1_m .. size = 1.50 MiB -> 0.16 MiB
99
- [ 41/ 255] blk.4.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
100
- [ 42/ 255] blk.4.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to iq2_xxs .. size = 6.00 MiB -> 0.77 MiB
101
- [ 43/ 255] blk.4.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to iq1_m .. size = 6.00 MiB -> 0.66 MiB
102
- [ 44/ 255] blk.4.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
103
- [ 45/ 255] blk.4.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to iq1_m .. size = 15.00 MiB -> 1.64 MiB
104
- [ 46/ 255] blk.4.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 15.00 MiB -> 1.64 MiB
105
- [ 47/ 255] blk.4.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
106
- [ 48/ 255] blk.4.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 15.00 MiB -> 1.64 MiB
107
- [ 49/ 255] blk.5.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to iq1_m .. size = 1.50 MiB -> 0.16 MiB
108
- [ 50/ 255] blk.5.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
109
- [ 51/ 255] blk.5.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to iq2_xxs .. size = 6.00 MiB -> 0.77 MiB
110
- [ 52/ 255] blk.5.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to iq1_m .. size = 6.00 MiB -> 0.66 MiB
111
- [ 53/ 255] blk.5.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
112
- [ 54/ 255] blk.5.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to iq1_m .. size = 15.00 MiB -> 1.64 MiB
113
- [ 55/ 255] blk.5.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 15.00 MiB -> 1.64 MiB
114
- [ 56/ 255] blk.5.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
115
- [ 57/ 255] blk.5.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 15.00 MiB -> 1.64 MiB
116
- [ 58/ 255] blk.6.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to iq1_m .. size = 1.50 MiB -> 0.16 MiB
117
- [ 59/ 255] blk.6.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
118
- [ 60/ 255] blk.6.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to iq2_xxs .. size = 6.00 MiB -> 0.77 MiB
119
- [ 61/ 255] blk.6.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to iq1_m .. size = 6.00 MiB -> 0.66 MiB
120
- [ 62/ 255] blk.6.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
121
- [ 63/ 255] blk.6.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to iq1_m .. size = 15.00 MiB -> 1.64 MiB
122
- [ 64/ 255] blk.6.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 15.00 MiB -> 1.64 MiB
123
- [ 65/ 255] blk.6.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
124
- [ 66/ 255] blk.6.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 15.00 MiB -> 1.64 MiB
125
- [ 67/ 255] blk.7.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to iq1_m .. size = 1.50 MiB -> 0.16 MiB
126
- [ 68/ 255] blk.7.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
127
- [ 69/ 255] blk.7.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to iq2_xxs .. size = 6.00 MiB -> 0.77 MiB
128
- [ 70/ 255] blk.7.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to iq1_m .. size = 6.00 MiB -> 0.66 MiB
129
- [ 71/ 255] blk.7.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
130
- [ 72/ 255] blk.7.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to iq1_m .. size = 15.00 MiB -> 1.64 MiB
131
- [ 73/ 255] blk.7.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 15.00 MiB -> 1.64 MiB
132
- [ 74/ 255] blk.7.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
133
- [ 75/ 255] blk.7.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 15.00 MiB -> 1.64 MiB
134
- [ 76/ 255] blk.8.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to iq1_m .. size = 1.50 MiB -> 0.16 MiB
135
- [ 77/ 255] blk.8.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
136
- [ 78/ 255] blk.8.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to iq2_xxs .. size = 6.00 MiB -> 0.77 MiB
137
- [ 79/ 255] blk.8.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to iq1_m .. size = 6.00 MiB -> 0.66 MiB
138
- [ 80/ 255] blk.8.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
139
- [ 81/ 255] blk.8.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to iq1_m .. size = 15.00 MiB -> 1.64 MiB
140
- [ 82/ 255] blk.8.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 15.00 MiB -> 1.64 MiB
141
- [ 83/ 255] blk.8.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
142
- [ 84/ 255] blk.8.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 15.00 MiB -> 1.64 MiB
143
- [ 85/ 255] blk.9.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to iq1_m .. size = 1.50 MiB -> 0.16 MiB
144
- [ 86/ 255] blk.9.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
145
- [ 87/ 255] blk.9.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to iq2_xxs .. size = 6.00 MiB -> 0.77 MiB
146
- [ 88/ 255] blk.9.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to iq1_m .. size = 6.00 MiB -> 0.66 MiB
147
- [ 89/ 255] blk.9.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
148
- [ 90/ 255] blk.9.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to iq1_m .. size = 15.00 MiB -> 1.64 MiB
149
- [ 91/ 255] blk.9.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 15.00 MiB -> 1.64 MiB
150
- [ 92/ 255] blk.9.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
151
- [ 93/ 255] blk.9.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 15.00 MiB -> 1.64 MiB
152
- [ 94/ 255] blk.10.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to iq1_m .. size = 1.50 MiB -> 0.16 MiB
153
- [ 95/ 255] blk.10.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
154
- [ 96/ 255] blk.10.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to iq2_xxs .. size = 6.00 MiB -> 0.77 MiB
155
- [ 97/ 255] blk.10.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to iq1_m .. size = 6.00 MiB -> 0.66 MiB
156
- [ 98/ 255] blk.10.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
157
- [ 99/ 255] blk.10.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to iq1_m .. size = 15.00 MiB -> 1.64 MiB
158
- [ 100/ 255] blk.10.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 15.00 MiB -> 1.64 MiB
159
- [ 101/ 255] blk.10.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
160
- [ 102/ 255] blk.10.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 15.00 MiB -> 1.64 MiB
161
- [ 103/ 255] blk.11.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to iq1_m .. size = 1.50 MiB -> 0.16 MiB
162
- [ 104/ 255] blk.11.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
163
- [ 105/ 255] blk.11.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to iq2_xxs .. size = 6.00 MiB -> 0.77 MiB
164
- [ 106/ 255] blk.11.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to iq1_m .. size = 6.00 MiB -> 0.66 MiB
165
- [ 107/ 255] blk.11.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
166
- [ 108/ 255] blk.11.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to iq1_m .. size = 15.00 MiB -> 1.64 MiB
167
- [ 109/ 255] blk.11.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 15.00 MiB -> 1.64 MiB
168
- [ 110/ 255] blk.11.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
169
- [ 111/ 255] blk.11.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 15.00 MiB -> 1.64 MiB
170
- [ 112/ 255] blk.12.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to iq1_m .. size = 1.50 MiB -> 0.16 MiB
171
- [ 113/ 255] blk.12.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
172
- [ 114/ 255] blk.12.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to iq2_xxs .. size = 6.00 MiB -> 0.77 MiB
173
- [ 115/ 255] blk.12.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to iq1_m .. size = 6.00 MiB -> 0.66 MiB
174
- [ 116/ 255] blk.12.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
175
- [ 117/ 255] blk.12.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to iq1_m .. size = 15.00 MiB -> 1.64 MiB
176
- [ 118/ 255] blk.12.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 15.00 MiB -> 1.64 MiB
177
- [ 119/ 255] blk.12.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
178
- [ 120/ 255] blk.12.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 15.00 MiB -> 1.64 MiB
179
- [ 121/ 255] blk.13.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to iq1_m .. size = 1.50 MiB -> 0.16 MiB
180
- [ 122/ 255] blk.13.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
181
- [ 123/ 255] blk.13.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to iq2_xxs .. size = 6.00 MiB -> 0.77 MiB
182
- [ 124/ 255] blk.13.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to iq1_m .. size = 6.00 MiB -> 0.66 MiB
183
- [ 125/ 255] blk.13.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
184
- [ 126/ 255] blk.13.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to iq1_m .. size = 15.00 MiB -> 1.64 MiB
185
- [ 127/ 255] blk.13.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 15.00 MiB -> 1.64 MiB
186
- [ 128/ 255] blk.13.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
187
- [ 129/ 255] blk.13.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 15.00 MiB -> 1.64 MiB
188
- [ 130/ 255] blk.14.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to iq1_m .. size = 1.50 MiB -> 0.16 MiB
189
- [ 131/ 255] blk.14.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
190
- [ 132/ 255] blk.14.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to iq2_xxs .. size = 6.00 MiB -> 0.77 MiB
191
- [ 133/ 255] blk.14.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to iq1_m .. size = 6.00 MiB -> 0.66 MiB
192
- [ 134/ 255] blk.14.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
193
- [ 135/ 255] blk.14.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to iq1_m .. size = 15.00 MiB -> 1.64 MiB
194
- [ 136/ 255] blk.14.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 15.00 MiB -> 1.64 MiB
195
- [ 137/ 255] blk.14.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
196
- [ 138/ 255] blk.14.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 15.00 MiB -> 1.64 MiB
197
- [ 139/ 255] blk.15.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to iq1_m .. size = 1.50 MiB -> 0.16 MiB
198
- [ 140/ 255] blk.15.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
199
- [ 141/ 255] blk.15.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to iq2_xxs .. size = 6.00 MiB -> 0.77 MiB
200
- [ 142/ 255] blk.15.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to iq1_m .. size = 6.00 MiB -> 0.66 MiB
201
- [ 143/ 255] blk.15.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
202
- [ 144/ 255] blk.15.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to iq1_m .. size = 15.00 MiB -> 1.64 MiB
203
- [ 145/ 255] blk.15.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 15.00 MiB -> 1.64 MiB
204
- [ 146/ 255] blk.15.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
205
- [ 147/ 255] blk.15.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 15.00 MiB -> 1.64 MiB
206
- [ 148/ 255] blk.16.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to iq1_m .. size = 1.50 MiB -> 0.16 MiB
207
- [ 149/ 255] blk.16.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
208
- [ 150/ 255] blk.16.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to iq2_xxs .. size = 6.00 MiB -> 0.77 MiB
209
- [ 151/ 255] blk.16.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to iq1_m .. size = 6.00 MiB -> 0.66 MiB
210
- [ 152/ 255] blk.16.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
211
- [ 153/ 255] blk.16.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to iq1_m .. size = 15.00 MiB -> 1.64 MiB
212
- [ 154/ 255] blk.16.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 15.00 MiB -> 1.64 MiB
213
- [ 155/ 255] blk.16.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
214
- [ 156/ 255] blk.16.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 15.00 MiB -> 1.64 MiB
215
- [ 157/ 255] blk.17.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to iq1_m .. size = 1.50 MiB -> 0.16 MiB
216
- [ 158/ 255] blk.17.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
217
- [ 159/ 255] blk.17.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to iq2_xxs .. size = 6.00 MiB -> 0.77 MiB
218
- [ 160/ 255] blk.17.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to iq1_m .. size = 6.00 MiB -> 0.66 MiB
219
- [ 161/ 255] blk.17.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
220
- [ 162/ 255] blk.17.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to iq1_m .. size = 15.00 MiB -> 1.64 MiB
221
- [ 163/ 255] blk.17.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 15.00 MiB -> 1.64 MiB
222
- [ 164/ 255] blk.17.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
223
- [ 165/ 255] blk.17.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 15.00 MiB -> 1.64 MiB
224
- [ 166/ 255] blk.18.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to iq1_m .. size = 1.50 MiB -> 0.16 MiB
225
- [ 167/ 255] blk.18.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
226
- [ 168/ 255] blk.18.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to iq2_xxs .. size = 6.00 MiB -> 0.77 MiB
227
- [ 169/ 255] blk.18.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to iq1_m .. size = 6.00 MiB -> 0.66 MiB
228
- [ 170/ 255] blk.18.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
229
- [ 171/ 255] blk.18.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to iq1_m .. size = 15.00 MiB -> 1.64 MiB
230
- [ 172/ 255] blk.18.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 15.00 MiB -> 1.64 MiB
231
- [ 173/ 255] blk.18.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
232
- [ 174/ 255] blk.18.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 15.00 MiB -> 1.64 MiB
233
- [ 175/ 255] blk.19.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to iq1_m .. size = 1.50 MiB -> 0.16 MiB
234
- [ 176/ 255] blk.19.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
235
- [ 177/ 255] blk.19.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to iq2_xxs .. size = 6.00 MiB -> 0.77 MiB
236
- [ 178/ 255] blk.19.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to iq1_m .. size = 6.00 MiB -> 0.66 MiB
237
- [ 179/ 255] blk.19.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
238
- [ 180/ 255] blk.19.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to iq1_m .. size = 15.00 MiB -> 1.64 MiB
239
- [ 181/ 255] blk.19.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 15.00 MiB -> 1.64 MiB
240
- [ 182/ 255] blk.19.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
241
- [ 183/ 255] blk.19.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 15.00 MiB -> 1.64 MiB
242
- [ 184/ 255] blk.20.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to iq1_m .. size = 1.50 MiB -> 0.16 MiB
243
- [ 185/ 255] blk.20.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
244
- [ 186/ 255] blk.20.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to iq2_xxs .. size = 6.00 MiB -> 0.77 MiB
245
- [ 187/ 255] blk.20.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to iq1_m .. size = 6.00 MiB -> 0.66 MiB
246
- [ 188/ 255] blk.20.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
247
- [ 189/ 255] blk.20.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to iq1_m .. size = 15.00 MiB -> 1.64 MiB
248
- [ 190/ 255] blk.20.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 15.00 MiB -> 1.64 MiB
249
- [ 191/ 255] blk.20.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
250
- [ 192/ 255] blk.20.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 15.00 MiB -> 1.64 MiB
251
- [ 193/ 255] blk.21.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to iq1_m .. size = 1.50 MiB -> 0.16 MiB
252
- [ 194/ 255] blk.21.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
253
- [ 195/ 255] blk.21.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to iq2_xxs .. size = 6.00 MiB -> 0.77 MiB
254
- [ 196/ 255] blk.21.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to iq1_m .. size = 6.00 MiB -> 0.66 MiB
255
- [ 197/ 255] blk.21.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
256
- [ 198/ 255] blk.21.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to iq1_m .. size = 15.00 MiB -> 1.64 MiB
257
- [ 199/ 255] blk.21.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 15.00 MiB -> 1.64 MiB
258
- [ 200/ 255] blk.21.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
259
- [ 201/ 255] blk.21.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 15.00 MiB -> 1.64 MiB
260
- [ 202/ 255] blk.22.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to iq1_m .. size = 1.50 MiB -> 0.16 MiB
261
- [ 203/ 255] blk.22.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
262
- [ 204/ 255] blk.22.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to iq2_xxs .. size = 6.00 MiB -> 0.77 MiB
263
- [ 205/ 255] blk.22.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to iq1_m .. size = 6.00 MiB -> 0.66 MiB
264
- [ 206/ 255] blk.22.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
265
- [ 207/ 255] blk.22.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to iq1_m .. size = 15.00 MiB -> 1.64 MiB
266
- [ 208/ 255] blk.22.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 15.00 MiB -> 1.64 MiB
267
- [ 209/ 255] blk.22.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
268
- [ 210/ 255] blk.22.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 15.00 MiB -> 1.64 MiB
269
- [ 211/ 255] blk.23.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to iq1_m .. size = 1.50 MiB -> 0.16 MiB
270
- [ 212/ 255] blk.23.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
271
- [ 213/ 255] blk.23.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to iq2_xxs .. size = 6.00 MiB -> 0.77 MiB
272
- [ 214/ 255] blk.23.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to iq1_m .. size = 6.00 MiB -> 0.66 MiB
273
- [ 215/ 255] blk.23.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
274
- [ 216/ 255] blk.23.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to iq1_m .. size = 15.00 MiB -> 1.64 MiB
275
- [ 217/ 255] blk.23.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 15.00 MiB -> 1.64 MiB
276
- [ 218/ 255] blk.23.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
277
- [ 219/ 255] blk.23.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 15.00 MiB -> 1.64 MiB
278
- [ 220/ 255] blk.24.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to iq1_m .. size = 1.50 MiB -> 0.16 MiB
279
- [ 221/ 255] blk.24.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
280
- [ 222/ 255] blk.24.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to iq2_xxs .. size = 6.00 MiB -> 0.77 MiB
281
- [ 223/ 255] blk.24.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to iq1_m .. size = 6.00 MiB -> 0.66 MiB
282
- [ 224/ 255] blk.24.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
283
- [ 225/ 255] blk.24.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to iq1_m .. size = 15.00 MiB -> 1.64 MiB
284
- [ 226/ 255] blk.24.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 15.00 MiB -> 1.64 MiB
285
- [ 227/ 255] blk.24.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
286
- [ 228/ 255] blk.24.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 15.00 MiB -> 1.64 MiB
287
- [ 229/ 255] blk.25.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to iq1_m .. size = 1.50 MiB -> 0.16 MiB
288
- [ 230/ 255] blk.25.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
289
- [ 231/ 255] blk.25.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to iq2_xxs .. size = 6.00 MiB -> 0.77 MiB
290
- [ 232/ 255] blk.25.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to iq1_m .. size = 6.00 MiB -> 0.66 MiB
291
- [ 233/ 255] blk.25.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
292
- [ 234/ 255] blk.25.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to iq1_m .. size = 15.00 MiB -> 1.64 MiB
293
- [ 235/ 255] blk.25.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 15.00 MiB -> 1.64 MiB
294
- [ 236/ 255] blk.25.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
295
- [ 237/ 255] blk.25.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 15.00 MiB -> 1.64 MiB
296
- [ 238/ 255] blk.26.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to iq1_m .. size = 1.50 MiB -> 0.16 MiB
297
- [ 239/ 255] blk.26.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
298
- [ 240/ 255] blk.26.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to iq2_xxs .. size = 6.00 MiB -> 0.77 MiB
299
- [ 241/ 255] blk.26.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to iq1_m .. size = 6.00 MiB -> 0.66 MiB
300
- [ 242/ 255] blk.26.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
301
- [ 243/ 255] blk.26.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to iq1_m .. size = 15.00 MiB -> 1.64 MiB
302
- [ 244/ 255] blk.26.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 15.00 MiB -> 1.64 MiB
303
- [ 245/ 255] blk.26.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
304
- [ 246/ 255] blk.26.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 15.00 MiB -> 1.64 MiB
305
- [ 247/ 255] blk.27.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to iq1_m .. size = 1.50 MiB -> 0.16 MiB
306
- [ 248/ 255] blk.27.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
307
- [ 249/ 255] blk.27.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to iq2_xxs .. size = 6.00 MiB -> 0.77 MiB
308
- [ 250/ 255] blk.27.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to iq1_m .. size = 6.00 MiB -> 0.66 MiB
309
- [ 251/ 255] blk.27.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
310
- [ 252/ 255] blk.27.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to iq1_m .. size = 15.00 MiB -> 1.64 MiB
311
- [ 253/ 255] blk.27.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 15.00 MiB -> 1.64 MiB
312
- [ 254/ 255] blk.27.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
313
- [ 255/ 255] blk.27.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq1_m .. size = 15.00 MiB -> 1.64 MiB
314
- llama_model_quantize_impl: model size = 2056.83 MiB (16.00 BPW)
315
- llama_model_quantize_impl: quant size = 292.64 MiB (2.28 BPW)
316
-
317
- llama_quantize: quantize time = 21151.08 ms
318
- llama_quantize: total time = 21151.08 ms
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
reproducibility/validation/quantize-logs/IQ2_XS.log DELETED
@@ -1,318 +0,0 @@
1
- ggml_cuda_init: failed to initialize CUDA: no CUDA-capable device is detected
2
- version: 0.3.0-dev (build 10671, commit 35999d101)
3
- built with GNU 13.3.0 for Linux x86_64
4
- llama_quantize: quantizing '/root/workspace/HF/source-official-bf16-k2-horizon-0.9b/K2-Horizon-0.9B-BF16.gguf' to '/root/workspace/HF/quantized-k2-horizon-0.9b/K2-Horizon-0.9B-IQ2_XS.gguf' as IQ2_XS using 128 threads
5
- llama_model_loader: loaded meta data with 41 key-value pairs and 255 tensors from /root/workspace/HF/source-official-bf16-k2-horizon-0.9b/K2-Horizon-0.9B-BF16.gguf (version GGUF V3 (latest))
6
- llama_model_loader: Dumping metadata keys/values. Note: KV overrides do not apply in this output.
7
- llama_model_loader: - kv 0: general.architecture str = k2-horizon
8
- llama_model_loader: - kv 1: general.type str = model
9
- llama_model_loader: - kv 2: general.name str = K2-Horizon-0.9B
10
- llama_model_loader: - kv 3: general.basename str = source-hf-k2-horizon
11
- llama_model_loader: - kv 4: general.size_label str = 0.9B
12
- llama_model_loader: - kv 5: general.license str = apache-2.0
13
- llama_model_loader: - kv 6: general.license.name str = internal-only
14
- llama_model_loader: - kv 7: general.license.link str = LICENSE
15
- llama_model_loader: - kv 8: general.tags arr[str,7] = ["k2-horizon", "0.9b", "dense", "reas...
16
- llama_model_loader: - kv 9: general.languages arr[str,2] = ["en", "zh"]
17
- llama_model_loader: - kv 10: k2-horizon.block_count u32 = 28
18
- llama_model_loader: - kv 11: k2-horizon.context_length u32 = 131072
19
- llama_model_loader: - kv 12: k2-horizon.embedding_length u32 = 1536
20
- llama_model_loader: - kv 13: k2-horizon.feed_forward_length u32 = 5120
21
- llama_model_loader: - kv 14: k2-horizon.attention.head_count u32 = 32
22
- llama_model_loader: - kv 15: k2-horizon.attention.head_count_kv u32 = 8
23
- llama_model_loader: - kv 16: k2-horizon.rope.scaling.type str = yarn
24
- llama_model_loader: - kv 17: k2-horizon.rope.scaling.factor f32 = 16.000000
25
- llama_model_loader: - kv 18: k2-horizon.rope.scaling.original_context_length u32 = 8192
26
- llama_model_loader: - kv 19: k2-horizon.rope.scaling.yarn_attn_factor f32 = 1.277259
27
- llama_model_loader: - kv 20: k2-horizon.rope.scaling.yarn_beta_fast f32 = 128.000000
28
- llama_model_loader: - kv 21: k2-horizon.rope.scaling.yarn_beta_slow f32 = 4.000000
29
- llama_model_loader: - kv 22: k2-horizon.rope.freq_base f32 = 1000000.000000
30
- llama_model_loader: - kv 23: k2-horizon.attention.layer_norm_rms_epsilon f32 = 0.000001
31
- llama_model_loader: - kv 24: k2-horizon.expert_count u32 = 0
32
- llama_model_loader: - kv 25: k2-horizon.expert_used_count u32 = 0
33
- llama_model_loader: - kv 26: k2-horizon.attention.key_length u32 = 64
34
- llama_model_loader: - kv 27: k2-horizon.attention.value_length u32 = 64
35
- llama_model_loader: - kv 28: general.file_type u32 = 32
36
- llama_model_loader: - kv 29: k2-horizon.attention.group_norm_groups u32 = 1
37
- llama_model_loader: - kv 30: k2-horizon.rope.dimension_count u32 = 64
38
- llama_model_loader: - kv 31: general.quantization_version u32 = 2
39
- llama_model_loader: - kv 32: tokenizer.ggml.model str = gpt2
40
- llama_model_loader: - kv 33: tokenizer.ggml.pre str = k2-horizon
41
- llama_model_loader: - kv 34: tokenizer.ggml.tokens arr[str,64256] = ["<|begin_of_text|>", "<|endoftext|>"...
42
- llama_model_loader: - kv 35: tokenizer.ggml.token_type arr[i32,64256] = [3, 3, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, ...
43
- llama_model_loader: - kv 36: tokenizer.ggml.merges arr[str,63742] = ["Δ  Δ ", "Ø Β§", "Γ™ Δ¦", "Γ  Β€", ...
44
- llama_model_loader: - kv 37: tokenizer.ggml.bos_token_id u32 = 0
45
- llama_model_loader: - kv 38: tokenizer.ggml.eos_token_id u32 = 1
46
- llama_model_loader: - kv 39: tokenizer.ggml.padding_token_id u32 = 64255
47
- llama_model_loader: - kv 40: tokenizer.chat_template str = {{- bos_token }}\n{%- if tool_presenta...
48
- llama_model_loader: - type f32: 57 tensors
49
- llama_model_loader: - type bf16: 198 tensors
50
-
51
- llama_model_quantize_impl: have importance matrix data with 196 entries
52
- [ 1/ 255] output.weight - [ 1536, 64256, 1, 1], type = bf16,
53
- ====== llama_model_quantize_impl: did not find weights for output.weight
54
- converting to q5_K .. load_imatrix: imatrix datasets=['calibration/wikitext-2-raw/wiki.train.raw', 'calibration/k2-horizon-0.9b/k2_horizon_en_zh_code_tool.txt']
55
- load_imatrix: loaded 196 importance matrix entries from /root/workspace/HF/calibration/k2-horizon-0.9b/k2_horizon_combined.imatrix.gguf computed on 99 chunks
56
- prepare_imatrix: have 196 importance matrix entries
57
- size = 188.25 MiB -> 64.71 MiB
58
- [ 2/ 255] output_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
59
- [ 3/ 255] token_embd.weight - [ 1536, 64256, 1, 1], type = bf16,
60
- ====== llama_model_quantize_impl: did not find weights for token_embd.weight
61
- converting to q2_K .. size = 188.25 MiB -> 30.88 MiB
62
- [ 4/ 255] blk.0.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to iq2_xs .. size = 1.50 MiB -> 0.22 MiB
63
- [ 5/ 255] blk.0.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
64
- [ 6/ 255] blk.0.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to iq2_xs .. size = 6.00 MiB -> 0.87 MiB
65
- [ 7/ 255] blk.0.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to iq2_xs .. size = 6.00 MiB -> 0.87 MiB
66
- [ 8/ 255] blk.0.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
67
- [ 9/ 255] blk.0.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q2_K .. size = 15.00 MiB -> 2.46 MiB
68
- [ 10/ 255] blk.0.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq2_xs .. size = 15.00 MiB -> 2.17 MiB
69
- [ 11/ 255] blk.0.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
70
- [ 12/ 255] blk.0.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq2_xs .. size = 15.00 MiB -> 2.17 MiB
71
- [ 13/ 255] blk.1.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to iq2_xs .. size = 1.50 MiB -> 0.22 MiB
72
- [ 14/ 255] blk.1.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
73
- [ 15/ 255] blk.1.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to iq2_xs .. size = 6.00 MiB -> 0.87 MiB
74
- [ 16/ 255] blk.1.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to iq2_xs .. size = 6.00 MiB -> 0.87 MiB
75
- [ 17/ 255] blk.1.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
76
- [ 18/ 255] blk.1.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q2_K .. size = 15.00 MiB -> 2.46 MiB
77
- [ 19/ 255] blk.1.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq2_xs .. size = 15.00 MiB -> 2.17 MiB
78
- [ 20/ 255] blk.1.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
79
- [ 21/ 255] blk.1.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq2_xs .. size = 15.00 MiB -> 2.17 MiB
80
- [ 22/ 255] blk.2.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to iq2_xs .. size = 1.50 MiB -> 0.22 MiB
81
- [ 23/ 255] blk.2.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
82
- [ 24/ 255] blk.2.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to iq2_xs .. size = 6.00 MiB -> 0.87 MiB
83
- [ 25/ 255] blk.2.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to iq2_xs .. size = 6.00 MiB -> 0.87 MiB
84
- [ 26/ 255] blk.2.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
85
- [ 27/ 255] blk.2.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q2_K .. size = 15.00 MiB -> 2.46 MiB
86
- [ 28/ 255] blk.2.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq2_xs .. size = 15.00 MiB -> 2.17 MiB
87
- [ 29/ 255] blk.2.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
88
- [ 30/ 255] blk.2.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq2_xs .. size = 15.00 MiB -> 2.17 MiB
89
- [ 31/ 255] blk.3.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to iq2_xs .. size = 1.50 MiB -> 0.22 MiB
90
- [ 32/ 255] blk.3.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
91
- [ 33/ 255] blk.3.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to iq2_xs .. size = 6.00 MiB -> 0.87 MiB
92
- [ 34/ 255] blk.3.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to iq2_xs .. size = 6.00 MiB -> 0.87 MiB
93
- [ 35/ 255] blk.3.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
94
- [ 36/ 255] blk.3.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to iq2_xs .. size = 15.00 MiB -> 2.17 MiB
95
- [ 37/ 255] blk.3.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq2_xs .. size = 15.00 MiB -> 2.17 MiB
96
- [ 38/ 255] blk.3.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
97
- [ 39/ 255] blk.3.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq2_xs .. size = 15.00 MiB -> 2.17 MiB
98
- [ 40/ 255] blk.4.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to iq2_xs .. size = 1.50 MiB -> 0.22 MiB
99
- [ 41/ 255] blk.4.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
100
- [ 42/ 255] blk.4.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to iq2_xs .. size = 6.00 MiB -> 0.87 MiB
101
- [ 43/ 255] blk.4.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to iq2_xs .. size = 6.00 MiB -> 0.87 MiB
102
- [ 44/ 255] blk.4.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
103
- [ 45/ 255] blk.4.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to iq2_xs .. size = 15.00 MiB -> 2.17 MiB
104
- [ 46/ 255] blk.4.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq2_xs .. size = 15.00 MiB -> 2.17 MiB
105
- [ 47/ 255] blk.4.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
106
- [ 48/ 255] blk.4.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq2_xs .. size = 15.00 MiB -> 2.17 MiB
107
- [ 49/ 255] blk.5.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to iq2_xs .. size = 1.50 MiB -> 0.22 MiB
108
- [ 50/ 255] blk.5.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
109
- [ 51/ 255] blk.5.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to iq2_xs .. size = 6.00 MiB -> 0.87 MiB
110
- [ 52/ 255] blk.5.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to iq2_xs .. size = 6.00 MiB -> 0.87 MiB
111
- [ 53/ 255] blk.5.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
112
- [ 54/ 255] blk.5.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to iq2_xs .. size = 15.00 MiB -> 2.17 MiB
113
- [ 55/ 255] blk.5.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq2_xs .. size = 15.00 MiB -> 2.17 MiB
114
- [ 56/ 255] blk.5.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
115
- [ 57/ 255] blk.5.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq2_xs .. size = 15.00 MiB -> 2.17 MiB
116
- [ 58/ 255] blk.6.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to iq2_xs .. size = 1.50 MiB -> 0.22 MiB
117
- [ 59/ 255] blk.6.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
118
- [ 60/ 255] blk.6.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to iq2_xs .. size = 6.00 MiB -> 0.87 MiB
119
- [ 61/ 255] blk.6.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to iq2_xs .. size = 6.00 MiB -> 0.87 MiB
120
- [ 62/ 255] blk.6.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
121
- [ 63/ 255] blk.6.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to iq2_xs .. size = 15.00 MiB -> 2.17 MiB
122
- [ 64/ 255] blk.6.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq2_xs .. size = 15.00 MiB -> 2.17 MiB
123
- [ 65/ 255] blk.6.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
124
- [ 66/ 255] blk.6.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq2_xs .. size = 15.00 MiB -> 2.17 MiB
125
- [ 67/ 255] blk.7.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to iq2_xs .. size = 1.50 MiB -> 0.22 MiB
126
- [ 68/ 255] blk.7.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
127
- [ 69/ 255] blk.7.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to iq2_xs .. size = 6.00 MiB -> 0.87 MiB
128
- [ 70/ 255] blk.7.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to iq2_xs .. size = 6.00 MiB -> 0.87 MiB
129
- [ 71/ 255] blk.7.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
130
- [ 72/ 255] blk.7.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to iq2_xs .. size = 15.00 MiB -> 2.17 MiB
131
- [ 73/ 255] blk.7.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq2_xs .. size = 15.00 MiB -> 2.17 MiB
132
- [ 74/ 255] blk.7.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
133
- [ 75/ 255] blk.7.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq2_xs .. size = 15.00 MiB -> 2.17 MiB
134
- [ 76/ 255] blk.8.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to iq2_xs .. size = 1.50 MiB -> 0.22 MiB
135
- [ 77/ 255] blk.8.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
136
- [ 78/ 255] blk.8.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to iq2_xs .. size = 6.00 MiB -> 0.87 MiB
137
- [ 79/ 255] blk.8.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to iq2_xs .. size = 6.00 MiB -> 0.87 MiB
138
- [ 80/ 255] blk.8.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
139
- [ 81/ 255] blk.8.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to iq2_xs .. size = 15.00 MiB -> 2.17 MiB
140
- [ 82/ 255] blk.8.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq2_xs .. size = 15.00 MiB -> 2.17 MiB
141
- [ 83/ 255] blk.8.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
142
- [ 84/ 255] blk.8.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq2_xs .. size = 15.00 MiB -> 2.17 MiB
143
- [ 85/ 255] blk.9.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to iq2_xs .. size = 1.50 MiB -> 0.22 MiB
144
- [ 86/ 255] blk.9.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
145
- [ 87/ 255] blk.9.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to iq2_xs .. size = 6.00 MiB -> 0.87 MiB
146
- [ 88/ 255] blk.9.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to iq2_xs .. size = 6.00 MiB -> 0.87 MiB
147
- [ 89/ 255] blk.9.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
148
- [ 90/ 255] blk.9.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to iq2_xs .. size = 15.00 MiB -> 2.17 MiB
149
- [ 91/ 255] blk.9.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq2_xs .. size = 15.00 MiB -> 2.17 MiB
150
- [ 92/ 255] blk.9.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
151
- [ 93/ 255] blk.9.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq2_xs .. size = 15.00 MiB -> 2.17 MiB
152
- [ 94/ 255] blk.10.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to iq2_xs .. size = 1.50 MiB -> 0.22 MiB
153
- [ 95/ 255] blk.10.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
154
- [ 96/ 255] blk.10.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to iq2_xs .. size = 6.00 MiB -> 0.87 MiB
155
- [ 97/ 255] blk.10.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to iq2_xs .. size = 6.00 MiB -> 0.87 MiB
156
- [ 98/ 255] blk.10.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
157
- [ 99/ 255] blk.10.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to iq2_xs .. size = 15.00 MiB -> 2.17 MiB
158
- [ 100/ 255] blk.10.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq2_xs .. size = 15.00 MiB -> 2.17 MiB
159
- [ 101/ 255] blk.10.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
160
- [ 102/ 255] blk.10.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq2_xs .. size = 15.00 MiB -> 2.17 MiB
161
- [ 103/ 255] blk.11.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to iq2_xs .. size = 1.50 MiB -> 0.22 MiB
162
- [ 104/ 255] blk.11.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
163
- [ 105/ 255] blk.11.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to iq2_xs .. size = 6.00 MiB -> 0.87 MiB
164
- [ 106/ 255] blk.11.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to iq2_xs .. size = 6.00 MiB -> 0.87 MiB
165
- [ 107/ 255] blk.11.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
166
- [ 108/ 255] blk.11.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to iq2_xs .. size = 15.00 MiB -> 2.17 MiB
167
- [ 109/ 255] blk.11.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq2_xs .. size = 15.00 MiB -> 2.17 MiB
168
- [ 110/ 255] blk.11.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
169
- [ 111/ 255] blk.11.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq2_xs .. size = 15.00 MiB -> 2.17 MiB
170
- [ 112/ 255] blk.12.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to iq2_xs .. size = 1.50 MiB -> 0.22 MiB
171
- [ 113/ 255] blk.12.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
172
- [ 114/ 255] blk.12.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to iq2_xs .. size = 6.00 MiB -> 0.87 MiB
173
- [ 115/ 255] blk.12.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to iq2_xs .. size = 6.00 MiB -> 0.87 MiB
174
- [ 116/ 255] blk.12.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
175
- [ 117/ 255] blk.12.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to iq2_xs .. size = 15.00 MiB -> 2.17 MiB
176
- [ 118/ 255] blk.12.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq2_xs .. size = 15.00 MiB -> 2.17 MiB
177
- [ 119/ 255] blk.12.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
178
- [ 120/ 255] blk.12.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq2_xs .. size = 15.00 MiB -> 2.17 MiB
179
- [ 121/ 255] blk.13.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to iq2_xs .. size = 1.50 MiB -> 0.22 MiB
180
- [ 122/ 255] blk.13.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
181
- [ 123/ 255] blk.13.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to iq2_xs .. size = 6.00 MiB -> 0.87 MiB
182
- [ 124/ 255] blk.13.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to iq2_xs .. size = 6.00 MiB -> 0.87 MiB
183
- [ 125/ 255] blk.13.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
184
- [ 126/ 255] blk.13.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to iq2_xs .. size = 15.00 MiB -> 2.17 MiB
185
- [ 127/ 255] blk.13.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq2_xs .. size = 15.00 MiB -> 2.17 MiB
186
- [ 128/ 255] blk.13.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
187
- [ 129/ 255] blk.13.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq2_xs .. size = 15.00 MiB -> 2.17 MiB
188
- [ 130/ 255] blk.14.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to iq2_xs .. size = 1.50 MiB -> 0.22 MiB
189
- [ 131/ 255] blk.14.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
190
- [ 132/ 255] blk.14.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to iq2_xs .. size = 6.00 MiB -> 0.87 MiB
191
- [ 133/ 255] blk.14.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to iq2_xs .. size = 6.00 MiB -> 0.87 MiB
192
- [ 134/ 255] blk.14.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
193
- [ 135/ 255] blk.14.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to iq2_xs .. size = 15.00 MiB -> 2.17 MiB
194
- [ 136/ 255] blk.14.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq2_xs .. size = 15.00 MiB -> 2.17 MiB
195
- [ 137/ 255] blk.14.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
196
- [ 138/ 255] blk.14.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq2_xs .. size = 15.00 MiB -> 2.17 MiB
197
- [ 139/ 255] blk.15.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to iq2_xs .. size = 1.50 MiB -> 0.22 MiB
198
- [ 140/ 255] blk.15.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
199
- [ 141/ 255] blk.15.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to iq2_xs .. size = 6.00 MiB -> 0.87 MiB
200
- [ 142/ 255] blk.15.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to iq2_xs .. size = 6.00 MiB -> 0.87 MiB
201
- [ 143/ 255] blk.15.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
202
- [ 144/ 255] blk.15.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to iq2_xs .. size = 15.00 MiB -> 2.17 MiB
203
- [ 145/ 255] blk.15.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq2_xs .. size = 15.00 MiB -> 2.17 MiB
204
- [ 146/ 255] blk.15.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
205
- [ 147/ 255] blk.15.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq2_xs .. size = 15.00 MiB -> 2.17 MiB
206
- [ 148/ 255] blk.16.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to iq2_xs .. size = 1.50 MiB -> 0.22 MiB
207
- [ 149/ 255] blk.16.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
208
- [ 150/ 255] blk.16.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to iq2_xs .. size = 6.00 MiB -> 0.87 MiB
209
- [ 151/ 255] blk.16.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to iq2_xs .. size = 6.00 MiB -> 0.87 MiB
210
- [ 152/ 255] blk.16.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
211
- [ 153/ 255] blk.16.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to iq2_xs .. size = 15.00 MiB -> 2.17 MiB
212
- [ 154/ 255] blk.16.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq2_xs .. size = 15.00 MiB -> 2.17 MiB
213
- [ 155/ 255] blk.16.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
214
- [ 156/ 255] blk.16.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq2_xs .. size = 15.00 MiB -> 2.17 MiB
215
- [ 157/ 255] blk.17.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to iq2_xs .. size = 1.50 MiB -> 0.22 MiB
216
- [ 158/ 255] blk.17.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
217
- [ 159/ 255] blk.17.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to iq2_xs .. size = 6.00 MiB -> 0.87 MiB
218
- [ 160/ 255] blk.17.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to iq2_xs .. size = 6.00 MiB -> 0.87 MiB
219
- [ 161/ 255] blk.17.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
220
- [ 162/ 255] blk.17.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to iq2_xs .. size = 15.00 MiB -> 2.17 MiB
221
- [ 163/ 255] blk.17.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq2_xs .. size = 15.00 MiB -> 2.17 MiB
222
- [ 164/ 255] blk.17.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
223
- [ 165/ 255] blk.17.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq2_xs .. size = 15.00 MiB -> 2.17 MiB
224
- [ 166/ 255] blk.18.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to iq2_xs .. size = 1.50 MiB -> 0.22 MiB
225
- [ 167/ 255] blk.18.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
226
- [ 168/ 255] blk.18.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to iq2_xs .. size = 6.00 MiB -> 0.87 MiB
227
- [ 169/ 255] blk.18.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to iq2_xs .. size = 6.00 MiB -> 0.87 MiB
228
- [ 170/ 255] blk.18.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
229
- [ 171/ 255] blk.18.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to iq2_xs .. size = 15.00 MiB -> 2.17 MiB
230
- [ 172/ 255] blk.18.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq2_xs .. size = 15.00 MiB -> 2.17 MiB
231
- [ 173/ 255] blk.18.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
232
- [ 174/ 255] blk.18.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq2_xs .. size = 15.00 MiB -> 2.17 MiB
233
- [ 175/ 255] blk.19.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to iq2_xs .. size = 1.50 MiB -> 0.22 MiB
234
- [ 176/ 255] blk.19.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
235
- [ 177/ 255] blk.19.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to iq2_xs .. size = 6.00 MiB -> 0.87 MiB
236
- [ 178/ 255] blk.19.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to iq2_xs .. size = 6.00 MiB -> 0.87 MiB
237
- [ 179/ 255] blk.19.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
238
- [ 180/ 255] blk.19.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to iq2_xs .. size = 15.00 MiB -> 2.17 MiB
239
- [ 181/ 255] blk.19.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq2_xs .. size = 15.00 MiB -> 2.17 MiB
240
- [ 182/ 255] blk.19.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
241
- [ 183/ 255] blk.19.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq2_xs .. size = 15.00 MiB -> 2.17 MiB
242
- [ 184/ 255] blk.20.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to iq2_xs .. size = 1.50 MiB -> 0.22 MiB
243
- [ 185/ 255] blk.20.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
244
- [ 186/ 255] blk.20.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to iq2_xs .. size = 6.00 MiB -> 0.87 MiB
245
- [ 187/ 255] blk.20.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to iq2_xs .. size = 6.00 MiB -> 0.87 MiB
246
- [ 188/ 255] blk.20.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
247
- [ 189/ 255] blk.20.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to iq2_xs .. size = 15.00 MiB -> 2.17 MiB
248
- [ 190/ 255] blk.20.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq2_xs .. size = 15.00 MiB -> 2.17 MiB
249
- [ 191/ 255] blk.20.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
250
- [ 192/ 255] blk.20.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq2_xs .. size = 15.00 MiB -> 2.17 MiB
251
- [ 193/ 255] blk.21.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to iq2_xs .. size = 1.50 MiB -> 0.22 MiB
252
- [ 194/ 255] blk.21.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
253
- [ 195/ 255] blk.21.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to iq2_xs .. size = 6.00 MiB -> 0.87 MiB
254
- [ 196/ 255] blk.21.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to iq2_xs .. size = 6.00 MiB -> 0.87 MiB
255
- [ 197/ 255] blk.21.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
256
- [ 198/ 255] blk.21.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to iq2_xs .. size = 15.00 MiB -> 2.17 MiB
257
- [ 199/ 255] blk.21.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq2_xs .. size = 15.00 MiB -> 2.17 MiB
258
- [ 200/ 255] blk.21.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
259
- [ 201/ 255] blk.21.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq2_xs .. size = 15.00 MiB -> 2.17 MiB
260
- [ 202/ 255] blk.22.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to iq2_xs .. size = 1.50 MiB -> 0.22 MiB
261
- [ 203/ 255] blk.22.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
262
- [ 204/ 255] blk.22.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to iq2_xs .. size = 6.00 MiB -> 0.87 MiB
263
- [ 205/ 255] blk.22.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to iq2_xs .. size = 6.00 MiB -> 0.87 MiB
264
- [ 206/ 255] blk.22.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
265
- [ 207/ 255] blk.22.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to iq2_xs .. size = 15.00 MiB -> 2.17 MiB
266
- [ 208/ 255] blk.22.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq2_xs .. size = 15.00 MiB -> 2.17 MiB
267
- [ 209/ 255] blk.22.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
268
- [ 210/ 255] blk.22.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq2_xs .. size = 15.00 MiB -> 2.17 MiB
269
- [ 211/ 255] blk.23.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to iq2_xs .. size = 1.50 MiB -> 0.22 MiB
270
- [ 212/ 255] blk.23.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
271
- [ 213/ 255] blk.23.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to iq2_xs .. size = 6.00 MiB -> 0.87 MiB
272
- [ 214/ 255] blk.23.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to iq2_xs .. size = 6.00 MiB -> 0.87 MiB
273
- [ 215/ 255] blk.23.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
274
- [ 216/ 255] blk.23.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to iq2_xs .. size = 15.00 MiB -> 2.17 MiB
275
- [ 217/ 255] blk.23.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq2_xs .. size = 15.00 MiB -> 2.17 MiB
276
- [ 218/ 255] blk.23.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
277
- [ 219/ 255] blk.23.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq2_xs .. size = 15.00 MiB -> 2.17 MiB
278
- [ 220/ 255] blk.24.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to iq2_xs .. size = 1.50 MiB -> 0.22 MiB
279
- [ 221/ 255] blk.24.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
280
- [ 222/ 255] blk.24.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to iq2_xs .. size = 6.00 MiB -> 0.87 MiB
281
- [ 223/ 255] blk.24.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to iq2_xs .. size = 6.00 MiB -> 0.87 MiB
282
- [ 224/ 255] blk.24.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
283
- [ 225/ 255] blk.24.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to iq2_xs .. size = 15.00 MiB -> 2.17 MiB
284
- [ 226/ 255] blk.24.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq2_xs .. size = 15.00 MiB -> 2.17 MiB
285
- [ 227/ 255] blk.24.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
286
- [ 228/ 255] blk.24.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq2_xs .. size = 15.00 MiB -> 2.17 MiB
287
- [ 229/ 255] blk.25.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to iq2_xs .. size = 1.50 MiB -> 0.22 MiB
288
- [ 230/ 255] blk.25.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
289
- [ 231/ 255] blk.25.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to iq2_xs .. size = 6.00 MiB -> 0.87 MiB
290
- [ 232/ 255] blk.25.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to iq2_xs .. size = 6.00 MiB -> 0.87 MiB
291
- [ 233/ 255] blk.25.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
292
- [ 234/ 255] blk.25.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to iq2_xs .. size = 15.00 MiB -> 2.17 MiB
293
- [ 235/ 255] blk.25.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq2_xs .. size = 15.00 MiB -> 2.17 MiB
294
- [ 236/ 255] blk.25.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
295
- [ 237/ 255] blk.25.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq2_xs .. size = 15.00 MiB -> 2.17 MiB
296
- [ 238/ 255] blk.26.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to iq2_xs .. size = 1.50 MiB -> 0.22 MiB
297
- [ 239/ 255] blk.26.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
298
- [ 240/ 255] blk.26.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to iq2_xs .. size = 6.00 MiB -> 0.87 MiB
299
- [ 241/ 255] blk.26.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to iq2_xs .. size = 6.00 MiB -> 0.87 MiB
300
- [ 242/ 255] blk.26.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
301
- [ 243/ 255] blk.26.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to iq2_xs .. size = 15.00 MiB -> 2.17 MiB
302
- [ 244/ 255] blk.26.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq2_xs .. size = 15.00 MiB -> 2.17 MiB
303
- [ 245/ 255] blk.26.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
304
- [ 246/ 255] blk.26.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq2_xs .. size = 15.00 MiB -> 2.17 MiB
305
- [ 247/ 255] blk.27.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to iq2_xs .. size = 1.50 MiB -> 0.22 MiB
306
- [ 248/ 255] blk.27.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
307
- [ 249/ 255] blk.27.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to iq2_xs .. size = 6.00 MiB -> 0.87 MiB
308
- [ 250/ 255] blk.27.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to iq2_xs .. size = 6.00 MiB -> 0.87 MiB
309
- [ 251/ 255] blk.27.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
310
- [ 252/ 255] blk.27.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to iq2_xs .. size = 15.00 MiB -> 2.17 MiB
311
- [ 253/ 255] blk.27.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq2_xs .. size = 15.00 MiB -> 2.17 MiB
312
- [ 254/ 255] blk.27.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
313
- [ 255/ 255] blk.27.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to iq2_xs .. size = 15.00 MiB -> 2.17 MiB
314
- llama_model_quantize_impl: model size = 2056.83 MiB (16.00 BPW)
315
- llama_model_quantize_impl: quant size = 345.36 MiB (2.69 BPW)
316
-
317
- llama_quantize: quantize time = 22938.39 ms
318
- llama_quantize: total time = 22938.39 ms
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
reproducibility/validation/quantize-logs/Q1_0.log DELETED
@@ -1,318 +0,0 @@
1
- ggml_cuda_init: failed to initialize CUDA: no CUDA-capable device is detected
2
- version: 0.3.0-dev (build 10671, commit 35999d101)
3
- built with GNU 13.3.0 for Linux x86_64
4
- llama_quantize: quantizing '/root/workspace/HF/source-official-bf16-k2-horizon-0.9b/K2-Horizon-0.9B-BF16.gguf' to '/root/workspace/HF/quantized-k2-horizon-0.9b/K2-Horizon-0.9B-Q1_0.gguf' as Q1_0 using 128 threads
5
- llama_model_loader: loaded meta data with 41 key-value pairs and 255 tensors from /root/workspace/HF/source-official-bf16-k2-horizon-0.9b/K2-Horizon-0.9B-BF16.gguf (version GGUF V3 (latest))
6
- llama_model_loader: Dumping metadata keys/values. Note: KV overrides do not apply in this output.
7
- llama_model_loader: - kv 0: general.architecture str = k2-horizon
8
- llama_model_loader: - kv 1: general.type str = model
9
- llama_model_loader: - kv 2: general.name str = K2-Horizon-0.9B
10
- llama_model_loader: - kv 3: general.basename str = source-hf-k2-horizon
11
- llama_model_loader: - kv 4: general.size_label str = 0.9B
12
- llama_model_loader: - kv 5: general.license str = apache-2.0
13
- llama_model_loader: - kv 6: general.license.name str = internal-only
14
- llama_model_loader: - kv 7: general.license.link str = LICENSE
15
- llama_model_loader: - kv 8: general.tags arr[str,7] = ["k2-horizon", "0.9b", "dense", "reas...
16
- llama_model_loader: - kv 9: general.languages arr[str,2] = ["en", "zh"]
17
- llama_model_loader: - kv 10: k2-horizon.block_count u32 = 28
18
- llama_model_loader: - kv 11: k2-horizon.context_length u32 = 131072
19
- llama_model_loader: - kv 12: k2-horizon.embedding_length u32 = 1536
20
- llama_model_loader: - kv 13: k2-horizon.feed_forward_length u32 = 5120
21
- llama_model_loader: - kv 14: k2-horizon.attention.head_count u32 = 32
22
- llama_model_loader: - kv 15: k2-horizon.attention.head_count_kv u32 = 8
23
- llama_model_loader: - kv 16: k2-horizon.rope.scaling.type str = yarn
24
- llama_model_loader: - kv 17: k2-horizon.rope.scaling.factor f32 = 16.000000
25
- llama_model_loader: - kv 18: k2-horizon.rope.scaling.original_context_length u32 = 8192
26
- llama_model_loader: - kv 19: k2-horizon.rope.scaling.yarn_attn_factor f32 = 1.277259
27
- llama_model_loader: - kv 20: k2-horizon.rope.scaling.yarn_beta_fast f32 = 128.000000
28
- llama_model_loader: - kv 21: k2-horizon.rope.scaling.yarn_beta_slow f32 = 4.000000
29
- llama_model_loader: - kv 22: k2-horizon.rope.freq_base f32 = 1000000.000000
30
- llama_model_loader: - kv 23: k2-horizon.attention.layer_norm_rms_epsilon f32 = 0.000001
31
- llama_model_loader: - kv 24: k2-horizon.expert_count u32 = 0
32
- llama_model_loader: - kv 25: k2-horizon.expert_used_count u32 = 0
33
- llama_model_loader: - kv 26: k2-horizon.attention.key_length u32 = 64
34
- llama_model_loader: - kv 27: k2-horizon.attention.value_length u32 = 64
35
- llama_model_loader: - kv 28: general.file_type u32 = 32
36
- llama_model_loader: - kv 29: k2-horizon.attention.group_norm_groups u32 = 1
37
- llama_model_loader: - kv 30: k2-horizon.rope.dimension_count u32 = 64
38
- llama_model_loader: - kv 31: general.quantization_version u32 = 2
39
- llama_model_loader: - kv 32: tokenizer.ggml.model str = gpt2
40
- llama_model_loader: - kv 33: tokenizer.ggml.pre str = k2-horizon
41
- llama_model_loader: - kv 34: tokenizer.ggml.tokens arr[str,64256] = ["<|begin_of_text|>", "<|endoftext|>"...
42
- llama_model_loader: - kv 35: tokenizer.ggml.token_type arr[i32,64256] = [3, 3, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, ...
43
- llama_model_loader: - kv 36: tokenizer.ggml.merges arr[str,63742] = ["Δ  Δ ", "Ø Β§", "Γ™ Δ¦", "Γ  Β€", ...
44
- llama_model_loader: - kv 37: tokenizer.ggml.bos_token_id u32 = 0
45
- llama_model_loader: - kv 38: tokenizer.ggml.eos_token_id u32 = 1
46
- llama_model_loader: - kv 39: tokenizer.ggml.padding_token_id u32 = 64255
47
- llama_model_loader: - kv 40: tokenizer.chat_template str = {{- bos_token }}\n{%- if tool_presenta...
48
- llama_model_loader: - type f32: 57 tensors
49
- llama_model_loader: - type bf16: 198 tensors
50
-
51
- llama_model_quantize_impl: have importance matrix data with 196 entries
52
- [ 1/ 255] output.weight - [ 1536, 64256, 1, 1], type = bf16,
53
- ====== llama_model_quantize_impl: did not find weights for output.weight
54
- converting to q6_K .. load_imatrix: imatrix datasets=['calibration/wikitext-2-raw/wiki.train.raw', 'calibration/k2-horizon-0.9b/k2_horizon_en_zh_code_tool.txt']
55
- load_imatrix: loaded 196 importance matrix entries from /root/workspace/HF/calibration/k2-horizon-0.9b/k2_horizon_combined.imatrix.gguf computed on 99 chunks
56
- prepare_imatrix: have 196 importance matrix entries
57
- size = 188.25 MiB -> 77.21 MiB
58
- [ 2/ 255] output_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
59
- [ 3/ 255] token_embd.weight - [ 1536, 64256, 1, 1], type = bf16,
60
- ====== llama_model_quantize_impl: did not find weights for token_embd.weight
61
- converting to q1_0 .. size = 188.25 MiB -> 13.24 MiB
62
- [ 4/ 255] blk.0.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q1_0 .. size = 1.50 MiB -> 0.11 MiB
63
- [ 5/ 255] blk.0.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
64
- [ 6/ 255] blk.0.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q1_0 .. size = 6.00 MiB -> 0.42 MiB
65
- [ 7/ 255] blk.0.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q1_0 .. size = 6.00 MiB -> 0.42 MiB
66
- [ 8/ 255] blk.0.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q1_0 .. size = 1.50 MiB -> 0.11 MiB
67
- [ 9/ 255] blk.0.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q1_0 .. size = 15.00 MiB -> 1.05 MiB
68
- [ 10/ 255] blk.0.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q1_0 .. size = 15.00 MiB -> 1.05 MiB
69
- [ 11/ 255] blk.0.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
70
- [ 12/ 255] blk.0.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q1_0 .. size = 15.00 MiB -> 1.05 MiB
71
- [ 13/ 255] blk.1.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q1_0 .. size = 1.50 MiB -> 0.11 MiB
72
- [ 14/ 255] blk.1.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
73
- [ 15/ 255] blk.1.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q1_0 .. size = 6.00 MiB -> 0.42 MiB
74
- [ 16/ 255] blk.1.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q1_0 .. size = 6.00 MiB -> 0.42 MiB
75
- [ 17/ 255] blk.1.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q1_0 .. size = 1.50 MiB -> 0.11 MiB
76
- [ 18/ 255] blk.1.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q1_0 .. size = 15.00 MiB -> 1.05 MiB
77
- [ 19/ 255] blk.1.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q1_0 .. size = 15.00 MiB -> 1.05 MiB
78
- [ 20/ 255] blk.1.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
79
- [ 21/ 255] blk.1.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q1_0 .. size = 15.00 MiB -> 1.05 MiB
80
- [ 22/ 255] blk.2.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q1_0 .. size = 1.50 MiB -> 0.11 MiB
81
- [ 23/ 255] blk.2.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
82
- [ 24/ 255] blk.2.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q1_0 .. size = 6.00 MiB -> 0.42 MiB
83
- [ 25/ 255] blk.2.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q1_0 .. size = 6.00 MiB -> 0.42 MiB
84
- [ 26/ 255] blk.2.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q1_0 .. size = 1.50 MiB -> 0.11 MiB
85
- [ 27/ 255] blk.2.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q1_0 .. size = 15.00 MiB -> 1.05 MiB
86
- [ 28/ 255] blk.2.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q1_0 .. size = 15.00 MiB -> 1.05 MiB
87
- [ 29/ 255] blk.2.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
88
- [ 30/ 255] blk.2.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q1_0 .. size = 15.00 MiB -> 1.05 MiB
89
- [ 31/ 255] blk.3.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q1_0 .. size = 1.50 MiB -> 0.11 MiB
90
- [ 32/ 255] blk.3.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
91
- [ 33/ 255] blk.3.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q1_0 .. size = 6.00 MiB -> 0.42 MiB
92
- [ 34/ 255] blk.3.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q1_0 .. size = 6.00 MiB -> 0.42 MiB
93
- [ 35/ 255] blk.3.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q1_0 .. size = 1.50 MiB -> 0.11 MiB
94
- [ 36/ 255] blk.3.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q1_0 .. size = 15.00 MiB -> 1.05 MiB
95
- [ 37/ 255] blk.3.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q1_0 .. size = 15.00 MiB -> 1.05 MiB
96
- [ 38/ 255] blk.3.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
97
- [ 39/ 255] blk.3.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q1_0 .. size = 15.00 MiB -> 1.05 MiB
98
- [ 40/ 255] blk.4.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q1_0 .. size = 1.50 MiB -> 0.11 MiB
99
- [ 41/ 255] blk.4.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
100
- [ 42/ 255] blk.4.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q1_0 .. size = 6.00 MiB -> 0.42 MiB
101
- [ 43/ 255] blk.4.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q1_0 .. size = 6.00 MiB -> 0.42 MiB
102
- [ 44/ 255] blk.4.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q1_0 .. size = 1.50 MiB -> 0.11 MiB
103
- [ 45/ 255] blk.4.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q1_0 .. size = 15.00 MiB -> 1.05 MiB
104
- [ 46/ 255] blk.4.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q1_0 .. size = 15.00 MiB -> 1.05 MiB
105
- [ 47/ 255] blk.4.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
106
- [ 48/ 255] blk.4.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q1_0 .. size = 15.00 MiB -> 1.05 MiB
107
- [ 49/ 255] blk.5.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q1_0 .. size = 1.50 MiB -> 0.11 MiB
108
- [ 50/ 255] blk.5.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
109
- [ 51/ 255] blk.5.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q1_0 .. size = 6.00 MiB -> 0.42 MiB
110
- [ 52/ 255] blk.5.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q1_0 .. size = 6.00 MiB -> 0.42 MiB
111
- [ 53/ 255] blk.5.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q1_0 .. size = 1.50 MiB -> 0.11 MiB
112
- [ 54/ 255] blk.5.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q1_0 .. size = 15.00 MiB -> 1.05 MiB
113
- [ 55/ 255] blk.5.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q1_0 .. size = 15.00 MiB -> 1.05 MiB
114
- [ 56/ 255] blk.5.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
115
- [ 57/ 255] blk.5.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q1_0 .. size = 15.00 MiB -> 1.05 MiB
116
- [ 58/ 255] blk.6.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q1_0 .. size = 1.50 MiB -> 0.11 MiB
117
- [ 59/ 255] blk.6.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
118
- [ 60/ 255] blk.6.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q1_0 .. size = 6.00 MiB -> 0.42 MiB
119
- [ 61/ 255] blk.6.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q1_0 .. size = 6.00 MiB -> 0.42 MiB
120
- [ 62/ 255] blk.6.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q1_0 .. size = 1.50 MiB -> 0.11 MiB
121
- [ 63/ 255] blk.6.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q1_0 .. size = 15.00 MiB -> 1.05 MiB
122
- [ 64/ 255] blk.6.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q1_0 .. size = 15.00 MiB -> 1.05 MiB
123
- [ 65/ 255] blk.6.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
124
- [ 66/ 255] blk.6.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q1_0 .. size = 15.00 MiB -> 1.05 MiB
125
- [ 67/ 255] blk.7.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q1_0 .. size = 1.50 MiB -> 0.11 MiB
126
- [ 68/ 255] blk.7.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
127
- [ 69/ 255] blk.7.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q1_0 .. size = 6.00 MiB -> 0.42 MiB
128
- [ 70/ 255] blk.7.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q1_0 .. size = 6.00 MiB -> 0.42 MiB
129
- [ 71/ 255] blk.7.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q1_0 .. size = 1.50 MiB -> 0.11 MiB
130
- [ 72/ 255] blk.7.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q1_0 .. size = 15.00 MiB -> 1.05 MiB
131
- [ 73/ 255] blk.7.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q1_0 .. size = 15.00 MiB -> 1.05 MiB
132
- [ 74/ 255] blk.7.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
133
- [ 75/ 255] blk.7.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q1_0 .. size = 15.00 MiB -> 1.05 MiB
134
- [ 76/ 255] blk.8.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q1_0 .. size = 1.50 MiB -> 0.11 MiB
135
- [ 77/ 255] blk.8.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
136
- [ 78/ 255] blk.8.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q1_0 .. size = 6.00 MiB -> 0.42 MiB
137
- [ 79/ 255] blk.8.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q1_0 .. size = 6.00 MiB -> 0.42 MiB
138
- [ 80/ 255] blk.8.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q1_0 .. size = 1.50 MiB -> 0.11 MiB
139
- [ 81/ 255] blk.8.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q1_0 .. size = 15.00 MiB -> 1.05 MiB
140
- [ 82/ 255] blk.8.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q1_0 .. size = 15.00 MiB -> 1.05 MiB
141
- [ 83/ 255] blk.8.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
142
- [ 84/ 255] blk.8.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q1_0 .. size = 15.00 MiB -> 1.05 MiB
143
- [ 85/ 255] blk.9.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q1_0 .. size = 1.50 MiB -> 0.11 MiB
144
- [ 86/ 255] blk.9.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
145
- [ 87/ 255] blk.9.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q1_0 .. size = 6.00 MiB -> 0.42 MiB
146
- [ 88/ 255] blk.9.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q1_0 .. size = 6.00 MiB -> 0.42 MiB
147
- [ 89/ 255] blk.9.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q1_0 .. size = 1.50 MiB -> 0.11 MiB
148
- [ 90/ 255] blk.9.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q1_0 .. size = 15.00 MiB -> 1.05 MiB
149
- [ 91/ 255] blk.9.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q1_0 .. size = 15.00 MiB -> 1.05 MiB
150
- [ 92/ 255] blk.9.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
151
- [ 93/ 255] blk.9.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q1_0 .. size = 15.00 MiB -> 1.05 MiB
152
- [ 94/ 255] blk.10.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q1_0 .. size = 1.50 MiB -> 0.11 MiB
153
- [ 95/ 255] blk.10.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
154
- [ 96/ 255] blk.10.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q1_0 .. size = 6.00 MiB -> 0.42 MiB
155
- [ 97/ 255] blk.10.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q1_0 .. size = 6.00 MiB -> 0.42 MiB
156
- [ 98/ 255] blk.10.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q1_0 .. size = 1.50 MiB -> 0.11 MiB
157
- [ 99/ 255] blk.10.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q1_0 .. size = 15.00 MiB -> 1.05 MiB
158
- [ 100/ 255] blk.10.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q1_0 .. size = 15.00 MiB -> 1.05 MiB
159
- [ 101/ 255] blk.10.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
160
- [ 102/ 255] blk.10.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q1_0 .. size = 15.00 MiB -> 1.05 MiB
161
- [ 103/ 255] blk.11.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q1_0 .. size = 1.50 MiB -> 0.11 MiB
162
- [ 104/ 255] blk.11.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
163
- [ 105/ 255] blk.11.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q1_0 .. size = 6.00 MiB -> 0.42 MiB
164
- [ 106/ 255] blk.11.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q1_0 .. size = 6.00 MiB -> 0.42 MiB
165
- [ 107/ 255] blk.11.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q1_0 .. size = 1.50 MiB -> 0.11 MiB
166
- [ 108/ 255] blk.11.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q1_0 .. size = 15.00 MiB -> 1.05 MiB
167
- [ 109/ 255] blk.11.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q1_0 .. size = 15.00 MiB -> 1.05 MiB
168
- [ 110/ 255] blk.11.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
169
- [ 111/ 255] blk.11.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q1_0 .. size = 15.00 MiB -> 1.05 MiB
170
- [ 112/ 255] blk.12.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q1_0 .. size = 1.50 MiB -> 0.11 MiB
171
- [ 113/ 255] blk.12.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
172
- [ 114/ 255] blk.12.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q1_0 .. size = 6.00 MiB -> 0.42 MiB
173
- [ 115/ 255] blk.12.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q1_0 .. size = 6.00 MiB -> 0.42 MiB
174
- [ 116/ 255] blk.12.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q1_0 .. size = 1.50 MiB -> 0.11 MiB
175
- [ 117/ 255] blk.12.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q1_0 .. size = 15.00 MiB -> 1.05 MiB
176
- [ 118/ 255] blk.12.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q1_0 .. size = 15.00 MiB -> 1.05 MiB
177
- [ 119/ 255] blk.12.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
178
- [ 120/ 255] blk.12.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q1_0 .. size = 15.00 MiB -> 1.05 MiB
179
- [ 121/ 255] blk.13.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q1_0 .. size = 1.50 MiB -> 0.11 MiB
180
- [ 122/ 255] blk.13.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
181
- [ 123/ 255] blk.13.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q1_0 .. size = 6.00 MiB -> 0.42 MiB
182
- [ 124/ 255] blk.13.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q1_0 .. size = 6.00 MiB -> 0.42 MiB
183
- [ 125/ 255] blk.13.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q1_0 .. size = 1.50 MiB -> 0.11 MiB
184
- [ 126/ 255] blk.13.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q1_0 .. size = 15.00 MiB -> 1.05 MiB
185
- [ 127/ 255] blk.13.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q1_0 .. size = 15.00 MiB -> 1.05 MiB
186
- [ 128/ 255] blk.13.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
187
- [ 129/ 255] blk.13.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q1_0 .. size = 15.00 MiB -> 1.05 MiB
188
- [ 130/ 255] blk.14.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q1_0 .. size = 1.50 MiB -> 0.11 MiB
189
- [ 131/ 255] blk.14.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
190
- [ 132/ 255] blk.14.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q1_0 .. size = 6.00 MiB -> 0.42 MiB
191
- [ 133/ 255] blk.14.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q1_0 .. size = 6.00 MiB -> 0.42 MiB
192
- [ 134/ 255] blk.14.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q1_0 .. size = 1.50 MiB -> 0.11 MiB
193
- [ 135/ 255] blk.14.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q1_0 .. size = 15.00 MiB -> 1.05 MiB
194
- [ 136/ 255] blk.14.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q1_0 .. size = 15.00 MiB -> 1.05 MiB
195
- [ 137/ 255] blk.14.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
196
- [ 138/ 255] blk.14.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q1_0 .. size = 15.00 MiB -> 1.05 MiB
197
- [ 139/ 255] blk.15.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q1_0 .. size = 1.50 MiB -> 0.11 MiB
198
- [ 140/ 255] blk.15.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
199
- [ 141/ 255] blk.15.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q1_0 .. size = 6.00 MiB -> 0.42 MiB
200
- [ 142/ 255] blk.15.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q1_0 .. size = 6.00 MiB -> 0.42 MiB
201
- [ 143/ 255] blk.15.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q1_0 .. size = 1.50 MiB -> 0.11 MiB
202
- [ 144/ 255] blk.15.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q1_0 .. size = 15.00 MiB -> 1.05 MiB
203
- [ 145/ 255] blk.15.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q1_0 .. size = 15.00 MiB -> 1.05 MiB
204
- [ 146/ 255] blk.15.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
205
- [ 147/ 255] blk.15.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q1_0 .. size = 15.00 MiB -> 1.05 MiB
206
- [ 148/ 255] blk.16.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q1_0 .. size = 1.50 MiB -> 0.11 MiB
207
- [ 149/ 255] blk.16.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
208
- [ 150/ 255] blk.16.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q1_0 .. size = 6.00 MiB -> 0.42 MiB
209
- [ 151/ 255] blk.16.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q1_0 .. size = 6.00 MiB -> 0.42 MiB
210
- [ 152/ 255] blk.16.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q1_0 .. size = 1.50 MiB -> 0.11 MiB
211
- [ 153/ 255] blk.16.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q1_0 .. size = 15.00 MiB -> 1.05 MiB
212
- [ 154/ 255] blk.16.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q1_0 .. size = 15.00 MiB -> 1.05 MiB
213
- [ 155/ 255] blk.16.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
214
- [ 156/ 255] blk.16.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q1_0 .. size = 15.00 MiB -> 1.05 MiB
215
- [ 157/ 255] blk.17.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q1_0 .. size = 1.50 MiB -> 0.11 MiB
216
- [ 158/ 255] blk.17.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
217
- [ 159/ 255] blk.17.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q1_0 .. size = 6.00 MiB -> 0.42 MiB
218
- [ 160/ 255] blk.17.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q1_0 .. size = 6.00 MiB -> 0.42 MiB
219
- [ 161/ 255] blk.17.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q1_0 .. size = 1.50 MiB -> 0.11 MiB
220
- [ 162/ 255] blk.17.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q1_0 .. size = 15.00 MiB -> 1.05 MiB
221
- [ 163/ 255] blk.17.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q1_0 .. size = 15.00 MiB -> 1.05 MiB
222
- [ 164/ 255] blk.17.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
223
- [ 165/ 255] blk.17.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q1_0 .. size = 15.00 MiB -> 1.05 MiB
224
- [ 166/ 255] blk.18.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q1_0 .. size = 1.50 MiB -> 0.11 MiB
225
- [ 167/ 255] blk.18.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
226
- [ 168/ 255] blk.18.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q1_0 .. size = 6.00 MiB -> 0.42 MiB
227
- [ 169/ 255] blk.18.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q1_0 .. size = 6.00 MiB -> 0.42 MiB
228
- [ 170/ 255] blk.18.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q1_0 .. size = 1.50 MiB -> 0.11 MiB
229
- [ 171/ 255] blk.18.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q1_0 .. size = 15.00 MiB -> 1.05 MiB
230
- [ 172/ 255] blk.18.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q1_0 .. size = 15.00 MiB -> 1.05 MiB
231
- [ 173/ 255] blk.18.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
232
- [ 174/ 255] blk.18.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q1_0 .. size = 15.00 MiB -> 1.05 MiB
233
- [ 175/ 255] blk.19.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q1_0 .. size = 1.50 MiB -> 0.11 MiB
234
- [ 176/ 255] blk.19.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
235
- [ 177/ 255] blk.19.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q1_0 .. size = 6.00 MiB -> 0.42 MiB
236
- [ 178/ 255] blk.19.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q1_0 .. size = 6.00 MiB -> 0.42 MiB
237
- [ 179/ 255] blk.19.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q1_0 .. size = 1.50 MiB -> 0.11 MiB
238
- [ 180/ 255] blk.19.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q1_0 .. size = 15.00 MiB -> 1.05 MiB
239
- [ 181/ 255] blk.19.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q1_0 .. size = 15.00 MiB -> 1.05 MiB
240
- [ 182/ 255] blk.19.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
241
- [ 183/ 255] blk.19.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q1_0 .. size = 15.00 MiB -> 1.05 MiB
242
- [ 184/ 255] blk.20.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q1_0 .. size = 1.50 MiB -> 0.11 MiB
243
- [ 185/ 255] blk.20.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
244
- [ 186/ 255] blk.20.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q1_0 .. size = 6.00 MiB -> 0.42 MiB
245
- [ 187/ 255] blk.20.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q1_0 .. size = 6.00 MiB -> 0.42 MiB
246
- [ 188/ 255] blk.20.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q1_0 .. size = 1.50 MiB -> 0.11 MiB
247
- [ 189/ 255] blk.20.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q1_0 .. size = 15.00 MiB -> 1.05 MiB
248
- [ 190/ 255] blk.20.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q1_0 .. size = 15.00 MiB -> 1.05 MiB
249
- [ 191/ 255] blk.20.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
250
- [ 192/ 255] blk.20.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q1_0 .. size = 15.00 MiB -> 1.05 MiB
251
- [ 193/ 255] blk.21.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q1_0 .. size = 1.50 MiB -> 0.11 MiB
252
- [ 194/ 255] blk.21.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
253
- [ 195/ 255] blk.21.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q1_0 .. size = 6.00 MiB -> 0.42 MiB
254
- [ 196/ 255] blk.21.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q1_0 .. size = 6.00 MiB -> 0.42 MiB
255
- [ 197/ 255] blk.21.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q1_0 .. size = 1.50 MiB -> 0.11 MiB
256
- [ 198/ 255] blk.21.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q1_0 .. size = 15.00 MiB -> 1.05 MiB
257
- [ 199/ 255] blk.21.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q1_0 .. size = 15.00 MiB -> 1.05 MiB
258
- [ 200/ 255] blk.21.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
259
- [ 201/ 255] blk.21.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q1_0 .. size = 15.00 MiB -> 1.05 MiB
260
- [ 202/ 255] blk.22.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q1_0 .. size = 1.50 MiB -> 0.11 MiB
261
- [ 203/ 255] blk.22.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
262
- [ 204/ 255] blk.22.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q1_0 .. size = 6.00 MiB -> 0.42 MiB
263
- [ 205/ 255] blk.22.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q1_0 .. size = 6.00 MiB -> 0.42 MiB
264
- [ 206/ 255] blk.22.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q1_0 .. size = 1.50 MiB -> 0.11 MiB
265
- [ 207/ 255] blk.22.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q1_0 .. size = 15.00 MiB -> 1.05 MiB
266
- [ 208/ 255] blk.22.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q1_0 .. size = 15.00 MiB -> 1.05 MiB
267
- [ 209/ 255] blk.22.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
268
- [ 210/ 255] blk.22.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q1_0 .. size = 15.00 MiB -> 1.05 MiB
269
- [ 211/ 255] blk.23.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q1_0 .. size = 1.50 MiB -> 0.11 MiB
270
- [ 212/ 255] blk.23.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
271
- [ 213/ 255] blk.23.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q1_0 .. size = 6.00 MiB -> 0.42 MiB
272
- [ 214/ 255] blk.23.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q1_0 .. size = 6.00 MiB -> 0.42 MiB
273
- [ 215/ 255] blk.23.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q1_0 .. size = 1.50 MiB -> 0.11 MiB
274
- [ 216/ 255] blk.23.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q1_0 .. size = 15.00 MiB -> 1.05 MiB
275
- [ 217/ 255] blk.23.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q1_0 .. size = 15.00 MiB -> 1.05 MiB
276
- [ 218/ 255] blk.23.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
277
- [ 219/ 255] blk.23.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q1_0 .. size = 15.00 MiB -> 1.05 MiB
278
- [ 220/ 255] blk.24.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q1_0 .. size = 1.50 MiB -> 0.11 MiB
279
- [ 221/ 255] blk.24.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
280
- [ 222/ 255] blk.24.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q1_0 .. size = 6.00 MiB -> 0.42 MiB
281
- [ 223/ 255] blk.24.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q1_0 .. size = 6.00 MiB -> 0.42 MiB
282
- [ 224/ 255] blk.24.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q1_0 .. size = 1.50 MiB -> 0.11 MiB
283
- [ 225/ 255] blk.24.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q1_0 .. size = 15.00 MiB -> 1.05 MiB
284
- [ 226/ 255] blk.24.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q1_0 .. size = 15.00 MiB -> 1.05 MiB
285
- [ 227/ 255] blk.24.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
286
- [ 228/ 255] blk.24.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q1_0 .. size = 15.00 MiB -> 1.05 MiB
287
- [ 229/ 255] blk.25.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q1_0 .. size = 1.50 MiB -> 0.11 MiB
288
- [ 230/ 255] blk.25.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
289
- [ 231/ 255] blk.25.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q1_0 .. size = 6.00 MiB -> 0.42 MiB
290
- [ 232/ 255] blk.25.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q1_0 .. size = 6.00 MiB -> 0.42 MiB
291
- [ 233/ 255] blk.25.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q1_0 .. size = 1.50 MiB -> 0.11 MiB
292
- [ 234/ 255] blk.25.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q1_0 .. size = 15.00 MiB -> 1.05 MiB
293
- [ 235/ 255] blk.25.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q1_0 .. size = 15.00 MiB -> 1.05 MiB
294
- [ 236/ 255] blk.25.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
295
- [ 237/ 255] blk.25.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q1_0 .. size = 15.00 MiB -> 1.05 MiB
296
- [ 238/ 255] blk.26.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q1_0 .. size = 1.50 MiB -> 0.11 MiB
297
- [ 239/ 255] blk.26.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
298
- [ 240/ 255] blk.26.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q1_0 .. size = 6.00 MiB -> 0.42 MiB
299
- [ 241/ 255] blk.26.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q1_0 .. size = 6.00 MiB -> 0.42 MiB
300
- [ 242/ 255] blk.26.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q1_0 .. size = 1.50 MiB -> 0.11 MiB
301
- [ 243/ 255] blk.26.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q1_0 .. size = 15.00 MiB -> 1.05 MiB
302
- [ 244/ 255] blk.26.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q1_0 .. size = 15.00 MiB -> 1.05 MiB
303
- [ 245/ 255] blk.26.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
304
- [ 246/ 255] blk.26.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q1_0 .. size = 15.00 MiB -> 1.05 MiB
305
- [ 247/ 255] blk.27.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q1_0 .. size = 1.50 MiB -> 0.11 MiB
306
- [ 248/ 255] blk.27.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
307
- [ 249/ 255] blk.27.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q1_0 .. size = 6.00 MiB -> 0.42 MiB
308
- [ 250/ 255] blk.27.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q1_0 .. size = 6.00 MiB -> 0.42 MiB
309
- [ 251/ 255] blk.27.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q1_0 .. size = 1.50 MiB -> 0.11 MiB
310
- [ 252/ 255] blk.27.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q1_0 .. size = 15.00 MiB -> 1.05 MiB
311
- [ 253/ 255] blk.27.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q1_0 .. size = 15.00 MiB -> 1.05 MiB
312
- [ 254/ 255] blk.27.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
313
- [ 255/ 255] blk.27.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q1_0 .. size = 15.00 MiB -> 1.05 MiB
314
- llama_model_quantize_impl: model size = 2056.83 MiB (16.00 BPW)
315
- llama_model_quantize_impl: quant size = 208.91 MiB (1.63 BPW)
316
-
317
- llama_quantize: quantize time = 9949.08 ms
318
- llama_quantize: total time = 9949.08 ms
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
reproducibility/validation/quantize-logs/Q2_K.log DELETED
@@ -1,318 +0,0 @@
1
- ggml_cuda_init: failed to initialize CUDA: no CUDA-capable device is detected
2
- version: 0.3.0-dev (build 10671, commit 35999d101)
3
- built with GNU 13.3.0 for Linux x86_64
4
- llama_quantize: quantizing '/root/workspace/HF/source-official-bf16-k2-horizon-0.9b/K2-Horizon-0.9B-BF16.gguf' to '/root/workspace/HF/quantized-k2-horizon-0.9b/K2-Horizon-0.9B-Q2_K.gguf' as Q2_K using 128 threads
5
- llama_model_loader: loaded meta data with 41 key-value pairs and 255 tensors from /root/workspace/HF/source-official-bf16-k2-horizon-0.9b/K2-Horizon-0.9B-BF16.gguf (version GGUF V3 (latest))
6
- llama_model_loader: Dumping metadata keys/values. Note: KV overrides do not apply in this output.
7
- llama_model_loader: - kv 0: general.architecture str = k2-horizon
8
- llama_model_loader: - kv 1: general.type str = model
9
- llama_model_loader: - kv 2: general.name str = K2-Horizon-0.9B
10
- llama_model_loader: - kv 3: general.basename str = source-hf-k2-horizon
11
- llama_model_loader: - kv 4: general.size_label str = 0.9B
12
- llama_model_loader: - kv 5: general.license str = apache-2.0
13
- llama_model_loader: - kv 6: general.license.name str = internal-only
14
- llama_model_loader: - kv 7: general.license.link str = LICENSE
15
- llama_model_loader: - kv 8: general.tags arr[str,7] = ["k2-horizon", "0.9b", "dense", "reas...
16
- llama_model_loader: - kv 9: general.languages arr[str,2] = ["en", "zh"]
17
- llama_model_loader: - kv 10: k2-horizon.block_count u32 = 28
18
- llama_model_loader: - kv 11: k2-horizon.context_length u32 = 131072
19
- llama_model_loader: - kv 12: k2-horizon.embedding_length u32 = 1536
20
- llama_model_loader: - kv 13: k2-horizon.feed_forward_length u32 = 5120
21
- llama_model_loader: - kv 14: k2-horizon.attention.head_count u32 = 32
22
- llama_model_loader: - kv 15: k2-horizon.attention.head_count_kv u32 = 8
23
- llama_model_loader: - kv 16: k2-horizon.rope.scaling.type str = yarn
24
- llama_model_loader: - kv 17: k2-horizon.rope.scaling.factor f32 = 16.000000
25
- llama_model_loader: - kv 18: k2-horizon.rope.scaling.original_context_length u32 = 8192
26
- llama_model_loader: - kv 19: k2-horizon.rope.scaling.yarn_attn_factor f32 = 1.277259
27
- llama_model_loader: - kv 20: k2-horizon.rope.scaling.yarn_beta_fast f32 = 128.000000
28
- llama_model_loader: - kv 21: k2-horizon.rope.scaling.yarn_beta_slow f32 = 4.000000
29
- llama_model_loader: - kv 22: k2-horizon.rope.freq_base f32 = 1000000.000000
30
- llama_model_loader: - kv 23: k2-horizon.attention.layer_norm_rms_epsilon f32 = 0.000001
31
- llama_model_loader: - kv 24: k2-horizon.expert_count u32 = 0
32
- llama_model_loader: - kv 25: k2-horizon.expert_used_count u32 = 0
33
- llama_model_loader: - kv 26: k2-horizon.attention.key_length u32 = 64
34
- llama_model_loader: - kv 27: k2-horizon.attention.value_length u32 = 64
35
- llama_model_loader: - kv 28: general.file_type u32 = 32
36
- llama_model_loader: - kv 29: k2-horizon.attention.group_norm_groups u32 = 1
37
- llama_model_loader: - kv 30: k2-horizon.rope.dimension_count u32 = 64
38
- llama_model_loader: - kv 31: general.quantization_version u32 = 2
39
- llama_model_loader: - kv 32: tokenizer.ggml.model str = gpt2
40
- llama_model_loader: - kv 33: tokenizer.ggml.pre str = k2-horizon
41
- llama_model_loader: - kv 34: tokenizer.ggml.tokens arr[str,64256] = ["<|begin_of_text|>", "<|endoftext|>"...
42
- llama_model_loader: - kv 35: tokenizer.ggml.token_type arr[i32,64256] = [3, 3, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, ...
43
- llama_model_loader: - kv 36: tokenizer.ggml.merges arr[str,63742] = ["Δ  Δ ", "Ø Β§", "Γ™ Δ¦", "Γ  Β€", ...
44
- llama_model_loader: - kv 37: tokenizer.ggml.bos_token_id u32 = 0
45
- llama_model_loader: - kv 38: tokenizer.ggml.eos_token_id u32 = 1
46
- llama_model_loader: - kv 39: tokenizer.ggml.padding_token_id u32 = 64255
47
- llama_model_loader: - kv 40: tokenizer.chat_template str = {{- bos_token }}\n{%- if tool_presenta...
48
- llama_model_loader: - type f32: 57 tensors
49
- llama_model_loader: - type bf16: 198 tensors
50
-
51
- llama_model_quantize_impl: have importance matrix data with 196 entries
52
- [ 1/ 255] output.weight - [ 1536, 64256, 1, 1], type = bf16,
53
- ====== llama_model_quantize_impl: did not find weights for output.weight
54
- converting to q6_K .. load_imatrix: imatrix datasets=['calibration/wikitext-2-raw/wiki.train.raw', 'calibration/k2-horizon-0.9b/k2_horizon_en_zh_code_tool.txt']
55
- load_imatrix: loaded 196 importance matrix entries from /root/workspace/HF/calibration/k2-horizon-0.9b/k2_horizon_combined.imatrix.gguf computed on 99 chunks
56
- prepare_imatrix: have 196 importance matrix entries
57
- size = 188.25 MiB -> 77.21 MiB
58
- [ 2/ 255] output_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
59
- [ 3/ 255] token_embd.weight - [ 1536, 64256, 1, 1], type = bf16,
60
- ====== llama_model_quantize_impl: did not find weights for token_embd.weight
61
- converting to q2_K .. size = 188.25 MiB -> 30.88 MiB
62
- [ 4/ 255] blk.0.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q2_K .. size = 1.50 MiB -> 0.25 MiB
63
- [ 5/ 255] blk.0.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
64
- [ 6/ 255] blk.0.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q3_K .. size = 6.00 MiB -> 1.29 MiB
65
- [ 7/ 255] blk.0.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q2_K .. size = 6.00 MiB -> 0.98 MiB
66
- [ 8/ 255] blk.0.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
67
- [ 9/ 255] blk.0.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q3_K .. size = 15.00 MiB -> 3.22 MiB
68
- [ 10/ 255] blk.0.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q2_K .. size = 15.00 MiB -> 2.46 MiB
69
- [ 11/ 255] blk.0.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
70
- [ 12/ 255] blk.0.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q2_K .. size = 15.00 MiB -> 2.46 MiB
71
- [ 13/ 255] blk.1.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q2_K .. size = 1.50 MiB -> 0.25 MiB
72
- [ 14/ 255] blk.1.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
73
- [ 15/ 255] blk.1.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q3_K .. size = 6.00 MiB -> 1.29 MiB
74
- [ 16/ 255] blk.1.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q2_K .. size = 6.00 MiB -> 0.98 MiB
75
- [ 17/ 255] blk.1.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
76
- [ 18/ 255] blk.1.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q3_K .. size = 15.00 MiB -> 3.22 MiB
77
- [ 19/ 255] blk.1.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q2_K .. size = 15.00 MiB -> 2.46 MiB
78
- [ 20/ 255] blk.1.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
79
- [ 21/ 255] blk.1.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q2_K .. size = 15.00 MiB -> 2.46 MiB
80
- [ 22/ 255] blk.2.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q2_K .. size = 1.50 MiB -> 0.25 MiB
81
- [ 23/ 255] blk.2.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
82
- [ 24/ 255] blk.2.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q3_K .. size = 6.00 MiB -> 1.29 MiB
83
- [ 25/ 255] blk.2.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q2_K .. size = 6.00 MiB -> 0.98 MiB
84
- [ 26/ 255] blk.2.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
85
- [ 27/ 255] blk.2.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q3_K .. size = 15.00 MiB -> 3.22 MiB
86
- [ 28/ 255] blk.2.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q2_K .. size = 15.00 MiB -> 2.46 MiB
87
- [ 29/ 255] blk.2.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
88
- [ 30/ 255] blk.2.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q2_K .. size = 15.00 MiB -> 2.46 MiB
89
- [ 31/ 255] blk.3.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q2_K .. size = 1.50 MiB -> 0.25 MiB
90
- [ 32/ 255] blk.3.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
91
- [ 33/ 255] blk.3.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q3_K .. size = 6.00 MiB -> 1.29 MiB
92
- [ 34/ 255] blk.3.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q2_K .. size = 6.00 MiB -> 0.98 MiB
93
- [ 35/ 255] blk.3.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
94
- [ 36/ 255] blk.3.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q3_K .. size = 15.00 MiB -> 3.22 MiB
95
- [ 37/ 255] blk.3.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q2_K .. size = 15.00 MiB -> 2.46 MiB
96
- [ 38/ 255] blk.3.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
97
- [ 39/ 255] blk.3.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q2_K .. size = 15.00 MiB -> 2.46 MiB
98
- [ 40/ 255] blk.4.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q2_K .. size = 1.50 MiB -> 0.25 MiB
99
- [ 41/ 255] blk.4.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
100
- [ 42/ 255] blk.4.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q3_K .. size = 6.00 MiB -> 1.29 MiB
101
- [ 43/ 255] blk.4.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q2_K .. size = 6.00 MiB -> 0.98 MiB
102
- [ 44/ 255] blk.4.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
103
- [ 45/ 255] blk.4.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q3_K .. size = 15.00 MiB -> 3.22 MiB
104
- [ 46/ 255] blk.4.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q2_K .. size = 15.00 MiB -> 2.46 MiB
105
- [ 47/ 255] blk.4.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
106
- [ 48/ 255] blk.4.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q2_K .. size = 15.00 MiB -> 2.46 MiB
107
- [ 49/ 255] blk.5.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q2_K .. size = 1.50 MiB -> 0.25 MiB
108
- [ 50/ 255] blk.5.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
109
- [ 51/ 255] blk.5.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q3_K .. size = 6.00 MiB -> 1.29 MiB
110
- [ 52/ 255] blk.5.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q2_K .. size = 6.00 MiB -> 0.98 MiB
111
- [ 53/ 255] blk.5.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
112
- [ 54/ 255] blk.5.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q3_K .. size = 15.00 MiB -> 3.22 MiB
113
- [ 55/ 255] blk.5.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q2_K .. size = 15.00 MiB -> 2.46 MiB
114
- [ 56/ 255] blk.5.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
115
- [ 57/ 255] blk.5.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q2_K .. size = 15.00 MiB -> 2.46 MiB
116
- [ 58/ 255] blk.6.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q2_K .. size = 1.50 MiB -> 0.25 MiB
117
- [ 59/ 255] blk.6.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
118
- [ 60/ 255] blk.6.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q3_K .. size = 6.00 MiB -> 1.29 MiB
119
- [ 61/ 255] blk.6.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q2_K .. size = 6.00 MiB -> 0.98 MiB
120
- [ 62/ 255] blk.6.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
121
- [ 63/ 255] blk.6.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q3_K .. size = 15.00 MiB -> 3.22 MiB
122
- [ 64/ 255] blk.6.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q2_K .. size = 15.00 MiB -> 2.46 MiB
123
- [ 65/ 255] blk.6.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
124
- [ 66/ 255] blk.6.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q2_K .. size = 15.00 MiB -> 2.46 MiB
125
- [ 67/ 255] blk.7.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q2_K .. size = 1.50 MiB -> 0.25 MiB
126
- [ 68/ 255] blk.7.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
127
- [ 69/ 255] blk.7.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q3_K .. size = 6.00 MiB -> 1.29 MiB
128
- [ 70/ 255] blk.7.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q2_K .. size = 6.00 MiB -> 0.98 MiB
129
- [ 71/ 255] blk.7.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
130
- [ 72/ 255] blk.7.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q3_K .. size = 15.00 MiB -> 3.22 MiB
131
- [ 73/ 255] blk.7.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q2_K .. size = 15.00 MiB -> 2.46 MiB
132
- [ 74/ 255] blk.7.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
133
- [ 75/ 255] blk.7.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q2_K .. size = 15.00 MiB -> 2.46 MiB
134
- [ 76/ 255] blk.8.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q2_K .. size = 1.50 MiB -> 0.25 MiB
135
- [ 77/ 255] blk.8.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
136
- [ 78/ 255] blk.8.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q3_K .. size = 6.00 MiB -> 1.29 MiB
137
- [ 79/ 255] blk.8.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q2_K .. size = 6.00 MiB -> 0.98 MiB
138
- [ 80/ 255] blk.8.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
139
- [ 81/ 255] blk.8.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q3_K .. size = 15.00 MiB -> 3.22 MiB
140
- [ 82/ 255] blk.8.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q2_K .. size = 15.00 MiB -> 2.46 MiB
141
- [ 83/ 255] blk.8.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
142
- [ 84/ 255] blk.8.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q2_K .. size = 15.00 MiB -> 2.46 MiB
143
- [ 85/ 255] blk.9.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q2_K .. size = 1.50 MiB -> 0.25 MiB
144
- [ 86/ 255] blk.9.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
145
- [ 87/ 255] blk.9.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q3_K .. size = 6.00 MiB -> 1.29 MiB
146
- [ 88/ 255] blk.9.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q2_K .. size = 6.00 MiB -> 0.98 MiB
147
- [ 89/ 255] blk.9.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
148
- [ 90/ 255] blk.9.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q3_K .. size = 15.00 MiB -> 3.22 MiB
149
- [ 91/ 255] blk.9.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q2_K .. size = 15.00 MiB -> 2.46 MiB
150
- [ 92/ 255] blk.9.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
151
- [ 93/ 255] blk.9.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q2_K .. size = 15.00 MiB -> 2.46 MiB
152
- [ 94/ 255] blk.10.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q2_K .. size = 1.50 MiB -> 0.25 MiB
153
- [ 95/ 255] blk.10.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
154
- [ 96/ 255] blk.10.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q3_K .. size = 6.00 MiB -> 1.29 MiB
155
- [ 97/ 255] blk.10.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q2_K .. size = 6.00 MiB -> 0.98 MiB
156
- [ 98/ 255] blk.10.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
157
- [ 99/ 255] blk.10.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q3_K .. size = 15.00 MiB -> 3.22 MiB
158
- [ 100/ 255] blk.10.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q2_K .. size = 15.00 MiB -> 2.46 MiB
159
- [ 101/ 255] blk.10.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
160
- [ 102/ 255] blk.10.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q2_K .. size = 15.00 MiB -> 2.46 MiB
161
- [ 103/ 255] blk.11.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q2_K .. size = 1.50 MiB -> 0.25 MiB
162
- [ 104/ 255] blk.11.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
163
- [ 105/ 255] blk.11.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q3_K .. size = 6.00 MiB -> 1.29 MiB
164
- [ 106/ 255] blk.11.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q2_K .. size = 6.00 MiB -> 0.98 MiB
165
- [ 107/ 255] blk.11.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
166
- [ 108/ 255] blk.11.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q3_K .. size = 15.00 MiB -> 3.22 MiB
167
- [ 109/ 255] blk.11.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q2_K .. size = 15.00 MiB -> 2.46 MiB
168
- [ 110/ 255] blk.11.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
169
- [ 111/ 255] blk.11.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q2_K .. size = 15.00 MiB -> 2.46 MiB
170
- [ 112/ 255] blk.12.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q2_K .. size = 1.50 MiB -> 0.25 MiB
171
- [ 113/ 255] blk.12.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
172
- [ 114/ 255] blk.12.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q3_K .. size = 6.00 MiB -> 1.29 MiB
173
- [ 115/ 255] blk.12.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q2_K .. size = 6.00 MiB -> 0.98 MiB
174
- [ 116/ 255] blk.12.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
175
- [ 117/ 255] blk.12.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q3_K .. size = 15.00 MiB -> 3.22 MiB
176
- [ 118/ 255] blk.12.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q2_K .. size = 15.00 MiB -> 2.46 MiB
177
- [ 119/ 255] blk.12.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
178
- [ 120/ 255] blk.12.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q2_K .. size = 15.00 MiB -> 2.46 MiB
179
- [ 121/ 255] blk.13.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q2_K .. size = 1.50 MiB -> 0.25 MiB
180
- [ 122/ 255] blk.13.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
181
- [ 123/ 255] blk.13.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q3_K .. size = 6.00 MiB -> 1.29 MiB
182
- [ 124/ 255] blk.13.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q2_K .. size = 6.00 MiB -> 0.98 MiB
183
- [ 125/ 255] blk.13.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
184
- [ 126/ 255] blk.13.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q3_K .. size = 15.00 MiB -> 3.22 MiB
185
- [ 127/ 255] blk.13.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q2_K .. size = 15.00 MiB -> 2.46 MiB
186
- [ 128/ 255] blk.13.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
187
- [ 129/ 255] blk.13.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q2_K .. size = 15.00 MiB -> 2.46 MiB
188
- [ 130/ 255] blk.14.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q2_K .. size = 1.50 MiB -> 0.25 MiB
189
- [ 131/ 255] blk.14.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
190
- [ 132/ 255] blk.14.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q3_K .. size = 6.00 MiB -> 1.29 MiB
191
- [ 133/ 255] blk.14.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q2_K .. size = 6.00 MiB -> 0.98 MiB
192
- [ 134/ 255] blk.14.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
193
- [ 135/ 255] blk.14.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q3_K .. size = 15.00 MiB -> 3.22 MiB
194
- [ 136/ 255] blk.14.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q2_K .. size = 15.00 MiB -> 2.46 MiB
195
- [ 137/ 255] blk.14.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
196
- [ 138/ 255] blk.14.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q2_K .. size = 15.00 MiB -> 2.46 MiB
197
- [ 139/ 255] blk.15.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q2_K .. size = 1.50 MiB -> 0.25 MiB
198
- [ 140/ 255] blk.15.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
199
- [ 141/ 255] blk.15.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q3_K .. size = 6.00 MiB -> 1.29 MiB
200
- [ 142/ 255] blk.15.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q2_K .. size = 6.00 MiB -> 0.98 MiB
201
- [ 143/ 255] blk.15.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
202
- [ 144/ 255] blk.15.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q3_K .. size = 15.00 MiB -> 3.22 MiB
203
- [ 145/ 255] blk.15.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q2_K .. size = 15.00 MiB -> 2.46 MiB
204
- [ 146/ 255] blk.15.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
205
- [ 147/ 255] blk.15.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q2_K .. size = 15.00 MiB -> 2.46 MiB
206
- [ 148/ 255] blk.16.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q2_K .. size = 1.50 MiB -> 0.25 MiB
207
- [ 149/ 255] blk.16.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
208
- [ 150/ 255] blk.16.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q3_K .. size = 6.00 MiB -> 1.29 MiB
209
- [ 151/ 255] blk.16.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q2_K .. size = 6.00 MiB -> 0.98 MiB
210
- [ 152/ 255] blk.16.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
211
- [ 153/ 255] blk.16.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q3_K .. size = 15.00 MiB -> 3.22 MiB
212
- [ 154/ 255] blk.16.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q2_K .. size = 15.00 MiB -> 2.46 MiB
213
- [ 155/ 255] blk.16.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
214
- [ 156/ 255] blk.16.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q2_K .. size = 15.00 MiB -> 2.46 MiB
215
- [ 157/ 255] blk.17.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q2_K .. size = 1.50 MiB -> 0.25 MiB
216
- [ 158/ 255] blk.17.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
217
- [ 159/ 255] blk.17.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q3_K .. size = 6.00 MiB -> 1.29 MiB
218
- [ 160/ 255] blk.17.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q2_K .. size = 6.00 MiB -> 0.98 MiB
219
- [ 161/ 255] blk.17.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
220
- [ 162/ 255] blk.17.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q3_K .. size = 15.00 MiB -> 3.22 MiB
221
- [ 163/ 255] blk.17.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q2_K .. size = 15.00 MiB -> 2.46 MiB
222
- [ 164/ 255] blk.17.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
223
- [ 165/ 255] blk.17.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q2_K .. size = 15.00 MiB -> 2.46 MiB
224
- [ 166/ 255] blk.18.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q2_K .. size = 1.50 MiB -> 0.25 MiB
225
- [ 167/ 255] blk.18.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
226
- [ 168/ 255] blk.18.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q3_K .. size = 6.00 MiB -> 1.29 MiB
227
- [ 169/ 255] blk.18.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q2_K .. size = 6.00 MiB -> 0.98 MiB
228
- [ 170/ 255] blk.18.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
229
- [ 171/ 255] blk.18.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q3_K .. size = 15.00 MiB -> 3.22 MiB
230
- [ 172/ 255] blk.18.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q2_K .. size = 15.00 MiB -> 2.46 MiB
231
- [ 173/ 255] blk.18.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
232
- [ 174/ 255] blk.18.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q2_K .. size = 15.00 MiB -> 2.46 MiB
233
- [ 175/ 255] blk.19.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q2_K .. size = 1.50 MiB -> 0.25 MiB
234
- [ 176/ 255] blk.19.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
235
- [ 177/ 255] blk.19.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q3_K .. size = 6.00 MiB -> 1.29 MiB
236
- [ 178/ 255] blk.19.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q2_K .. size = 6.00 MiB -> 0.98 MiB
237
- [ 179/ 255] blk.19.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
238
- [ 180/ 255] blk.19.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q3_K .. size = 15.00 MiB -> 3.22 MiB
239
- [ 181/ 255] blk.19.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q2_K .. size = 15.00 MiB -> 2.46 MiB
240
- [ 182/ 255] blk.19.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
241
- [ 183/ 255] blk.19.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q2_K .. size = 15.00 MiB -> 2.46 MiB
242
- [ 184/ 255] blk.20.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q2_K .. size = 1.50 MiB -> 0.25 MiB
243
- [ 185/ 255] blk.20.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
244
- [ 186/ 255] blk.20.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q3_K .. size = 6.00 MiB -> 1.29 MiB
245
- [ 187/ 255] blk.20.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q2_K .. size = 6.00 MiB -> 0.98 MiB
246
- [ 188/ 255] blk.20.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
247
- [ 189/ 255] blk.20.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q3_K .. size = 15.00 MiB -> 3.22 MiB
248
- [ 190/ 255] blk.20.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q2_K .. size = 15.00 MiB -> 2.46 MiB
249
- [ 191/ 255] blk.20.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
250
- [ 192/ 255] blk.20.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q2_K .. size = 15.00 MiB -> 2.46 MiB
251
- [ 193/ 255] blk.21.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q2_K .. size = 1.50 MiB -> 0.25 MiB
252
- [ 194/ 255] blk.21.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
253
- [ 195/ 255] blk.21.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q3_K .. size = 6.00 MiB -> 1.29 MiB
254
- [ 196/ 255] blk.21.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q2_K .. size = 6.00 MiB -> 0.98 MiB
255
- [ 197/ 255] blk.21.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
256
- [ 198/ 255] blk.21.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q3_K .. size = 15.00 MiB -> 3.22 MiB
257
- [ 199/ 255] blk.21.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q2_K .. size = 15.00 MiB -> 2.46 MiB
258
- [ 200/ 255] blk.21.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
259
- [ 201/ 255] blk.21.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q2_K .. size = 15.00 MiB -> 2.46 MiB
260
- [ 202/ 255] blk.22.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q2_K .. size = 1.50 MiB -> 0.25 MiB
261
- [ 203/ 255] blk.22.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
262
- [ 204/ 255] blk.22.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q3_K .. size = 6.00 MiB -> 1.29 MiB
263
- [ 205/ 255] blk.22.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q2_K .. size = 6.00 MiB -> 0.98 MiB
264
- [ 206/ 255] blk.22.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
265
- [ 207/ 255] blk.22.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q3_K .. size = 15.00 MiB -> 3.22 MiB
266
- [ 208/ 255] blk.22.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q2_K .. size = 15.00 MiB -> 2.46 MiB
267
- [ 209/ 255] blk.22.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
268
- [ 210/ 255] blk.22.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q2_K .. size = 15.00 MiB -> 2.46 MiB
269
- [ 211/ 255] blk.23.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q2_K .. size = 1.50 MiB -> 0.25 MiB
270
- [ 212/ 255] blk.23.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
271
- [ 213/ 255] blk.23.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q3_K .. size = 6.00 MiB -> 1.29 MiB
272
- [ 214/ 255] blk.23.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q2_K .. size = 6.00 MiB -> 0.98 MiB
273
- [ 215/ 255] blk.23.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
274
- [ 216/ 255] blk.23.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q3_K .. size = 15.00 MiB -> 3.22 MiB
275
- [ 217/ 255] blk.23.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q2_K .. size = 15.00 MiB -> 2.46 MiB
276
- [ 218/ 255] blk.23.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
277
- [ 219/ 255] blk.23.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q2_K .. size = 15.00 MiB -> 2.46 MiB
278
- [ 220/ 255] blk.24.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q2_K .. size = 1.50 MiB -> 0.25 MiB
279
- [ 221/ 255] blk.24.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
280
- [ 222/ 255] blk.24.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q3_K .. size = 6.00 MiB -> 1.29 MiB
281
- [ 223/ 255] blk.24.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q2_K .. size = 6.00 MiB -> 0.98 MiB
282
- [ 224/ 255] blk.24.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
283
- [ 225/ 255] blk.24.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q3_K .. size = 15.00 MiB -> 3.22 MiB
284
- [ 226/ 255] blk.24.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q2_K .. size = 15.00 MiB -> 2.46 MiB
285
- [ 227/ 255] blk.24.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
286
- [ 228/ 255] blk.24.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q2_K .. size = 15.00 MiB -> 2.46 MiB
287
- [ 229/ 255] blk.25.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q2_K .. size = 1.50 MiB -> 0.25 MiB
288
- [ 230/ 255] blk.25.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
289
- [ 231/ 255] blk.25.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q3_K .. size = 6.00 MiB -> 1.29 MiB
290
- [ 232/ 255] blk.25.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q2_K .. size = 6.00 MiB -> 0.98 MiB
291
- [ 233/ 255] blk.25.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
292
- [ 234/ 255] blk.25.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q3_K .. size = 15.00 MiB -> 3.22 MiB
293
- [ 235/ 255] blk.25.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q2_K .. size = 15.00 MiB -> 2.46 MiB
294
- [ 236/ 255] blk.25.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
295
- [ 237/ 255] blk.25.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q2_K .. size = 15.00 MiB -> 2.46 MiB
296
- [ 238/ 255] blk.26.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q2_K .. size = 1.50 MiB -> 0.25 MiB
297
- [ 239/ 255] blk.26.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
298
- [ 240/ 255] blk.26.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q3_K .. size = 6.00 MiB -> 1.29 MiB
299
- [ 241/ 255] blk.26.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q2_K .. size = 6.00 MiB -> 0.98 MiB
300
- [ 242/ 255] blk.26.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
301
- [ 243/ 255] blk.26.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q3_K .. size = 15.00 MiB -> 3.22 MiB
302
- [ 244/ 255] blk.26.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q2_K .. size = 15.00 MiB -> 2.46 MiB
303
- [ 245/ 255] blk.26.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
304
- [ 246/ 255] blk.26.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q2_K .. size = 15.00 MiB -> 2.46 MiB
305
- [ 247/ 255] blk.27.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q2_K .. size = 1.50 MiB -> 0.25 MiB
306
- [ 248/ 255] blk.27.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
307
- [ 249/ 255] blk.27.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q3_K .. size = 6.00 MiB -> 1.29 MiB
308
- [ 250/ 255] blk.27.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q2_K .. size = 6.00 MiB -> 0.98 MiB
309
- [ 251/ 255] blk.27.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
310
- [ 252/ 255] blk.27.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q3_K .. size = 15.00 MiB -> 3.22 MiB
311
- [ 253/ 255] blk.27.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q2_K .. size = 15.00 MiB -> 2.46 MiB
312
- [ 254/ 255] blk.27.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
313
- [ 255/ 255] blk.27.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q2_K .. size = 15.00 MiB -> 2.46 MiB
314
- llama_model_quantize_impl: model size = 2056.83 MiB (16.00 BPW)
315
- llama_model_quantize_impl: quant size = 418.84 MiB (3.26 BPW)
316
-
317
- llama_quantize: quantize time = 12375.14 ms
318
- llama_quantize: total time = 12375.14 ms
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
reproducibility/validation/quantize-logs/Q3_K_M.log DELETED
@@ -1,318 +0,0 @@
1
- ggml_cuda_init: failed to initialize CUDA: no CUDA-capable device is detected
2
- version: 0.3.0-dev (build 10671, commit 35999d101)
3
- built with GNU 13.3.0 for Linux x86_64
4
- llama_quantize: quantizing '/root/workspace/HF/source-official-bf16-k2-horizon-0.9b/K2-Horizon-0.9B-BF16.gguf' to '/root/workspace/HF/quantized-k2-horizon-0.9b/K2-Horizon-0.9B-Q3_K_M.gguf' as Q3_K_M using 128 threads
5
- llama_model_loader: loaded meta data with 41 key-value pairs and 255 tensors from /root/workspace/HF/source-official-bf16-k2-horizon-0.9b/K2-Horizon-0.9B-BF16.gguf (version GGUF V3 (latest))
6
- llama_model_loader: Dumping metadata keys/values. Note: KV overrides do not apply in this output.
7
- llama_model_loader: - kv 0: general.architecture str = k2-horizon
8
- llama_model_loader: - kv 1: general.type str = model
9
- llama_model_loader: - kv 2: general.name str = K2-Horizon-0.9B
10
- llama_model_loader: - kv 3: general.basename str = source-hf-k2-horizon
11
- llama_model_loader: - kv 4: general.size_label str = 0.9B
12
- llama_model_loader: - kv 5: general.license str = apache-2.0
13
- llama_model_loader: - kv 6: general.license.name str = internal-only
14
- llama_model_loader: - kv 7: general.license.link str = LICENSE
15
- llama_model_loader: - kv 8: general.tags arr[str,7] = ["k2-horizon", "0.9b", "dense", "reas...
16
- llama_model_loader: - kv 9: general.languages arr[str,2] = ["en", "zh"]
17
- llama_model_loader: - kv 10: k2-horizon.block_count u32 = 28
18
- llama_model_loader: - kv 11: k2-horizon.context_length u32 = 131072
19
- llama_model_loader: - kv 12: k2-horizon.embedding_length u32 = 1536
20
- llama_model_loader: - kv 13: k2-horizon.feed_forward_length u32 = 5120
21
- llama_model_loader: - kv 14: k2-horizon.attention.head_count u32 = 32
22
- llama_model_loader: - kv 15: k2-horizon.attention.head_count_kv u32 = 8
23
- llama_model_loader: - kv 16: k2-horizon.rope.scaling.type str = yarn
24
- llama_model_loader: - kv 17: k2-horizon.rope.scaling.factor f32 = 16.000000
25
- llama_model_loader: - kv 18: k2-horizon.rope.scaling.original_context_length u32 = 8192
26
- llama_model_loader: - kv 19: k2-horizon.rope.scaling.yarn_attn_factor f32 = 1.277259
27
- llama_model_loader: - kv 20: k2-horizon.rope.scaling.yarn_beta_fast f32 = 128.000000
28
- llama_model_loader: - kv 21: k2-horizon.rope.scaling.yarn_beta_slow f32 = 4.000000
29
- llama_model_loader: - kv 22: k2-horizon.rope.freq_base f32 = 1000000.000000
30
- llama_model_loader: - kv 23: k2-horizon.attention.layer_norm_rms_epsilon f32 = 0.000001
31
- llama_model_loader: - kv 24: k2-horizon.expert_count u32 = 0
32
- llama_model_loader: - kv 25: k2-horizon.expert_used_count u32 = 0
33
- llama_model_loader: - kv 26: k2-horizon.attention.key_length u32 = 64
34
- llama_model_loader: - kv 27: k2-horizon.attention.value_length u32 = 64
35
- llama_model_loader: - kv 28: general.file_type u32 = 32
36
- llama_model_loader: - kv 29: k2-horizon.attention.group_norm_groups u32 = 1
37
- llama_model_loader: - kv 30: k2-horizon.rope.dimension_count u32 = 64
38
- llama_model_loader: - kv 31: general.quantization_version u32 = 2
39
- llama_model_loader: - kv 32: tokenizer.ggml.model str = gpt2
40
- llama_model_loader: - kv 33: tokenizer.ggml.pre str = k2-horizon
41
- llama_model_loader: - kv 34: tokenizer.ggml.tokens arr[str,64256] = ["<|begin_of_text|>", "<|endoftext|>"...
42
- llama_model_loader: - kv 35: tokenizer.ggml.token_type arr[i32,64256] = [3, 3, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, ...
43
- llama_model_loader: - kv 36: tokenizer.ggml.merges arr[str,63742] = ["Δ  Δ ", "Ø Β§", "Γ™ Δ¦", "Γ  Β€", ...
44
- llama_model_loader: - kv 37: tokenizer.ggml.bos_token_id u32 = 0
45
- llama_model_loader: - kv 38: tokenizer.ggml.eos_token_id u32 = 1
46
- llama_model_loader: - kv 39: tokenizer.ggml.padding_token_id u32 = 64255
47
- llama_model_loader: - kv 40: tokenizer.chat_template str = {{- bos_token }}\n{%- if tool_presenta...
48
- llama_model_loader: - type f32: 57 tensors
49
- llama_model_loader: - type bf16: 198 tensors
50
-
51
- llama_model_quantize_impl: have importance matrix data with 196 entries
52
- [ 1/ 255] output.weight - [ 1536, 64256, 1, 1], type = bf16,
53
- ====== llama_model_quantize_impl: did not find weights for output.weight
54
- converting to q6_K .. load_imatrix: imatrix datasets=['calibration/wikitext-2-raw/wiki.train.raw', 'calibration/k2-horizon-0.9b/k2_horizon_en_zh_code_tool.txt']
55
- load_imatrix: loaded 196 importance matrix entries from /root/workspace/HF/calibration/k2-horizon-0.9b/k2_horizon_combined.imatrix.gguf computed on 99 chunks
56
- prepare_imatrix: have 196 importance matrix entries
57
- size = 188.25 MiB -> 77.21 MiB
58
- [ 2/ 255] output_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
59
- [ 3/ 255] token_embd.weight - [ 1536, 64256, 1, 1], type = bf16,
60
- ====== llama_model_quantize_impl: did not find weights for token_embd.weight
61
- converting to q3_K .. size = 188.25 MiB -> 40.44 MiB
62
- [ 4/ 255] blk.0.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q3_K .. size = 1.50 MiB -> 0.32 MiB
63
- [ 5/ 255] blk.0.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
64
- [ 6/ 255] blk.0.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q4_K .. size = 6.00 MiB -> 1.69 MiB
65
- [ 7/ 255] blk.0.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q3_K .. size = 6.00 MiB -> 1.29 MiB
66
- [ 8/ 255] blk.0.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q5_K .. size = 1.50 MiB -> 0.52 MiB
67
- [ 9/ 255] blk.0.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q5_K .. size = 15.00 MiB -> 5.16 MiB
68
- [ 10/ 255] blk.0.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q3_K .. size = 15.00 MiB -> 3.22 MiB
69
- [ 11/ 255] blk.0.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
70
- [ 12/ 255] blk.0.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q3_K .. size = 15.00 MiB -> 3.22 MiB
71
- [ 13/ 255] blk.1.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q3_K .. size = 1.50 MiB -> 0.32 MiB
72
- [ 14/ 255] blk.1.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
73
- [ 15/ 255] blk.1.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q4_K .. size = 6.00 MiB -> 1.69 MiB
74
- [ 16/ 255] blk.1.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q3_K .. size = 6.00 MiB -> 1.29 MiB
75
- [ 17/ 255] blk.1.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q5_K .. size = 1.50 MiB -> 0.52 MiB
76
- [ 18/ 255] blk.1.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q4_K .. size = 15.00 MiB -> 4.22 MiB
77
- [ 19/ 255] blk.1.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q3_K .. size = 15.00 MiB -> 3.22 MiB
78
- [ 20/ 255] blk.1.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
79
- [ 21/ 255] blk.1.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q3_K .. size = 15.00 MiB -> 3.22 MiB
80
- [ 22/ 255] blk.2.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q3_K .. size = 1.50 MiB -> 0.32 MiB
81
- [ 23/ 255] blk.2.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
82
- [ 24/ 255] blk.2.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q4_K .. size = 6.00 MiB -> 1.69 MiB
83
- [ 25/ 255] blk.2.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q3_K .. size = 6.00 MiB -> 1.29 MiB
84
- [ 26/ 255] blk.2.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
85
- [ 27/ 255] blk.2.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q4_K .. size = 15.00 MiB -> 4.22 MiB
86
- [ 28/ 255] blk.2.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q3_K .. size = 15.00 MiB -> 3.22 MiB
87
- [ 29/ 255] blk.2.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
88
- [ 30/ 255] blk.2.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q3_K .. size = 15.00 MiB -> 3.22 MiB
89
- [ 31/ 255] blk.3.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q3_K .. size = 1.50 MiB -> 0.32 MiB
90
- [ 32/ 255] blk.3.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
91
- [ 33/ 255] blk.3.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q4_K .. size = 6.00 MiB -> 1.69 MiB
92
- [ 34/ 255] blk.3.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q3_K .. size = 6.00 MiB -> 1.29 MiB
93
- [ 35/ 255] blk.3.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
94
- [ 36/ 255] blk.3.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q4_K .. size = 15.00 MiB -> 4.22 MiB
95
- [ 37/ 255] blk.3.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q3_K .. size = 15.00 MiB -> 3.22 MiB
96
- [ 38/ 255] blk.3.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
97
- [ 39/ 255] blk.3.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q3_K .. size = 15.00 MiB -> 3.22 MiB
98
- [ 40/ 255] blk.4.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q3_K .. size = 1.50 MiB -> 0.32 MiB
99
- [ 41/ 255] blk.4.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
100
- [ 42/ 255] blk.4.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q4_K .. size = 6.00 MiB -> 1.69 MiB
101
- [ 43/ 255] blk.4.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q3_K .. size = 6.00 MiB -> 1.29 MiB
102
- [ 44/ 255] blk.4.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
103
- [ 45/ 255] blk.4.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q4_K .. size = 15.00 MiB -> 4.22 MiB
104
- [ 46/ 255] blk.4.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q3_K .. size = 15.00 MiB -> 3.22 MiB
105
- [ 47/ 255] blk.4.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
106
- [ 48/ 255] blk.4.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q3_K .. size = 15.00 MiB -> 3.22 MiB
107
- [ 49/ 255] blk.5.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q3_K .. size = 1.50 MiB -> 0.32 MiB
108
- [ 50/ 255] blk.5.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
109
- [ 51/ 255] blk.5.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q4_K .. size = 6.00 MiB -> 1.69 MiB
110
- [ 52/ 255] blk.5.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q3_K .. size = 6.00 MiB -> 1.29 MiB
111
- [ 53/ 255] blk.5.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
112
- [ 54/ 255] blk.5.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q4_K .. size = 15.00 MiB -> 4.22 MiB
113
- [ 55/ 255] blk.5.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q3_K .. size = 15.00 MiB -> 3.22 MiB
114
- [ 56/ 255] blk.5.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
115
- [ 57/ 255] blk.5.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q3_K .. size = 15.00 MiB -> 3.22 MiB
116
- [ 58/ 255] blk.6.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q3_K .. size = 1.50 MiB -> 0.32 MiB
117
- [ 59/ 255] blk.6.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
118
- [ 60/ 255] blk.6.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q4_K .. size = 6.00 MiB -> 1.69 MiB
119
- [ 61/ 255] blk.6.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q3_K .. size = 6.00 MiB -> 1.29 MiB
120
- [ 62/ 255] blk.6.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
121
- [ 63/ 255] blk.6.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q4_K .. size = 15.00 MiB -> 4.22 MiB
122
- [ 64/ 255] blk.6.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q3_K .. size = 15.00 MiB -> 3.22 MiB
123
- [ 65/ 255] blk.6.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
124
- [ 66/ 255] blk.6.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q3_K .. size = 15.00 MiB -> 3.22 MiB
125
- [ 67/ 255] blk.7.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q3_K .. size = 1.50 MiB -> 0.32 MiB
126
- [ 68/ 255] blk.7.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
127
- [ 69/ 255] blk.7.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q4_K .. size = 6.00 MiB -> 1.69 MiB
128
- [ 70/ 255] blk.7.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q3_K .. size = 6.00 MiB -> 1.29 MiB
129
- [ 71/ 255] blk.7.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
130
- [ 72/ 255] blk.7.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q4_K .. size = 15.00 MiB -> 4.22 MiB
131
- [ 73/ 255] blk.7.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q3_K .. size = 15.00 MiB -> 3.22 MiB
132
- [ 74/ 255] blk.7.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
133
- [ 75/ 255] blk.7.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q3_K .. size = 15.00 MiB -> 3.22 MiB
134
- [ 76/ 255] blk.8.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q3_K .. size = 1.50 MiB -> 0.32 MiB
135
- [ 77/ 255] blk.8.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
136
- [ 78/ 255] blk.8.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q4_K .. size = 6.00 MiB -> 1.69 MiB
137
- [ 79/ 255] blk.8.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q3_K .. size = 6.00 MiB -> 1.29 MiB
138
- [ 80/ 255] blk.8.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
139
- [ 81/ 255] blk.8.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q4_K .. size = 15.00 MiB -> 4.22 MiB
140
- [ 82/ 255] blk.8.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q3_K .. size = 15.00 MiB -> 3.22 MiB
141
- [ 83/ 255] blk.8.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
142
- [ 84/ 255] blk.8.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q3_K .. size = 15.00 MiB -> 3.22 MiB
143
- [ 85/ 255] blk.9.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q3_K .. size = 1.50 MiB -> 0.32 MiB
144
- [ 86/ 255] blk.9.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
145
- [ 87/ 255] blk.9.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q4_K .. size = 6.00 MiB -> 1.69 MiB
146
- [ 88/ 255] blk.9.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q3_K .. size = 6.00 MiB -> 1.29 MiB
147
- [ 89/ 255] blk.9.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
148
- [ 90/ 255] blk.9.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q4_K .. size = 15.00 MiB -> 4.22 MiB
149
- [ 91/ 255] blk.9.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q3_K .. size = 15.00 MiB -> 3.22 MiB
150
- [ 92/ 255] blk.9.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
151
- [ 93/ 255] blk.9.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q3_K .. size = 15.00 MiB -> 3.22 MiB
152
- [ 94/ 255] blk.10.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q3_K .. size = 1.50 MiB -> 0.32 MiB
153
- [ 95/ 255] blk.10.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
154
- [ 96/ 255] blk.10.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q4_K .. size = 6.00 MiB -> 1.69 MiB
155
- [ 97/ 255] blk.10.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q3_K .. size = 6.00 MiB -> 1.29 MiB
156
- [ 98/ 255] blk.10.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
157
- [ 99/ 255] blk.10.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q4_K .. size = 15.00 MiB -> 4.22 MiB
158
- [ 100/ 255] blk.10.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q3_K .. size = 15.00 MiB -> 3.22 MiB
159
- [ 101/ 255] blk.10.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
160
- [ 102/ 255] blk.10.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q3_K .. size = 15.00 MiB -> 3.22 MiB
161
- [ 103/ 255] blk.11.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q3_K .. size = 1.50 MiB -> 0.32 MiB
162
- [ 104/ 255] blk.11.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
163
- [ 105/ 255] blk.11.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q4_K .. size = 6.00 MiB -> 1.69 MiB
164
- [ 106/ 255] blk.11.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q3_K .. size = 6.00 MiB -> 1.29 MiB
165
- [ 107/ 255] blk.11.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
166
- [ 108/ 255] blk.11.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q4_K .. size = 15.00 MiB -> 4.22 MiB
167
- [ 109/ 255] blk.11.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q3_K .. size = 15.00 MiB -> 3.22 MiB
168
- [ 110/ 255] blk.11.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
169
- [ 111/ 255] blk.11.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q3_K .. size = 15.00 MiB -> 3.22 MiB
170
- [ 112/ 255] blk.12.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q3_K .. size = 1.50 MiB -> 0.32 MiB
171
- [ 113/ 255] blk.12.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
172
- [ 114/ 255] blk.12.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q4_K .. size = 6.00 MiB -> 1.69 MiB
173
- [ 115/ 255] blk.12.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q3_K .. size = 6.00 MiB -> 1.29 MiB
174
- [ 116/ 255] blk.12.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
175
- [ 117/ 255] blk.12.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q4_K .. size = 15.00 MiB -> 4.22 MiB
176
- [ 118/ 255] blk.12.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q3_K .. size = 15.00 MiB -> 3.22 MiB
177
- [ 119/ 255] blk.12.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
178
- [ 120/ 255] blk.12.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q3_K .. size = 15.00 MiB -> 3.22 MiB
179
- [ 121/ 255] blk.13.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q3_K .. size = 1.50 MiB -> 0.32 MiB
180
- [ 122/ 255] blk.13.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
181
- [ 123/ 255] blk.13.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q4_K .. size = 6.00 MiB -> 1.69 MiB
182
- [ 124/ 255] blk.13.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q3_K .. size = 6.00 MiB -> 1.29 MiB
183
- [ 125/ 255] blk.13.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
184
- [ 126/ 255] blk.13.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q4_K .. size = 15.00 MiB -> 4.22 MiB
185
- [ 127/ 255] blk.13.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q3_K .. size = 15.00 MiB -> 3.22 MiB
186
- [ 128/ 255] blk.13.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
187
- [ 129/ 255] blk.13.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q3_K .. size = 15.00 MiB -> 3.22 MiB
188
- [ 130/ 255] blk.14.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q3_K .. size = 1.50 MiB -> 0.32 MiB
189
- [ 131/ 255] blk.14.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
190
- [ 132/ 255] blk.14.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q4_K .. size = 6.00 MiB -> 1.69 MiB
191
- [ 133/ 255] blk.14.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q3_K .. size = 6.00 MiB -> 1.29 MiB
192
- [ 134/ 255] blk.14.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
193
- [ 135/ 255] blk.14.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q4_K .. size = 15.00 MiB -> 4.22 MiB
194
- [ 136/ 255] blk.14.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q3_K .. size = 15.00 MiB -> 3.22 MiB
195
- [ 137/ 255] blk.14.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
196
- [ 138/ 255] blk.14.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q3_K .. size = 15.00 MiB -> 3.22 MiB
197
- [ 139/ 255] blk.15.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q3_K .. size = 1.50 MiB -> 0.32 MiB
198
- [ 140/ 255] blk.15.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
199
- [ 141/ 255] blk.15.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q4_K .. size = 6.00 MiB -> 1.69 MiB
200
- [ 142/ 255] blk.15.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q3_K .. size = 6.00 MiB -> 1.29 MiB
201
- [ 143/ 255] blk.15.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
202
- [ 144/ 255] blk.15.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q4_K .. size = 15.00 MiB -> 4.22 MiB
203
- [ 145/ 255] blk.15.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q3_K .. size = 15.00 MiB -> 3.22 MiB
204
- [ 146/ 255] blk.15.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
205
- [ 147/ 255] blk.15.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q3_K .. size = 15.00 MiB -> 3.22 MiB
206
- [ 148/ 255] blk.16.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q3_K .. size = 1.50 MiB -> 0.32 MiB
207
- [ 149/ 255] blk.16.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
208
- [ 150/ 255] blk.16.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q4_K .. size = 6.00 MiB -> 1.69 MiB
209
- [ 151/ 255] blk.16.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q3_K .. size = 6.00 MiB -> 1.29 MiB
210
- [ 152/ 255] blk.16.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
211
- [ 153/ 255] blk.16.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q4_K .. size = 15.00 MiB -> 4.22 MiB
212
- [ 154/ 255] blk.16.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q3_K .. size = 15.00 MiB -> 3.22 MiB
213
- [ 155/ 255] blk.16.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
214
- [ 156/ 255] blk.16.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q3_K .. size = 15.00 MiB -> 3.22 MiB
215
- [ 157/ 255] blk.17.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q3_K .. size = 1.50 MiB -> 0.32 MiB
216
- [ 158/ 255] blk.17.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
217
- [ 159/ 255] blk.17.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q4_K .. size = 6.00 MiB -> 1.69 MiB
218
- [ 160/ 255] blk.17.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q3_K .. size = 6.00 MiB -> 1.29 MiB
219
- [ 161/ 255] blk.17.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
220
- [ 162/ 255] blk.17.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q4_K .. size = 15.00 MiB -> 4.22 MiB
221
- [ 163/ 255] blk.17.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q3_K .. size = 15.00 MiB -> 3.22 MiB
222
- [ 164/ 255] blk.17.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
223
- [ 165/ 255] blk.17.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q3_K .. size = 15.00 MiB -> 3.22 MiB
224
- [ 166/ 255] blk.18.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q3_K .. size = 1.50 MiB -> 0.32 MiB
225
- [ 167/ 255] blk.18.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
226
- [ 168/ 255] blk.18.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q4_K .. size = 6.00 MiB -> 1.69 MiB
227
- [ 169/ 255] blk.18.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q3_K .. size = 6.00 MiB -> 1.29 MiB
228
- [ 170/ 255] blk.18.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
229
- [ 171/ 255] blk.18.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q4_K .. size = 15.00 MiB -> 4.22 MiB
230
- [ 172/ 255] blk.18.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q3_K .. size = 15.00 MiB -> 3.22 MiB
231
- [ 173/ 255] blk.18.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
232
- [ 174/ 255] blk.18.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q3_K .. size = 15.00 MiB -> 3.22 MiB
233
- [ 175/ 255] blk.19.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q3_K .. size = 1.50 MiB -> 0.32 MiB
234
- [ 176/ 255] blk.19.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
235
- [ 177/ 255] blk.19.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q4_K .. size = 6.00 MiB -> 1.69 MiB
236
- [ 178/ 255] blk.19.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q3_K .. size = 6.00 MiB -> 1.29 MiB
237
- [ 179/ 255] blk.19.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
238
- [ 180/ 255] blk.19.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q4_K .. size = 15.00 MiB -> 4.22 MiB
239
- [ 181/ 255] blk.19.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q3_K .. size = 15.00 MiB -> 3.22 MiB
240
- [ 182/ 255] blk.19.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
241
- [ 183/ 255] blk.19.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q3_K .. size = 15.00 MiB -> 3.22 MiB
242
- [ 184/ 255] blk.20.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q3_K .. size = 1.50 MiB -> 0.32 MiB
243
- [ 185/ 255] blk.20.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
244
- [ 186/ 255] blk.20.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q4_K .. size = 6.00 MiB -> 1.69 MiB
245
- [ 187/ 255] blk.20.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q3_K .. size = 6.00 MiB -> 1.29 MiB
246
- [ 188/ 255] blk.20.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
247
- [ 189/ 255] blk.20.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q4_K .. size = 15.00 MiB -> 4.22 MiB
248
- [ 190/ 255] blk.20.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q3_K .. size = 15.00 MiB -> 3.22 MiB
249
- [ 191/ 255] blk.20.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
250
- [ 192/ 255] blk.20.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q3_K .. size = 15.00 MiB -> 3.22 MiB
251
- [ 193/ 255] blk.21.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q3_K .. size = 1.50 MiB -> 0.32 MiB
252
- [ 194/ 255] blk.21.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
253
- [ 195/ 255] blk.21.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q4_K .. size = 6.00 MiB -> 1.69 MiB
254
- [ 196/ 255] blk.21.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q3_K .. size = 6.00 MiB -> 1.29 MiB
255
- [ 197/ 255] blk.21.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
256
- [ 198/ 255] blk.21.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q4_K .. size = 15.00 MiB -> 4.22 MiB
257
- [ 199/ 255] blk.21.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q3_K .. size = 15.00 MiB -> 3.22 MiB
258
- [ 200/ 255] blk.21.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
259
- [ 201/ 255] blk.21.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q3_K .. size = 15.00 MiB -> 3.22 MiB
260
- [ 202/ 255] blk.22.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q3_K .. size = 1.50 MiB -> 0.32 MiB
261
- [ 203/ 255] blk.22.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
262
- [ 204/ 255] blk.22.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q4_K .. size = 6.00 MiB -> 1.69 MiB
263
- [ 205/ 255] blk.22.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q3_K .. size = 6.00 MiB -> 1.29 MiB
264
- [ 206/ 255] blk.22.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
265
- [ 207/ 255] blk.22.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q4_K .. size = 15.00 MiB -> 4.22 MiB
266
- [ 208/ 255] blk.22.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q3_K .. size = 15.00 MiB -> 3.22 MiB
267
- [ 209/ 255] blk.22.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
268
- [ 210/ 255] blk.22.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q3_K .. size = 15.00 MiB -> 3.22 MiB
269
- [ 211/ 255] blk.23.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q3_K .. size = 1.50 MiB -> 0.32 MiB
270
- [ 212/ 255] blk.23.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
271
- [ 213/ 255] blk.23.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q4_K .. size = 6.00 MiB -> 1.69 MiB
272
- [ 214/ 255] blk.23.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q3_K .. size = 6.00 MiB -> 1.29 MiB
273
- [ 215/ 255] blk.23.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
274
- [ 216/ 255] blk.23.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q4_K .. size = 15.00 MiB -> 4.22 MiB
275
- [ 217/ 255] blk.23.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q3_K .. size = 15.00 MiB -> 3.22 MiB
276
- [ 218/ 255] blk.23.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
277
- [ 219/ 255] blk.23.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q3_K .. size = 15.00 MiB -> 3.22 MiB
278
- [ 220/ 255] blk.24.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q3_K .. size = 1.50 MiB -> 0.32 MiB
279
- [ 221/ 255] blk.24.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
280
- [ 222/ 255] blk.24.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q4_K .. size = 6.00 MiB -> 1.69 MiB
281
- [ 223/ 255] blk.24.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q3_K .. size = 6.00 MiB -> 1.29 MiB
282
- [ 224/ 255] blk.24.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
283
- [ 225/ 255] blk.24.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q4_K .. size = 15.00 MiB -> 4.22 MiB
284
- [ 226/ 255] blk.24.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q3_K .. size = 15.00 MiB -> 3.22 MiB
285
- [ 227/ 255] blk.24.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
286
- [ 228/ 255] blk.24.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q3_K .. size = 15.00 MiB -> 3.22 MiB
287
- [ 229/ 255] blk.25.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q3_K .. size = 1.50 MiB -> 0.32 MiB
288
- [ 230/ 255] blk.25.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
289
- [ 231/ 255] blk.25.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q4_K .. size = 6.00 MiB -> 1.69 MiB
290
- [ 232/ 255] blk.25.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q3_K .. size = 6.00 MiB -> 1.29 MiB
291
- [ 233/ 255] blk.25.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
292
- [ 234/ 255] blk.25.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q4_K .. size = 15.00 MiB -> 4.22 MiB
293
- [ 235/ 255] blk.25.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q3_K .. size = 15.00 MiB -> 3.22 MiB
294
- [ 236/ 255] blk.25.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
295
- [ 237/ 255] blk.25.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q3_K .. size = 15.00 MiB -> 3.22 MiB
296
- [ 238/ 255] blk.26.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q3_K .. size = 1.50 MiB -> 0.32 MiB
297
- [ 239/ 255] blk.26.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
298
- [ 240/ 255] blk.26.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q4_K .. size = 6.00 MiB -> 1.69 MiB
299
- [ 241/ 255] blk.26.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q3_K .. size = 6.00 MiB -> 1.29 MiB
300
- [ 242/ 255] blk.26.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
301
- [ 243/ 255] blk.26.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q4_K .. size = 15.00 MiB -> 4.22 MiB
302
- [ 244/ 255] blk.26.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q3_K .. size = 15.00 MiB -> 3.22 MiB
303
- [ 245/ 255] blk.26.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
304
- [ 246/ 255] blk.26.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q3_K .. size = 15.00 MiB -> 3.22 MiB
305
- [ 247/ 255] blk.27.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q3_K .. size = 1.50 MiB -> 0.32 MiB
306
- [ 248/ 255] blk.27.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
307
- [ 249/ 255] blk.27.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q4_K .. size = 6.00 MiB -> 1.69 MiB
308
- [ 250/ 255] blk.27.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q3_K .. size = 6.00 MiB -> 1.29 MiB
309
- [ 251/ 255] blk.27.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
310
- [ 252/ 255] blk.27.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q4_K .. size = 15.00 MiB -> 4.22 MiB
311
- [ 253/ 255] blk.27.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q3_K .. size = 15.00 MiB -> 3.22 MiB
312
- [ 254/ 255] blk.27.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
313
- [ 255/ 255] blk.27.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q3_K .. size = 15.00 MiB -> 3.22 MiB
314
- llama_model_quantize_impl: model size = 2056.83 MiB (16.00 BPW)
315
- llama_model_quantize_impl: quant size = 521.89 MiB (4.06 BPW)
316
-
317
- llama_quantize: quantize time = 12134.65 ms
318
- llama_quantize: total time = 12134.65 ms
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
reproducibility/validation/quantize-logs/Q4_K_M.log DELETED
@@ -1,318 +0,0 @@
1
- ggml_cuda_init: failed to initialize CUDA: no CUDA-capable device is detected
2
- version: 0.3.0-dev (build 10671, commit 35999d101)
3
- built with GNU 13.3.0 for Linux x86_64
4
- llama_quantize: quantizing '/root/workspace/HF/source-official-bf16-k2-horizon-0.9b/K2-Horizon-0.9B-BF16.gguf' to '/root/workspace/HF/quantized-k2-horizon-0.9b/K2-Horizon-0.9B-Q4_K_M.gguf' as Q4_K_M using 128 threads
5
- llama_model_loader: loaded meta data with 41 key-value pairs and 255 tensors from /root/workspace/HF/source-official-bf16-k2-horizon-0.9b/K2-Horizon-0.9B-BF16.gguf (version GGUF V3 (latest))
6
- llama_model_loader: Dumping metadata keys/values. Note: KV overrides do not apply in this output.
7
- llama_model_loader: - kv 0: general.architecture str = k2-horizon
8
- llama_model_loader: - kv 1: general.type str = model
9
- llama_model_loader: - kv 2: general.name str = K2-Horizon-0.9B
10
- llama_model_loader: - kv 3: general.basename str = source-hf-k2-horizon
11
- llama_model_loader: - kv 4: general.size_label str = 0.9B
12
- llama_model_loader: - kv 5: general.license str = apache-2.0
13
- llama_model_loader: - kv 6: general.license.name str = internal-only
14
- llama_model_loader: - kv 7: general.license.link str = LICENSE
15
- llama_model_loader: - kv 8: general.tags arr[str,7] = ["k2-horizon", "0.9b", "dense", "reas...
16
- llama_model_loader: - kv 9: general.languages arr[str,2] = ["en", "zh"]
17
- llama_model_loader: - kv 10: k2-horizon.block_count u32 = 28
18
- llama_model_loader: - kv 11: k2-horizon.context_length u32 = 131072
19
- llama_model_loader: - kv 12: k2-horizon.embedding_length u32 = 1536
20
- llama_model_loader: - kv 13: k2-horizon.feed_forward_length u32 = 5120
21
- llama_model_loader: - kv 14: k2-horizon.attention.head_count u32 = 32
22
- llama_model_loader: - kv 15: k2-horizon.attention.head_count_kv u32 = 8
23
- llama_model_loader: - kv 16: k2-horizon.rope.scaling.type str = yarn
24
- llama_model_loader: - kv 17: k2-horizon.rope.scaling.factor f32 = 16.000000
25
- llama_model_loader: - kv 18: k2-horizon.rope.scaling.original_context_length u32 = 8192
26
- llama_model_loader: - kv 19: k2-horizon.rope.scaling.yarn_attn_factor f32 = 1.277259
27
- llama_model_loader: - kv 20: k2-horizon.rope.scaling.yarn_beta_fast f32 = 128.000000
28
- llama_model_loader: - kv 21: k2-horizon.rope.scaling.yarn_beta_slow f32 = 4.000000
29
- llama_model_loader: - kv 22: k2-horizon.rope.freq_base f32 = 1000000.000000
30
- llama_model_loader: - kv 23: k2-horizon.attention.layer_norm_rms_epsilon f32 = 0.000001
31
- llama_model_loader: - kv 24: k2-horizon.expert_count u32 = 0
32
- llama_model_loader: - kv 25: k2-horizon.expert_used_count u32 = 0
33
- llama_model_loader: - kv 26: k2-horizon.attention.key_length u32 = 64
34
- llama_model_loader: - kv 27: k2-horizon.attention.value_length u32 = 64
35
- llama_model_loader: - kv 28: general.file_type u32 = 32
36
- llama_model_loader: - kv 29: k2-horizon.attention.group_norm_groups u32 = 1
37
- llama_model_loader: - kv 30: k2-horizon.rope.dimension_count u32 = 64
38
- llama_model_loader: - kv 31: general.quantization_version u32 = 2
39
- llama_model_loader: - kv 32: tokenizer.ggml.model str = gpt2
40
- llama_model_loader: - kv 33: tokenizer.ggml.pre str = k2-horizon
41
- llama_model_loader: - kv 34: tokenizer.ggml.tokens arr[str,64256] = ["<|begin_of_text|>", "<|endoftext|>"...
42
- llama_model_loader: - kv 35: tokenizer.ggml.token_type arr[i32,64256] = [3, 3, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, ...
43
- llama_model_loader: - kv 36: tokenizer.ggml.merges arr[str,63742] = ["Δ  Δ ", "Ø Β§", "Γ™ Δ¦", "Γ  Β€", ...
44
- llama_model_loader: - kv 37: tokenizer.ggml.bos_token_id u32 = 0
45
- llama_model_loader: - kv 38: tokenizer.ggml.eos_token_id u32 = 1
46
- llama_model_loader: - kv 39: tokenizer.ggml.padding_token_id u32 = 64255
47
- llama_model_loader: - kv 40: tokenizer.chat_template str = {{- bos_token }}\n{%- if tool_presenta...
48
- llama_model_loader: - type f32: 57 tensors
49
- llama_model_loader: - type bf16: 198 tensors
50
-
51
- llama_model_quantize_impl: have importance matrix data with 196 entries
52
- [ 1/ 255] output.weight - [ 1536, 64256, 1, 1], type = bf16,
53
- ====== llama_model_quantize_impl: did not find weights for output.weight
54
- converting to q6_K .. load_imatrix: imatrix datasets=['calibration/wikitext-2-raw/wiki.train.raw', 'calibration/k2-horizon-0.9b/k2_horizon_en_zh_code_tool.txt']
55
- load_imatrix: loaded 196 importance matrix entries from /root/workspace/HF/calibration/k2-horizon-0.9b/k2_horizon_combined.imatrix.gguf computed on 99 chunks
56
- prepare_imatrix: have 196 importance matrix entries
57
- size = 188.25 MiB -> 77.21 MiB
58
- [ 2/ 255] output_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
59
- [ 3/ 255] token_embd.weight - [ 1536, 64256, 1, 1], type = bf16,
60
- ====== llama_model_quantize_impl: did not find weights for token_embd.weight
61
- converting to q4_K .. size = 188.25 MiB -> 52.95 MiB
62
- [ 4/ 255] blk.0.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
63
- [ 5/ 255] blk.0.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
64
- [ 6/ 255] blk.0.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q4_K .. size = 6.00 MiB -> 1.69 MiB
65
- [ 7/ 255] blk.0.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q4_K .. size = 6.00 MiB -> 1.69 MiB
66
- [ 8/ 255] blk.0.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q6_K .. size = 1.50 MiB -> 0.62 MiB
67
- [ 9/ 255] blk.0.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
68
- [ 10/ 255] blk.0.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q4_K .. size = 15.00 MiB -> 4.22 MiB
69
- [ 11/ 255] blk.0.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
70
- [ 12/ 255] blk.0.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q4_K .. size = 15.00 MiB -> 4.22 MiB
71
- [ 13/ 255] blk.1.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
72
- [ 14/ 255] blk.1.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
73
- [ 15/ 255] blk.1.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q4_K .. size = 6.00 MiB -> 1.69 MiB
74
- [ 16/ 255] blk.1.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q4_K .. size = 6.00 MiB -> 1.69 MiB
75
- [ 17/ 255] blk.1.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q6_K .. size = 1.50 MiB -> 0.62 MiB
76
- [ 18/ 255] blk.1.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
77
- [ 19/ 255] blk.1.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q4_K .. size = 15.00 MiB -> 4.22 MiB
78
- [ 20/ 255] blk.1.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
79
- [ 21/ 255] blk.1.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q4_K .. size = 15.00 MiB -> 4.22 MiB
80
- [ 22/ 255] blk.2.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
81
- [ 23/ 255] blk.2.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
82
- [ 24/ 255] blk.2.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q4_K .. size = 6.00 MiB -> 1.69 MiB
83
- [ 25/ 255] blk.2.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q4_K .. size = 6.00 MiB -> 1.69 MiB
84
- [ 26/ 255] blk.2.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q6_K .. size = 1.50 MiB -> 0.62 MiB
85
- [ 27/ 255] blk.2.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
86
- [ 28/ 255] blk.2.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q4_K .. size = 15.00 MiB -> 4.22 MiB
87
- [ 29/ 255] blk.2.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
88
- [ 30/ 255] blk.2.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q4_K .. size = 15.00 MiB -> 4.22 MiB
89
- [ 31/ 255] blk.3.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
90
- [ 32/ 255] blk.3.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
91
- [ 33/ 255] blk.3.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q4_K .. size = 6.00 MiB -> 1.69 MiB
92
- [ 34/ 255] blk.3.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q4_K .. size = 6.00 MiB -> 1.69 MiB
93
- [ 35/ 255] blk.3.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
94
- [ 36/ 255] blk.3.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q4_K .. size = 15.00 MiB -> 4.22 MiB
95
- [ 37/ 255] blk.3.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q4_K .. size = 15.00 MiB -> 4.22 MiB
96
- [ 38/ 255] blk.3.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
97
- [ 39/ 255] blk.3.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q4_K .. size = 15.00 MiB -> 4.22 MiB
98
- [ 40/ 255] blk.4.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
99
- [ 41/ 255] blk.4.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
100
- [ 42/ 255] blk.4.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q4_K .. size = 6.00 MiB -> 1.69 MiB
101
- [ 43/ 255] blk.4.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q4_K .. size = 6.00 MiB -> 1.69 MiB
102
- [ 44/ 255] blk.4.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
103
- [ 45/ 255] blk.4.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q4_K .. size = 15.00 MiB -> 4.22 MiB
104
- [ 46/ 255] blk.4.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q4_K .. size = 15.00 MiB -> 4.22 MiB
105
- [ 47/ 255] blk.4.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
106
- [ 48/ 255] blk.4.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q4_K .. size = 15.00 MiB -> 4.22 MiB
107
- [ 49/ 255] blk.5.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
108
- [ 50/ 255] blk.5.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
109
- [ 51/ 255] blk.5.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q4_K .. size = 6.00 MiB -> 1.69 MiB
110
- [ 52/ 255] blk.5.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q4_K .. size = 6.00 MiB -> 1.69 MiB
111
- [ 53/ 255] blk.5.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q6_K .. size = 1.50 MiB -> 0.62 MiB
112
- [ 54/ 255] blk.5.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
113
- [ 55/ 255] blk.5.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q4_K .. size = 15.00 MiB -> 4.22 MiB
114
- [ 56/ 255] blk.5.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
115
- [ 57/ 255] blk.5.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q4_K .. size = 15.00 MiB -> 4.22 MiB
116
- [ 58/ 255] blk.6.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
117
- [ 59/ 255] blk.6.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
118
- [ 60/ 255] blk.6.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q4_K .. size = 6.00 MiB -> 1.69 MiB
119
- [ 61/ 255] blk.6.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q4_K .. size = 6.00 MiB -> 1.69 MiB
120
- [ 62/ 255] blk.6.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
121
- [ 63/ 255] blk.6.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q4_K .. size = 15.00 MiB -> 4.22 MiB
122
- [ 64/ 255] blk.6.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q4_K .. size = 15.00 MiB -> 4.22 MiB
123
- [ 65/ 255] blk.6.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
124
- [ 66/ 255] blk.6.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q4_K .. size = 15.00 MiB -> 4.22 MiB
125
- [ 67/ 255] blk.7.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
126
- [ 68/ 255] blk.7.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
127
- [ 69/ 255] blk.7.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q4_K .. size = 6.00 MiB -> 1.69 MiB
128
- [ 70/ 255] blk.7.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q4_K .. size = 6.00 MiB -> 1.69 MiB
129
- [ 71/ 255] blk.7.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
130
- [ 72/ 255] blk.7.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q4_K .. size = 15.00 MiB -> 4.22 MiB
131
- [ 73/ 255] blk.7.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q4_K .. size = 15.00 MiB -> 4.22 MiB
132
- [ 74/ 255] blk.7.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
133
- [ 75/ 255] blk.7.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q4_K .. size = 15.00 MiB -> 4.22 MiB
134
- [ 76/ 255] blk.8.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
135
- [ 77/ 255] blk.8.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
136
- [ 78/ 255] blk.8.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q4_K .. size = 6.00 MiB -> 1.69 MiB
137
- [ 79/ 255] blk.8.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q4_K .. size = 6.00 MiB -> 1.69 MiB
138
- [ 80/ 255] blk.8.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q6_K .. size = 1.50 MiB -> 0.62 MiB
139
- [ 81/ 255] blk.8.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
140
- [ 82/ 255] blk.8.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q4_K .. size = 15.00 MiB -> 4.22 MiB
141
- [ 83/ 255] blk.8.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
142
- [ 84/ 255] blk.8.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q4_K .. size = 15.00 MiB -> 4.22 MiB
143
- [ 85/ 255] blk.9.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
144
- [ 86/ 255] blk.9.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
145
- [ 87/ 255] blk.9.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q4_K .. size = 6.00 MiB -> 1.69 MiB
146
- [ 88/ 255] blk.9.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q4_K .. size = 6.00 MiB -> 1.69 MiB
147
- [ 89/ 255] blk.9.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
148
- [ 90/ 255] blk.9.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q4_K .. size = 15.00 MiB -> 4.22 MiB
149
- [ 91/ 255] blk.9.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q4_K .. size = 15.00 MiB -> 4.22 MiB
150
- [ 92/ 255] blk.9.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
151
- [ 93/ 255] blk.9.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q4_K .. size = 15.00 MiB -> 4.22 MiB
152
- [ 94/ 255] blk.10.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
153
- [ 95/ 255] blk.10.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
154
- [ 96/ 255] blk.10.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q4_K .. size = 6.00 MiB -> 1.69 MiB
155
- [ 97/ 255] blk.10.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q4_K .. size = 6.00 MiB -> 1.69 MiB
156
- [ 98/ 255] blk.10.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
157
- [ 99/ 255] blk.10.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q4_K .. size = 15.00 MiB -> 4.22 MiB
158
- [ 100/ 255] blk.10.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q4_K .. size = 15.00 MiB -> 4.22 MiB
159
- [ 101/ 255] blk.10.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
160
- [ 102/ 255] blk.10.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q4_K .. size = 15.00 MiB -> 4.22 MiB
161
- [ 103/ 255] blk.11.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
162
- [ 104/ 255] blk.11.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
163
- [ 105/ 255] blk.11.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q4_K .. size = 6.00 MiB -> 1.69 MiB
164
- [ 106/ 255] blk.11.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q4_K .. size = 6.00 MiB -> 1.69 MiB
165
- [ 107/ 255] blk.11.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q6_K .. size = 1.50 MiB -> 0.62 MiB
166
- [ 108/ 255] blk.11.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
167
- [ 109/ 255] blk.11.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q4_K .. size = 15.00 MiB -> 4.22 MiB
168
- [ 110/ 255] blk.11.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
169
- [ 111/ 255] blk.11.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q4_K .. size = 15.00 MiB -> 4.22 MiB
170
- [ 112/ 255] blk.12.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
171
- [ 113/ 255] blk.12.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
172
- [ 114/ 255] blk.12.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q4_K .. size = 6.00 MiB -> 1.69 MiB
173
- [ 115/ 255] blk.12.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q4_K .. size = 6.00 MiB -> 1.69 MiB
174
- [ 116/ 255] blk.12.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
175
- [ 117/ 255] blk.12.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q4_K .. size = 15.00 MiB -> 4.22 MiB
176
- [ 118/ 255] blk.12.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q4_K .. size = 15.00 MiB -> 4.22 MiB
177
- [ 119/ 255] blk.12.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
178
- [ 120/ 255] blk.12.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q4_K .. size = 15.00 MiB -> 4.22 MiB
179
- [ 121/ 255] blk.13.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
180
- [ 122/ 255] blk.13.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
181
- [ 123/ 255] blk.13.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q4_K .. size = 6.00 MiB -> 1.69 MiB
182
- [ 124/ 255] blk.13.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q4_K .. size = 6.00 MiB -> 1.69 MiB
183
- [ 125/ 255] blk.13.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
184
- [ 126/ 255] blk.13.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q4_K .. size = 15.00 MiB -> 4.22 MiB
185
- [ 127/ 255] blk.13.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q4_K .. size = 15.00 MiB -> 4.22 MiB
186
- [ 128/ 255] blk.13.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
187
- [ 129/ 255] blk.13.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q4_K .. size = 15.00 MiB -> 4.22 MiB
188
- [ 130/ 255] blk.14.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
189
- [ 131/ 255] blk.14.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
190
- [ 132/ 255] blk.14.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q4_K .. size = 6.00 MiB -> 1.69 MiB
191
- [ 133/ 255] blk.14.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q4_K .. size = 6.00 MiB -> 1.69 MiB
192
- [ 134/ 255] blk.14.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q6_K .. size = 1.50 MiB -> 0.62 MiB
193
- [ 135/ 255] blk.14.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
194
- [ 136/ 255] blk.14.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q4_K .. size = 15.00 MiB -> 4.22 MiB
195
- [ 137/ 255] blk.14.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
196
- [ 138/ 255] blk.14.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q4_K .. size = 15.00 MiB -> 4.22 MiB
197
- [ 139/ 255] blk.15.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
198
- [ 140/ 255] blk.15.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
199
- [ 141/ 255] blk.15.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q4_K .. size = 6.00 MiB -> 1.69 MiB
200
- [ 142/ 255] blk.15.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q4_K .. size = 6.00 MiB -> 1.69 MiB
201
- [ 143/ 255] blk.15.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
202
- [ 144/ 255] blk.15.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q4_K .. size = 15.00 MiB -> 4.22 MiB
203
- [ 145/ 255] blk.15.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q4_K .. size = 15.00 MiB -> 4.22 MiB
204
- [ 146/ 255] blk.15.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
205
- [ 147/ 255] blk.15.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q4_K .. size = 15.00 MiB -> 4.22 MiB
206
- [ 148/ 255] blk.16.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
207
- [ 149/ 255] blk.16.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
208
- [ 150/ 255] blk.16.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q4_K .. size = 6.00 MiB -> 1.69 MiB
209
- [ 151/ 255] blk.16.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q4_K .. size = 6.00 MiB -> 1.69 MiB
210
- [ 152/ 255] blk.16.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
211
- [ 153/ 255] blk.16.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q4_K .. size = 15.00 MiB -> 4.22 MiB
212
- [ 154/ 255] blk.16.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q4_K .. size = 15.00 MiB -> 4.22 MiB
213
- [ 155/ 255] blk.16.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
214
- [ 156/ 255] blk.16.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q4_K .. size = 15.00 MiB -> 4.22 MiB
215
- [ 157/ 255] blk.17.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
216
- [ 158/ 255] blk.17.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
217
- [ 159/ 255] blk.17.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q4_K .. size = 6.00 MiB -> 1.69 MiB
218
- [ 160/ 255] blk.17.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q4_K .. size = 6.00 MiB -> 1.69 MiB
219
- [ 161/ 255] blk.17.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q6_K .. size = 1.50 MiB -> 0.62 MiB
220
- [ 162/ 255] blk.17.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
221
- [ 163/ 255] blk.17.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q4_K .. size = 15.00 MiB -> 4.22 MiB
222
- [ 164/ 255] blk.17.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
223
- [ 165/ 255] blk.17.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q4_K .. size = 15.00 MiB -> 4.22 MiB
224
- [ 166/ 255] blk.18.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
225
- [ 167/ 255] blk.18.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
226
- [ 168/ 255] blk.18.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q4_K .. size = 6.00 MiB -> 1.69 MiB
227
- [ 169/ 255] blk.18.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q4_K .. size = 6.00 MiB -> 1.69 MiB
228
- [ 170/ 255] blk.18.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
229
- [ 171/ 255] blk.18.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q4_K .. size = 15.00 MiB -> 4.22 MiB
230
- [ 172/ 255] blk.18.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q4_K .. size = 15.00 MiB -> 4.22 MiB
231
- [ 173/ 255] blk.18.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
232
- [ 174/ 255] blk.18.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q4_K .. size = 15.00 MiB -> 4.22 MiB
233
- [ 175/ 255] blk.19.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
234
- [ 176/ 255] blk.19.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
235
- [ 177/ 255] blk.19.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q4_K .. size = 6.00 MiB -> 1.69 MiB
236
- [ 178/ 255] blk.19.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q4_K .. size = 6.00 MiB -> 1.69 MiB
237
- [ 179/ 255] blk.19.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
238
- [ 180/ 255] blk.19.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q4_K .. size = 15.00 MiB -> 4.22 MiB
239
- [ 181/ 255] blk.19.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q4_K .. size = 15.00 MiB -> 4.22 MiB
240
- [ 182/ 255] blk.19.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
241
- [ 183/ 255] blk.19.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q4_K .. size = 15.00 MiB -> 4.22 MiB
242
- [ 184/ 255] blk.20.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
243
- [ 185/ 255] blk.20.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
244
- [ 186/ 255] blk.20.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q4_K .. size = 6.00 MiB -> 1.69 MiB
245
- [ 187/ 255] blk.20.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q4_K .. size = 6.00 MiB -> 1.69 MiB
246
- [ 188/ 255] blk.20.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q6_K .. size = 1.50 MiB -> 0.62 MiB
247
- [ 189/ 255] blk.20.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
248
- [ 190/ 255] blk.20.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q4_K .. size = 15.00 MiB -> 4.22 MiB
249
- [ 191/ 255] blk.20.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
250
- [ 192/ 255] blk.20.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q4_K .. size = 15.00 MiB -> 4.22 MiB
251
- [ 193/ 255] blk.21.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
252
- [ 194/ 255] blk.21.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
253
- [ 195/ 255] blk.21.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q4_K .. size = 6.00 MiB -> 1.69 MiB
254
- [ 196/ 255] blk.21.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q4_K .. size = 6.00 MiB -> 1.69 MiB
255
- [ 197/ 255] blk.21.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
256
- [ 198/ 255] blk.21.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q4_K .. size = 15.00 MiB -> 4.22 MiB
257
- [ 199/ 255] blk.21.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q4_K .. size = 15.00 MiB -> 4.22 MiB
258
- [ 200/ 255] blk.21.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
259
- [ 201/ 255] blk.21.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q4_K .. size = 15.00 MiB -> 4.22 MiB
260
- [ 202/ 255] blk.22.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
261
- [ 203/ 255] blk.22.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
262
- [ 204/ 255] blk.22.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q4_K .. size = 6.00 MiB -> 1.69 MiB
263
- [ 205/ 255] blk.22.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q4_K .. size = 6.00 MiB -> 1.69 MiB
264
- [ 206/ 255] blk.22.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
265
- [ 207/ 255] blk.22.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q4_K .. size = 15.00 MiB -> 4.22 MiB
266
- [ 208/ 255] blk.22.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q4_K .. size = 15.00 MiB -> 4.22 MiB
267
- [ 209/ 255] blk.22.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
268
- [ 210/ 255] blk.22.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q4_K .. size = 15.00 MiB -> 4.22 MiB
269
- [ 211/ 255] blk.23.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
270
- [ 212/ 255] blk.23.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
271
- [ 213/ 255] blk.23.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q4_K .. size = 6.00 MiB -> 1.69 MiB
272
- [ 214/ 255] blk.23.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q4_K .. size = 6.00 MiB -> 1.69 MiB
273
- [ 215/ 255] blk.23.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q6_K .. size = 1.50 MiB -> 0.62 MiB
274
- [ 216/ 255] blk.23.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
275
- [ 217/ 255] blk.23.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q4_K .. size = 15.00 MiB -> 4.22 MiB
276
- [ 218/ 255] blk.23.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
277
- [ 219/ 255] blk.23.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q4_K .. size = 15.00 MiB -> 4.22 MiB
278
- [ 220/ 255] blk.24.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
279
- [ 221/ 255] blk.24.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
280
- [ 222/ 255] blk.24.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q4_K .. size = 6.00 MiB -> 1.69 MiB
281
- [ 223/ 255] blk.24.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q4_K .. size = 6.00 MiB -> 1.69 MiB
282
- [ 224/ 255] blk.24.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q6_K .. size = 1.50 MiB -> 0.62 MiB
283
- [ 225/ 255] blk.24.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
284
- [ 226/ 255] blk.24.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q4_K .. size = 15.00 MiB -> 4.22 MiB
285
- [ 227/ 255] blk.24.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
286
- [ 228/ 255] blk.24.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q4_K .. size = 15.00 MiB -> 4.22 MiB
287
- [ 229/ 255] blk.25.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
288
- [ 230/ 255] blk.25.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
289
- [ 231/ 255] blk.25.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q4_K .. size = 6.00 MiB -> 1.69 MiB
290
- [ 232/ 255] blk.25.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q4_K .. size = 6.00 MiB -> 1.69 MiB
291
- [ 233/ 255] blk.25.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q6_K .. size = 1.50 MiB -> 0.62 MiB
292
- [ 234/ 255] blk.25.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
293
- [ 235/ 255] blk.25.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q4_K .. size = 15.00 MiB -> 4.22 MiB
294
- [ 236/ 255] blk.25.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
295
- [ 237/ 255] blk.25.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q4_K .. size = 15.00 MiB -> 4.22 MiB
296
- [ 238/ 255] blk.26.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
297
- [ 239/ 255] blk.26.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
298
- [ 240/ 255] blk.26.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q4_K .. size = 6.00 MiB -> 1.69 MiB
299
- [ 241/ 255] blk.26.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q4_K .. size = 6.00 MiB -> 1.69 MiB
300
- [ 242/ 255] blk.26.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q6_K .. size = 1.50 MiB -> 0.62 MiB
301
- [ 243/ 255] blk.26.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
302
- [ 244/ 255] blk.26.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q4_K .. size = 15.00 MiB -> 4.22 MiB
303
- [ 245/ 255] blk.26.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
304
- [ 246/ 255] blk.26.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q4_K .. size = 15.00 MiB -> 4.22 MiB
305
- [ 247/ 255] blk.27.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q4_K .. size = 1.50 MiB -> 0.42 MiB
306
- [ 248/ 255] blk.27.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
307
- [ 249/ 255] blk.27.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q4_K .. size = 6.00 MiB -> 1.69 MiB
308
- [ 250/ 255] blk.27.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q4_K .. size = 6.00 MiB -> 1.69 MiB
309
- [ 251/ 255] blk.27.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q6_K .. size = 1.50 MiB -> 0.62 MiB
310
- [ 252/ 255] blk.27.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
311
- [ 253/ 255] blk.27.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q4_K .. size = 15.00 MiB -> 4.22 MiB
312
- [ 254/ 255] blk.27.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
313
- [ 255/ 255] blk.27.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q4_K .. size = 15.00 MiB -> 4.22 MiB
314
- llama_model_quantize_impl: model size = 2056.83 MiB (16.00 BPW)
315
- llama_model_quantize_impl: quant size = 632.77 MiB (4.92 BPW)
316
-
317
- llama_quantize: quantize time = 12638.64 ms
318
- llama_quantize: total time = 12638.64 ms
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
reproducibility/validation/quantize-logs/Q5_K_M.log DELETED
@@ -1,318 +0,0 @@
1
- ggml_cuda_init: failed to initialize CUDA: no CUDA-capable device is detected
2
- version: 0.3.0-dev (build 10671, commit 35999d101)
3
- built with GNU 13.3.0 for Linux x86_64
4
- llama_quantize: quantizing '/root/workspace/HF/source-official-bf16-k2-horizon-0.9b/K2-Horizon-0.9B-BF16.gguf' to '/root/workspace/HF/quantized-k2-horizon-0.9b/K2-Horizon-0.9B-Q5_K_M.gguf' as Q5_K_M using 128 threads
5
- llama_model_loader: loaded meta data with 41 key-value pairs and 255 tensors from /root/workspace/HF/source-official-bf16-k2-horizon-0.9b/K2-Horizon-0.9B-BF16.gguf (version GGUF V3 (latest))
6
- llama_model_loader: Dumping metadata keys/values. Note: KV overrides do not apply in this output.
7
- llama_model_loader: - kv 0: general.architecture str = k2-horizon
8
- llama_model_loader: - kv 1: general.type str = model
9
- llama_model_loader: - kv 2: general.name str = K2-Horizon-0.9B
10
- llama_model_loader: - kv 3: general.basename str = source-hf-k2-horizon
11
- llama_model_loader: - kv 4: general.size_label str = 0.9B
12
- llama_model_loader: - kv 5: general.license str = apache-2.0
13
- llama_model_loader: - kv 6: general.license.name str = internal-only
14
- llama_model_loader: - kv 7: general.license.link str = LICENSE
15
- llama_model_loader: - kv 8: general.tags arr[str,7] = ["k2-horizon", "0.9b", "dense", "reas...
16
- llama_model_loader: - kv 9: general.languages arr[str,2] = ["en", "zh"]
17
- llama_model_loader: - kv 10: k2-horizon.block_count u32 = 28
18
- llama_model_loader: - kv 11: k2-horizon.context_length u32 = 131072
19
- llama_model_loader: - kv 12: k2-horizon.embedding_length u32 = 1536
20
- llama_model_loader: - kv 13: k2-horizon.feed_forward_length u32 = 5120
21
- llama_model_loader: - kv 14: k2-horizon.attention.head_count u32 = 32
22
- llama_model_loader: - kv 15: k2-horizon.attention.head_count_kv u32 = 8
23
- llama_model_loader: - kv 16: k2-horizon.rope.scaling.type str = yarn
24
- llama_model_loader: - kv 17: k2-horizon.rope.scaling.factor f32 = 16.000000
25
- llama_model_loader: - kv 18: k2-horizon.rope.scaling.original_context_length u32 = 8192
26
- llama_model_loader: - kv 19: k2-horizon.rope.scaling.yarn_attn_factor f32 = 1.277259
27
- llama_model_loader: - kv 20: k2-horizon.rope.scaling.yarn_beta_fast f32 = 128.000000
28
- llama_model_loader: - kv 21: k2-horizon.rope.scaling.yarn_beta_slow f32 = 4.000000
29
- llama_model_loader: - kv 22: k2-horizon.rope.freq_base f32 = 1000000.000000
30
- llama_model_loader: - kv 23: k2-horizon.attention.layer_norm_rms_epsilon f32 = 0.000001
31
- llama_model_loader: - kv 24: k2-horizon.expert_count u32 = 0
32
- llama_model_loader: - kv 25: k2-horizon.expert_used_count u32 = 0
33
- llama_model_loader: - kv 26: k2-horizon.attention.key_length u32 = 64
34
- llama_model_loader: - kv 27: k2-horizon.attention.value_length u32 = 64
35
- llama_model_loader: - kv 28: general.file_type u32 = 32
36
- llama_model_loader: - kv 29: k2-horizon.attention.group_norm_groups u32 = 1
37
- llama_model_loader: - kv 30: k2-horizon.rope.dimension_count u32 = 64
38
- llama_model_loader: - kv 31: general.quantization_version u32 = 2
39
- llama_model_loader: - kv 32: tokenizer.ggml.model str = gpt2
40
- llama_model_loader: - kv 33: tokenizer.ggml.pre str = k2-horizon
41
- llama_model_loader: - kv 34: tokenizer.ggml.tokens arr[str,64256] = ["<|begin_of_text|>", "<|endoftext|>"...
42
- llama_model_loader: - kv 35: tokenizer.ggml.token_type arr[i32,64256] = [3, 3, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, ...
43
- llama_model_loader: - kv 36: tokenizer.ggml.merges arr[str,63742] = ["Δ  Δ ", "Ø Β§", "Γ™ Δ¦", "Γ  Β€", ...
44
- llama_model_loader: - kv 37: tokenizer.ggml.bos_token_id u32 = 0
45
- llama_model_loader: - kv 38: tokenizer.ggml.eos_token_id u32 = 1
46
- llama_model_loader: - kv 39: tokenizer.ggml.padding_token_id u32 = 64255
47
- llama_model_loader: - kv 40: tokenizer.chat_template str = {{- bos_token }}\n{%- if tool_presenta...
48
- llama_model_loader: - type f32: 57 tensors
49
- llama_model_loader: - type bf16: 198 tensors
50
-
51
- llama_model_quantize_impl: have importance matrix data with 196 entries
52
- [ 1/ 255] output.weight - [ 1536, 64256, 1, 1], type = bf16,
53
- ====== llama_model_quantize_impl: did not find weights for output.weight
54
- converting to q6_K .. load_imatrix: imatrix datasets=['calibration/wikitext-2-raw/wiki.train.raw', 'calibration/k2-horizon-0.9b/k2_horizon_en_zh_code_tool.txt']
55
- load_imatrix: loaded 196 importance matrix entries from /root/workspace/HF/calibration/k2-horizon-0.9b/k2_horizon_combined.imatrix.gguf computed on 99 chunks
56
- prepare_imatrix: have 196 importance matrix entries
57
- size = 188.25 MiB -> 77.21 MiB
58
- [ 2/ 255] output_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
59
- [ 3/ 255] token_embd.weight - [ 1536, 64256, 1, 1], type = bf16,
60
- ====== llama_model_quantize_impl: did not find weights for token_embd.weight
61
- converting to q5_K .. size = 188.25 MiB -> 64.71 MiB
62
- [ 4/ 255] blk.0.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q5_K .. size = 1.50 MiB -> 0.52 MiB
63
- [ 5/ 255] blk.0.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
64
- [ 6/ 255] blk.0.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q5_K .. size = 6.00 MiB -> 2.06 MiB
65
- [ 7/ 255] blk.0.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q5_K .. size = 6.00 MiB -> 2.06 MiB
66
- [ 8/ 255] blk.0.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q6_K .. size = 1.50 MiB -> 0.62 MiB
67
- [ 9/ 255] blk.0.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
68
- [ 10/ 255] blk.0.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q5_K .. size = 15.00 MiB -> 5.16 MiB
69
- [ 11/ 255] blk.0.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
70
- [ 12/ 255] blk.0.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q5_K .. size = 15.00 MiB -> 5.16 MiB
71
- [ 13/ 255] blk.1.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q5_K .. size = 1.50 MiB -> 0.52 MiB
72
- [ 14/ 255] blk.1.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
73
- [ 15/ 255] blk.1.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q5_K .. size = 6.00 MiB -> 2.06 MiB
74
- [ 16/ 255] blk.1.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q5_K .. size = 6.00 MiB -> 2.06 MiB
75
- [ 17/ 255] blk.1.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q6_K .. size = 1.50 MiB -> 0.62 MiB
76
- [ 18/ 255] blk.1.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
77
- [ 19/ 255] blk.1.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q5_K .. size = 15.00 MiB -> 5.16 MiB
78
- [ 20/ 255] blk.1.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
79
- [ 21/ 255] blk.1.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q5_K .. size = 15.00 MiB -> 5.16 MiB
80
- [ 22/ 255] blk.2.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q5_K .. size = 1.50 MiB -> 0.52 MiB
81
- [ 23/ 255] blk.2.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
82
- [ 24/ 255] blk.2.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q5_K .. size = 6.00 MiB -> 2.06 MiB
83
- [ 25/ 255] blk.2.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q5_K .. size = 6.00 MiB -> 2.06 MiB
84
- [ 26/ 255] blk.2.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q6_K .. size = 1.50 MiB -> 0.62 MiB
85
- [ 27/ 255] blk.2.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
86
- [ 28/ 255] blk.2.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q5_K .. size = 15.00 MiB -> 5.16 MiB
87
- [ 29/ 255] blk.2.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
88
- [ 30/ 255] blk.2.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q5_K .. size = 15.00 MiB -> 5.16 MiB
89
- [ 31/ 255] blk.3.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q5_K .. size = 1.50 MiB -> 0.52 MiB
90
- [ 32/ 255] blk.3.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
91
- [ 33/ 255] blk.3.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q5_K .. size = 6.00 MiB -> 2.06 MiB
92
- [ 34/ 255] blk.3.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q5_K .. size = 6.00 MiB -> 2.06 MiB
93
- [ 35/ 255] blk.3.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q5_K .. size = 1.50 MiB -> 0.52 MiB
94
- [ 36/ 255] blk.3.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q5_K .. size = 15.00 MiB -> 5.16 MiB
95
- [ 37/ 255] blk.3.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q5_K .. size = 15.00 MiB -> 5.16 MiB
96
- [ 38/ 255] blk.3.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
97
- [ 39/ 255] blk.3.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q5_K .. size = 15.00 MiB -> 5.16 MiB
98
- [ 40/ 255] blk.4.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q5_K .. size = 1.50 MiB -> 0.52 MiB
99
- [ 41/ 255] blk.4.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
100
- [ 42/ 255] blk.4.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q5_K .. size = 6.00 MiB -> 2.06 MiB
101
- [ 43/ 255] blk.4.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q5_K .. size = 6.00 MiB -> 2.06 MiB
102
- [ 44/ 255] blk.4.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q5_K .. size = 1.50 MiB -> 0.52 MiB
103
- [ 45/ 255] blk.4.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q5_K .. size = 15.00 MiB -> 5.16 MiB
104
- [ 46/ 255] blk.4.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q5_K .. size = 15.00 MiB -> 5.16 MiB
105
- [ 47/ 255] blk.4.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
106
- [ 48/ 255] blk.4.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q5_K .. size = 15.00 MiB -> 5.16 MiB
107
- [ 49/ 255] blk.5.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q5_K .. size = 1.50 MiB -> 0.52 MiB
108
- [ 50/ 255] blk.5.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
109
- [ 51/ 255] blk.5.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q5_K .. size = 6.00 MiB -> 2.06 MiB
110
- [ 52/ 255] blk.5.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q5_K .. size = 6.00 MiB -> 2.06 MiB
111
- [ 53/ 255] blk.5.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q6_K .. size = 1.50 MiB -> 0.62 MiB
112
- [ 54/ 255] blk.5.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
113
- [ 55/ 255] blk.5.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q5_K .. size = 15.00 MiB -> 5.16 MiB
114
- [ 56/ 255] blk.5.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
115
- [ 57/ 255] blk.5.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q5_K .. size = 15.00 MiB -> 5.16 MiB
116
- [ 58/ 255] blk.6.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q5_K .. size = 1.50 MiB -> 0.52 MiB
117
- [ 59/ 255] blk.6.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
118
- [ 60/ 255] blk.6.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q5_K .. size = 6.00 MiB -> 2.06 MiB
119
- [ 61/ 255] blk.6.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q5_K .. size = 6.00 MiB -> 2.06 MiB
120
- [ 62/ 255] blk.6.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q5_K .. size = 1.50 MiB -> 0.52 MiB
121
- [ 63/ 255] blk.6.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q5_K .. size = 15.00 MiB -> 5.16 MiB
122
- [ 64/ 255] blk.6.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q5_K .. size = 15.00 MiB -> 5.16 MiB
123
- [ 65/ 255] blk.6.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
124
- [ 66/ 255] blk.6.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q5_K .. size = 15.00 MiB -> 5.16 MiB
125
- [ 67/ 255] blk.7.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q5_K .. size = 1.50 MiB -> 0.52 MiB
126
- [ 68/ 255] blk.7.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
127
- [ 69/ 255] blk.7.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q5_K .. size = 6.00 MiB -> 2.06 MiB
128
- [ 70/ 255] blk.7.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q5_K .. size = 6.00 MiB -> 2.06 MiB
129
- [ 71/ 255] blk.7.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q5_K .. size = 1.50 MiB -> 0.52 MiB
130
- [ 72/ 255] blk.7.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q5_K .. size = 15.00 MiB -> 5.16 MiB
131
- [ 73/ 255] blk.7.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q5_K .. size = 15.00 MiB -> 5.16 MiB
132
- [ 74/ 255] blk.7.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
133
- [ 75/ 255] blk.7.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q5_K .. size = 15.00 MiB -> 5.16 MiB
134
- [ 76/ 255] blk.8.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q5_K .. size = 1.50 MiB -> 0.52 MiB
135
- [ 77/ 255] blk.8.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
136
- [ 78/ 255] blk.8.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q5_K .. size = 6.00 MiB -> 2.06 MiB
137
- [ 79/ 255] blk.8.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q5_K .. size = 6.00 MiB -> 2.06 MiB
138
- [ 80/ 255] blk.8.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q6_K .. size = 1.50 MiB -> 0.62 MiB
139
- [ 81/ 255] blk.8.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
140
- [ 82/ 255] blk.8.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q5_K .. size = 15.00 MiB -> 5.16 MiB
141
- [ 83/ 255] blk.8.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
142
- [ 84/ 255] blk.8.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q5_K .. size = 15.00 MiB -> 5.16 MiB
143
- [ 85/ 255] blk.9.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q5_K .. size = 1.50 MiB -> 0.52 MiB
144
- [ 86/ 255] blk.9.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
145
- [ 87/ 255] blk.9.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q5_K .. size = 6.00 MiB -> 2.06 MiB
146
- [ 88/ 255] blk.9.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q5_K .. size = 6.00 MiB -> 2.06 MiB
147
- [ 89/ 255] blk.9.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q5_K .. size = 1.50 MiB -> 0.52 MiB
148
- [ 90/ 255] blk.9.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q5_K .. size = 15.00 MiB -> 5.16 MiB
149
- [ 91/ 255] blk.9.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q5_K .. size = 15.00 MiB -> 5.16 MiB
150
- [ 92/ 255] blk.9.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
151
- [ 93/ 255] blk.9.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q5_K .. size = 15.00 MiB -> 5.16 MiB
152
- [ 94/ 255] blk.10.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q5_K .. size = 1.50 MiB -> 0.52 MiB
153
- [ 95/ 255] blk.10.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
154
- [ 96/ 255] blk.10.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q5_K .. size = 6.00 MiB -> 2.06 MiB
155
- [ 97/ 255] blk.10.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q5_K .. size = 6.00 MiB -> 2.06 MiB
156
- [ 98/ 255] blk.10.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q5_K .. size = 1.50 MiB -> 0.52 MiB
157
- [ 99/ 255] blk.10.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q5_K .. size = 15.00 MiB -> 5.16 MiB
158
- [ 100/ 255] blk.10.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q5_K .. size = 15.00 MiB -> 5.16 MiB
159
- [ 101/ 255] blk.10.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
160
- [ 102/ 255] blk.10.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q5_K .. size = 15.00 MiB -> 5.16 MiB
161
- [ 103/ 255] blk.11.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q5_K .. size = 1.50 MiB -> 0.52 MiB
162
- [ 104/ 255] blk.11.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
163
- [ 105/ 255] blk.11.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q5_K .. size = 6.00 MiB -> 2.06 MiB
164
- [ 106/ 255] blk.11.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q5_K .. size = 6.00 MiB -> 2.06 MiB
165
- [ 107/ 255] blk.11.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q6_K .. size = 1.50 MiB -> 0.62 MiB
166
- [ 108/ 255] blk.11.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
167
- [ 109/ 255] blk.11.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q5_K .. size = 15.00 MiB -> 5.16 MiB
168
- [ 110/ 255] blk.11.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
169
- [ 111/ 255] blk.11.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q5_K .. size = 15.00 MiB -> 5.16 MiB
170
- [ 112/ 255] blk.12.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q5_K .. size = 1.50 MiB -> 0.52 MiB
171
- [ 113/ 255] blk.12.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
172
- [ 114/ 255] blk.12.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q5_K .. size = 6.00 MiB -> 2.06 MiB
173
- [ 115/ 255] blk.12.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q5_K .. size = 6.00 MiB -> 2.06 MiB
174
- [ 116/ 255] blk.12.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q5_K .. size = 1.50 MiB -> 0.52 MiB
175
- [ 117/ 255] blk.12.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q5_K .. size = 15.00 MiB -> 5.16 MiB
176
- [ 118/ 255] blk.12.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q5_K .. size = 15.00 MiB -> 5.16 MiB
177
- [ 119/ 255] blk.12.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
178
- [ 120/ 255] blk.12.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q5_K .. size = 15.00 MiB -> 5.16 MiB
179
- [ 121/ 255] blk.13.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q5_K .. size = 1.50 MiB -> 0.52 MiB
180
- [ 122/ 255] blk.13.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
181
- [ 123/ 255] blk.13.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q5_K .. size = 6.00 MiB -> 2.06 MiB
182
- [ 124/ 255] blk.13.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q5_K .. size = 6.00 MiB -> 2.06 MiB
183
- [ 125/ 255] blk.13.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q5_K .. size = 1.50 MiB -> 0.52 MiB
184
- [ 126/ 255] blk.13.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q5_K .. size = 15.00 MiB -> 5.16 MiB
185
- [ 127/ 255] blk.13.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q5_K .. size = 15.00 MiB -> 5.16 MiB
186
- [ 128/ 255] blk.13.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
187
- [ 129/ 255] blk.13.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q5_K .. size = 15.00 MiB -> 5.16 MiB
188
- [ 130/ 255] blk.14.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q5_K .. size = 1.50 MiB -> 0.52 MiB
189
- [ 131/ 255] blk.14.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
190
- [ 132/ 255] blk.14.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q5_K .. size = 6.00 MiB -> 2.06 MiB
191
- [ 133/ 255] blk.14.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q5_K .. size = 6.00 MiB -> 2.06 MiB
192
- [ 134/ 255] blk.14.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q6_K .. size = 1.50 MiB -> 0.62 MiB
193
- [ 135/ 255] blk.14.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
194
- [ 136/ 255] blk.14.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q5_K .. size = 15.00 MiB -> 5.16 MiB
195
- [ 137/ 255] blk.14.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
196
- [ 138/ 255] blk.14.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q5_K .. size = 15.00 MiB -> 5.16 MiB
197
- [ 139/ 255] blk.15.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q5_K .. size = 1.50 MiB -> 0.52 MiB
198
- [ 140/ 255] blk.15.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
199
- [ 141/ 255] blk.15.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q5_K .. size = 6.00 MiB -> 2.06 MiB
200
- [ 142/ 255] blk.15.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q5_K .. size = 6.00 MiB -> 2.06 MiB
201
- [ 143/ 255] blk.15.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q5_K .. size = 1.50 MiB -> 0.52 MiB
202
- [ 144/ 255] blk.15.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q5_K .. size = 15.00 MiB -> 5.16 MiB
203
- [ 145/ 255] blk.15.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q5_K .. size = 15.00 MiB -> 5.16 MiB
204
- [ 146/ 255] blk.15.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
205
- [ 147/ 255] blk.15.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q5_K .. size = 15.00 MiB -> 5.16 MiB
206
- [ 148/ 255] blk.16.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q5_K .. size = 1.50 MiB -> 0.52 MiB
207
- [ 149/ 255] blk.16.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
208
- [ 150/ 255] blk.16.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q5_K .. size = 6.00 MiB -> 2.06 MiB
209
- [ 151/ 255] blk.16.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q5_K .. size = 6.00 MiB -> 2.06 MiB
210
- [ 152/ 255] blk.16.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q5_K .. size = 1.50 MiB -> 0.52 MiB
211
- [ 153/ 255] blk.16.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q5_K .. size = 15.00 MiB -> 5.16 MiB
212
- [ 154/ 255] blk.16.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q5_K .. size = 15.00 MiB -> 5.16 MiB
213
- [ 155/ 255] blk.16.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
214
- [ 156/ 255] blk.16.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q5_K .. size = 15.00 MiB -> 5.16 MiB
215
- [ 157/ 255] blk.17.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q5_K .. size = 1.50 MiB -> 0.52 MiB
216
- [ 158/ 255] blk.17.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
217
- [ 159/ 255] blk.17.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q5_K .. size = 6.00 MiB -> 2.06 MiB
218
- [ 160/ 255] blk.17.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q5_K .. size = 6.00 MiB -> 2.06 MiB
219
- [ 161/ 255] blk.17.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q6_K .. size = 1.50 MiB -> 0.62 MiB
220
- [ 162/ 255] blk.17.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
221
- [ 163/ 255] blk.17.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q5_K .. size = 15.00 MiB -> 5.16 MiB
222
- [ 164/ 255] blk.17.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
223
- [ 165/ 255] blk.17.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q5_K .. size = 15.00 MiB -> 5.16 MiB
224
- [ 166/ 255] blk.18.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q5_K .. size = 1.50 MiB -> 0.52 MiB
225
- [ 167/ 255] blk.18.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
226
- [ 168/ 255] blk.18.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q5_K .. size = 6.00 MiB -> 2.06 MiB
227
- [ 169/ 255] blk.18.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q5_K .. size = 6.00 MiB -> 2.06 MiB
228
- [ 170/ 255] blk.18.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q5_K .. size = 1.50 MiB -> 0.52 MiB
229
- [ 171/ 255] blk.18.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q5_K .. size = 15.00 MiB -> 5.16 MiB
230
- [ 172/ 255] blk.18.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q5_K .. size = 15.00 MiB -> 5.16 MiB
231
- [ 173/ 255] blk.18.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
232
- [ 174/ 255] blk.18.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q5_K .. size = 15.00 MiB -> 5.16 MiB
233
- [ 175/ 255] blk.19.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q5_K .. size = 1.50 MiB -> 0.52 MiB
234
- [ 176/ 255] blk.19.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
235
- [ 177/ 255] blk.19.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q5_K .. size = 6.00 MiB -> 2.06 MiB
236
- [ 178/ 255] blk.19.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q5_K .. size = 6.00 MiB -> 2.06 MiB
237
- [ 179/ 255] blk.19.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q5_K .. size = 1.50 MiB -> 0.52 MiB
238
- [ 180/ 255] blk.19.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q5_K .. size = 15.00 MiB -> 5.16 MiB
239
- [ 181/ 255] blk.19.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q5_K .. size = 15.00 MiB -> 5.16 MiB
240
- [ 182/ 255] blk.19.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
241
- [ 183/ 255] blk.19.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q5_K .. size = 15.00 MiB -> 5.16 MiB
242
- [ 184/ 255] blk.20.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q5_K .. size = 1.50 MiB -> 0.52 MiB
243
- [ 185/ 255] blk.20.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
244
- [ 186/ 255] blk.20.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q5_K .. size = 6.00 MiB -> 2.06 MiB
245
- [ 187/ 255] blk.20.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q5_K .. size = 6.00 MiB -> 2.06 MiB
246
- [ 188/ 255] blk.20.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q6_K .. size = 1.50 MiB -> 0.62 MiB
247
- [ 189/ 255] blk.20.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
248
- [ 190/ 255] blk.20.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q5_K .. size = 15.00 MiB -> 5.16 MiB
249
- [ 191/ 255] blk.20.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
250
- [ 192/ 255] blk.20.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q5_K .. size = 15.00 MiB -> 5.16 MiB
251
- [ 193/ 255] blk.21.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q5_K .. size = 1.50 MiB -> 0.52 MiB
252
- [ 194/ 255] blk.21.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
253
- [ 195/ 255] blk.21.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q5_K .. size = 6.00 MiB -> 2.06 MiB
254
- [ 196/ 255] blk.21.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q5_K .. size = 6.00 MiB -> 2.06 MiB
255
- [ 197/ 255] blk.21.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q5_K .. size = 1.50 MiB -> 0.52 MiB
256
- [ 198/ 255] blk.21.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q5_K .. size = 15.00 MiB -> 5.16 MiB
257
- [ 199/ 255] blk.21.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q5_K .. size = 15.00 MiB -> 5.16 MiB
258
- [ 200/ 255] blk.21.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
259
- [ 201/ 255] blk.21.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q5_K .. size = 15.00 MiB -> 5.16 MiB
260
- [ 202/ 255] blk.22.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q5_K .. size = 1.50 MiB -> 0.52 MiB
261
- [ 203/ 255] blk.22.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
262
- [ 204/ 255] blk.22.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q5_K .. size = 6.00 MiB -> 2.06 MiB
263
- [ 205/ 255] blk.22.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q5_K .. size = 6.00 MiB -> 2.06 MiB
264
- [ 206/ 255] blk.22.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q5_K .. size = 1.50 MiB -> 0.52 MiB
265
- [ 207/ 255] blk.22.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q5_K .. size = 15.00 MiB -> 5.16 MiB
266
- [ 208/ 255] blk.22.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q5_K .. size = 15.00 MiB -> 5.16 MiB
267
- [ 209/ 255] blk.22.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
268
- [ 210/ 255] blk.22.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q5_K .. size = 15.00 MiB -> 5.16 MiB
269
- [ 211/ 255] blk.23.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q5_K .. size = 1.50 MiB -> 0.52 MiB
270
- [ 212/ 255] blk.23.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
271
- [ 213/ 255] blk.23.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q5_K .. size = 6.00 MiB -> 2.06 MiB
272
- [ 214/ 255] blk.23.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q5_K .. size = 6.00 MiB -> 2.06 MiB
273
- [ 215/ 255] blk.23.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q6_K .. size = 1.50 MiB -> 0.62 MiB
274
- [ 216/ 255] blk.23.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
275
- [ 217/ 255] blk.23.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q5_K .. size = 15.00 MiB -> 5.16 MiB
276
- [ 218/ 255] blk.23.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
277
- [ 219/ 255] blk.23.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q5_K .. size = 15.00 MiB -> 5.16 MiB
278
- [ 220/ 255] blk.24.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q5_K .. size = 1.50 MiB -> 0.52 MiB
279
- [ 221/ 255] blk.24.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
280
- [ 222/ 255] blk.24.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q5_K .. size = 6.00 MiB -> 2.06 MiB
281
- [ 223/ 255] blk.24.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q5_K .. size = 6.00 MiB -> 2.06 MiB
282
- [ 224/ 255] blk.24.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q6_K .. size = 1.50 MiB -> 0.62 MiB
283
- [ 225/ 255] blk.24.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
284
- [ 226/ 255] blk.24.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q5_K .. size = 15.00 MiB -> 5.16 MiB
285
- [ 227/ 255] blk.24.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
286
- [ 228/ 255] blk.24.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q5_K .. size = 15.00 MiB -> 5.16 MiB
287
- [ 229/ 255] blk.25.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q5_K .. size = 1.50 MiB -> 0.52 MiB
288
- [ 230/ 255] blk.25.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
289
- [ 231/ 255] blk.25.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q5_K .. size = 6.00 MiB -> 2.06 MiB
290
- [ 232/ 255] blk.25.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q5_K .. size = 6.00 MiB -> 2.06 MiB
291
- [ 233/ 255] blk.25.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q6_K .. size = 1.50 MiB -> 0.62 MiB
292
- [ 234/ 255] blk.25.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
293
- [ 235/ 255] blk.25.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q5_K .. size = 15.00 MiB -> 5.16 MiB
294
- [ 236/ 255] blk.25.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
295
- [ 237/ 255] blk.25.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q5_K .. size = 15.00 MiB -> 5.16 MiB
296
- [ 238/ 255] blk.26.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q5_K .. size = 1.50 MiB -> 0.52 MiB
297
- [ 239/ 255] blk.26.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
298
- [ 240/ 255] blk.26.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q5_K .. size = 6.00 MiB -> 2.06 MiB
299
- [ 241/ 255] blk.26.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q5_K .. size = 6.00 MiB -> 2.06 MiB
300
- [ 242/ 255] blk.26.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q6_K .. size = 1.50 MiB -> 0.62 MiB
301
- [ 243/ 255] blk.26.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
302
- [ 244/ 255] blk.26.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q5_K .. size = 15.00 MiB -> 5.16 MiB
303
- [ 245/ 255] blk.26.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
304
- [ 246/ 255] blk.26.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q5_K .. size = 15.00 MiB -> 5.16 MiB
305
- [ 247/ 255] blk.27.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q5_K .. size = 1.50 MiB -> 0.52 MiB
306
- [ 248/ 255] blk.27.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
307
- [ 249/ 255] blk.27.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q5_K .. size = 6.00 MiB -> 2.06 MiB
308
- [ 250/ 255] blk.27.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q5_K .. size = 6.00 MiB -> 2.06 MiB
309
- [ 251/ 255] blk.27.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q6_K .. size = 1.50 MiB -> 0.62 MiB
310
- [ 252/ 255] blk.27.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
311
- [ 253/ 255] blk.27.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q5_K .. size = 15.00 MiB -> 5.16 MiB
312
- [ 254/ 255] blk.27.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
313
- [ 255/ 255] blk.27.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q5_K .. size = 15.00 MiB -> 5.16 MiB
314
- llama_model_quantize_impl: model size = 2056.83 MiB (16.00 BPW)
315
- llama_model_quantize_impl: quant size = 735.10 MiB (5.72 BPW)
316
-
317
- llama_quantize: quantize time = 13244.49 ms
318
- llama_quantize: total time = 13244.49 ms
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
reproducibility/validation/quantize-logs/Q6_K.log DELETED
@@ -1,318 +0,0 @@
1
- ggml_cuda_init: failed to initialize CUDA: no CUDA-capable device is detected
2
- version: 0.3.0-dev (build 10671, commit 35999d101)
3
- built with GNU 13.3.0 for Linux x86_64
4
- llama_quantize: quantizing '/root/workspace/HF/source-official-bf16-k2-horizon-0.9b/K2-Horizon-0.9B-BF16.gguf' to '/root/workspace/HF/quantized-k2-horizon-0.9b/K2-Horizon-0.9B-Q6_K.gguf' as Q6_K using 128 threads
5
- llama_model_loader: loaded meta data with 41 key-value pairs and 255 tensors from /root/workspace/HF/source-official-bf16-k2-horizon-0.9b/K2-Horizon-0.9B-BF16.gguf (version GGUF V3 (latest))
6
- llama_model_loader: Dumping metadata keys/values. Note: KV overrides do not apply in this output.
7
- llama_model_loader: - kv 0: general.architecture str = k2-horizon
8
- llama_model_loader: - kv 1: general.type str = model
9
- llama_model_loader: - kv 2: general.name str = K2-Horizon-0.9B
10
- llama_model_loader: - kv 3: general.basename str = source-hf-k2-horizon
11
- llama_model_loader: - kv 4: general.size_label str = 0.9B
12
- llama_model_loader: - kv 5: general.license str = apache-2.0
13
- llama_model_loader: - kv 6: general.license.name str = internal-only
14
- llama_model_loader: - kv 7: general.license.link str = LICENSE
15
- llama_model_loader: - kv 8: general.tags arr[str,7] = ["k2-horizon", "0.9b", "dense", "reas...
16
- llama_model_loader: - kv 9: general.languages arr[str,2] = ["en", "zh"]
17
- llama_model_loader: - kv 10: k2-horizon.block_count u32 = 28
18
- llama_model_loader: - kv 11: k2-horizon.context_length u32 = 131072
19
- llama_model_loader: - kv 12: k2-horizon.embedding_length u32 = 1536
20
- llama_model_loader: - kv 13: k2-horizon.feed_forward_length u32 = 5120
21
- llama_model_loader: - kv 14: k2-horizon.attention.head_count u32 = 32
22
- llama_model_loader: - kv 15: k2-horizon.attention.head_count_kv u32 = 8
23
- llama_model_loader: - kv 16: k2-horizon.rope.scaling.type str = yarn
24
- llama_model_loader: - kv 17: k2-horizon.rope.scaling.factor f32 = 16.000000
25
- llama_model_loader: - kv 18: k2-horizon.rope.scaling.original_context_length u32 = 8192
26
- llama_model_loader: - kv 19: k2-horizon.rope.scaling.yarn_attn_factor f32 = 1.277259
27
- llama_model_loader: - kv 20: k2-horizon.rope.scaling.yarn_beta_fast f32 = 128.000000
28
- llama_model_loader: - kv 21: k2-horizon.rope.scaling.yarn_beta_slow f32 = 4.000000
29
- llama_model_loader: - kv 22: k2-horizon.rope.freq_base f32 = 1000000.000000
30
- llama_model_loader: - kv 23: k2-horizon.attention.layer_norm_rms_epsilon f32 = 0.000001
31
- llama_model_loader: - kv 24: k2-horizon.expert_count u32 = 0
32
- llama_model_loader: - kv 25: k2-horizon.expert_used_count u32 = 0
33
- llama_model_loader: - kv 26: k2-horizon.attention.key_length u32 = 64
34
- llama_model_loader: - kv 27: k2-horizon.attention.value_length u32 = 64
35
- llama_model_loader: - kv 28: general.file_type u32 = 32
36
- llama_model_loader: - kv 29: k2-horizon.attention.group_norm_groups u32 = 1
37
- llama_model_loader: - kv 30: k2-horizon.rope.dimension_count u32 = 64
38
- llama_model_loader: - kv 31: general.quantization_version u32 = 2
39
- llama_model_loader: - kv 32: tokenizer.ggml.model str = gpt2
40
- llama_model_loader: - kv 33: tokenizer.ggml.pre str = k2-horizon
41
- llama_model_loader: - kv 34: tokenizer.ggml.tokens arr[str,64256] = ["<|begin_of_text|>", "<|endoftext|>"...
42
- llama_model_loader: - kv 35: tokenizer.ggml.token_type arr[i32,64256] = [3, 3, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, ...
43
- llama_model_loader: - kv 36: tokenizer.ggml.merges arr[str,63742] = ["Δ  Δ ", "Ø Β§", "Γ™ Δ¦", "Γ  Β€", ...
44
- llama_model_loader: - kv 37: tokenizer.ggml.bos_token_id u32 = 0
45
- llama_model_loader: - kv 38: tokenizer.ggml.eos_token_id u32 = 1
46
- llama_model_loader: - kv 39: tokenizer.ggml.padding_token_id u32 = 64255
47
- llama_model_loader: - kv 40: tokenizer.chat_template str = {{- bos_token }}\n{%- if tool_presenta...
48
- llama_model_loader: - type f32: 57 tensors
49
- llama_model_loader: - type bf16: 198 tensors
50
-
51
- llama_model_quantize_impl: have importance matrix data with 196 entries
52
- [ 1/ 255] output.weight - [ 1536, 64256, 1, 1], type = bf16,
53
- ====== llama_model_quantize_impl: did not find weights for output.weight
54
- converting to q6_K .. load_imatrix: imatrix datasets=['calibration/wikitext-2-raw/wiki.train.raw', 'calibration/k2-horizon-0.9b/k2_horizon_en_zh_code_tool.txt']
55
- load_imatrix: loaded 196 importance matrix entries from /root/workspace/HF/calibration/k2-horizon-0.9b/k2_horizon_combined.imatrix.gguf computed on 99 chunks
56
- prepare_imatrix: have 196 importance matrix entries
57
- size = 188.25 MiB -> 77.21 MiB
58
- [ 2/ 255] output_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
59
- [ 3/ 255] token_embd.weight - [ 1536, 64256, 1, 1], type = bf16,
60
- ====== llama_model_quantize_impl: did not find weights for token_embd.weight
61
- converting to q6_K .. size = 188.25 MiB -> 77.21 MiB
62
- [ 4/ 255] blk.0.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q6_K .. size = 1.50 MiB -> 0.62 MiB
63
- [ 5/ 255] blk.0.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
64
- [ 6/ 255] blk.0.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q6_K .. size = 6.00 MiB -> 2.46 MiB
65
- [ 7/ 255] blk.0.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q6_K .. size = 6.00 MiB -> 2.46 MiB
66
- [ 8/ 255] blk.0.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q6_K .. size = 1.50 MiB -> 0.62 MiB
67
- [ 9/ 255] blk.0.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
68
- [ 10/ 255] blk.0.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
69
- [ 11/ 255] blk.0.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
70
- [ 12/ 255] blk.0.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
71
- [ 13/ 255] blk.1.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q6_K .. size = 1.50 MiB -> 0.62 MiB
72
- [ 14/ 255] blk.1.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
73
- [ 15/ 255] blk.1.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q6_K .. size = 6.00 MiB -> 2.46 MiB
74
- [ 16/ 255] blk.1.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q6_K .. size = 6.00 MiB -> 2.46 MiB
75
- [ 17/ 255] blk.1.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q6_K .. size = 1.50 MiB -> 0.62 MiB
76
- [ 18/ 255] blk.1.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
77
- [ 19/ 255] blk.1.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
78
- [ 20/ 255] blk.1.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
79
- [ 21/ 255] blk.1.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
80
- [ 22/ 255] blk.2.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q6_K .. size = 1.50 MiB -> 0.62 MiB
81
- [ 23/ 255] blk.2.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
82
- [ 24/ 255] blk.2.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q6_K .. size = 6.00 MiB -> 2.46 MiB
83
- [ 25/ 255] blk.2.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q6_K .. size = 6.00 MiB -> 2.46 MiB
84
- [ 26/ 255] blk.2.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q6_K .. size = 1.50 MiB -> 0.62 MiB
85
- [ 27/ 255] blk.2.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
86
- [ 28/ 255] blk.2.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
87
- [ 29/ 255] blk.2.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
88
- [ 30/ 255] blk.2.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
89
- [ 31/ 255] blk.3.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q6_K .. size = 1.50 MiB -> 0.62 MiB
90
- [ 32/ 255] blk.3.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
91
- [ 33/ 255] blk.3.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q6_K .. size = 6.00 MiB -> 2.46 MiB
92
- [ 34/ 255] blk.3.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q6_K .. size = 6.00 MiB -> 2.46 MiB
93
- [ 35/ 255] blk.3.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q6_K .. size = 1.50 MiB -> 0.62 MiB
94
- [ 36/ 255] blk.3.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
95
- [ 37/ 255] blk.3.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
96
- [ 38/ 255] blk.3.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
97
- [ 39/ 255] blk.3.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
98
- [ 40/ 255] blk.4.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q6_K .. size = 1.50 MiB -> 0.62 MiB
99
- [ 41/ 255] blk.4.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
100
- [ 42/ 255] blk.4.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q6_K .. size = 6.00 MiB -> 2.46 MiB
101
- [ 43/ 255] blk.4.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q6_K .. size = 6.00 MiB -> 2.46 MiB
102
- [ 44/ 255] blk.4.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q6_K .. size = 1.50 MiB -> 0.62 MiB
103
- [ 45/ 255] blk.4.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
104
- [ 46/ 255] blk.4.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
105
- [ 47/ 255] blk.4.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
106
- [ 48/ 255] blk.4.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
107
- [ 49/ 255] blk.5.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q6_K .. size = 1.50 MiB -> 0.62 MiB
108
- [ 50/ 255] blk.5.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
109
- [ 51/ 255] blk.5.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q6_K .. size = 6.00 MiB -> 2.46 MiB
110
- [ 52/ 255] blk.5.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q6_K .. size = 6.00 MiB -> 2.46 MiB
111
- [ 53/ 255] blk.5.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q6_K .. size = 1.50 MiB -> 0.62 MiB
112
- [ 54/ 255] blk.5.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
113
- [ 55/ 255] blk.5.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
114
- [ 56/ 255] blk.5.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
115
- [ 57/ 255] blk.5.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
116
- [ 58/ 255] blk.6.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q6_K .. size = 1.50 MiB -> 0.62 MiB
117
- [ 59/ 255] blk.6.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
118
- [ 60/ 255] blk.6.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q6_K .. size = 6.00 MiB -> 2.46 MiB
119
- [ 61/ 255] blk.6.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q6_K .. size = 6.00 MiB -> 2.46 MiB
120
- [ 62/ 255] blk.6.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q6_K .. size = 1.50 MiB -> 0.62 MiB
121
- [ 63/ 255] blk.6.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
122
- [ 64/ 255] blk.6.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
123
- [ 65/ 255] blk.6.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
124
- [ 66/ 255] blk.6.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
125
- [ 67/ 255] blk.7.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q6_K .. size = 1.50 MiB -> 0.62 MiB
126
- [ 68/ 255] blk.7.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
127
- [ 69/ 255] blk.7.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q6_K .. size = 6.00 MiB -> 2.46 MiB
128
- [ 70/ 255] blk.7.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q6_K .. size = 6.00 MiB -> 2.46 MiB
129
- [ 71/ 255] blk.7.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q6_K .. size = 1.50 MiB -> 0.62 MiB
130
- [ 72/ 255] blk.7.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
131
- [ 73/ 255] blk.7.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
132
- [ 74/ 255] blk.7.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
133
- [ 75/ 255] blk.7.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
134
- [ 76/ 255] blk.8.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q6_K .. size = 1.50 MiB -> 0.62 MiB
135
- [ 77/ 255] blk.8.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
136
- [ 78/ 255] blk.8.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q6_K .. size = 6.00 MiB -> 2.46 MiB
137
- [ 79/ 255] blk.8.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q6_K .. size = 6.00 MiB -> 2.46 MiB
138
- [ 80/ 255] blk.8.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q6_K .. size = 1.50 MiB -> 0.62 MiB
139
- [ 81/ 255] blk.8.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
140
- [ 82/ 255] blk.8.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
141
- [ 83/ 255] blk.8.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
142
- [ 84/ 255] blk.8.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
143
- [ 85/ 255] blk.9.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q6_K .. size = 1.50 MiB -> 0.62 MiB
144
- [ 86/ 255] blk.9.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
145
- [ 87/ 255] blk.9.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q6_K .. size = 6.00 MiB -> 2.46 MiB
146
- [ 88/ 255] blk.9.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q6_K .. size = 6.00 MiB -> 2.46 MiB
147
- [ 89/ 255] blk.9.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q6_K .. size = 1.50 MiB -> 0.62 MiB
148
- [ 90/ 255] blk.9.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
149
- [ 91/ 255] blk.9.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
150
- [ 92/ 255] blk.9.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
151
- [ 93/ 255] blk.9.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
152
- [ 94/ 255] blk.10.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q6_K .. size = 1.50 MiB -> 0.62 MiB
153
- [ 95/ 255] blk.10.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
154
- [ 96/ 255] blk.10.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q6_K .. size = 6.00 MiB -> 2.46 MiB
155
- [ 97/ 255] blk.10.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q6_K .. size = 6.00 MiB -> 2.46 MiB
156
- [ 98/ 255] blk.10.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q6_K .. size = 1.50 MiB -> 0.62 MiB
157
- [ 99/ 255] blk.10.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
158
- [ 100/ 255] blk.10.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
159
- [ 101/ 255] blk.10.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
160
- [ 102/ 255] blk.10.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
161
- [ 103/ 255] blk.11.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q6_K .. size = 1.50 MiB -> 0.62 MiB
162
- [ 104/ 255] blk.11.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
163
- [ 105/ 255] blk.11.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q6_K .. size = 6.00 MiB -> 2.46 MiB
164
- [ 106/ 255] blk.11.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q6_K .. size = 6.00 MiB -> 2.46 MiB
165
- [ 107/ 255] blk.11.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q6_K .. size = 1.50 MiB -> 0.62 MiB
166
- [ 108/ 255] blk.11.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
167
- [ 109/ 255] blk.11.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
168
- [ 110/ 255] blk.11.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
169
- [ 111/ 255] blk.11.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
170
- [ 112/ 255] blk.12.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q6_K .. size = 1.50 MiB -> 0.62 MiB
171
- [ 113/ 255] blk.12.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
172
- [ 114/ 255] blk.12.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q6_K .. size = 6.00 MiB -> 2.46 MiB
173
- [ 115/ 255] blk.12.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q6_K .. size = 6.00 MiB -> 2.46 MiB
174
- [ 116/ 255] blk.12.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q6_K .. size = 1.50 MiB -> 0.62 MiB
175
- [ 117/ 255] blk.12.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
176
- [ 118/ 255] blk.12.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
177
- [ 119/ 255] blk.12.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
178
- [ 120/ 255] blk.12.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
179
- [ 121/ 255] blk.13.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q6_K .. size = 1.50 MiB -> 0.62 MiB
180
- [ 122/ 255] blk.13.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
181
- [ 123/ 255] blk.13.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q6_K .. size = 6.00 MiB -> 2.46 MiB
182
- [ 124/ 255] blk.13.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q6_K .. size = 6.00 MiB -> 2.46 MiB
183
- [ 125/ 255] blk.13.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q6_K .. size = 1.50 MiB -> 0.62 MiB
184
- [ 126/ 255] blk.13.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
185
- [ 127/ 255] blk.13.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
186
- [ 128/ 255] blk.13.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
187
- [ 129/ 255] blk.13.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
188
- [ 130/ 255] blk.14.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q6_K .. size = 1.50 MiB -> 0.62 MiB
189
- [ 131/ 255] blk.14.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
190
- [ 132/ 255] blk.14.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q6_K .. size = 6.00 MiB -> 2.46 MiB
191
- [ 133/ 255] blk.14.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q6_K .. size = 6.00 MiB -> 2.46 MiB
192
- [ 134/ 255] blk.14.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q6_K .. size = 1.50 MiB -> 0.62 MiB
193
- [ 135/ 255] blk.14.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
194
- [ 136/ 255] blk.14.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
195
- [ 137/ 255] blk.14.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
196
- [ 138/ 255] blk.14.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
197
- [ 139/ 255] blk.15.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q6_K .. size = 1.50 MiB -> 0.62 MiB
198
- [ 140/ 255] blk.15.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
199
- [ 141/ 255] blk.15.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q6_K .. size = 6.00 MiB -> 2.46 MiB
200
- [ 142/ 255] blk.15.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q6_K .. size = 6.00 MiB -> 2.46 MiB
201
- [ 143/ 255] blk.15.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q6_K .. size = 1.50 MiB -> 0.62 MiB
202
- [ 144/ 255] blk.15.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
203
- [ 145/ 255] blk.15.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
204
- [ 146/ 255] blk.15.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
205
- [ 147/ 255] blk.15.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
206
- [ 148/ 255] blk.16.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q6_K .. size = 1.50 MiB -> 0.62 MiB
207
- [ 149/ 255] blk.16.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
208
- [ 150/ 255] blk.16.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q6_K .. size = 6.00 MiB -> 2.46 MiB
209
- [ 151/ 255] blk.16.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q6_K .. size = 6.00 MiB -> 2.46 MiB
210
- [ 152/ 255] blk.16.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q6_K .. size = 1.50 MiB -> 0.62 MiB
211
- [ 153/ 255] blk.16.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
212
- [ 154/ 255] blk.16.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
213
- [ 155/ 255] blk.16.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
214
- [ 156/ 255] blk.16.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
215
- [ 157/ 255] blk.17.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q6_K .. size = 1.50 MiB -> 0.62 MiB
216
- [ 158/ 255] blk.17.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
217
- [ 159/ 255] blk.17.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q6_K .. size = 6.00 MiB -> 2.46 MiB
218
- [ 160/ 255] blk.17.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q6_K .. size = 6.00 MiB -> 2.46 MiB
219
- [ 161/ 255] blk.17.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q6_K .. size = 1.50 MiB -> 0.62 MiB
220
- [ 162/ 255] blk.17.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
221
- [ 163/ 255] blk.17.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
222
- [ 164/ 255] blk.17.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
223
- [ 165/ 255] blk.17.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
224
- [ 166/ 255] blk.18.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q6_K .. size = 1.50 MiB -> 0.62 MiB
225
- [ 167/ 255] blk.18.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
226
- [ 168/ 255] blk.18.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q6_K .. size = 6.00 MiB -> 2.46 MiB
227
- [ 169/ 255] blk.18.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q6_K .. size = 6.00 MiB -> 2.46 MiB
228
- [ 170/ 255] blk.18.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q6_K .. size = 1.50 MiB -> 0.62 MiB
229
- [ 171/ 255] blk.18.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
230
- [ 172/ 255] blk.18.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
231
- [ 173/ 255] blk.18.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
232
- [ 174/ 255] blk.18.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
233
- [ 175/ 255] blk.19.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q6_K .. size = 1.50 MiB -> 0.62 MiB
234
- [ 176/ 255] blk.19.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
235
- [ 177/ 255] blk.19.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q6_K .. size = 6.00 MiB -> 2.46 MiB
236
- [ 178/ 255] blk.19.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q6_K .. size = 6.00 MiB -> 2.46 MiB
237
- [ 179/ 255] blk.19.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q6_K .. size = 1.50 MiB -> 0.62 MiB
238
- [ 180/ 255] blk.19.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
239
- [ 181/ 255] blk.19.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
240
- [ 182/ 255] blk.19.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
241
- [ 183/ 255] blk.19.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
242
- [ 184/ 255] blk.20.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q6_K .. size = 1.50 MiB -> 0.62 MiB
243
- [ 185/ 255] blk.20.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
244
- [ 186/ 255] blk.20.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q6_K .. size = 6.00 MiB -> 2.46 MiB
245
- [ 187/ 255] blk.20.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q6_K .. size = 6.00 MiB -> 2.46 MiB
246
- [ 188/ 255] blk.20.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q6_K .. size = 1.50 MiB -> 0.62 MiB
247
- [ 189/ 255] blk.20.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
248
- [ 190/ 255] blk.20.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
249
- [ 191/ 255] blk.20.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
250
- [ 192/ 255] blk.20.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
251
- [ 193/ 255] blk.21.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q6_K .. size = 1.50 MiB -> 0.62 MiB
252
- [ 194/ 255] blk.21.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
253
- [ 195/ 255] blk.21.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q6_K .. size = 6.00 MiB -> 2.46 MiB
254
- [ 196/ 255] blk.21.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q6_K .. size = 6.00 MiB -> 2.46 MiB
255
- [ 197/ 255] blk.21.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q6_K .. size = 1.50 MiB -> 0.62 MiB
256
- [ 198/ 255] blk.21.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
257
- [ 199/ 255] blk.21.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
258
- [ 200/ 255] blk.21.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
259
- [ 201/ 255] blk.21.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
260
- [ 202/ 255] blk.22.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q6_K .. size = 1.50 MiB -> 0.62 MiB
261
- [ 203/ 255] blk.22.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
262
- [ 204/ 255] blk.22.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q6_K .. size = 6.00 MiB -> 2.46 MiB
263
- [ 205/ 255] blk.22.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q6_K .. size = 6.00 MiB -> 2.46 MiB
264
- [ 206/ 255] blk.22.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q6_K .. size = 1.50 MiB -> 0.62 MiB
265
- [ 207/ 255] blk.22.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
266
- [ 208/ 255] blk.22.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
267
- [ 209/ 255] blk.22.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
268
- [ 210/ 255] blk.22.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
269
- [ 211/ 255] blk.23.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q6_K .. size = 1.50 MiB -> 0.62 MiB
270
- [ 212/ 255] blk.23.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
271
- [ 213/ 255] blk.23.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q6_K .. size = 6.00 MiB -> 2.46 MiB
272
- [ 214/ 255] blk.23.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q6_K .. size = 6.00 MiB -> 2.46 MiB
273
- [ 215/ 255] blk.23.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q6_K .. size = 1.50 MiB -> 0.62 MiB
274
- [ 216/ 255] blk.23.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
275
- [ 217/ 255] blk.23.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
276
- [ 218/ 255] blk.23.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
277
- [ 219/ 255] blk.23.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
278
- [ 220/ 255] blk.24.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q6_K .. size = 1.50 MiB -> 0.62 MiB
279
- [ 221/ 255] blk.24.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
280
- [ 222/ 255] blk.24.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q6_K .. size = 6.00 MiB -> 2.46 MiB
281
- [ 223/ 255] blk.24.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q6_K .. size = 6.00 MiB -> 2.46 MiB
282
- [ 224/ 255] blk.24.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q6_K .. size = 1.50 MiB -> 0.62 MiB
283
- [ 225/ 255] blk.24.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
284
- [ 226/ 255] blk.24.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
285
- [ 227/ 255] blk.24.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
286
- [ 228/ 255] blk.24.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
287
- [ 229/ 255] blk.25.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q6_K .. size = 1.50 MiB -> 0.62 MiB
288
- [ 230/ 255] blk.25.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
289
- [ 231/ 255] blk.25.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q6_K .. size = 6.00 MiB -> 2.46 MiB
290
- [ 232/ 255] blk.25.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q6_K .. size = 6.00 MiB -> 2.46 MiB
291
- [ 233/ 255] blk.25.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q6_K .. size = 1.50 MiB -> 0.62 MiB
292
- [ 234/ 255] blk.25.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
293
- [ 235/ 255] blk.25.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
294
- [ 236/ 255] blk.25.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
295
- [ 237/ 255] blk.25.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
296
- [ 238/ 255] blk.26.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q6_K .. size = 1.50 MiB -> 0.62 MiB
297
- [ 239/ 255] blk.26.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
298
- [ 240/ 255] blk.26.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q6_K .. size = 6.00 MiB -> 2.46 MiB
299
- [ 241/ 255] blk.26.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q6_K .. size = 6.00 MiB -> 2.46 MiB
300
- [ 242/ 255] blk.26.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q6_K .. size = 1.50 MiB -> 0.62 MiB
301
- [ 243/ 255] blk.26.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
302
- [ 244/ 255] blk.26.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
303
- [ 245/ 255] blk.26.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
304
- [ 246/ 255] blk.26.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
305
- [ 247/ 255] blk.27.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q6_K .. size = 1.50 MiB -> 0.62 MiB
306
- [ 248/ 255] blk.27.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
307
- [ 249/ 255] blk.27.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q6_K .. size = 6.00 MiB -> 2.46 MiB
308
- [ 250/ 255] blk.27.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q6_K .. size = 6.00 MiB -> 2.46 MiB
309
- [ 251/ 255] blk.27.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q6_K .. size = 1.50 MiB -> 0.62 MiB
310
- [ 252/ 255] blk.27.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
311
- [ 253/ 255] blk.27.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
312
- [ 254/ 255] blk.27.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
313
- [ 255/ 255] blk.27.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q6_K .. size = 15.00 MiB -> 6.15 MiB
314
- llama_model_quantize_impl: model size = 2056.83 MiB (16.00 BPW)
315
- llama_model_quantize_impl: quant size = 843.82 MiB (6.56 BPW)
316
-
317
- llama_quantize: quantize time = 11849.92 ms
318
- llama_quantize: total time = 11849.92 ms
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
reproducibility/validation/quantize-logs/Q8_0.log DELETED
@@ -1,309 +0,0 @@
1
- ggml_cuda_init: failed to initialize CUDA: no CUDA-capable device is detected
2
- version: 0.3.0-dev (build 10671, commit 35999d101)
3
- built with GNU 13.3.0 for Linux x86_64
4
- llama_quantize: quantizing '/root/workspace/HF/source-official-bf16-k2-horizon-0.9b/K2-Horizon-0.9B-BF16.gguf' to '/root/workspace/HF/quantized-k2-horizon-0.9b/K2-Horizon-0.9B-Q8_0.gguf' as Q8_0 using 128 threads
5
- llama_model_loader: loaded meta data with 41 key-value pairs and 255 tensors from /root/workspace/HF/source-official-bf16-k2-horizon-0.9b/K2-Horizon-0.9B-BF16.gguf (version GGUF V3 (latest))
6
- llama_model_loader: Dumping metadata keys/values. Note: KV overrides do not apply in this output.
7
- llama_model_loader: - kv 0: general.architecture str = k2-horizon
8
- llama_model_loader: - kv 1: general.type str = model
9
- llama_model_loader: - kv 2: general.name str = K2-Horizon-0.9B
10
- llama_model_loader: - kv 3: general.basename str = source-hf-k2-horizon
11
- llama_model_loader: - kv 4: general.size_label str = 0.9B
12
- llama_model_loader: - kv 5: general.license str = apache-2.0
13
- llama_model_loader: - kv 6: general.license.name str = internal-only
14
- llama_model_loader: - kv 7: general.license.link str = LICENSE
15
- llama_model_loader: - kv 8: general.tags arr[str,7] = ["k2-horizon", "0.9b", "dense", "reas...
16
- llama_model_loader: - kv 9: general.languages arr[str,2] = ["en", "zh"]
17
- llama_model_loader: - kv 10: k2-horizon.block_count u32 = 28
18
- llama_model_loader: - kv 11: k2-horizon.context_length u32 = 131072
19
- llama_model_loader: - kv 12: k2-horizon.embedding_length u32 = 1536
20
- llama_model_loader: - kv 13: k2-horizon.feed_forward_length u32 = 5120
21
- llama_model_loader: - kv 14: k2-horizon.attention.head_count u32 = 32
22
- llama_model_loader: - kv 15: k2-horizon.attention.head_count_kv u32 = 8
23
- llama_model_loader: - kv 16: k2-horizon.rope.scaling.type str = yarn
24
- llama_model_loader: - kv 17: k2-horizon.rope.scaling.factor f32 = 16.000000
25
- llama_model_loader: - kv 18: k2-horizon.rope.scaling.original_context_length u32 = 8192
26
- llama_model_loader: - kv 19: k2-horizon.rope.scaling.yarn_attn_factor f32 = 1.277259
27
- llama_model_loader: - kv 20: k2-horizon.rope.scaling.yarn_beta_fast f32 = 128.000000
28
- llama_model_loader: - kv 21: k2-horizon.rope.scaling.yarn_beta_slow f32 = 4.000000
29
- llama_model_loader: - kv 22: k2-horizon.rope.freq_base f32 = 1000000.000000
30
- llama_model_loader: - kv 23: k2-horizon.attention.layer_norm_rms_epsilon f32 = 0.000001
31
- llama_model_loader: - kv 24: k2-horizon.expert_count u32 = 0
32
- llama_model_loader: - kv 25: k2-horizon.expert_used_count u32 = 0
33
- llama_model_loader: - kv 26: k2-horizon.attention.key_length u32 = 64
34
- llama_model_loader: - kv 27: k2-horizon.attention.value_length u32 = 64
35
- llama_model_loader: - kv 28: general.file_type u32 = 32
36
- llama_model_loader: - kv 29: k2-horizon.attention.group_norm_groups u32 = 1
37
- llama_model_loader: - kv 30: k2-horizon.rope.dimension_count u32 = 64
38
- llama_model_loader: - kv 31: general.quantization_version u32 = 2
39
- llama_model_loader: - kv 32: tokenizer.ggml.model str = gpt2
40
- llama_model_loader: - kv 33: tokenizer.ggml.pre str = k2-horizon
41
- llama_model_loader: - kv 34: tokenizer.ggml.tokens arr[str,64256] = ["<|begin_of_text|>", "<|endoftext|>"...
42
- llama_model_loader: - kv 35: tokenizer.ggml.token_type arr[i32,64256] = [3, 3, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, ...
43
- llama_model_loader: - kv 36: tokenizer.ggml.merges arr[str,63742] = ["Δ  Δ ", "Ø Β§", "Γ™ Δ¦", "Γ  Β€", ...
44
- llama_model_loader: - kv 37: tokenizer.ggml.bos_token_id u32 = 0
45
- llama_model_loader: - kv 38: tokenizer.ggml.eos_token_id u32 = 1
46
- llama_model_loader: - kv 39: tokenizer.ggml.padding_token_id u32 = 64255
47
- llama_model_loader: - kv 40: tokenizer.chat_template str = {{- bos_token }}\n{%- if tool_presenta...
48
- llama_model_loader: - type f32: 57 tensors
49
- llama_model_loader: - type bf16: 198 tensors
50
- [ 1/ 255] output.weight - [ 1536, 64256, 1, 1], type = bf16, converting to q8_0 .. size = 188.25 MiB -> 100.01 MiB
51
- [ 2/ 255] output_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
52
- [ 3/ 255] token_embd.weight - [ 1536, 64256, 1, 1], type = bf16, converting to q8_0 .. size = 188.25 MiB -> 100.01 MiB
53
- [ 4/ 255] blk.0.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q8_0 .. size = 1.50 MiB -> 0.80 MiB
54
- [ 5/ 255] blk.0.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
55
- [ 6/ 255] blk.0.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q8_0 .. size = 6.00 MiB -> 3.19 MiB
56
- [ 7/ 255] blk.0.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q8_0 .. size = 6.00 MiB -> 3.19 MiB
57
- [ 8/ 255] blk.0.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q8_0 .. size = 1.50 MiB -> 0.80 MiB
58
- [ 9/ 255] blk.0.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q8_0 .. size = 15.00 MiB -> 7.97 MiB
59
- [ 10/ 255] blk.0.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 15.00 MiB -> 7.97 MiB
60
- [ 11/ 255] blk.0.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
61
- [ 12/ 255] blk.0.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 15.00 MiB -> 7.97 MiB
62
- [ 13/ 255] blk.1.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q8_0 .. size = 1.50 MiB -> 0.80 MiB
63
- [ 14/ 255] blk.1.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
64
- [ 15/ 255] blk.1.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q8_0 .. size = 6.00 MiB -> 3.19 MiB
65
- [ 16/ 255] blk.1.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q8_0 .. size = 6.00 MiB -> 3.19 MiB
66
- [ 17/ 255] blk.1.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q8_0 .. size = 1.50 MiB -> 0.80 MiB
67
- [ 18/ 255] blk.1.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q8_0 .. size = 15.00 MiB -> 7.97 MiB
68
- [ 19/ 255] blk.1.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 15.00 MiB -> 7.97 MiB
69
- [ 20/ 255] blk.1.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
70
- [ 21/ 255] blk.1.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 15.00 MiB -> 7.97 MiB
71
- [ 22/ 255] blk.2.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q8_0 .. size = 1.50 MiB -> 0.80 MiB
72
- [ 23/ 255] blk.2.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
73
- [ 24/ 255] blk.2.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q8_0 .. size = 6.00 MiB -> 3.19 MiB
74
- [ 25/ 255] blk.2.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q8_0 .. size = 6.00 MiB -> 3.19 MiB
75
- [ 26/ 255] blk.2.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q8_0 .. size = 1.50 MiB -> 0.80 MiB
76
- [ 27/ 255] blk.2.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q8_0 .. size = 15.00 MiB -> 7.97 MiB
77
- [ 28/ 255] blk.2.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 15.00 MiB -> 7.97 MiB
78
- [ 29/ 255] blk.2.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
79
- [ 30/ 255] blk.2.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 15.00 MiB -> 7.97 MiB
80
- [ 31/ 255] blk.3.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q8_0 .. size = 1.50 MiB -> 0.80 MiB
81
- [ 32/ 255] blk.3.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
82
- [ 33/ 255] blk.3.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q8_0 .. size = 6.00 MiB -> 3.19 MiB
83
- [ 34/ 255] blk.3.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q8_0 .. size = 6.00 MiB -> 3.19 MiB
84
- [ 35/ 255] blk.3.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q8_0 .. size = 1.50 MiB -> 0.80 MiB
85
- [ 36/ 255] blk.3.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q8_0 .. size = 15.00 MiB -> 7.97 MiB
86
- [ 37/ 255] blk.3.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 15.00 MiB -> 7.97 MiB
87
- [ 38/ 255] blk.3.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
88
- [ 39/ 255] blk.3.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 15.00 MiB -> 7.97 MiB
89
- [ 40/ 255] blk.4.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q8_0 .. size = 1.50 MiB -> 0.80 MiB
90
- [ 41/ 255] blk.4.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
91
- [ 42/ 255] blk.4.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q8_0 .. size = 6.00 MiB -> 3.19 MiB
92
- [ 43/ 255] blk.4.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q8_0 .. size = 6.00 MiB -> 3.19 MiB
93
- [ 44/ 255] blk.4.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q8_0 .. size = 1.50 MiB -> 0.80 MiB
94
- [ 45/ 255] blk.4.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q8_0 .. size = 15.00 MiB -> 7.97 MiB
95
- [ 46/ 255] blk.4.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 15.00 MiB -> 7.97 MiB
96
- [ 47/ 255] blk.4.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
97
- [ 48/ 255] blk.4.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 15.00 MiB -> 7.97 MiB
98
- [ 49/ 255] blk.5.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q8_0 .. size = 1.50 MiB -> 0.80 MiB
99
- [ 50/ 255] blk.5.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
100
- [ 51/ 255] blk.5.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q8_0 .. size = 6.00 MiB -> 3.19 MiB
101
- [ 52/ 255] blk.5.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q8_0 .. size = 6.00 MiB -> 3.19 MiB
102
- [ 53/ 255] blk.5.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q8_0 .. size = 1.50 MiB -> 0.80 MiB
103
- [ 54/ 255] blk.5.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q8_0 .. size = 15.00 MiB -> 7.97 MiB
104
- [ 55/ 255] blk.5.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 15.00 MiB -> 7.97 MiB
105
- [ 56/ 255] blk.5.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
106
- [ 57/ 255] blk.5.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 15.00 MiB -> 7.97 MiB
107
- [ 58/ 255] blk.6.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q8_0 .. size = 1.50 MiB -> 0.80 MiB
108
- [ 59/ 255] blk.6.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
109
- [ 60/ 255] blk.6.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q8_0 .. size = 6.00 MiB -> 3.19 MiB
110
- [ 61/ 255] blk.6.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q8_0 .. size = 6.00 MiB -> 3.19 MiB
111
- [ 62/ 255] blk.6.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q8_0 .. size = 1.50 MiB -> 0.80 MiB
112
- [ 63/ 255] blk.6.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q8_0 .. size = 15.00 MiB -> 7.97 MiB
113
- [ 64/ 255] blk.6.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 15.00 MiB -> 7.97 MiB
114
- [ 65/ 255] blk.6.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
115
- [ 66/ 255] blk.6.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 15.00 MiB -> 7.97 MiB
116
- [ 67/ 255] blk.7.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q8_0 .. size = 1.50 MiB -> 0.80 MiB
117
- [ 68/ 255] blk.7.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
118
- [ 69/ 255] blk.7.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q8_0 .. size = 6.00 MiB -> 3.19 MiB
119
- [ 70/ 255] blk.7.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q8_0 .. size = 6.00 MiB -> 3.19 MiB
120
- [ 71/ 255] blk.7.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q8_0 .. size = 1.50 MiB -> 0.80 MiB
121
- [ 72/ 255] blk.7.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q8_0 .. size = 15.00 MiB -> 7.97 MiB
122
- [ 73/ 255] blk.7.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 15.00 MiB -> 7.97 MiB
123
- [ 74/ 255] blk.7.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
124
- [ 75/ 255] blk.7.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 15.00 MiB -> 7.97 MiB
125
- [ 76/ 255] blk.8.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q8_0 .. size = 1.50 MiB -> 0.80 MiB
126
- [ 77/ 255] blk.8.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
127
- [ 78/ 255] blk.8.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q8_0 .. size = 6.00 MiB -> 3.19 MiB
128
- [ 79/ 255] blk.8.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q8_0 .. size = 6.00 MiB -> 3.19 MiB
129
- [ 80/ 255] blk.8.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q8_0 .. size = 1.50 MiB -> 0.80 MiB
130
- [ 81/ 255] blk.8.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q8_0 .. size = 15.00 MiB -> 7.97 MiB
131
- [ 82/ 255] blk.8.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 15.00 MiB -> 7.97 MiB
132
- [ 83/ 255] blk.8.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
133
- [ 84/ 255] blk.8.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 15.00 MiB -> 7.97 MiB
134
- [ 85/ 255] blk.9.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q8_0 .. size = 1.50 MiB -> 0.80 MiB
135
- [ 86/ 255] blk.9.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
136
- [ 87/ 255] blk.9.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q8_0 .. size = 6.00 MiB -> 3.19 MiB
137
- [ 88/ 255] blk.9.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q8_0 .. size = 6.00 MiB -> 3.19 MiB
138
- [ 89/ 255] blk.9.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q8_0 .. size = 1.50 MiB -> 0.80 MiB
139
- [ 90/ 255] blk.9.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q8_0 .. size = 15.00 MiB -> 7.97 MiB
140
- [ 91/ 255] blk.9.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 15.00 MiB -> 7.97 MiB
141
- [ 92/ 255] blk.9.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
142
- [ 93/ 255] blk.9.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 15.00 MiB -> 7.97 MiB
143
- [ 94/ 255] blk.10.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q8_0 .. size = 1.50 MiB -> 0.80 MiB
144
- [ 95/ 255] blk.10.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
145
- [ 96/ 255] blk.10.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q8_0 .. size = 6.00 MiB -> 3.19 MiB
146
- [ 97/ 255] blk.10.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q8_0 .. size = 6.00 MiB -> 3.19 MiB
147
- [ 98/ 255] blk.10.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q8_0 .. size = 1.50 MiB -> 0.80 MiB
148
- [ 99/ 255] blk.10.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q8_0 .. size = 15.00 MiB -> 7.97 MiB
149
- [ 100/ 255] blk.10.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 15.00 MiB -> 7.97 MiB
150
- [ 101/ 255] blk.10.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
151
- [ 102/ 255] blk.10.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 15.00 MiB -> 7.97 MiB
152
- [ 103/ 255] blk.11.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q8_0 .. size = 1.50 MiB -> 0.80 MiB
153
- [ 104/ 255] blk.11.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
154
- [ 105/ 255] blk.11.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q8_0 .. size = 6.00 MiB -> 3.19 MiB
155
- [ 106/ 255] blk.11.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q8_0 .. size = 6.00 MiB -> 3.19 MiB
156
- [ 107/ 255] blk.11.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q8_0 .. size = 1.50 MiB -> 0.80 MiB
157
- [ 108/ 255] blk.11.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q8_0 .. size = 15.00 MiB -> 7.97 MiB
158
- [ 109/ 255] blk.11.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 15.00 MiB -> 7.97 MiB
159
- [ 110/ 255] blk.11.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
160
- [ 111/ 255] blk.11.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 15.00 MiB -> 7.97 MiB
161
- [ 112/ 255] blk.12.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q8_0 .. size = 1.50 MiB -> 0.80 MiB
162
- [ 113/ 255] blk.12.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
163
- [ 114/ 255] blk.12.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q8_0 .. size = 6.00 MiB -> 3.19 MiB
164
- [ 115/ 255] blk.12.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q8_0 .. size = 6.00 MiB -> 3.19 MiB
165
- [ 116/ 255] blk.12.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q8_0 .. size = 1.50 MiB -> 0.80 MiB
166
- [ 117/ 255] blk.12.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q8_0 .. size = 15.00 MiB -> 7.97 MiB
167
- [ 118/ 255] blk.12.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 15.00 MiB -> 7.97 MiB
168
- [ 119/ 255] blk.12.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
169
- [ 120/ 255] blk.12.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 15.00 MiB -> 7.97 MiB
170
- [ 121/ 255] blk.13.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q8_0 .. size = 1.50 MiB -> 0.80 MiB
171
- [ 122/ 255] blk.13.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
172
- [ 123/ 255] blk.13.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q8_0 .. size = 6.00 MiB -> 3.19 MiB
173
- [ 124/ 255] blk.13.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q8_0 .. size = 6.00 MiB -> 3.19 MiB
174
- [ 125/ 255] blk.13.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q8_0 .. size = 1.50 MiB -> 0.80 MiB
175
- [ 126/ 255] blk.13.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q8_0 .. size = 15.00 MiB -> 7.97 MiB
176
- [ 127/ 255] blk.13.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 15.00 MiB -> 7.97 MiB
177
- [ 128/ 255] blk.13.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
178
- [ 129/ 255] blk.13.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 15.00 MiB -> 7.97 MiB
179
- [ 130/ 255] blk.14.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q8_0 .. size = 1.50 MiB -> 0.80 MiB
180
- [ 131/ 255] blk.14.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
181
- [ 132/ 255] blk.14.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q8_0 .. size = 6.00 MiB -> 3.19 MiB
182
- [ 133/ 255] blk.14.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q8_0 .. size = 6.00 MiB -> 3.19 MiB
183
- [ 134/ 255] blk.14.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q8_0 .. size = 1.50 MiB -> 0.80 MiB
184
- [ 135/ 255] blk.14.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q8_0 .. size = 15.00 MiB -> 7.97 MiB
185
- [ 136/ 255] blk.14.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 15.00 MiB -> 7.97 MiB
186
- [ 137/ 255] blk.14.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
187
- [ 138/ 255] blk.14.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 15.00 MiB -> 7.97 MiB
188
- [ 139/ 255] blk.15.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q8_0 .. size = 1.50 MiB -> 0.80 MiB
189
- [ 140/ 255] blk.15.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
190
- [ 141/ 255] blk.15.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q8_0 .. size = 6.00 MiB -> 3.19 MiB
191
- [ 142/ 255] blk.15.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q8_0 .. size = 6.00 MiB -> 3.19 MiB
192
- [ 143/ 255] blk.15.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q8_0 .. size = 1.50 MiB -> 0.80 MiB
193
- [ 144/ 255] blk.15.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q8_0 .. size = 15.00 MiB -> 7.97 MiB
194
- [ 145/ 255] blk.15.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 15.00 MiB -> 7.97 MiB
195
- [ 146/ 255] blk.15.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
196
- [ 147/ 255] blk.15.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 15.00 MiB -> 7.97 MiB
197
- [ 148/ 255] blk.16.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q8_0 .. size = 1.50 MiB -> 0.80 MiB
198
- [ 149/ 255] blk.16.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
199
- [ 150/ 255] blk.16.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q8_0 .. size = 6.00 MiB -> 3.19 MiB
200
- [ 151/ 255] blk.16.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q8_0 .. size = 6.00 MiB -> 3.19 MiB
201
- [ 152/ 255] blk.16.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q8_0 .. size = 1.50 MiB -> 0.80 MiB
202
- [ 153/ 255] blk.16.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q8_0 .. size = 15.00 MiB -> 7.97 MiB
203
- [ 154/ 255] blk.16.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 15.00 MiB -> 7.97 MiB
204
- [ 155/ 255] blk.16.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
205
- [ 156/ 255] blk.16.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 15.00 MiB -> 7.97 MiB
206
- [ 157/ 255] blk.17.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q8_0 .. size = 1.50 MiB -> 0.80 MiB
207
- [ 158/ 255] blk.17.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
208
- [ 159/ 255] blk.17.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q8_0 .. size = 6.00 MiB -> 3.19 MiB
209
- [ 160/ 255] blk.17.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q8_0 .. size = 6.00 MiB -> 3.19 MiB
210
- [ 161/ 255] blk.17.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q8_0 .. size = 1.50 MiB -> 0.80 MiB
211
- [ 162/ 255] blk.17.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q8_0 .. size = 15.00 MiB -> 7.97 MiB
212
- [ 163/ 255] blk.17.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 15.00 MiB -> 7.97 MiB
213
- [ 164/ 255] blk.17.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
214
- [ 165/ 255] blk.17.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 15.00 MiB -> 7.97 MiB
215
- [ 166/ 255] blk.18.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q8_0 .. size = 1.50 MiB -> 0.80 MiB
216
- [ 167/ 255] blk.18.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
217
- [ 168/ 255] blk.18.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q8_0 .. size = 6.00 MiB -> 3.19 MiB
218
- [ 169/ 255] blk.18.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q8_0 .. size = 6.00 MiB -> 3.19 MiB
219
- [ 170/ 255] blk.18.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q8_0 .. size = 1.50 MiB -> 0.80 MiB
220
- [ 171/ 255] blk.18.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q8_0 .. size = 15.00 MiB -> 7.97 MiB
221
- [ 172/ 255] blk.18.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 15.00 MiB -> 7.97 MiB
222
- [ 173/ 255] blk.18.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
223
- [ 174/ 255] blk.18.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 15.00 MiB -> 7.97 MiB
224
- [ 175/ 255] blk.19.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q8_0 .. size = 1.50 MiB -> 0.80 MiB
225
- [ 176/ 255] blk.19.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
226
- [ 177/ 255] blk.19.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q8_0 .. size = 6.00 MiB -> 3.19 MiB
227
- [ 178/ 255] blk.19.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q8_0 .. size = 6.00 MiB -> 3.19 MiB
228
- [ 179/ 255] blk.19.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q8_0 .. size = 1.50 MiB -> 0.80 MiB
229
- [ 180/ 255] blk.19.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q8_0 .. size = 15.00 MiB -> 7.97 MiB
230
- [ 181/ 255] blk.19.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 15.00 MiB -> 7.97 MiB
231
- [ 182/ 255] blk.19.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
232
- [ 183/ 255] blk.19.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 15.00 MiB -> 7.97 MiB
233
- [ 184/ 255] blk.20.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q8_0 .. size = 1.50 MiB -> 0.80 MiB
234
- [ 185/ 255] blk.20.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
235
- [ 186/ 255] blk.20.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q8_0 .. size = 6.00 MiB -> 3.19 MiB
236
- [ 187/ 255] blk.20.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q8_0 .. size = 6.00 MiB -> 3.19 MiB
237
- [ 188/ 255] blk.20.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q8_0 .. size = 1.50 MiB -> 0.80 MiB
238
- [ 189/ 255] blk.20.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q8_0 .. size = 15.00 MiB -> 7.97 MiB
239
- [ 190/ 255] blk.20.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 15.00 MiB -> 7.97 MiB
240
- [ 191/ 255] blk.20.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
241
- [ 192/ 255] blk.20.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 15.00 MiB -> 7.97 MiB
242
- [ 193/ 255] blk.21.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q8_0 .. size = 1.50 MiB -> 0.80 MiB
243
- [ 194/ 255] blk.21.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
244
- [ 195/ 255] blk.21.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q8_0 .. size = 6.00 MiB -> 3.19 MiB
245
- [ 196/ 255] blk.21.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q8_0 .. size = 6.00 MiB -> 3.19 MiB
246
- [ 197/ 255] blk.21.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q8_0 .. size = 1.50 MiB -> 0.80 MiB
247
- [ 198/ 255] blk.21.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q8_0 .. size = 15.00 MiB -> 7.97 MiB
248
- [ 199/ 255] blk.21.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 15.00 MiB -> 7.97 MiB
249
- [ 200/ 255] blk.21.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
250
- [ 201/ 255] blk.21.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 15.00 MiB -> 7.97 MiB
251
- [ 202/ 255] blk.22.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q8_0 .. size = 1.50 MiB -> 0.80 MiB
252
- [ 203/ 255] blk.22.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
253
- [ 204/ 255] blk.22.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q8_0 .. size = 6.00 MiB -> 3.19 MiB
254
- [ 205/ 255] blk.22.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q8_0 .. size = 6.00 MiB -> 3.19 MiB
255
- [ 206/ 255] blk.22.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q8_0 .. size = 1.50 MiB -> 0.80 MiB
256
- [ 207/ 255] blk.22.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q8_0 .. size = 15.00 MiB -> 7.97 MiB
257
- [ 208/ 255] blk.22.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 15.00 MiB -> 7.97 MiB
258
- [ 209/ 255] blk.22.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
259
- [ 210/ 255] blk.22.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 15.00 MiB -> 7.97 MiB
260
- [ 211/ 255] blk.23.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q8_0 .. size = 1.50 MiB -> 0.80 MiB
261
- [ 212/ 255] blk.23.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
262
- [ 213/ 255] blk.23.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q8_0 .. size = 6.00 MiB -> 3.19 MiB
263
- [ 214/ 255] blk.23.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q8_0 .. size = 6.00 MiB -> 3.19 MiB
264
- [ 215/ 255] blk.23.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q8_0 .. size = 1.50 MiB -> 0.80 MiB
265
- [ 216/ 255] blk.23.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q8_0 .. size = 15.00 MiB -> 7.97 MiB
266
- [ 217/ 255] blk.23.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 15.00 MiB -> 7.97 MiB
267
- [ 218/ 255] blk.23.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
268
- [ 219/ 255] blk.23.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 15.00 MiB -> 7.97 MiB
269
- [ 220/ 255] blk.24.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q8_0 .. size = 1.50 MiB -> 0.80 MiB
270
- [ 221/ 255] blk.24.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
271
- [ 222/ 255] blk.24.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q8_0 .. size = 6.00 MiB -> 3.19 MiB
272
- [ 223/ 255] blk.24.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q8_0 .. size = 6.00 MiB -> 3.19 MiB
273
- [ 224/ 255] blk.24.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q8_0 .. size = 1.50 MiB -> 0.80 MiB
274
- [ 225/ 255] blk.24.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q8_0 .. size = 15.00 MiB -> 7.97 MiB
275
- [ 226/ 255] blk.24.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 15.00 MiB -> 7.97 MiB
276
- [ 227/ 255] blk.24.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
277
- [ 228/ 255] blk.24.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 15.00 MiB -> 7.97 MiB
278
- [ 229/ 255] blk.25.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q8_0 .. size = 1.50 MiB -> 0.80 MiB
279
- [ 230/ 255] blk.25.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
280
- [ 231/ 255] blk.25.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q8_0 .. size = 6.00 MiB -> 3.19 MiB
281
- [ 232/ 255] blk.25.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q8_0 .. size = 6.00 MiB -> 3.19 MiB
282
- [ 233/ 255] blk.25.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q8_0 .. size = 1.50 MiB -> 0.80 MiB
283
- [ 234/ 255] blk.25.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q8_0 .. size = 15.00 MiB -> 7.97 MiB
284
- [ 235/ 255] blk.25.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 15.00 MiB -> 7.97 MiB
285
- [ 236/ 255] blk.25.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
286
- [ 237/ 255] blk.25.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 15.00 MiB -> 7.97 MiB
287
- [ 238/ 255] blk.26.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q8_0 .. size = 1.50 MiB -> 0.80 MiB
288
- [ 239/ 255] blk.26.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
289
- [ 240/ 255] blk.26.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q8_0 .. size = 6.00 MiB -> 3.19 MiB
290
- [ 241/ 255] blk.26.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q8_0 .. size = 6.00 MiB -> 3.19 MiB
291
- [ 242/ 255] blk.26.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q8_0 .. size = 1.50 MiB -> 0.80 MiB
292
- [ 243/ 255] blk.26.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q8_0 .. size = 15.00 MiB -> 7.97 MiB
293
- [ 244/ 255] blk.26.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 15.00 MiB -> 7.97 MiB
294
- [ 245/ 255] blk.26.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
295
- [ 246/ 255] blk.26.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 15.00 MiB -> 7.97 MiB
296
- [ 247/ 255] blk.27.attn_k.weight - [ 1536, 512, 1, 1], type = bf16, converting to q8_0 .. size = 1.50 MiB -> 0.80 MiB
297
- [ 248/ 255] blk.27.attn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
298
- [ 249/ 255] blk.27.attn_output.weight - [ 2048, 1536, 1, 1], type = bf16, converting to q8_0 .. size = 6.00 MiB -> 3.19 MiB
299
- [ 250/ 255] blk.27.attn_q.weight - [ 1536, 2048, 1, 1], type = bf16, converting to q8_0 .. size = 6.00 MiB -> 3.19 MiB
300
- [ 251/ 255] blk.27.attn_v.weight - [ 1536, 512, 1, 1], type = bf16, converting to q8_0 .. size = 1.50 MiB -> 0.80 MiB
301
- [ 252/ 255] blk.27.ffn_down.weight - [ 5120, 1536, 1, 1], type = bf16, converting to q8_0 .. size = 15.00 MiB -> 7.97 MiB
302
- [ 253/ 255] blk.27.ffn_gate.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 15.00 MiB -> 7.97 MiB
303
- [ 254/ 255] blk.27.ffn_norm.weight - [ 1536, 1, 1, 1], type = f32, size = 0.006 MiB
304
- [ 255/ 255] blk.27.ffn_up.weight - [ 1536, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 15.00 MiB -> 7.97 MiB
305
- llama_model_quantize_impl: model size = 2056.83 MiB (16.00 BPW)
306
- llama_model_quantize_impl: quant size = 1092.85 MiB (8.50 BPW)
307
-
308
- llama_quantize: quantize time = 11529.23 ms
309
- llama_quantize: total time = 11529.23 ms
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
reproducibility/validation/smoke-test-gpu1.tsv DELETED
@@ -1,10 +0,0 @@
1
- model status log
2
- K2-Horizon-0.9B-Q8_0.gguf PASS /root/workspace/HF/reports/k2-horizon-0.9b/smoke-tests-gpu1/K2-Horizon-0.9B-Q8_0.log
3
- K2-Horizon-0.9B-Q6_K.gguf PASS /root/workspace/HF/reports/k2-horizon-0.9b/smoke-tests-gpu1/K2-Horizon-0.9B-Q6_K.log
4
- K2-Horizon-0.9B-Q5_K_M.gguf PASS /root/workspace/HF/reports/k2-horizon-0.9b/smoke-tests-gpu1/K2-Horizon-0.9B-Q5_K_M.log
5
- K2-Horizon-0.9B-Q4_K_M.gguf PASS /root/workspace/HF/reports/k2-horizon-0.9b/smoke-tests-gpu1/K2-Horizon-0.9B-Q4_K_M.log
6
- K2-Horizon-0.9B-Q3_K_M.gguf PASS /root/workspace/HF/reports/k2-horizon-0.9b/smoke-tests-gpu1/K2-Horizon-0.9B-Q3_K_M.log
7
- K2-Horizon-0.9B-Q2_K.gguf PASS /root/workspace/HF/reports/k2-horizon-0.9b/smoke-tests-gpu1/K2-Horizon-0.9B-Q2_K.log
8
- K2-Horizon-0.9B-IQ2_XS.gguf PASS /root/workspace/HF/reports/k2-horizon-0.9b/smoke-tests-gpu1/K2-Horizon-0.9B-IQ2_XS.log
9
- K2-Horizon-0.9B-IQ1_M.gguf PASS /root/workspace/HF/reports/k2-horizon-0.9b/smoke-tests-gpu1/K2-Horizon-0.9B-IQ1_M.log
10
- K2-Horizon-0.9B-Q1_0.gguf PASS /root/workspace/HF/reports/k2-horizon-0.9b/smoke-tests-gpu1/K2-Horizon-0.9B-Q1_0.log
 
 
 
 
 
 
 
 
 
 
 
reproducibility/validation/smoke-tests-gpu1/K2-Horizon-0.9B-IQ1_M.log DELETED
@@ -1,56 +0,0 @@
1
-
2
-
3
- Loading model...
4
-
5
- β–„β–„ β–„β–„
6
- β–ˆβ–ˆ β–ˆβ–ˆ
7
- β–ˆβ–ˆ β–ˆβ–ˆ β–€β–€β–ˆβ–„ β–ˆβ–ˆβ–ˆβ–„β–ˆβ–ˆβ–ˆβ–„ β–€β–€β–ˆβ–„ β–„β–ˆβ–ˆβ–ˆβ–ˆ β–ˆβ–ˆβ–ˆβ–ˆβ–„ β–ˆβ–ˆβ–ˆβ–ˆβ–„
8
- β–ˆβ–ˆ β–ˆβ–ˆ β–„β–ˆβ–€β–ˆβ–ˆ β–ˆβ–ˆ β–ˆβ–ˆ β–ˆβ–ˆ β–„β–ˆβ–€β–ˆβ–ˆ β–ˆβ–ˆ β–ˆβ–ˆ β–ˆβ–ˆ β–ˆβ–ˆ β–ˆβ–ˆ
9
- β–ˆβ–ˆ β–ˆβ–ˆ β–€β–ˆβ–„β–ˆβ–ˆ β–ˆβ–ˆ β–ˆβ–ˆ β–ˆβ–ˆ β–€β–ˆβ–„β–ˆβ–ˆ β–ˆβ–ˆ β–€β–ˆβ–ˆβ–ˆβ–ˆ β–ˆβ–ˆβ–ˆβ–ˆβ–€ β–ˆβ–ˆβ–ˆβ–ˆβ–€
10
- β–ˆβ–ˆ β–ˆβ–ˆ
11
- β–€β–€ β–€β–€
12
-
13
- build : b10671-35999d101
14
- model : /root/workspace/HF/quantized-k2-horizon-0.9b/K2-Horizon-0.9B-IQ1_M.gguf
15
- ftype : IQ1_M - 1.75 bpw
16
- modalities : text
17
-
18
- available commands:
19
- /exit or Ctrl+C stop or exit
20
- /regen regenerate the last response
21
- /clear clear the chat history
22
- /read <file> add a text file
23
- /glob <pattern> add text files using globbing pattern
24
-
25
-
26
-
27
- > Answer in one sentence: what does a GGUF file contain?
28
- =
29
-
30
- Answer: "GNU file" contains?
31
-
32
- (...):
33
- )
34
-
35
- (...)
36
-
37
- The initial answer was being generated.
38
-
39
- (...)
40
-
41
-
42
-
43
- ```
44
-
45
- (...)
46
-
47
-
48
-
49
- ``````
50
-
51
- ````````````````````````````````````````````
52
-
53
- [ Prompt: 241.4 t/s | Generation: 320.7 t/s ]
54
-
55
-
56
- Exiting...
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
reproducibility/validation/smoke-tests-gpu1/K2-Horizon-0.9B-IQ2_XS.log DELETED
@@ -1,33 +0,0 @@
1
-
2
-
3
- Loading model...
4
-
5
- β–„β–„ β–„β–„
6
- β–ˆβ–ˆ β–ˆβ–ˆ
7
- β–ˆβ–ˆ β–ˆβ–ˆ β–€β–€β–ˆβ–„ β–ˆβ–ˆβ–ˆβ–„β–ˆβ–ˆβ–ˆβ–„ β–€β–€β–ˆβ–„ β–„β–ˆβ–ˆβ–ˆβ–ˆ β–ˆβ–ˆβ–ˆβ–ˆβ–„ β–ˆβ–ˆβ–ˆβ–ˆβ–„
8
- β–ˆβ–ˆ β–ˆβ–ˆ β–„β–ˆβ–€β–ˆβ–ˆ β–ˆβ–ˆ β–ˆβ–ˆ β–ˆβ–ˆ β–„β–ˆβ–€β–ˆβ–ˆ β–ˆβ–ˆ β–ˆβ–ˆ β–ˆβ–ˆ β–ˆβ–ˆ β–ˆβ–ˆ
9
- β–ˆβ–ˆ β–ˆβ–ˆ β–€β–ˆβ–„β–ˆβ–ˆ β–ˆβ–ˆ β–ˆβ–ˆ β–ˆβ–ˆ β–€β–ˆβ–„β–ˆβ–ˆ β–ˆβ–ˆ β–€β–ˆβ–ˆβ–ˆβ–ˆ β–ˆβ–ˆβ–ˆβ–ˆβ–€ β–ˆβ–ˆβ–ˆβ–ˆβ–€
10
- β–ˆβ–ˆ β–ˆβ–ˆ
11
- β–€β–€ β–€β–€
12
-
13
- build : b10671-35999d101
14
- model : /root/workspace/HF/quantized-k2-horizon-0.9b/K2-Horizon-0.9B-IQ2_XS.gguf
15
- ftype : IQ2_XS - 2.3125 bpw
16
- modalities : text
17
-
18
- available commands:
19
- /exit or Ctrl+C stop or exit
20
- /regen regenerate the last response
21
- /clear clear the chat history
22
- /read <file> add a text file
23
- /glob <pattern> add text files using globbing pattern
24
-
25
-
26
-
27
- > Answer in one sentence: what does a GGUF file contain?
28
- The user is asking for a concise answer to a question about GGUF files. Let me think about what this refers to. GGUF is a type of file format used in the Google Folder (GluF) library, which is a type of library used in the Google Folder (Glu
29
-
30
- [ Prompt: 1179.2 t/s | Generation: 303.3 t/s ]
31
-
32
-
33
- Exiting...
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
reproducibility/validation/smoke-tests-gpu1/K2-Horizon-0.9B-Q1_0.log DELETED
@@ -1,33 +0,0 @@
1
-
2
-
3
- Loading model...
4
-
5
- β–„β–„ β–„β–„
6
- β–ˆβ–ˆ β–ˆβ–ˆ
7
- β–ˆβ–ˆ β–ˆβ–ˆ β–€β–€β–ˆβ–„ β–ˆβ–ˆβ–ˆβ–„β–ˆβ–ˆβ–ˆβ–„ β–€β–€β–ˆβ–„ β–„β–ˆβ–ˆβ–ˆβ–ˆ β–ˆβ–ˆβ–ˆβ–ˆβ–„ β–ˆβ–ˆβ–ˆβ–ˆβ–„
8
- β–ˆβ–ˆ β–ˆβ–ˆ β–„β–ˆβ–€β–ˆβ–ˆ β–ˆβ–ˆ β–ˆβ–ˆ β–ˆβ–ˆ β–„β–ˆβ–€β–ˆβ–ˆ β–ˆβ–ˆ β–ˆβ–ˆ β–ˆβ–ˆ β–ˆβ–ˆ β–ˆβ–ˆ
9
- β–ˆβ–ˆ β–ˆβ–ˆ β–€β–ˆβ–„β–ˆβ–ˆ β–ˆβ–ˆ β–ˆβ–ˆ β–ˆβ–ˆ β–€β–ˆβ–„β–ˆβ–ˆ β–ˆβ–ˆ β–€β–ˆβ–ˆβ–ˆβ–ˆ β–ˆβ–ˆβ–ˆβ–ˆβ–€ β–ˆβ–ˆβ–ˆβ–ˆβ–€
10
- β–ˆβ–ˆ β–ˆβ–ˆ
11
- β–€β–€ β–€β–€
12
-
13
- build : b10671-35999d101
14
- model : /root/workspace/HF/quantized-k2-horizon-0.9b/K2-Horizon-0.9B-Q1_0.gguf
15
- ftype : Q1_0
16
- modalities : text
17
-
18
- available commands:
19
- /exit or Ctrl+C stop or exit
20
- /regen regenerate the last response
21
- /clear clear the chat history
22
- /read <file> add a text file
23
- /glob <pattern> add text files using globbing pattern
24
-
25
-
26
-
27
- > Answer in one sentence: what does a GGUF file contain?
28
- .g.g.g.g.g.g.g.g.g.g.g.g.g.g.g.g.g.g.g.g.g.g.g.g.g.g.g.g.g.g.g.g.g.g.g.g.g.g.g.g.g.g.g.g.g.g.g.g.g.g.g.g.g.g.g.g.g.g.g.g.g.g.g.g
29
-
30
- [ Prompt: 1935.7 t/s | Generation: 407.4 t/s ]
31
-
32
-
33
- Exiting...
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
reproducibility/validation/smoke-tests-gpu1/K2-Horizon-0.9B-Q2_K.log DELETED
@@ -1,33 +0,0 @@
1
-
2
-
3
- Loading model...
4
-
5
- β–„β–„ β–„β–„
6
- β–ˆβ–ˆ β–ˆβ–ˆ
7
- β–ˆβ–ˆ β–ˆβ–ˆ β–€β–€β–ˆβ–„ β–ˆβ–ˆβ–ˆβ–„β–ˆβ–ˆβ–ˆβ–„ β–€β–€β–ˆβ–„ β–„β–ˆβ–ˆβ–ˆβ–ˆ β–ˆβ–ˆβ–ˆβ–ˆβ–„ β–ˆβ–ˆβ–ˆβ–ˆβ–„
8
- β–ˆβ–ˆ β–ˆβ–ˆ β–„β–ˆβ–€β–ˆβ–ˆ β–ˆβ–ˆ β–ˆβ–ˆ β–ˆβ–ˆ β–„β–ˆβ–€β–ˆβ–ˆ β–ˆβ–ˆ β–ˆβ–ˆ β–ˆβ–ˆ β–ˆβ–ˆ β–ˆβ–ˆ
9
- β–ˆβ–ˆ β–ˆβ–ˆ β–€β–ˆβ–„β–ˆβ–ˆ β–ˆβ–ˆ β–ˆβ–ˆ β–ˆβ–ˆ β–€β–ˆβ–„β–ˆβ–ˆ β–ˆβ–ˆ β–€β–ˆβ–ˆβ–ˆβ–ˆ β–ˆβ–ˆβ–ˆβ–ˆβ–€ β–ˆβ–ˆβ–ˆβ–ˆβ–€
10
- β–ˆβ–ˆ β–ˆβ–ˆ
11
- β–€β–€ β–€β–€
12
-
13
- build : b10671-35999d101
14
- model : /root/workspace/HF/quantized-k2-horizon-0.9b/K2-Horizon-0.9B-Q2_K.gguf
15
- ftype : Q2_K - Medium
16
- modalities : text
17
-
18
- available commands:
19
- /exit or Ctrl+C stop or exit
20
- /regen regenerate the last response
21
- /clear clear the chat history
22
- /read <file> add a text file
23
- /glob <pattern> add text files using globbing pattern
24
-
25
-
26
-
27
- > Answer in one sentence: what does a GGUF file contain?
28
- The user asks: "Answer in one sentence: what does a GGUF file contain?" They want a single sentence answer. The question: "what does a GGUF file contain?" GGUF stands for "GitHub Gist" or "GitHub Gist file"? Actually GGUF
29
-
30
- [ Prompt: 1188.2 t/s | Generation: 304.3 t/s ]
31
-
32
-
33
- Exiting...
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
reproducibility/validation/smoke-tests-gpu1/K2-Horizon-0.9B-Q3_K_M.log DELETED
@@ -1,33 +0,0 @@
1
-
2
-
3
- Loading model...
4
-
5
- β–„β–„ β–„β–„
6
- β–ˆβ–ˆ β–ˆβ–ˆ
7
- β–ˆβ–ˆ β–ˆβ–ˆ β–€β–€β–ˆβ–„ β–ˆβ–ˆβ–ˆβ–„β–ˆβ–ˆβ–ˆβ–„ β–€β–€β–ˆβ–„ β–„β–ˆβ–ˆβ–ˆβ–ˆ β–ˆβ–ˆβ–ˆβ–ˆβ–„ β–ˆβ–ˆβ–ˆβ–ˆβ–„
8
- β–ˆβ–ˆ β–ˆβ–ˆ β–„β–ˆβ–€β–ˆβ–ˆ β–ˆβ–ˆ β–ˆβ–ˆ β–ˆβ–ˆ β–„β–ˆβ–€β–ˆβ–ˆ β–ˆβ–ˆ β–ˆβ–ˆ β–ˆβ–ˆ β–ˆβ–ˆ β–ˆβ–ˆ
9
- β–ˆβ–ˆ β–ˆβ–ˆ β–€β–ˆβ–„β–ˆβ–ˆ β–ˆβ–ˆ β–ˆβ–ˆ β–ˆβ–ˆ β–€β–ˆβ–„β–ˆβ–ˆ β–ˆβ–ˆ β–€β–ˆβ–ˆβ–ˆβ–ˆ β–ˆβ–ˆβ–ˆβ–ˆβ–€ β–ˆβ–ˆβ–ˆβ–ˆβ–€
10
- β–ˆβ–ˆ β–ˆβ–ˆ
11
- β–€β–€ β–€β–€
12
-
13
- build : b10671-35999d101
14
- model : /root/workspace/HF/quantized-k2-horizon-0.9b/K2-Horizon-0.9B-Q3_K_M.gguf
15
- ftype : Q3_K - Medium
16
- modalities : text
17
-
18
- available commands:
19
- /exit or Ctrl+C stop or exit
20
- /regen regenerate the last response
21
- /clear clear the chat history
22
- /read <file> add a text file
23
- /glob <pattern> add text files using globbing pattern
24
-
25
-
26
-
27
- > Answer in one sentence: what does a GGUF file contain?
28
- The user asks: "Answer in one sentence: what does a GGUF file contain?" They want a one-sentence answer. GGUF stands for "GPU GIF"? Actually GGUF is a format for GIF files? Wait, GGUF is a format for "GPU G
29
-
30
- [ Prompt: 1203.4 t/s | Generation: 247.5 t/s ]
31
-
32
-
33
- Exiting...
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
reproducibility/validation/smoke-tests-gpu1/K2-Horizon-0.9B-Q4_K_M.log DELETED
@@ -1,33 +0,0 @@
1
-
2
-
3
- Loading model...
4
-
5
- β–„β–„ β–„β–„
6
- β–ˆβ–ˆ β–ˆβ–ˆ
7
- β–ˆβ–ˆ β–ˆβ–ˆ β–€β–€β–ˆβ–„ β–ˆβ–ˆβ–ˆβ–„β–ˆβ–ˆβ–ˆβ–„ β–€β–€β–ˆβ–„ β–„β–ˆβ–ˆβ–ˆβ–ˆ β–ˆβ–ˆβ–ˆβ–ˆβ–„ β–ˆβ–ˆβ–ˆβ–ˆβ–„
8
- β–ˆβ–ˆ β–ˆβ–ˆ β–„β–ˆβ–€β–ˆβ–ˆ β–ˆβ–ˆ β–ˆβ–ˆ β–ˆβ–ˆ β–„β–ˆβ–€β–ˆβ–ˆ β–ˆβ–ˆ β–ˆβ–ˆ β–ˆβ–ˆ β–ˆβ–ˆ β–ˆβ–ˆ
9
- β–ˆβ–ˆ β–ˆβ–ˆ β–€β–ˆβ–„β–ˆβ–ˆ β–ˆβ–ˆ β–ˆβ–ˆ β–ˆβ–ˆ β–€β–ˆβ–„β–ˆβ–ˆ β–ˆβ–ˆ β–€β–ˆβ–ˆβ–ˆβ–ˆ β–ˆβ–ˆβ–ˆβ–ˆβ–€ β–ˆβ–ˆβ–ˆβ–ˆβ–€
10
- β–ˆβ–ˆ β–ˆβ–ˆ
11
- β–€β–€ β–€β–€
12
-
13
- build : b10671-35999d101
14
- model : /root/workspace/HF/quantized-k2-horizon-0.9b/K2-Horizon-0.9B-Q4_K_M.gguf
15
- ftype : Q4_K - Medium
16
- modalities : text
17
-
18
- available commands:
19
- /exit or Ctrl+C stop or exit
20
- /regen regenerate the last response
21
- /clear clear the chat history
22
- /read <file> add a text file
23
- /glob <pattern> add text files using globbing pattern
24
-
25
-
26
-
27
- > Answer in one sentence: what does a GGUF file contain?
28
- The user asks: "Answer in one sentence: what does a GGUF file contain?" They want a concise answer: a GGUF file contains a model's weights and metadata, typically a JSON file with model architecture, tokenizer, etc. But they ask "what does a GGUF file contain
29
-
30
- [ Prompt: 1477.1 t/s | Generation: 325.6 t/s ]
31
-
32
-
33
- Exiting...
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
reproducibility/validation/smoke-tests-gpu1/K2-Horizon-0.9B-Q5_K_M.log DELETED
@@ -1,33 +0,0 @@
1
-
2
-
3
- Loading model...
4
-
5
- β–„β–„ β–„β–„
6
- β–ˆβ–ˆ β–ˆβ–ˆ
7
- β–ˆβ–ˆ β–ˆβ–ˆ β–€β–€β–ˆβ–„ β–ˆβ–ˆβ–ˆβ–„β–ˆβ–ˆβ–ˆβ–„ β–€β–€β–ˆβ–„ β–„β–ˆβ–ˆβ–ˆβ–ˆ β–ˆβ–ˆβ–ˆβ–ˆβ–„ β–ˆβ–ˆβ–ˆβ–ˆβ–„
8
- β–ˆβ–ˆ β–ˆβ–ˆ β–„β–ˆβ–€β–ˆβ–ˆ β–ˆβ–ˆ β–ˆβ–ˆ β–ˆβ–ˆ β–„β–ˆβ–€β–ˆβ–ˆ β–ˆβ–ˆ β–ˆβ–ˆ β–ˆβ–ˆ β–ˆβ–ˆ β–ˆβ–ˆ
9
- β–ˆβ–ˆ β–ˆβ–ˆ β–€β–ˆβ–„β–ˆβ–ˆ β–ˆβ–ˆ β–ˆβ–ˆ β–ˆβ–ˆ β–€β–ˆβ–„β–ˆβ–ˆ β–ˆβ–ˆ β–€β–ˆβ–ˆβ–ˆβ–ˆ β–ˆβ–ˆβ–ˆβ–ˆβ–€ β–ˆβ–ˆβ–ˆβ–ˆβ–€
10
- β–ˆβ–ˆ β–ˆβ–ˆ
11
- β–€β–€ β–€β–€
12
-
13
- build : b10671-35999d101
14
- model : /root/workspace/HF/quantized-k2-horizon-0.9b/K2-Horizon-0.9B-Q5_K_M.gguf
15
- ftype : Q5_K - Medium
16
- modalities : text
17
-
18
- available commands:
19
- /exit or Ctrl+C stop or exit
20
- /regen regenerate the last response
21
- /clear clear the chat history
22
- /read <file> add a text file
23
- /glob <pattern> add text files using globbing pattern
24
-
25
-
26
-
27
- > Answer in one sentence: what does a GGUF file contain?
28
- The user asks: "Answer in one sentence: what does a GGUF file contain?" They want a concise answer: a GGUF file contains a model's weights and metadata, typically a quantized version of a model, used for inference. So answer: "A GGUF file contains the quant
29
-
30
- [ Prompt: 1422.8 t/s | Generation: 295.2 t/s ]
31
-
32
-
33
- Exiting...
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
reproducibility/validation/smoke-tests-gpu1/K2-Horizon-0.9B-Q6_K.log DELETED
@@ -1,33 +0,0 @@
1
-
2
-
3
- Loading model...
4
-
5
- β–„β–„ β–„β–„
6
- β–ˆβ–ˆ β–ˆβ–ˆ
7
- β–ˆβ–ˆ β–ˆβ–ˆ β–€β–€β–ˆβ–„ β–ˆβ–ˆβ–ˆβ–„β–ˆβ–ˆβ–ˆβ–„ β–€β–€β–ˆβ–„ β–„β–ˆβ–ˆβ–ˆβ–ˆ β–ˆβ–ˆβ–ˆβ–ˆβ–„ β–ˆβ–ˆβ–ˆβ–ˆβ–„
8
- β–ˆβ–ˆ β–ˆβ–ˆ β–„β–ˆβ–€β–ˆβ–ˆ β–ˆβ–ˆ β–ˆβ–ˆ β–ˆβ–ˆ β–„β–ˆβ–€β–ˆβ–ˆ β–ˆβ–ˆ β–ˆβ–ˆ β–ˆβ–ˆ β–ˆβ–ˆ β–ˆβ–ˆ
9
- β–ˆβ–ˆ β–ˆβ–ˆ β–€β–ˆβ–„β–ˆβ–ˆ β–ˆβ–ˆ β–ˆβ–ˆ β–ˆβ–ˆ β–€β–ˆβ–„β–ˆβ–ˆ β–ˆβ–ˆ β–€β–ˆβ–ˆβ–ˆβ–ˆ β–ˆβ–ˆβ–ˆβ–ˆβ–€ β–ˆβ–ˆβ–ˆβ–ˆβ–€
10
- β–ˆβ–ˆ β–ˆβ–ˆ
11
- β–€β–€ β–€β–€
12
-
13
- build : b10671-35999d101
14
- model : /root/workspace/HF/quantized-k2-horizon-0.9b/K2-Horizon-0.9B-Q6_K.gguf
15
- ftype : Q6_K
16
- modalities : text
17
-
18
- available commands:
19
- /exit or Ctrl+C stop or exit
20
- /regen regenerate the last response
21
- /clear clear the chat history
22
- /read <file> add a text file
23
- /glob <pattern> add text files using globbing pattern
24
-
25
-
26
-
27
- > Answer in one sentence: what does a GGUF file contain?
28
- The user asks: "Answer in one sentence: what does a GGUF file contain?" They want a concise answer: a GGUF file contains a model's weights and metadata, typically a quantized version of a model, used for inference. So answer: "A GGUF file contains the quant
29
-
30
- [ Prompt: 1664.8 t/s | Generation: 289.9 t/s ]
31
-
32
-
33
- Exiting...
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
reproducibility/validation/smoke-tests-gpu1/K2-Horizon-0.9B-Q8_0.log DELETED
@@ -1,33 +0,0 @@
1
-
2
-
3
- Loading model...
4
-
5
- β–„β–„ β–„β–„
6
- β–ˆβ–ˆ β–ˆβ–ˆ
7
- β–ˆβ–ˆ β–ˆβ–ˆ β–€β–€β–ˆβ–„ β–ˆβ–ˆβ–ˆβ–„β–ˆβ–ˆβ–ˆβ–„ β–€β–€β–ˆβ–„ β–„β–ˆβ–ˆβ–ˆβ–ˆ β–ˆβ–ˆβ–ˆβ–ˆβ–„ β–ˆβ–ˆβ–ˆβ–ˆβ–„
8
- β–ˆβ–ˆ β–ˆβ–ˆ β–„β–ˆβ–€β–ˆβ–ˆ β–ˆβ–ˆ β–ˆβ–ˆ β–ˆβ–ˆ β–„β–ˆβ–€β–ˆβ–ˆ β–ˆβ–ˆ β–ˆβ–ˆ β–ˆβ–ˆ β–ˆβ–ˆ β–ˆβ–ˆ
9
- β–ˆβ–ˆ β–ˆβ–ˆ β–€β–ˆβ–„β–ˆβ–ˆ β–ˆβ–ˆ β–ˆβ–ˆ β–ˆβ–ˆ β–€β–ˆβ–„β–ˆβ–ˆ β–ˆβ–ˆ β–€β–ˆβ–ˆβ–ˆβ–ˆ β–ˆβ–ˆβ–ˆβ–ˆβ–€ β–ˆβ–ˆβ–ˆβ–ˆβ–€
10
- β–ˆβ–ˆ β–ˆβ–ˆ
11
- β–€β–€ β–€β–€
12
-
13
- build : b10671-35999d101
14
- model : /root/workspace/HF/quantized-k2-horizon-0.9b/K2-Horizon-0.9B-Q8_0.gguf
15
- ftype : Q8_0
16
- modalities : text
17
-
18
- available commands:
19
- /exit or Ctrl+C stop or exit
20
- /regen regenerate the last response
21
- /clear clear the chat history
22
- /read <file> add a text file
23
- /glob <pattern> add text files using globbing pattern
24
-
25
-
26
-
27
- > Answer in one sentence: what does a GGUF file contain?
28
- The user asks: "Answer in one sentence: what does a GGUF file contain?" They want a concise answer: a GGUF file contains a model's weights and metadata, typically a quantized version of a model, used for inference. So answer: "A GGUF file contains the quant
29
-
30
- [ Prompt: 1714.1 t/s | Generation: 265.1 t/s ]
31
-
32
-
33
- Exiting...