--- base_model: IFM/K2-Horizon-7B base_model_relation: quantized license: apache-2.0 language: - en library_name: gguf pipeline_tag: text-generation tags: - gguf - llama.cpp - k2-horizon - long-context - 512k-context - dense --- > [!IMPORTANT] > **Compatibility:** These GGUF files require a `llama.cpp` build with K2 Horizon architecture support. Until upstream support lands, use the [MBZUAI-IFM fork](https://github.com/MBZUAI-IFM/llama.cpp/tree/model/K2Horizon). # K2-Horizon-7B GGUF GGUF quantizations of [IFM/K2-Horizon-7B](https://huggingface.co/IFM/K2-Horizon-7B), a 7B dense decoder-only model for reasoning, coding, long-context work, and tool use. The source checkpoint supports a native context length of **524,288 tokens (512K)**. ## Benchmarks ![K2-Horizon-7B benchmark results](assets/k2-horizon-7b-benchmarks.png) *Benchmark results reported by IFM for the original K2-Horizon-7B checkpoint.* ## GGUF files | Quantization | File | Size | | --- | --- | ---: | | Q4_0 | [K2-Horizon-7B-Q4_0.gguf](K2-Horizon-7B-Q4_0.gguf) | 5.34 GB | | Q4_K_M | [K2-Horizon-7B-Q4_K_M.gguf](K2-Horizon-7B-Q4_K_M.gguf) | 5.59 GB | | Q4_K_M Selective | [K2-Horizon-7B-Q4_K_M-Selective.gguf](K2-Horizon-7B-Q4_K_M-Selective.gguf) | 5.96 GB | | Q5_K_M | [K2-Horizon-7B-Q5_K_M.gguf](K2-Horizon-7B-Q5_K_M.gguf) | 6.47 GB | | Q6_K | [K2-Horizon-7B-Q6_K.gguf](K2-Horizon-7B-Q6_K.gguf) | 7.39 GB | | Q8_0 | [K2-Horizon-7B-Q8_0.gguf](K2-Horizon-7B-Q8_0.gguf) | 9.57 GB | The files are text-only GGUFs; no vision projector is required. The selective variant uses a Q4_K_M baseline with attention Q/K/V/O projection tensors kept at Q6_K; it is a manual tensor-selective build and does not use an importance matrix. SHA-256 checksums are provided in [`SHA256SUMS.txt`](SHA256SUMS.txt). ## Chat template Each GGUF embeds the llama.cpp-compatible chat template. [`chat_template.jinja`](chat_template.jinja) is a matching external copy for tools that require one. The original source template is retained as [`chat_template.upstream.jinja`](chat_template.upstream.jinja) for runtimes with full Jinja support. `xml` is the default tool-call format. Use `--chat-template-kwargs` to select `json` or `xml_typed` when required. ## Usage Use the [IFM K2 Horizon llama.cpp fork](https://github.com/MBZUAI-IFM/llama.cpp/tree/model/K2Horizon). The example below uses a practical 128K context; `-c 524288` can be used when the available memory is sufficient. ```bash llama-cli \ -m K2-Horizon-7B-Q4_K_M.gguf \ -c 131072 --jinja \ --temp 1.0 --top-p 0.95 ``` ## Source - Source model: [IFM/K2-Horizon-7B](https://huggingface.co/IFM/K2-Horizon-7B) - Source revision: [`2c9659a`](https://huggingface.co/IFM/K2-Horizon-7B/commit/2c9659a84c4eea6f9f60462221fe762c8c84d75c) - Source license: [Apache-2.0](https://www.apache.org/licenses/LICENSE-2.0)