---
pipeline_tag: text-generation
library_name: transformers
model_name: K2-Horizon-7B
language:
- en
license: apache-2.0
datasets:
- IFM/K2-Horizon-Pretrain-Data
- IFM/K2-Horizon-Midtrain-Data
tags:
- k2-horizon
- 7b
- dense
- open-weights
- ifm
---
# K2-Horizon-7B
K2-Horizon-7B is the medium dense member of the K2-Horizon family: a 7B-core decoder-only model with a 512K context window.
## K2-Horizon-7B Highlights
- **Strong dense baseline.** A 7B-class dense model evaluated across agentic, coding, long-context, and reasoning benchmarks.
- **512K context.** Native 524,288-token context from the midtraining stages onward.
- **Intermediate checkpoints.** Intermediate checkpoints are released so capability changes can be studied across training rather than at a single checkpoint.
- **Fully open.** Training data and recipe, training code, and evaluation resources are public.
## Benchmark Results
The chart at the top of this card shows K2-Horizon-7B against selected reference models. The table below lists every comparison model used in the figure.
### Full Results
| | Reference models ยท weak to strong |
|---|
| Benchmark | K2-Horizon-7B | Reference 1 | Reference 2 | Reference 3 |
|---|
| Math |
HMMT Feb 2026 Competition mathematics | 73.3 | Gemma 4-12B 63.1 | Qwen3.5-9B 65.7 | Granite 4.2-8B 66.5 |
| Coding |
SWE-bench Verified Software engineering | 70.6 | Gemma 4-12B 30.6 | Granite 4.2-8B 47.7 | Qwen3.5-9B 50.8 |
| Scientific Reasoning |
HLE Expert-level reasoning | 18.6 | Granite 4.2-8B 9.7 | Qwen3.5-9B 14.9 | Gemma 4-12B 15.7 |
| Coding |
SciCode Scientific coding | 31.6 | Qwen3.5-9B 27.5 | Mistral Small 4 28.0 | Granite 4.2-8B 30.4 |
| General |
LCR Long-context reasoning | 68.0 | Granite 4.2-8B 43.3 | Gemma 4-12B 61.7 | Qwen3.5-9B 65.3 |
| Coding |
Terminal-Bench 2.1 Agentic terminal use | 39.1 | Granite 4.2-8B 18.4 | Gemma 4-12B 27.3 | Qwen3.5-9B 29.2 |
| Agents |
tau3-Banking Agentic tool use | 25.8 | Qwen3.5-9B 7.0 | Granite 4.2-8B 7.6 | Muse Glimmer-30B 24.0 |
BrowseComp Web browsing | 59.0 | DeepSeek V4 Flash-0423 53.5 | GPT-5 54.9 | LongCat Flash Thinking-2601 56.6 |
Scores in %. Bold marks the best score in each row. BrowseComp: our model uses the Discard-all@95k context-length protocol proposed in the DeepSeek-V3.2 technical report; comparison models may use different harnesses.
## Quickstart
### Serving
vLLM, recipe at [recipes.vllm.ai/IFM](https://recipes.vllm.ai/IFM):
```shell
vllm serve IFM/K2-Horizon-7B \
--trust-remote-code \
--dtype bfloat16 \
--max-model-len 131072 \
--tensor-parallel-size 1 \
--reasoning-parser k2_horizon \
--enable-auto-tool-choice \
--tool-call-parser k2_horizon
```
SGLang, this is the recipe validated in the [SGLang K2 Horizon cookbook](https://docs.sglang.io/cookbook/autoregressive/IFM/K2-Horizon):
```shell
sglang serve \
--model-path IFM/K2-Horizon-7B \
--revision 69ada542b68fe13d767479db2ab9421baff88681 \
--tp 1 \
--dtype bfloat16 \
--attention-backend fa3 \
--reasoning-parser k2_horizon \
--host 0.0.0.0 \
--port 30000
```
### API Usage
> [!Tip]
> Recommended settings: `reasoning_effort="high"`, `temperature=1.0`, `top_p=0.95`, and at least 32,768 output tokens.
> Reasoning depth is selected per request through `chat_template_kwargs`. Thinking is returned in `reasoning_content` and the answer in `content`.
```python
from openai import OpenAI
client = OpenAI(base_url="http://localhost:30000/v1", api_key="EMPTY")
response = client.chat.completions.create(
model="IFM/K2-Horizon-7B",
messages=[{"role": "user", "content": "Explain the result step by step."}],
temperature=1.0,
top_p=0.95,
max_tokens=32768,
extra_body={"chat_template_kwargs": {"reasoning_effort": "high", "tool_call_format": "xml"}},
)
message = response.choices[0].message
print("Reasoning:", getattr(message, "reasoning_content", None))
print("Answer:", message.content)
```
Our model supports multiple tool calls formats, which can be changed with `chat_template_kwargs`. The supported values are `json`, `xml`, and `xml_typed` . The default is `xml`. Keep `--tool-call-parser k2_horizon` enabled to parse the selected format.
### Transformers
Validated with Transformers 5.15.0, PyTorch 2.13.0, Safetensors 0.8.0.
```python
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "IFM/K2-Horizon-7B"
tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
model_id, device_map="auto", dtype="bfloat16", low_cpu_mem_usage=True, trust_remote_code=True
)
inputs = tokenizer("Explain why long-context evaluation is difficult.", return_tensors="pt").to(model.device)
inputs.pop("token_type_ids", None)
outputs = model.generate(**inputs, max_new_tokens=32768, temperature=1.0, top_p=0.95, do_sample=True)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
```
## Best Practices
1. **Reasoning effort: always `high`.** All reported results use high reasoning effort. Pass `{"chat_template_kwargs": {"reasoning_effort": "high"}}` on every request; `medium` and `low` trade accuracy for speed and are not recommended for evaluation.
2. **Sampling parameters.** `temperature=1.0`, `top_p=0.95`.
3. **Output length.** Allow at least 32,768 output tokens so reasoning is never cut off. Truncated reasoning is a failed response, not a shorter one.
4. **Serving.** Use the validated SGLang recipe above: BF16, TP=1, FlashAttention-3. Full recipes for every K2-Horizon size, with measured H200 latency and throughput, are in the [SGLang cookbook](https://docs.sglang.io/cookbook/autoregressive/IFM/K2-Horizon).
5. **Parsers.** Enable the `k2_horizon` reasoning parser for chat, and add the `k2_horizon` tool-call parser for agent use. Leave both off for plain completion-style generation.
6. **Revisions.** Pin a revision tag when reproducibility matters. `main` is the default checkpoint; `base_final` and the `mid_*_final` tags identify training stages.
## Citation
```bibtex
@misc{k2horizon2026,
title = {Introducing K2 Horizon: Frontier Performance, Radically Open},
author = {{IFM Team}},
year = {2026},
url = {https://ifm.ai/blog/k2/},
}
```