---
pipeline_tag: text-generation
library_name: transformers
model_name: K2-Horizon-0.9B
language:
- en
- zh
license: apache-2.0
license_name: internal-only
license_link: LICENSE
tags:
- k2-horizon
- 0.9b
- dense
- reasoning
- knowledge-distillation
- ifm
---
# K2-Horizon-0.9B
K2-Horizon-0.9B is the compact dense member of the K2-Horizon family: a 0.9B-class decoder-only model with a 128K context window.
## K2-Horizon-0.9B Highlights
- **Compact reasoning model.** A 0.9B-class dense model evaluated across mathematics, coding, science, and tool-use benchmarks.
- **128K context.** Supports up to 131,072 tokens with YaRN RoPE scaling.
- **Multi-teacher distillation.** Trained with domain teachers for math and code, STEM, and instruction following.
- **Fully open.** Training data/recipe and the training code will be made public.
## Benchmark Results
The chart at the top of this card shows K2-Horizon-0.9B against selected reference models. The table below lists every comparison model used in the figure.
### Full Results
| Reference models |
|---|
| K2-Horizon-0.9B | Qwen3.5-0.8B | OpenBMB-1B | Qwen3.5-2B |
|---|
| # Params | 0.9B | 0.8B | 1B | 2B |
| # Activated params | 0.9B | 0.8B | 1B | 2B |
| Architecture | Dense | Dense | Dense | Dense |
| Math |
AIME 2025 Competition mathematics | 41.7 | 1.0 | 40.4 | 34.2 |
AIME 2026 Competition mathematics | 48.5 | 0.2 | 40.4 | 38.8 |
HMMT Feb 2026 Competition mathematics | 25.8 | 0.6 | 23.3 | 22.7 |
| Scientific Reasoning |
GPQA Diamond Graduate-level science QA | 27.3 | 11.9 | 26.3 | 54.9 |
| Coding |
HumanEval+ Code generation | 79.9 | 16.5 | 65.2 | 75.6 |
MBPP+ Code generation | 68.0 | 35.4 | 60.6 | 67.7 |
LiveCodeBench v6 Competitive coding | 37.4 | 6.6 | 33.5 | 29.8 |
| Agents |
BFCL v4 Function calling | 28.0 | 25.3 | 25.2 | 43.6 |
Scores in %. Bold highlights K2-Horizon-0.9B; Qwen3.5-2B is included as a larger reference model. Protocol and provenance details are in the [Technical Appendix](APPENDIX.md#evaluation).
## Quickstart
### Serving
vLLM (source at [PR #53806](https://github.com/vllm-project/vllm/pull/53806), commit `d9fd5f11`):
```shell
vllm serve IFM/K2-Horizon-0.9B \
--trust-remote-code \
--dtype bfloat16 \
--max-model-len 131072 \
--hf-overrides '{"rope_parameters":{rope_type: yarn, factor: 16, original_max_position_embeddings: 8192, rope_theta: 1000000, beta_fast: 128, beta_slow: 4}' \
--gpu-memory-utilization 0.85 \
--tensor-parallel-size 1 \
--reasoning-parser k2_horizon \
--enable-auto-tool-choice \
--tool-call-parser k2_horizon
```
SGLang, from a source checkout that includes [sgl-project/sglang#37654](https://github.com/sgl-project/sglang/pull/37654). This is the recipe validated in the [SGLang K2 Horizon cookbook](https://docs.sglang.io/cookbook/autoregressive/IFM/K2-Horizon):
```shell
sglang serve \
--model-path IFM/K2-Horizon-0.9B \
--revision 9b9ec1f7e17f62ed218df542687a144116219d84 \
--tp 1 \
--dtype bfloat16 \
--attention-backend fa3 \
--reasoning-parser k2_horizon \
--host 0.0.0.0 \
--port 30000
```
### API Usage
> [!Tip]
> Recommended settings: `reasoning_effort="high"`, `temperature=0.6`, `top_p=0.95`, and at least 32,768 output tokens.
> Reasoning depth is selected per request through `chat_template_kwargs`. Thinking is returned in `reasoning_content` and the answer in `content`.
```python
from openai import OpenAI
client = OpenAI(base_url="http://localhost:30000/v1", api_key="EMPTY")
response = client.chat.completions.create(
model="IFM/K2-Horizon-0.9B",
messages=[{"role": "user", "content": "Explain the result step by step."}],
temperature=0.6,
top_p=0.95,
max_tokens=32768,
extra_body={"chat_template_kwargs": {"reasoning_effort": "high"}},
)
message = response.choices[0].message
print("Reasoning:", getattr(message, "reasoning_content", None))
print("Answer:", message.content)
```
### Transformers
Validated with Transformers 5.15.0, PyTorch 2.13.0, Safetensors 0.8.0.
```python
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "IFM/K2-Horizon-0.9B"
tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
model_id, device_map="auto", dtype="bfloat16", low_cpu_mem_usage=True, trust_remote_code=True
)
inputs = tokenizer("Explain why long-context evaluation is difficult.", return_tensors="pt").to(model.device)
inputs.pop("token_type_ids", None)
outputs = model.generate(**inputs, max_new_tokens=32768, temperature=1.0, top_p=0.95, do_sample=True)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
```
## Best Practices
1. **Reasoning effort: always `high`.** All reported results use high reasoning effort. Pass `{"chat_template_kwargs": {"reasoning_effort": "high"}}` on every request; `medium` and `low` trade accuracy for speed and are not recommended for evaluation.
2. **Sampling parameters.** `temperature=0.6`, `top_p=0.95`.
3. **Output length.** Allow at least 32,768 output tokens so reasoning is never cut off. Truncated reasoning is a failed response, not a shorter one.
4. **Serving.** Use the validated SGLang recipe above: BF16, TP=1, FlashAttention-3. Full recipes for every K2-Horizon size, with measured H200 latency and throughput, are in the [SGLang cookbook](https://docs.sglang.io/cookbook/autoregressive/IFM/K2-Horizon).
5. **Parsers.** Enable the `k2_horizon` reasoning parser for chat, and add the `k2_horizon` tool-call parser for agent use. Leave both off for plain completion-style generation.
6. **Revisions.** `main` is the MOPD release checkpoint; `mid1_75k` and `mid2_47k` preserve the context-extension stages.
## Citation
```bibtex
@misc{k2horizon2026,
title = {Introducing K2 Horizon: Frontier Performance, Radically Open},
author = {{IFM Team}},
year = {2026},
url = {https://ifm.ai/blog/k2/},
}
```