--- base_model: IFM/K2-Horizon-MoVA-36B-A4B tags: - mlx - apple-silicon - text-generation - oQ license: apache-2.0 --- # K2-Horizon-MoVA-36B-A4B MLX-6bit 6-bit uniform quantization conversion of [IFM/K2-Horizon-MoVA-36B-A4B](https://huggingface.co/IFM/K2-Horizon-MoVA-36B-A4B), a sparse Mixture-of-Experts model with Mixture-of-Values attention (36B total / 4B active parameters, 512K context). **Upstream model:** [IFM/K2-Horizon-MoVA-36B-A4B](https://huggingface.co/IFM/K2-Horizon-MoVA-36B-A4B) by the IFM Team, released under Apache-2.0. **Conversion:** Quantized to MLX format using [Hermes Agent](https://hermes-agent.nousresearch.com) with `mlx-lm` and `oMLX`. ## Quickstart ```bash pip install -U mlx-lm python3 -m mlx_lm.generate --model hermitdave/K2-Horizon-MoVA-36B-A4B-MLX-6bit --prompt "Explain why long-context evaluation is difficult." --max-tokens 512 --temp 1.0 --top-p 0.95 ``` ## Reasoning K2-Horizon is a reasoning model. Always use `reasoning_effort="high"` for best results: ```python from openai import OpenAI client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY") response = client.chat.completions.create( model="hermitdave/K2-Horizon-MoVA-36B-A4B-MLX-6bit", messages=[{"role": "user", "content": "Explain quantum entanglement."}], extra_body={"chat_template_kwargs": {"reasoning_effort": "high"}}, ) print("Reasoning:", getattr(response.choices[0].message, "reasoning_content", None)) print("Answer:", response.choices[0].message.content) ``` ## Benchmark Results | Benchmark | K2-Horizon-MoVA-36B-A4B | |-----------|------------------------| | tau3-Banking (Agentic tool use) | **26.8** | | Terminal-Bench 2.1 (Agentic terminal use) | **58.6** | | GPQA Diamond (Graduate-level science QA) | 80.8 | | AA-LCR (Long-context reasoning) | 66.3 | Scores in %. See [model card](https://huggingface.co/IFM/K2-Horizon-MoVA-36B-A4B) for full results. ## oMLX Patch K2-Horizon requires oMLX v0.6.4+ with the [K2-Horizon support patch (PR #3441)](https://github.com/jundot/omlx/pull/3441). This patch adds: - `k2_horizon` model type support - Reasoning content handling (`` tags) - Tool call parsing (plain text and XML formats) - Multi-turn conversation support Without this patch, oMLX will refuse to load K2-Horizon models with `ValueError: Model type k2_horizon not supported`. ## Chat Template K2-Horizon uses IFM's custom chat template with reasoning and tool calling support. Key tags: | Tag | Purpose | |-----|---------| | ``, ``, `` | Thinking blocks | | `<\|ifm\|im_start|>`, `<\|ifm\|im_end\|>` | Message delimiters | | ``, ``, `` | Tool call structure | All tags are automatically stripped by oMLX before responses reach users. ## Citation ```bibtex @misc{k2horizon2026, title = {Introducing K2 Horizon: Frontier Performance, Radically Open}, author = {{IFM Team}}, year = {2026}, url = {https://ifm.ai/blog/k2/}, } ``` ## License Apache-2.0 (same as upstream).