How to use from the
Use from the
MLX library
# Make sure mlx-lm is installed
# pip install --upgrade mlx-lm

# Generate text with mlx-lm
from mlx_lm import load, generate

model, tokenizer = load("hermitdave/K2-Horizon-MoVA-36B-A4B-MLX-6bit")

prompt = "Write a story about Einstein"
messages = [{"role": "user", "content": prompt}]
prompt = tokenizer.apply_chat_template(
    messages, add_generation_prompt=True
)

text = generate(model, tokenizer, prompt=prompt, verbose=True)

K2-Horizon-MoVA-36B-A4B MLX-6bit

6-bit uniform quantization conversion of IFM/K2-Horizon-MoVA-36B-A4B, a sparse Mixture-of-Experts model with Mixture-of-Values attention (36B total / 4B active parameters, 512K context).

Upstream model: IFM/K2-Horizon-MoVA-36B-A4B by the IFM Team, released under Apache-2.0.

Conversion: Quantized to MLX format using Hermes Agent with mlx-lm and oMLX.

Quickstart

pip install -U mlx-lm

python3 -m mlx_lm.generate   --model hermitdave/K2-Horizon-MoVA-36B-A4B-MLX-6bit   --prompt "Explain why long-context evaluation is difficult."   --max-tokens 512 --temp 1.0 --top-p 0.95

Reasoning

K2-Horizon is a reasoning model. Always use reasoning_effort="high" for best results:

from openai import OpenAI

client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")
response = client.chat.completions.create(
    model="hermitdave/K2-Horizon-MoVA-36B-A4B-MLX-6bit",
    messages=[{"role": "user", "content": "Explain quantum entanglement."}],
    extra_body={"chat_template_kwargs": {"reasoning_effort": "high"}},
)
print("Reasoning:", getattr(response.choices[0].message, "reasoning_content", None))
print("Answer:", response.choices[0].message.content)

Benchmark Results

Benchmark K2-Horizon-MoVA-36B-A4B
tau3-Banking (Agentic tool use) 26.8
Terminal-Bench 2.1 (Agentic terminal use) 58.6
GPQA Diamond (Graduate-level science QA) 80.8
AA-LCR (Long-context reasoning) 66.3

Scores in %. See model card for full results.

oMLX Patch

K2-Horizon requires oMLX v0.6.4+ with the K2-Horizon support patch (PR #3441). This patch adds:

  • k2_horizon model type support
  • Reasoning content handling (<ifm|think> tags)
  • Tool call parsing (plain text and XML formats)
  • Multi-turn conversation support

Without this patch, oMLX will refuse to load K2-Horizon models with ValueError: Model type k2_horizon not supported.

Chat Template

K2-Horizon uses IFM's custom chat template with reasoning and tool calling support. Key tags:

Tag Purpose
<ifm|think>, <ifm|think_fast>, <ifm|think_faster> Thinking blocks
`<|ifm|im_start >, <|ifm|im_end|>`
<ifm|tool_call>, <ifm|arg_key>, <ifm|arg_value> Tool call structure

All tags are automatically stripped by oMLX before responses reach users.

Citation

@misc{k2horizon2026,
  title  = {Introducing K2 Horizon: Frontier Performance, Radically Open},
  author = {{IFM Team}},
  year   = {2026},
  url    = {https://ifm.ai/blog/k2/},
}

License

Apache-2.0 (same as upstream).

Downloads last month
337
Safetensors
Model size
37B params
Tensor type
U32
·
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

6-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for hermitdave/K2-Horizon-MoVA-36B-A4B-MLX-6bit

Quantized
(18)
this model