Quick answer: K2-Horizon-MoVA-36B-A4B is IFM's sparse Mixture-of-Experts model using Mixture-of-Values attention (MoVA) — 36B total parameters, only 4B active per token, with a native 512K-token context window. It outscores open-weight dense and MoE models up to 15x its active-parameter size on agentic and reasoning benchmarks, and is competitive with closed frontier models. Open weights, Apache 2.0.
Where K2-Horizon-MoVA-36B-A4B leads
Where it lags
Best for: cost-sensitive deployments needing strong agentic tool-use and terminal capability without the inference cost of a 30B+ active-parameter model.
K2-Horizon-MoVA-36B-A4B is the efficiency-focused member of IFM's K2-Horizon family. It uses a novel Mixture-of-Values attention (MoVA) mechanism paired with sparse MoE routing to store 36B parameters while activating just 4B per token — a ~9x sparsity ratio. IFM's own framing is that at this activated-parameter budget it competes with models up to 15x larger, positioning it as an efficiency play for agentic and reasoning workloads rather than a raw-parameter-count flagship.
| Field | Value |
|---|---|
| Organization | IFM |
| Parameters | 36B total / 4B active |
| Context window | 524,288 tokens (512K) |
| Architecture | Sparse MoE with Mixture-of-Values attention (MoVA) |
| License | Apache 2.0 |
| Release date | 2026-09 |
| Modality | Text |
| Input (per 1M tokens) | Output (per 1M tokens) | |
|---|---|---|
| Open weights | — | — |
Open weights: free to download and self-host. Validated SGLang recipe: 2x H200, TP=2, EP=2.
K2-Horizon-MoVA-36B-A4B has a 524,288-token (512K) context window, native from the midtraining stages onward — enabling long-document and long-horizon-agent workloads at a fraction of the active-parameter cost of its dense-context peers.
| Benchmark | Score | Source | Date |
|---|---|---|---|
| τ³ Banking | 26.8 | IFM model card | 2026-09 |
| TerminalBench 2.1 | 58.6 | IFM model card | 2026-09 |
| SciCode | 38.9 | IFM model card | 2026-09 |
| Humanity's Last Exam (no tools) | 25.2 | IFM model card | 2026-09 |
| GPQA Diamond | 80.8 | IFM model card | 2026-09 |
| CritPt | 2.1 | IFM model card | 2026-09 |
| AA-LCR | 66.3 | IFM model card | 2026-09 |
AA-Omniscience Accuracy (18.8) and Non-Hallucination (69.2) were also reported but use a different metric structure than the platform's combined AA-Omniscience Index — pending mapping decision, not yet added. Scores above are reported by IFM and shown for context; not Benchgen measurements.
from openai import OpenAI
client = OpenAI(base_url="http://localhost:30000/v1", api_key="EMPTY")
response = client.chat.completions.create(
model="IFM/K2-Horizon-MoVA-36B-A4B",
messages=[{"role": "user", "content": "Explain the result step by step."}],
temperature=1.0,
top_p=0.95,
extra_body={"chat_template_kwargs": {"reasoning_effort": "high"}},
)
print(response.choices[0].message.content)Specs and scores sourced from IFM's official Hugging Face model card; third-party benchmark scores attributed inline. Last updated 2026-09-03.