Quick answer: K2-Horizon-32B is IFM's large dense model in the K2-Horizon family — 32B parameters, all activated, with a native 512K-token context window. This is the Stage 1 training checkpoint; results are for stage 1 only, with stage 2 results expected later. Open weights, Apache 2.0.
Where K2-Horizon-32B leads
Where it lags
Best for: teams wanting to track K2-Horizon's dense-model trajectory before the Stage 2 checkpoint lands; not yet the strongest dense option in its size class.
K2-Horizon-32B is the large dense member of IFM's K2-Horizon family — a 32B-parameter decoder-only model evaluated on the same benchmark suite as its MoE siblings for apples-to-apples comparison. IFM explicitly labels current published results as "Stage 1" of the final model's training, with Stage 2 results to follow — a transparency choice that lets outside observers track capability changes across a checkpoint's training rather than judging only a finished model.
| Field | Value |
|---|---|
| Organization | IFM |
| Parameters | 32B (dense, all activated) |
| Context window | 524,288 tokens (512K) |
| Architecture | Dense decoder-only |
| License | Apache 2.0 |
| Release date | 2026-09 (Stage 1 checkpoint) |
| Modality | Text |
| Input (per 1M tokens) | Output (per 1M tokens) | |
|---|---|---|
| Open weights | — | — |
Open weights: free to download and self-host. Validated SGLang recipe: 2x H200, TP=2.
K2-Horizon-32B has a 524,288-token (512K) context window, native from the midtraining stages onward.
| Benchmark | Score | Source | Date |
|---|---|---|---|
| τ³ Banking | 22.5 | IFM model card | 2026-09 |
| TerminalBench 2.1 | 36.6 | IFM model card | 2026-09 |
| SciCode | 30.2 | IFM model card | 2026-09 |
| Humanity's Last Exam (no tools) | 22.8 | IFM model card | 2026-09 |
| GPQA Diamond | 82.3 | IFM model card | 2026-09 |
| CritPt | 1.4 | IFM model card | 2026-09 |
| AA-LCR | 65.3 | IFM model card | 2026-09 |
AA-Omniscience Accuracy (16.8) and Non-Hallucination (58.3) were also reported but use a different metric structure than the platform's combined AA-Omniscience Index — pending mapping decision, not yet added. Scores above are reported by IFM (Stage 1 checkpoint) and shown for context; not Benchgen measurements.
from openai import OpenAI
client = OpenAI(base_url="http://localhost:30000/v1", api_key="EMPTY")
response = client.chat.completions.create(
model="IFM/K2-Horizon-32B",
messages=[{"role": "user", "content": "Explain the result step by step."}],
temperature=1.0,
top_p=0.95,
extra_body={"chat_template_kwargs": {"reasoning_effort": "high"}},
)
print(response.choices[0].message.content)Specs and scores sourced from IFM's official Hugging Face model card; third-party benchmark scores attributed inline. Last updated 2026-09-03.