Quick answer: K2-Horizon-3.7B is IFM's small dense model — 3.7B parameters with a native 512K-token context window. It leads its comparison set (Qwen3.5-4B, G9v3-3B, Granite 4.2-3B, Nemotron 3 Nano-4B) on SWE-bench Verified (68.6%) and HMMT Feb 2026 (70.5%), though it trails on GPQA Diamond and BFCL v4 function calling. Open weights, Apache 2.0.
Where K2-Horizon-3.7B leads
Where it lags
Best for: lightweight, self-hosted coding-agent workloads where SWE-bench-style task completion matters more than broad science QA.
K2-Horizon-3.7B is the small dense model in IFM's K2-Horizon family, evaluated on the same agentic, coding, and reasoning benchmark suite as its larger siblings. Like the 7B model, it ships with intermediate checkpoints and public training data/recipe/code. Its benchmark profile — strong on SWE-bench and competition math, comparatively weaker on GPQA and function-calling — suggests targeted tuning toward coding-agent tasks rather than uniform improvement across all domains.
| Field | Value |
|---|---|
| Organization | IFM |
| Parameters | 3.7B (dense) |
| Context window | 524,288 tokens (512K) |
| Architecture | Dense decoder-only |
| License | Apache 2.0 |
| Release date | 2026-09 |
| Modality | Text |
| Input (per 1M tokens) | Output (per 1M tokens) | |
|---|---|---|
| Open weights | — | — |
Open weights: free to download and self-host. Validated SGLang/vLLM recipe: single-GPU (TP=1).
K2-Horizon-3.7B has a 524,288-token (512K) context window, native from the midtraining stages onward.
| Benchmark | Score | Source | Date |
|---|---|---|---|
| HMMT 2026 | 70.5 | IFM model card | 2026-09 |
| SWE-bench Verified | 68.6 | IFM model card | 2026-09 |
| GPQA Diamond | 65.4 | IFM model card | 2026-09 |
| Humanity's Last Exam | 12.9 | IFM model card | 2026-09 |
| SciCode | 25.9 | IFM model card | 2026-09 |
| TerminalBench 2.1 | 25.1 | IFM model card | 2026-09 |
| τ³ Banking | 17.7 | IFM model card | 2026-09 |
| BFCL v4 | 50.9 | IFM model card | 2026-09 |
Scores above are reported by IFM and shown for context; not Benchgen measurements. Baseline protocols for comparison models may differ.
from openai import OpenAI
client = OpenAI(base_url="http://localhost:30000/v1", api_key="EMPTY")
response = client.chat.completions.create(
model="IFM/K2-Horizon-3.7B",
messages=[{"role": "user", "content": "Explain the result step by step."}],
temperature=1.0,
top_p=0.95,
max_tokens=32768,
extra_body={"chat_template_kwargs": {"reasoning_effort": "high"}},
)
print(response.choices[0].message.content)Specs and scores sourced from IFM's official Hugging Face model card; third-party benchmark scores attributed inline. Last updated 2026-09-03.