Quick answer: K2-Horizon-7B is IFM's medium dense model — a 7B-core decoder-only model with a native 512K-token context window. It beats Gemma 4-12B, Qwen3.5-9B, and IBM Granite 4.2-8B on SWE-bench Verified (68.4%), HMMT Feb 2026 (73.3%), and Terminal-Bench 2.1 (39.1%), despite being smaller than several of its comparison models. Open weights, Apache 2.0.
Where K2-Horizon-7B leads
Where it lags
Best for: teams wanting strong coding/agentic capability (especially SWE-bench-style tasks) in a compact, self-hostable 7B footprint.
K2-Horizon-7B is the medium-sized dense model in IFM's K2-Horizon family — a 7B-core decoder-only architecture evaluated across agentic, coding, long-context, and reasoning benchmarks. It ships with intermediate checkpoints released alongside the final model, plus public training data, recipe, training code, and evaluation resources — a notably high transparency bar for a model this size. Its comparative strength on SWE-bench Verified and HMMT suggests it was tuned with real coding and competition-math workloads in mind rather than only broad instruction-following.
| Field | Value |
|---|---|
| Organization | IFM |
| Parameters | 7B (dense) |
| Context window | 524,288 tokens (512K) |
| Architecture | Dense decoder-only |
| License | Apache 2.0 |
| Release date | 2026-09 |
| Modality | Text |
| Input (per 1M tokens) | Output (per 1M tokens) | |
|---|---|---|
| Open weights | — | — |
Open weights: free to download and self-host. Validated SGLang/vLLM recipe: single-GPU (TP=1).
K2-Horizon-7B has a 524,288-token (512K) context window, native from the midtraining stages onward — long-context reasoning (AA-LCR) scores 68.0%, ahead of same-class peers.
| Benchmark | Score | Source | Date |
|---|---|---|---|
| HMMT 2026 | 73.3 | IFM model card | 2026-09 |
| SWE-bench Verified | 68.4 | IFM model card | 2026-09 |
| Humanity's Last Exam | 18.6 | IFM model card | 2026-09 |
| SciCode | 31.6 | IFM model card | 2026-09 |
| AA-LCR | 68.0 | IFM model card | 2026-09 |
| TerminalBench 2.1 | 39.1 | IFM model card | 2026-09 |
| τ³ Banking | 25.8 | IFM model card | 2026-09 |
| BrowseComp | 59.0 | IFM model card | 2026-09 |
Scores above are reported by IFM and shown for context; not Benchgen measurements. BrowseComp: IFM used the Discard-all@95k context-length protocol from the DeepSeek-V3.2 technical report; comparison models may use different harnesses.
from openai import OpenAI
client = OpenAI(base_url="http://localhost:30000/v1", api_key="EMPTY")
response = client.chat.completions.create(
model="IFM/K2-Horizon-7B",
messages=[{"role": "user", "content": "Explain the result step by step."}],
temperature=1.0,
top_p=0.95,
max_tokens=32768,
extra_body={"chat_template_kwargs": {"reasoning_effort": "high"}},
)
print(response.choices[0].message.content)Specs and scores sourced from IFM's official Hugging Face model card; third-party benchmark scores attributed inline. Last updated 2026-09-03.