Benchgen
Models/ifm/

K2 Horizon 7B

DraftPublic

Model Details

K2-Horizon-7B

Organization Context Pricing License Modality Released

Quick answer: K2-Horizon-7B is IFM's medium dense model — a 7B-core decoder-only model with a native 512K-token context window. It beats Gemma 4-12B, Qwen3.5-9B, and IBM Granite 4.2-8B on SWE-bench Verified (68.4%), HMMT Feb 2026 (73.3%), and Terminal-Bench 2.1 (39.1%), despite being smaller than several of its comparison models. Open weights, Apache 2.0.

At a Glance

Where K2-Horizon-7B leads

  • Strong SWE-bench Verified score (68.4%) — well ahead of larger dense peers like Gemma 4-12B (30.6%) and Qwen3.5-9B (50.8%)
  • Leads its comparison set on HMMT Feb 2026 competition math (73.3%) and long-context reasoning (AA-LCR, 68.0%)
  • 512K native context in a 7B-class model

Where it lags

  • BrowseComp (59.0%) is close to but not ahead of GPT-5 (54.9%) and LongCat Flash Thinking (56.6%) — a tighter race than other benchmarks
  • Smaller model, so absolute scores on frontier-class benchmarks (HLE: 18.6%) remain modest in absolute terms even where relatively strong

Best for: teams wanting strong coding/agentic capability (especially SWE-bench-style tasks) in a compact, self-hostable 7B footprint.

What K2-Horizon-7B Is

K2-Horizon-7B is the medium-sized dense model in IFM's K2-Horizon family — a 7B-core decoder-only architecture evaluated across agentic, coding, long-context, and reasoning benchmarks. It ships with intermediate checkpoints released alongside the final model, plus public training data, recipe, training code, and evaluation resources — a notably high transparency bar for a model this size. Its comparative strength on SWE-bench Verified and HMMT suggests it was tuned with real coding and competition-math workloads in mind rather than only broad instruction-following.

Specifications

FieldValue
OrganizationIFM
Parameters7B (dense)
Context window524,288 tokens (512K)
ArchitectureDense decoder-only
LicenseApache 2.0
Release date2026-09
ModalityText

Pricing

Input (per 1M tokens)Output (per 1M tokens)
Open weights——

Open weights: free to download and self-host. Validated SGLang/vLLM recipe: single-GPU (TP=1).

Context Window

K2-Horizon-7B has a 524,288-token (512K) context window, native from the midtraining stages onward — long-context reasoning (AA-LCR) scores 68.0%, ahead of same-class peers.

Public Benchmark Scores

Scores above are reported by IFM and shown for context; not Benchgen measurements. BrowseComp: IFM used the Discard-all@95k context-length protocol from the DeepSeek-V3.2 technical report; comparison models may use different harnesses.

Use K2-Horizon-7B via API

from openai import OpenAI

client = OpenAI(base_url="http://localhost:30000/v1", api_key="EMPTY")

response = client.chat.completions.create(
    model="IFM/K2-Horizon-7B",
    messages=[{"role": "user", "content": "Explain the result step by step."}],
    temperature=1.0,
    top_p=0.95,
    max_tokens=32768,
    extra_body={"chat_template_kwargs": {"reasoning_effort": "high"}},
)
print(response.choices[0].message.content)

Frequently Asked Questions

What is K2-Horizon-7B? It's IFM's medium dense model in the K2-Horizon family: 7B parameters, 512K-token context, with strong SWE-bench Verified and HMMT scores relative to its size.
What is K2-Horizon-7B's context window? 524,288 tokens (512K).
How much does K2-Horizon-7B cost? Open weights — free to download; cost is hosting/inference only.
Is K2-Horizon-7B open source? Yes, released under the Apache 2.0 license.

Specs and scores sourced from IFM's official Hugging Face model card; third-party benchmark scores attributed inline. Last updated 2026-09-03.