Benchgen
Models/ifm/

K2 Horizon 3.7B

DraftPublic

Model Details

K2-Horizon-3.7B

Organization Context Pricing License Modality Released

Quick answer: K2-Horizon-3.7B is IFM's small dense model — 3.7B parameters with a native 512K-token context window. It leads its comparison set (Qwen3.5-4B, G9v3-3B, Granite 4.2-3B, Nemotron 3 Nano-4B) on SWE-bench Verified (68.6%) and HMMT Feb 2026 (70.5%), though it trails on GPQA Diamond and BFCL v4 function calling. Open weights, Apache 2.0.

At a Glance

Where K2-Horizon-3.7B leads

  • Dominant SWE-bench Verified score (68.6%) versus all four comparison models (next best: Qwen3.5-4B at 41.2%)
  • Strong competition-math result on HMMT Feb 2026 (70.5%), well ahead of Qwen3.5-4B (61.6%)
  • 512K context in a sub-4B model

Where it lags

  • GPQA Diamond (65.4%) trails Qwen3.5-4B (77.1%)
  • BFCL v4 function calling (50.9%) trails Qwen3.5-4B (55.7%) and Granite 4.2-3B (50.8%) narrowly
  • τ³-Banking data is thin for comparison models (two show no reported score)

Best for: lightweight, self-hosted coding-agent workloads where SWE-bench-style task completion matters more than broad science QA.

What K2-Horizon-3.7B Is

K2-Horizon-3.7B is the small dense model in IFM's K2-Horizon family, evaluated on the same agentic, coding, and reasoning benchmark suite as its larger siblings. Like the 7B model, it ships with intermediate checkpoints and public training data/recipe/code. Its benchmark profile — strong on SWE-bench and competition math, comparatively weaker on GPQA and function-calling — suggests targeted tuning toward coding-agent tasks rather than uniform improvement across all domains.

Specifications

FieldValue
OrganizationIFM
Parameters3.7B (dense)
Context window524,288 tokens (512K)
ArchitectureDense decoder-only
LicenseApache 2.0
Release date2026-09
ModalityText

Pricing

Input (per 1M tokens)Output (per 1M tokens)
Open weights——

Open weights: free to download and self-host. Validated SGLang/vLLM recipe: single-GPU (TP=1).

Context Window

K2-Horizon-3.7B has a 524,288-token (512K) context window, native from the midtraining stages onward.

Public Benchmark Scores

Scores above are reported by IFM and shown for context; not Benchgen measurements. Baseline protocols for comparison models may differ.

Use K2-Horizon-3.7B via API

from openai import OpenAI

client = OpenAI(base_url="http://localhost:30000/v1", api_key="EMPTY")

response = client.chat.completions.create(
    model="IFM/K2-Horizon-3.7B",
    messages=[{"role": "user", "content": "Explain the result step by step."}],
    temperature=1.0,
    top_p=0.95,
    max_tokens=32768,
    extra_body={"chat_template_kwargs": {"reasoning_effort": "high"}},
)
print(response.choices[0].message.content)

Frequently Asked Questions

What is K2-Horizon-3.7B? It's IFM's small dense model in the K2-Horizon family: 3.7B parameters, 512K-token context, with a standout SWE-bench Verified score for its size class.
What is K2-Horizon-3.7B's context window? 524,288 tokens (512K).
How much does K2-Horizon-3.7B cost? Open weights — free to download; cost is hosting/inference only.
Is K2-Horizon-3.7B open source? Yes, released under the Apache 2.0 license.

Specs and scores sourced from IFM's official Hugging Face model card; third-party benchmark scores attributed inline. Last updated 2026-09-03.