Quick answer: K2-Horizon-375B-A23B is IFM's flagship open-weight Mixture-of-Experts model — 375B total parameters, 23B active per token, with a native 512K-token context window. On agentic tool-use, terminal, and long-horizon workflow benchmarks it matches or beats open-weight MoE models up to 2.6x its size and is competitive with closed frontier models. Open weights, Apache 2.0 license.
Where K2-Horizon-375B-A23B leads
Where it lags
Best for: teams wanting an open-weight model with frontier-adjacent agentic tool-use and terminal-use capability at a fraction of the active-parameter cost of larger MoE peers.
K2-Horizon-375B-A23B is the largest model in IFM's K2-Horizon family, a sparse Mixture-of-Experts architecture storing 375B parameters while activating only 23B per token. It is built for long-horizon agentic workflows — tool use, terminal interaction, and multi-step task completion — rather than pure static-benchmark chasing. The model ships with the final training checkpoint; intermediate checkpoints, training data, and training code are planned for future release, an unusually high transparency bar for a model at this scale.
| Field | Value |
|---|---|
| Organization | IFM |
| Parameters | 375B total / 23B active |
| Context window | 524,288 tokens (512K) |
| Architecture | Sparse Mixture-of-Experts (MoE) |
| License | Apache 2.0 |
| Release date | 2026-09 |
| Modality | Text |
| Input (per 1M tokens) | Output (per 1M tokens) | |
|---|---|---|
| Open weights | — | — |
Open weights: free to download and self-host; cost is hosting/inference only. Validated serving recipes exist for vLLM and SGLang (8x H200, tensor-parallel 8).
K2-Horizon-375B-A23B has a 524,288-token (512K) context window — roughly 1,300+ pages of text in a single request. This is native from the midtraining stages onward, supporting long-document and long-horizon-agent workloads without external context-extension tricks.
| Benchmark | Score | Source | Date |
|---|---|---|---|
| GDPVal-AA v2 | 1,441 (Elo) | IFM model card | 2026-09 |
| τ³ Banking | 34.0 | IFM model card | 2026-09 |
| TerminalBench 2.1 | 70.2 | IFM model card | 2026-09 |
| SciCode | 42.7 | IFM model card | 2026-09 |
| Humanity's Last Exam (no tools) | 32.0 | IFM model card | 2026-09 |
| GPQA Diamond | 87.3 | IFM model card | 2026-09 |
| CritPt | 8.6 | IFM model card | 2026-09 |
| AA-LCR | 76.0 | IFM model card | 2026-09 |
| Toolathlon-Verified | 65.3 | IFM model card | 2026-09 |
| AutomationBench (Public subset) | 25.3 | IFM model card | 2026-09 |
| APEX-Agents (pass@1, text subset) | 24.8 | IFM model card | 2026-09 |
| MCPMark-Verified | 67.7 | IFM model card | 2026-09 |
| BrowseComp | 72.8 | IFM model card | 2026-09 |
| WildClawBench (English text subset) | 50.9 | IFM model card | 2026-09 |
| SWE Bench Pro (strict, no internet) | 42.6 | IFM model card | 2026-09 |
AA-Omniscience Accuracy (23.0) and Non-Hallucination (74.7) were also reported but use a different metric structure than the platform's combined AA-Omniscience Index — pending mapping decision, not yet added. SWE-Atlas-QnA (48.4, strict) reported but pending exact-benchmark mapping confirmation. Scores above are reported by IFM and shown for context; not Benchgen measurements. GDPVal-AA is the Elo rating; other scores are % unless noted.
from openai import OpenAI
client = OpenAI(base_url="http://localhost:30000/v1", api_key="EMPTY")
response = client.chat.completions.create(
model="IFM/K2-Horizon-375B-A23B",
messages=[{"role": "user", "content": "Explain the result step by step."}],
temperature=1.0,
top_p=0.95,
extra_body={"chat_template_kwargs": {"reasoning_effort": "high"}},
)
print(response.choices[0].message.content)Specs and scores sourced from IFM's official Hugging Face model card; third-party benchmark scores attributed inline. Last updated 2026-09-03.