Benchgen
Models/ifm/

K2 Horizon MoVA 36B A4B

DraftPublic

Model Details

K2-Horizon-MoVA-36B-A4B

Organization Context Pricing License Modality Released

Quick answer: K2-Horizon-MoVA-36B-A4B is IFM's sparse Mixture-of-Experts model using Mixture-of-Values attention (MoVA) — 36B total parameters, only 4B active per token, with a native 512K-token context window. It outscores open-weight dense and MoE models up to 15x its active-parameter size on agentic and reasoning benchmarks, and is competitive with closed frontier models. Open weights, Apache 2.0.

At a Glance

Where K2-Horizon-MoVA-36B-A4B leads

  • Best-in-class efficiency: frontier-adjacent agentic scores at only 4B active parameters
  • Strong Terminal-Bench 2.1 and τ³-Banking scores versus much larger open peers
  • 512K native context; fully open weights

Where it lags

  • GPQA Diamond and AA-Omniscience trail larger dense/MoE peers like Gemma 4 31B and Nemotron 3 Ultra
  • CritPt (frontier physics) remains a weak point across the whole K2-Horizon family

Best for: cost-sensitive deployments needing strong agentic tool-use and terminal capability without the inference cost of a 30B+ active-parameter model.

What K2-Horizon-MoVA-36B-A4B Is

K2-Horizon-MoVA-36B-A4B is the efficiency-focused member of IFM's K2-Horizon family. It uses a novel Mixture-of-Values attention (MoVA) mechanism paired with sparse MoE routing to store 36B parameters while activating just 4B per token — a ~9x sparsity ratio. IFM's own framing is that at this activated-parameter budget it competes with models up to 15x larger, positioning it as an efficiency play for agentic and reasoning workloads rather than a raw-parameter-count flagship.

Specifications

FieldValue
OrganizationIFM
Parameters36B total / 4B active
Context window524,288 tokens (512K)
ArchitectureSparse MoE with Mixture-of-Values attention (MoVA)
LicenseApache 2.0
Release date2026-09
ModalityText

Pricing

Input (per 1M tokens)Output (per 1M tokens)
Open weights——

Open weights: free to download and self-host. Validated SGLang recipe: 2x H200, TP=2, EP=2.

Context Window

K2-Horizon-MoVA-36B-A4B has a 524,288-token (512K) context window, native from the midtraining stages onward — enabling long-document and long-horizon-agent workloads at a fraction of the active-parameter cost of its dense-context peers.

Public Benchmark Scores

AA-Omniscience Accuracy (18.8) and Non-Hallucination (69.2) were also reported but use a different metric structure than the platform's combined AA-Omniscience Index — pending mapping decision, not yet added. Scores above are reported by IFM and shown for context; not Benchgen measurements.

Use K2-Horizon-MoVA-36B-A4B via API

from openai import OpenAI

client = OpenAI(base_url="http://localhost:30000/v1", api_key="EMPTY")

response = client.chat.completions.create(
    model="IFM/K2-Horizon-MoVA-36B-A4B",
    messages=[{"role": "user", "content": "Explain the result step by step."}],
    temperature=1.0,
    top_p=0.95,
    extra_body={"chat_template_kwargs": {"reasoning_effort": "high"}},
)
print(response.choices[0].message.content)

Frequently Asked Questions

What is K2-Horizon-MoVA-36B-A4B? It's IFM's efficiency-focused Mixture-of-Experts model using Mixture-of-Values attention (MoVA): 36B total parameters, only 4B activated per token, with a 512K-token context window.
What is K2-Horizon-MoVA-36B-A4B's context window? 524,288 tokens (512K).
How much does K2-Horizon-MoVA-36B-A4B cost? Open weights — free to download; cost is hosting/inference only.
Is K2-Horizon-MoVA-36B-A4B open source? Yes, released under the Apache 2.0 license.

Specs and scores sourced from IFM's official Hugging Face model card; third-party benchmark scores attributed inline. Last updated 2026-09-03.