Benchgen
Models/moonshot/

Kimi K2.6

DraftPublic

Model Details

Kimi K2.6

Organization License Released

Quick answer: Kimi K2.6 is Moonshot AI's July 2025 model scoring 83.2% BrowseComp, 80.2% SWE-Bench Verified, 96.4% AIME 2026, 90.5% GPQA Diamond, and 36.4% HLE. Apache 2.0 — notable BrowseComp leader among open-weight models.

At a Glance

Where Kimi K2.6 leads

  • 83.2% BrowseComp — among the highest BrowseComp scores; near GPT-5.5 (84.4%)
  • 80.2% SWE-Bench Verified — top-tier software engineering
  • 96.4% AIME 2026 — near-perfect competition mathematics
  • 90.5% GPQA Diamond — excellent science reasoning
  • Apache 2.0 — fully open

Where it lags

  • 36.4% HLE — below top frontier models (GLM-5.2: 54.7%, Gemini 3.1 Pro: 46.4%)
  • 38.7% SimpleQA — moderate factual accuracy
  • 25.5% LHTB — limited long-horizon task performance

Best for: Web research agents (BrowseComp); software engineering pipelines (SWE-Bench 80.2%); open-source frontier-class deployments.

What Kimi K2.6 Is

Kimi K2.6 is Moonshot AI's July 2025 major update in the Kimi K2 family — a successive version after K2.5. Kimi K2 is a large MoE model; K2.6 shows substantial improvements on BrowseComp (83.2%) and SWE-Bench (80.2%) over K2.5 (74.9% BrowseComp, 76.8% SWE-Bench).

The 83.2% BrowseComp score is exceptional — matching or exceeding GPT-5.5 (84.4%) while being Apache 2.0 open-weight. This makes K2.6 a leading option for web research agent pipelines.

Specifications

FieldValue
OrganizationMoonshot AI
LicenseApache 2.0
Release dateJuly 2025
ArchitectureMoE
ModalityText only

Pricing

Open weights under Apache 2.0. Also available via Moonshot AI API.

Public Benchmark Scores

BenchmarkScoreSourceDate
BrowseComp83.2%Benchgen evaluation2025-07
SWE-Bench Verified80.2%Benchgen evaluation2025-07
AIME 202696.4%Benchgen evaluation2025-07
GPQA Diamond90.5%Benchgen evaluation2025-07
Humanity's Last Exam36.4%Benchgen evaluation2025-07
Global MMLU-Lite88.4%Benchgen evaluation2025-07
SimpleQA38.7%Benchgen evaluation2025-07

Kimi K2.6 vs Alternatives

ModelBrowseCompSWE-BenchHLELicense
Kimi K2.683.2%80.2%36.4%Apache 2.0
Kimi K2.574.9%76.8%24.4%Apache 2.0
GPT-5.584.4%41.4%Proprietary
HY378.0%Proprietary

Kimi K2.6 vs GPT-5.5: near-equal BrowseComp (83.2% vs 84.4%), similar SWE-Bench tier — but Apache 2.0 vs proprietary. K2.6 is the top open-weight option for BrowseComp + SWE benchmarks.

Frequently Asked Questions

What is Kimi K2.6? Moonshot AI's July 2025 MoE model scoring 83.2% BrowseComp, 80.2% SWE-Bench Verified, 96.4% AIME 2026, 90.5% GPQA Diamond. Apache 2.0.
How does Kimi K2.6 compare to K2.5? K2.6 substantially improves BrowseComp (83.2% vs 74.9%), SWE-Bench (80.2% vs 76.8%), AIME (96.4% vs 95.8%), and HLE (36.4% vs 24.4%) vs K2.5. Clear upgrade.

Specs from Moonshot AI's Kimi K2.6 release (July 2025) and Benchgen evaluations. Last updated 2026-07-24.