Quick answer: Kimi K2.6 is Moonshot AI's July 2025 model scoring 83.2% BrowseComp, 80.2% SWE-Bench Verified, 96.4% AIME 2026, 90.5% GPQA Diamond, and 36.4% HLE. Apache 2.0 — notable BrowseComp leader among open-weight models.
Where Kimi K2.6 leads
Where it lags
Best for: Web research agents (BrowseComp); software engineering pipelines (SWE-Bench 80.2%); open-source frontier-class deployments.
Kimi K2.6 is Moonshot AI's July 2025 major update in the Kimi K2 family — a successive version after K2.5. Kimi K2 is a large MoE model; K2.6 shows substantial improvements on BrowseComp (83.2%) and SWE-Bench (80.2%) over K2.5 (74.9% BrowseComp, 76.8% SWE-Bench).
The 83.2% BrowseComp score is exceptional — matching or exceeding GPT-5.5 (84.4%) while being Apache 2.0 open-weight. This makes K2.6 a leading option for web research agent pipelines.
| Field | Value |
|---|---|
| Organization | Moonshot AI |
| License | Apache 2.0 |
| Release date | July 2025 |
| Architecture | MoE |
| Modality | Text only |
Open weights under Apache 2.0. Also available via Moonshot AI API.
| Benchmark | Score | Source | Date |
|---|---|---|---|
| BrowseComp | 83.2% | Benchgen evaluation | 2025-07 |
| SWE-Bench Verified | 80.2% | Benchgen evaluation | 2025-07 |
| AIME 2026 | 96.4% | Benchgen evaluation | 2025-07 |
| GPQA Diamond | 90.5% | Benchgen evaluation | 2025-07 |
| Humanity's Last Exam | 36.4% | Benchgen evaluation | 2025-07 |
| Global MMLU-Lite | 88.4% | Benchgen evaluation | 2025-07 |
| SimpleQA | 38.7% | Benchgen evaluation | 2025-07 |
| Model | BrowseComp | SWE-Bench | HLE | License |
|---|---|---|---|---|
| Kimi K2.6 | 83.2% | 80.2% | 36.4% | Apache 2.0 |
| Kimi K2.5 | 74.9% | 76.8% | 24.4% | Apache 2.0 |
| GPT-5.5 | 84.4% | — | 41.4% | Proprietary |
| HY3 | — | 78.0% | — | Proprietary |
Kimi K2.6 vs GPT-5.5: near-equal BrowseComp (83.2% vs 84.4%), similar SWE-Bench tier — but Apache 2.0 vs proprietary. K2.6 is the top open-weight option for BrowseComp + SWE benchmarks.
Specs from Moonshot AI's Kimi K2.6 release (July 2025) and Benchgen evaluations. Last updated 2026-07-24.