Quick answer: Kimi K2 Thinking 0905 is Moonshot AI's September 2025 reasoning checkpoint of Kimi K2, scoring 60.2% BrowseComp, 51.0% HLE, 84.6% MMLU-Pro, 47.1% TerminalBench, and 44.8% SciCode. Apache 2.0 — notable as an open-weight model with these frontier-class reasoning scores.
Where Kimi K2 Thinking 0905 leads
Where it lags
Best for: Open-source frontier-class reasoning; web research pipelines; academic knowledge tasks; teams needing Apache 2.0 models with o3-tier reasoning.
Kimi K2 Thinking 0905 is the thinking (chain-of-thought reasoning) variant of Moonshot AI's Kimi K2 model family, at the September 5, 2025 checkpoint. Kimi K2 is a large MoE model; the "Thinking" variant extends it with a reasoning mode similar to DeepSeek-R1 and QwQ-32B.
The BrowseComp score (60.2%) is notable — this benchmark tests multi-step web research ability, and 60% is in the frontier tier. Combined with 51.0% HLE, this is one of the strongest open-weight reasoning models from mid-2025.
| Field | Value |
|---|---|
| Organization | Moonshot AI |
| License | Apache 2.0 |
| Release date | July 2025 (checkpoint: Sept 2025) |
| Architecture | MoE with chain-of-thought reasoning |
| Modality | Text only |
Open weights under Apache 2.0 — self-host at no license cost. Also available via Moonshot AI API.
| Benchmark | Score | Source | Date |
|---|---|---|---|
| BrowseComp | 60.2% | Benchgen evaluation | 2025-09 |
| Humanity's Last Exam | 51.0% | Benchgen evaluation | 2025-09 |
| MMLU-Pro | 84.6% | Benchgen evaluation | 2025-09 |
| TerminalBench | 47.1% | Benchgen evaluation | 2025-09 |
| SciCode | 44.8% | Benchgen evaluation | 2025-09 |
| Model | HLE | BrowseComp | MMLU-Pro | License |
|---|---|---|---|---|
| Kimi K2 Thinking 0905 | 51.0% | 60.2% | 84.6% | Apache 2.0 |
| Kimi K2 Instruct 0905 | — | — | 81.1% | Apache 2.0 |
| GLM-5.2 | 54.7% | — | — | Proprietary |
| GPT-5.5 | 41.4% | 84.4% | — | Proprietary |
Kimi K2 Thinking 0905 leads open-weight models on HLE (51.0%) and BrowseComp (60.2%). It beats GPT-5.5 on HLE (51.0% vs 41.4%) while remaining Apache 2.0. GLM-5.2 leads on HLE (54.7%) but is proprietary.
Specs from Moonshot AI's Kimi K2 Thinking 0905 release and Benchgen evaluations. Last updated 2026-07-24.
This model isn’t on any benchmark leaderboard yet.