Quick answer: Kimi K2.5 is Moonshot AI's July 2025 model scoring 74.9% BrowseComp, 76.8% SWE-Bench Verified, 95.8% AIME 2026, 87.6% GPQA Diamond, and 87.1% MMLU-Pro. Apache 2.0 — strong baseline before K2.6 release.
Where Kimi K2.5 leads
Where it lags
Best for: Teams locked to a K2.5 checkpoint; open-source deployments requiring strong BrowseComp/SWE balance before K2.6.
Kimi K2.5 is the July 2025 checkpoint of Moonshot AI's Kimi K2 MoE model, predating the K2.6 release. It shows the strong capabilities of the K2 family before the K2.6 improvements: 74.9% BrowseComp, 76.8% SWE-Bench, and 95.8% AIME represent a highly competitive set of open-weight benchmarks.
The model is superseded by K2.6 for most use cases. K2.5 may be preferred where a stable, vetted checkpoint is required or where K2.6's additional LHTB/safety scores have not yet been validated.
| Field | Value |
|---|---|
| Organization | Moonshot AI |
| License | Apache 2.0 |
| Release date | July 2025 |
| Architecture | MoE |
| Modality | Text only |
Open weights under Apache 2.0. Also available via Moonshot AI API.
| Benchmark | Score | Source | Date |
|---|---|---|---|
| BrowseComp | 74.9% | Benchgen evaluation | 2025-07 |
| SWE-Bench Verified | 76.8% | Benchgen evaluation | 2025-07 |
| AIME 2026 | 95.8% | Benchgen evaluation | 2025-07 |
| GPQA Diamond | 87.6% | Benchgen evaluation | 2025-07 |
| Humanity's Last Exam | 24.4% | Benchgen evaluation | 2025-07 |
| MMLU-Pro | 87.1% | Benchgen evaluation | 2025-07 |
| SimpleQA | 36.9% | Benchgen evaluation | 2025-07 |
| Model | BrowseComp | SWE-Bench | MMLU-Pro | License |
|---|---|---|---|---|
| Kimi K2.5 | 74.9% | 76.8% | 87.1% | Apache 2.0 |
| Kimi K2.6 | 83.2% | 80.2% | — | Apache 2.0 |
| Kimi K2 Thinking 0905 | 60.2% | — | 84.6% | Apache 2.0 |
| DeepSeek-R1 0528 | — | — | 85% | MIT |
Kimi K2.5 vs K2.6: K2.6 is better on all metrics. Use K2.5 only when K2.6 is not available or when a specific checkpoint is required.
Specs from Moonshot AI's Kimi K2.5 release (July 2025) and Benchgen evaluations. Last updated 2026-07-24.
This model isn’t on any benchmark leaderboard yet.