Quick answer: Kimi K2 0905 is Moonshot AI's September 2025 checkpoint of Kimi K2, scoring 94.5% HumanEval, 84.9% GPQA Diamond, 95.0% GSM8K, 89.1% MATH, and 53.7% LiveCodeBench. Apache 2.0.
Where Kimi K2 0905 leads
Where it lags
Best for: High-quality HumanEval/math pipelines in Apache 2.0; general Kimi K2 ecosystem deployments; September 2025 stable checkpoint.
Kimi K2 0905 is the September 5, 2025 checkpoint of Moonshot AI's Kimi K2 instruct/base model — a specific checkpoint between the K2.5 and K2.6 releases. The 0905 suffix (September 5) indicates a mid-series update.
With 94.5% HumanEval and 95.0% GSM8K, this checkpoint shows strong fundamental math and code quality. The 53.7% LiveCodeBench reflects practical coding ability, though below the later K2.6 SWE-Bench metric.
| Field | Value |
|---|---|
| Organization | Moonshot AI |
| License | Apache 2.0 |
| Release date | July 2025 (checkpoint: Sept 5, 2025) |
| Architecture | MoE |
| Modality | Text only |
Open weights under Apache 2.0. Also available via Moonshot AI API.
| Benchmark | Score | Source | Date |
|---|---|---|---|
| HumanEval | 94.5% | Benchgen evaluation | 2025-09 |
| GSM8K | 95.0% | Benchgen evaluation | 2025-09 |
| MATH | 89.1% | Benchgen evaluation | 2025-09 |
| GPQA Diamond | 84.9% | Benchgen evaluation | 2025-09 |
| LiveCodeBench | 53.7% | Benchgen evaluation | 2025-09 |
| Model | HumanEval | GPQA Diamond | LiveCodeBench | License |
|---|---|---|---|---|
| Kimi K2 0905 | 94.5% | 84.9% | 53.7% | Apache 2.0 |
| Kimi K2.5 | — | 87.6% | — | Apache 2.0 |
| Kimi K2 Instruct 0905 | — | — | — | Apache 2.0 |
| Qwen2.5 Coder 32B Instruct | 92.7% | — | — | Apache 2.0 |
Kimi K2 0905 vs Qwen2.5 Coder 32B: higher HumanEval (94.5% vs 92.7%) with GPQA coverage as a bonus. For Kimi-ecosystem general coding with STEM reasoning: K2 0905.
Specs from Moonshot AI's Kimi K2 0905 checkpoint and Benchgen evaluations. Last updated 2026-07-24.
This model isn’t on any benchmark leaderboard yet.