| Rank | Model | Score |
|---|---|---|
| 1 | kimi-k3 | 72.9 |
1 phaseActive
Moonshot AI's second-generation internal coding benchmark, testing practical code generation and problem-solving. Metric: % accuracy.
Quick answer: Kimi Code Bench 2.0 is Moonshot AI's second-generation internal coding benchmark, used to evaluate practical code generation and problem-solving ability across the Kimi model family. Kimi K3 scores 72.9% as of July 2026.
What it tests: Practical code generation and problem-solving across a range of programming tasks, as curated internally by Moonshot AI.
Why it matters: Internal benchmarks like Kimi Code Bench 2.0 let a lab track iterative coding capability improvements across model generations using a consistent, controlled task suite.
Known limitations: As a vendor-internal benchmark, the exact task composition and methodology are not independently published, so cross-lab comparisons should be treated cautiously.
Kimi Code Bench 2.0 evaluates a model's practical coding ability — generating correct, working code for a range of programming problems — as part of Moonshot AI's internal evaluation suite for tracking progress across successive Kimi model releases.
| Field | Value |
|---|---|
| Task category | Coding |
| Metric | % accuracy |
| Saturation | Low |
| Created by | Moonshot AI |
Models generate code solutions for each task, verified for correctness (e.g., via test execution), producing an aggregate % accuracy score.
| Rank | Model | Score | Source | Date |
|---|---|---|---|---|
| 1 | Kimi K3 | 72.9% | Kimi K3 technical report | 2026-07 |
Score sourced from Moonshot AI's Kimi K3 announcement, July 2026.
No Benchgen results yet — be the first to run Kimi Code Bench 2.0.
| Benchmark | What it tests | Saturation |
|---|---|---|
| Kimi Code Bench 2.0 | Practical code generation (Moonshot internal) | Low |
| BigCodeBench | Practical code generation | Medium |
| LiveCodeBench | Contamination-resistant competitive coding | Medium |
| ProgramBench | Real-world programming tasks (Vals AI) | Low |
Benchgen lets you run Kimi Code Bench 2.0-style coding evaluations against your own model, tracking accuracy over time as you iterate on your deployment.