Benchgen

Kimi Code Bench 2.0 — Results

RankModelScore
1kimi-k372.9
K

Kimi Code Bench 2.0

1 phaseActive

Moonshot AI's second-generation internal coding benchmark, testing practical code generation and problem-solving. Metric: % accuracy.

Overview

Kimi Code Bench 2.0

Category Metric Saturation Created

Quick answer: Kimi Code Bench 2.0 is Moonshot AI's second-generation internal coding benchmark, used to evaluate practical code generation and problem-solving ability across the Kimi model family. Kimi K3 scores 72.9% as of July 2026.

At a Glance

What it tests: Practical code generation and problem-solving across a range of programming tasks, as curated internally by Moonshot AI.

Why it matters: Internal benchmarks like Kimi Code Bench 2.0 let a lab track iterative coding capability improvements across model generations using a consistent, controlled task suite.

Known limitations: As a vendor-internal benchmark, the exact task composition and methodology are not independently published, so cross-lab comparisons should be treated cautiously.

What Kimi Code Bench 2.0 Measures

Kimi Code Bench 2.0 evaluates a model's practical coding ability — generating correct, working code for a range of programming problems — as part of Moonshot AI's internal evaluation suite for tracking progress across successive Kimi model releases.

Benchmark Specifications

FieldValue
Task categoryCoding
Metric% accuracy
SaturationLow
Created byMoonshot AI

How Kimi Code Bench 2.0 Is Scored

Models generate code solutions for each task, verified for correctness (e.g., via test execution), producing an aggregate % accuracy score.

State-of-the-Art Results

RankModelScoreSourceDate
1Kimi K372.9%Kimi K3 technical report2026-07

Score sourced from Moonshot AI's Kimi K3 announcement, July 2026.

Kimi Code Bench 2.0 on Benchgen

No Benchgen results yet — be the first to run Kimi Code Bench 2.0.

Kimi Code Bench 2.0 vs Other Benchmarks

BenchmarkWhat it testsSaturation
Kimi Code Bench 2.0Practical code generation (Moonshot internal)Low
BigCodeBenchPractical code generationMedium
LiveCodeBenchContamination-resistant competitive codingMedium
ProgramBenchReal-world programming tasks (Vals AI)Low

Run Kimi Code Bench 2.0 on Your Model

Benchgen lets you run Kimi Code Bench 2.0-style coding evaluations against your own model, tracking accuracy over time as you iterate on your deployment.

Frequently Asked Questions

What is Kimi Code Bench 2.0? Kimi Code Bench 2.0 is Moonshot AI's second-generation internal coding benchmark, used to track coding capability across successive Kimi model releases.
What does a good score look like on Kimi Code Bench 2.0? Kimi K3 reports 72.9% as of July 2026, its own frontier-level result on this internal benchmark.
Who created Kimi Code Bench 2.0? Kimi Code Bench 2.0 was created by Moonshot AI as part of its internal model evaluation suite.