Benchgen

ProgramBench — Results

RankModelScore
1kimi-k377.8
P

ProgramBench

1 phaseActive

Vals AI's real-world programming task benchmark, evaluating end-to-end code generation and correctness. Metric: % accuracy.

Overview

ProgramBench

Category Metric Saturation Created

Website

Quick answer: ProgramBench is a real-world programming benchmark maintained by Vals AI, an independent model-evaluation platform. It tests models on practical coding tasks with an emphasis on correctness and real-world applicability rather than isolated algorithmic puzzles. Kimi K3 scores 77.8% as of July 2026.

At a Glance

What it tests: A model's ability to produce correct, working code for practical programming tasks, evaluated by an independent third-party leaderboard.

Why it matters: Vals AI runs standardized, independently-administered evaluations across many frontier models, making ProgramBench a useful cross-vendor reference point that isn't self-reported by any single lab.

Known limitations: As a third-party leaderboard benchmark, detailed task composition and scoring methodology are less publicly documented than academic benchmarks with published papers.

What ProgramBench Measures

ProgramBench evaluates models on real-world programming tasks curated and administered by Vals AI, an independent AI evaluation platform. The benchmark emphasizes practical, applied coding correctness rather than narrow algorithmic puzzle-solving, and results are published on Vals AI's public leaderboard alongside many other frontier models for direct comparison.

Benchmark Specifications

FieldValue
Task categoryCoding
Metric% accuracy
SaturationLow
Created byVals AI
Leaderboardvals.ai/benchmarks/programbench

How ProgramBench Is Scored

Models generate solutions for each programming task, which are evaluated for correctness (e.g., via test execution or reference-solution comparison), producing an aggregate % accuracy score published on the Vals AI leaderboard.

State-of-the-Art Results

RankModelScoreSourceDate
1Kimi K377.8%Kimi K3 technical report2026-07

Score sourced from Moonshot AI's Kimi K3 announcement, July 2026.

ProgramBench on Benchgen

No Benchgen results yet — be the first to run ProgramBench.

ProgramBench vs Other Benchmarks

BenchmarkWhat it testsSaturation
ProgramBenchReal-world programming tasks (Vals AI)Low
BigCodeBenchPractical code generationMedium
LiveCodeBenchContamination-resistant competitive codingMedium
DeepSWESoftware engineering agent tasksLow

Run ProgramBench on Your Model

Benchgen lets you run ProgramBench-style evaluations against your own model, tracking coding accuracy over time as you fine-tune or update your deployment.

Frequently Asked Questions

What is ProgramBench? ProgramBench is a real-world programming benchmark maintained by Vals AI, an independent AI model evaluation platform, testing practical coding correctness across frontier models.
What does a good score look like on ProgramBench? Kimi K3 reports 77.8% as of July 2026, a strong result among current frontier coding models on this independently-administered benchmark.
Who created ProgramBench? ProgramBench is created and maintained by Vals AI, an independent AI evaluation company that publishes standardized benchmark results across many frontier models.