| Rank | Model | Score |
|---|---|---|
| 1 | kimi-k3 | 77.8 |
1 phaseActive
Vals AI's real-world programming task benchmark, evaluating end-to-end code generation and correctness. Metric: % accuracy.
Quick answer: ProgramBench is a real-world programming benchmark maintained by Vals AI, an independent model-evaluation platform. It tests models on practical coding tasks with an emphasis on correctness and real-world applicability rather than isolated algorithmic puzzles. Kimi K3 scores 77.8% as of July 2026.
What it tests: A model's ability to produce correct, working code for practical programming tasks, evaluated by an independent third-party leaderboard.
Why it matters: Vals AI runs standardized, independently-administered evaluations across many frontier models, making ProgramBench a useful cross-vendor reference point that isn't self-reported by any single lab.
Known limitations: As a third-party leaderboard benchmark, detailed task composition and scoring methodology are less publicly documented than academic benchmarks with published papers.
ProgramBench evaluates models on real-world programming tasks curated and administered by Vals AI, an independent AI evaluation platform. The benchmark emphasizes practical, applied coding correctness rather than narrow algorithmic puzzle-solving, and results are published on Vals AI's public leaderboard alongside many other frontier models for direct comparison.
| Field | Value |
|---|---|
| Task category | Coding |
| Metric | % accuracy |
| Saturation | Low |
| Created by | Vals AI |
| Leaderboard | vals.ai/benchmarks/programbench |
Models generate solutions for each programming task, which are evaluated for correctness (e.g., via test execution or reference-solution comparison), producing an aggregate % accuracy score published on the Vals AI leaderboard.
| Rank | Model | Score | Source | Date |
|---|---|---|---|---|
| 1 | Kimi K3 | 77.8% | Kimi K3 technical report | 2026-07 |
Score sourced from Moonshot AI's Kimi K3 announcement, July 2026.
No Benchgen results yet — be the first to run ProgramBench.
| Benchmark | What it tests | Saturation |
|---|---|---|
| ProgramBench | Real-world programming tasks (Vals AI) | Low |
| BigCodeBench | Practical code generation | Medium |
| LiveCodeBench | Contamination-resistant competitive coding | Medium |
| DeepSWE | Software engineering agent tasks | Low |
Benchgen lets you run ProgramBench-style evaluations against your own model, tracking coding accuracy over time as you fine-tune or update your deployment.