Benchgen

LiveCodeBench Pro — Results

RankModelScore
1fugu-ultra90.8
2fugu87.8

LiveCodeBench Pro

1 phaseActive

Hard-problem subset of LiveCodeBench — focuses on the most difficult contamination-free competitive programming tasks. Fugu Ultra leads at 90.8%.

Overview

LiveCodeBench Pro

Category Metric Difficulty Saturation

Paper GitHub

Quick answer: LiveCodeBench Pro is the hard-problem subset of LiveCodeBench, restricting evaluation to the most difficult contamination-free competitive programming problems. It provides a sharper discrimination signal between frontier systems at the top of the leaderboard. Fugu Ultra scores 90.8% and Fugu scores 87.8% as of June 2026.

At a Glance

What it tests: Challenging competitive programming problem solving — the hardest subset of fresh contest problems from LeetCode, AtCoder, and Codeforces.

Why it matters: At the frontier, many models cluster at 85–93% on standard LiveCodeBench, making it hard to differentiate top systems. The "Pro" harder subset spreads scores and provides a more informative ranking.

Known limitations: Difficulty is defined by contest rating, which may not directly map to practical software engineering complexity. Results depend on the time window of problems evaluated.

What LiveCodeBench Pro Measures

LiveCodeBench Pro applies the same contamination-free protocol as LiveCodeBench — collecting problems from major competitive programming platforms after model training cutoffs — but filters for the hardest-rated problems. A pass@1 metric is used: the model must solve the problem correctly on its first attempt, with no retries.

This harder subset is particularly useful for benchmarking multi-agent systems and reasoning-augmented models that exceed 85% on the standard benchmark, where standard LiveCodeBench provides less discriminating signal.

Benchmark Specifications

FieldValue
Task categoryCoding — competitive programming (hard)
Metricpass@1
DifficultyHard (top difficulty tier of LiveCodeBench)
SaturationLow
Related benchmarkLiveCodeBench

State-of-the-Art Results

RankModelScoreSourceDate
1Fugu Ultra90.8%Sakana Fugu technical report2026-06
2Fugu87.8%Sakana Fugu technical report2026-06
3Fable 5 / Mythos Preview (max)84.8%Sakana Fugu technical report2026-06

Scores sourced from Sakana AI's Fugu technical report, June 2026.