Benchgen

SWE-bench Pro — Results

RankModelScore
1fugu-ultra73.7
2qwen3-8-max67.7
3hy4-preview65.7
4nex-n2-5-max65.7
5ornith-1-5-397b65.1
6gpt-5-6-sol64.6
7claude-opus-4-764.3
8gpt-5-6-terra63.4
9claude-sonnet-563.2
10gpt-5-6-luna62.7
11qwen3-8-flash-next62.5
12glm-5-262.1
13qwen3-8-27b61.7
14nex-n2-5-pro61.2
15dots3-note-preview61
16ornith-1-5-35b-a3b59.6
17laguna-s-2-159.4
18fugu59
19gpt-5-558.6
20kimi-k2-658.6
21claude-sonnet-4-658.1
22seed-2-1-pro57.5
23mimo-v2-5-pro57.2
24seed-2-1-turbo57
25inkling54.3

SWE-bench Pro

1 phaseActive

Harder variant of SWE-bench for coding agents — tests real-world GitHub issue resolution on more challenging tasks than SWE-bench Verified.

Overview

SWE-bench Pro

Category Metric Saturation

Quick answer: SWE-bench Pro is a harder-difficulty variant of the SWE-bench coding agent benchmark, designed to test model performance on more challenging real-world GitHub issue resolution tasks than SWE-bench Verified. It is evaluated using the mini-swe-agent scaffold. Fugu Ultra achieves 73.7% as of June 2026.

At a Glance

What it tests: Whether a coding agent can resolve harder, real-world GitHub issues by producing patches that pass all relevant unit tests.

Why it matters: SWE-bench Verified is approaching saturation at the top end; SWE-bench Pro provides a more discriminating signal for frontier and multi-agent coding systems.

Known limitations: Evaluated with mini-swe-agent scaffold — scores reflect both model capability and the specific scaffolding. Cross-benchmark methodology details are provided in the Sakana Fugu technical report.

What SWE-bench Pro Measures

SWE-bench Pro follows the same core protocol as SWE-bench Verified: a model receives a GitHub repository at a specific commit and the text of a real issue, and must produce a patch that causes all relevant tests to pass. The "Pro" designation selects for harder tasks, targeting the segment of GitHub issues where current frontier models still struggle — providing a more informative separation between top-tier systems.

Scores are reported as the percentage of tasks resolved, using the mini-swe-agent as the agent scaffold. This lightweight scaffolding provides a more reproducible comparison point than custom multi-agent pipelines.

Benchmark Specifications

FieldValue
Task categoryCoding
Metric% resolved (patch passes all unit tests)
Scaffoldmini-swe-agent
SaturationLow
Related benchmarkSWE-bench Verified

State-of-the-Art Results

RankModelScoreSourceDate
1Fugu Ultra73.7%Sakana Fugu technical report2026-06
2Fable 5 / Mythos Preview (max)69.2%Sakana Fugu technical report2026-06
3Fugu59.0%Sakana Fugu technical report2026-06

All scores above use the mini-swe-agent scaffold, as reported by Sakana AI in the Fugu technical report.