| Rank | Model | Score |
|---|---|---|
| 1 | fugu-ultra | 73.7 |
| 2 | qwen3-8-max | 67.7 |
| 3 | hy4-preview | 65.7 |
| 4 | nex-n2-5-max | 65.7 |
| 5 | ornith-1-5-397b | 65.1 |
| 6 | gpt-5-6-sol | 64.6 |
| 7 | claude-opus-4-7 | 64.3 |
| 8 | gpt-5-6-terra | 63.4 |
| 9 | claude-sonnet-5 | 63.2 |
| 10 | gpt-5-6-luna | 62.7 |
| 11 | qwen3-8-flash-next | 62.5 |
| 12 | glm-5-2 | 62.1 |
| 13 | qwen3-8-27b | 61.7 |
| 14 | nex-n2-5-pro | 61.2 |
| 15 | dots3-note-preview | 61 |
| 16 | ornith-1-5-35b-a3b | 59.6 |
| 17 | laguna-s-2-1 | 59.4 |
| 18 | fugu | 59 |
| 19 | gpt-5-5 | 58.6 |
| 20 | kimi-k2-6 | 58.6 |
| 21 | claude-sonnet-4-6 | 58.1 |
| 22 | seed-2-1-pro | 57.5 |
| 23 | mimo-v2-5-pro | 57.2 |
| 24 | seed-2-1-turbo | 57 |
| 25 | inkling | 54.3 |
1 phaseActive
Harder variant of SWE-bench for coding agents — tests real-world GitHub issue resolution on more challenging tasks than SWE-bench Verified.
Quick answer: SWE-bench Pro is a harder-difficulty variant of the SWE-bench coding agent benchmark, designed to test model performance on more challenging real-world GitHub issue resolution tasks than SWE-bench Verified. It is evaluated using the mini-swe-agent scaffold. Fugu Ultra achieves 73.7% as of June 2026.
What it tests: Whether a coding agent can resolve harder, real-world GitHub issues by producing patches that pass all relevant unit tests.
Why it matters: SWE-bench Verified is approaching saturation at the top end; SWE-bench Pro provides a more discriminating signal for frontier and multi-agent coding systems.
Known limitations: Evaluated with mini-swe-agent scaffold — scores reflect both model capability and the specific scaffolding. Cross-benchmark methodology details are provided in the Sakana Fugu technical report.
SWE-bench Pro follows the same core protocol as SWE-bench Verified: a model receives a GitHub repository at a specific commit and the text of a real issue, and must produce a patch that causes all relevant tests to pass. The "Pro" designation selects for harder tasks, targeting the segment of GitHub issues where current frontier models still struggle — providing a more informative separation between top-tier systems.
Scores are reported as the percentage of tasks resolved, using the mini-swe-agent as the agent scaffold. This lightweight scaffolding provides a more reproducible comparison point than custom multi-agent pipelines.
| Field | Value |
|---|---|
| Task category | Coding |
| Metric | % resolved (patch passes all unit tests) |
| Scaffold | mini-swe-agent |
| Saturation | Low |
| Related benchmark | SWE-bench Verified |
| Rank | Model | Score | Source | Date |
|---|---|---|---|---|
| 1 | Fugu Ultra | 73.7% | Sakana Fugu technical report | 2026-06 |
| 2 | Fable 5 / Mythos Preview (max) | 69.2% | Sakana Fugu technical report | 2026-06 |
| 3 | Fugu | 59.0% | Sakana Fugu technical report | 2026-06 |
All scores above use the mini-swe-agent scaffold, as reported by Sakana AI in the Fugu technical report.