| Rank | Model | Score |
|---|---|---|
| 1 | kimi-k3 | 42 |
1 phaseActive
Long-horizon, endurance-style software engineering agent benchmark testing sustained multi-step coding workflows. Metric: % task success.
Quick answer: SWE-Marathon is a long-horizon, endurance-style software engineering benchmark that tests whether coding agents can sustain correct, coherent work across extended multi-step development tasks — rather than short, isolated bug fixes. Kimi K3 scores 42.0% as of July 2026.
What it tests: A coding agent's ability to maintain correctness and coherence across long-running, multi-session software engineering workflows.
Why it matters: Many coding benchmarks test single-shot fixes. As agents are deployed for hours-long autonomous coding sessions, endurance-style benchmarks like SWE-Marathon better predict real-world reliability.
Known limitations: As a newer benchmark with limited public documentation, task composition and scoring methodology details are not as thoroughly published as established academic benchmarks.
SWE-Marathon evaluates coding agents on extended, multi-step software engineering tasks that require sustained focus and coherent long-horizon planning, in contrast to single-turn bug-fix benchmarks. It aims to surface degradation in agent performance that only appears over long task horizons — such as context drift, plan abandonment, or compounding errors.
| Field | Value |
|---|---|
| Task category | Coding agent / long-horizon |
| Metric | % task success |
| Saturation | Low |
| Created by | SWE-Marathon project |
| Website | swe-marathon.org |
Agents attempt extended multi-step coding tasks, and outcomes are scored on % task success based on whether the final code state satisfies the task's verification criteria after the full multi-step workflow.
| Rank | Model | Score | Source | Date |
|---|---|---|---|---|
| 1 | Kimi K3 | 42.0% | Kimi K3 technical report | 2026-07 |
Score sourced from Moonshot AI's Kimi K3 announcement, July 2026.
No Benchgen results yet — be the first to run SWE-Marathon.
| Benchmark | What it tests | Saturation |
|---|---|---|
| SWE-Marathon | Long-horizon, endurance software engineering | Low |
| DeepSWE | Real-world software engineering tasks | Low |
| SWE-Bench Pro | Professional-grade SWE agent tasks | Low |
| FrontierSWE | Frontier-level software engineering tasks | Low |
Benchgen lets you run SWE-Marathon against your own coding agent, tracking long-horizon task success and identifying where agents drift or degrade over extended sessions.