| Rank | Model | Score |
|---|---|---|
| 1 | kimi-k3 | 81.2 |
| 2 | qwen3-8-max | 73.5 |
1 phaseActive
Frontier-level software engineering agent benchmark designed to stay unsaturated as coding agents improve. Metric: % task success.
Quick answer: FrontierSWE is a software engineering benchmark designed to test frontier-level coding agents on difficult, real-world development tasks that remain challenging even for the strongest available models. Kimi K3 scores 81.2% as of July 2026.
What it tests: A coding agent's ability to resolve advanced, real-world software engineering tasks that push beyond the difficulty ceiling of more saturated coding benchmarks.
Why it matters: As benchmarks like SWE-Bench Verified approach saturation for top models, FrontierSWE provides continued headroom to differentiate frontier coding agents.
Known limitations: As a newer benchmark, public documentation of exact task sourcing and difficulty calibration is limited.
FrontierSWE evaluates coding agents against a curated pool of advanced software engineering tasks intended to remain difficult even for current frontier models. It follows the broader trend of "frontier" benchmarks — deliberately targeting the upper edge of model capability to keep leaderboards meaningful as older benchmarks saturate.
| Field | Value |
|---|---|
| Task category | Coding agent |
| Metric | % task success |
| Saturation | Low |
| Created by | FrontierSWE project |
| Website | frontierswe.com |
Agents attempt each software engineering task, with correctness verified through automated checks (e.g., test suite execution), producing an aggregate % task success rate.
| Rank | Model | Score | Source | Date |
|---|---|---|---|---|
| 1 | Kimi K3 | 81.2% | Kimi K3 technical report | 2026-07 |
Score sourced from Moonshot AI's Kimi K3 announcement, July 2026.
No Benchgen results yet — be the first to run FrontierSWE.
| Benchmark | What it tests | Saturation |
|---|---|---|
| FrontierSWE | Frontier-level software engineering tasks | Low |
| SWE-Marathon | Long-horizon software engineering | Low |
| DeepSWE | Real-world software engineering tasks | Low |
| SWE-Bench Pro | Professional-grade SWE agent tasks | Low |
Benchgen lets you run FrontierSWE against your own coding agent, tracking task success rate over time as models and harnesses evolve.