Benchgen

FrontierSWE — Results

RankModelScore
1kimi-k381.2
2qwen3-8-max73.5
F

FrontierSWE

1 phaseActive

Frontier-level software engineering agent benchmark designed to stay unsaturated as coding agents improve. Metric: % task success.

Overview

FrontierSWE

Category Metric Saturation Created

Website

Quick answer: FrontierSWE is a software engineering benchmark designed to test frontier-level coding agents on difficult, real-world development tasks that remain challenging even for the strongest available models. Kimi K3 scores 81.2% as of July 2026.

At a Glance

What it tests: A coding agent's ability to resolve advanced, real-world software engineering tasks that push beyond the difficulty ceiling of more saturated coding benchmarks.

Why it matters: As benchmarks like SWE-Bench Verified approach saturation for top models, FrontierSWE provides continued headroom to differentiate frontier coding agents.

Known limitations: As a newer benchmark, public documentation of exact task sourcing and difficulty calibration is limited.

What FrontierSWE Measures

FrontierSWE evaluates coding agents against a curated pool of advanced software engineering tasks intended to remain difficult even for current frontier models. It follows the broader trend of "frontier" benchmarks — deliberately targeting the upper edge of model capability to keep leaderboards meaningful as older benchmarks saturate.

Benchmark Specifications

FieldValue
Task categoryCoding agent
Metric% task success
SaturationLow
Created byFrontierSWE project
Websitefrontierswe.com

How FrontierSWE Is Scored

Agents attempt each software engineering task, with correctness verified through automated checks (e.g., test suite execution), producing an aggregate % task success rate.

State-of-the-Art Results

RankModelScoreSourceDate
1Kimi K381.2%Kimi K3 technical report2026-07

Score sourced from Moonshot AI's Kimi K3 announcement, July 2026.

FrontierSWE on Benchgen

No Benchgen results yet — be the first to run FrontierSWE.

FrontierSWE vs Other Benchmarks

BenchmarkWhat it testsSaturation
FrontierSWEFrontier-level software engineering tasksLow
SWE-MarathonLong-horizon software engineeringLow
DeepSWEReal-world software engineering tasksLow
SWE-Bench ProProfessional-grade SWE agent tasksLow

Run FrontierSWE on Your Model

Benchgen lets you run FrontierSWE against your own coding agent, tracking task success rate over time as models and harnesses evolve.

Frequently Asked Questions

What is FrontierSWE? FrontierSWE is a software engineering benchmark designed to test frontier-level coding agents on advanced, real-world development tasks that stay difficult even for the strongest models.
What does a good score look like on FrontierSWE? Kimi K3 reports 81.2% as of July 2026, indicating strong frontier-level coding agent performance on this benchmark's task pool.
Who created FrontierSWE? FrontierSWE is maintained as an independent benchmark project (frontierswe.com) focused on frontier-level software engineering agent evaluation.