Benchgen

DeepSWE — Results

RankModelScore
1kimi-k367.5
2grok-4-665.9
3muse-spark-1-259.3
4qwen3-8-max56.6
D

DeepSWE

1 phaseActive

Real-world software engineering agent benchmark curated by DataCurve, testing end-to-end coding-agent workflows. Metric: % task success.

Overview

DeepSWE

Category Metric Saturation Created

Dataset

Quick answer: DeepSWE is a real-world software engineering agent benchmark curated by DataCurve, testing end-to-end coding agent workflows against verified ground-truth solutions. It's positioned alongside benchmarks like SWE-Bench as a harder, more diverse test of agentic coding capability. Kimi K3 scores 67.5% as of July 2026.

At a Glance

What it tests: An AI coding agent's ability to resolve real-world software engineering tasks end-to-end — understanding a codebase, implementing a fix or feature, and verifying correctness.

Why it matters: As SWE-Bench-style benchmarks approach saturation for frontier models, DeepSWE provides an independently curated task pool from DataCurve to cross-validate agentic coding performance.

Known limitations: As a newer, vendor-curated benchmark, public documentation of methodology and task composition is limited compared to established academic benchmarks like SWE-Bench.

What DeepSWE Measures

DeepSWE evaluates coding agents on real-world software engineering tasks curated by DataCurve, a data-curation company focused on agentic and coding evaluation datasets. Tasks are designed to mirror the type of work a software engineer performs day-to-day — bug fixes, feature implementation, and refactors — verified against reference solutions or test suites.

Benchmark Specifications

FieldValue
Task categoryCoding agent
Metric% task success
SaturationLow
Created byDataCurve
Datasetdatacurve/deep-swe on HuggingFace
Projectdeepswe.datacurve.ai

How DeepSWE Is Scored

Agents attempt each software engineering task and are scored on % of tasks resolved correctly, typically verified via automated test execution against the target repository.

State-of-the-Art Results

RankModelScoreSourceDate
1Kimi K367.5%Kimi K3 technical report2026-07

Score sourced from Moonshot AI's Kimi K3 announcement, July 2026.

DeepSWE on Benchgen

No Benchgen results yet — be the first to run DeepSWE.

DeepSWE vs Other Benchmarks

BenchmarkWhat it testsSaturation
DeepSWEReal-world software engineering agent tasksLow
SWE-Bench ProProfessional-grade SWE agent tasksLow
SWE-Bench VerifiedVerified real-world GitHub issue resolutionMedium
SWE-MarathonLong-horizon software engineeringLow

Run DeepSWE on Your Model

Benchgen lets you run DeepSWE against your own coding agent, tracking task success rate over time to catch regressions from model or harness updates.

Frequently Asked Questions

What is DeepSWE? DeepSWE is a real-world software engineering agent benchmark curated by DataCurve, testing coding agents on end-to-end development tasks verified against reference solutions.
What does a good score look like on DeepSWE? Kimi K3 reports 67.5% as of July 2026; scores in this range are considered strong for current frontier coding agents on real-world software engineering workflows.
Who created DeepSWE? DeepSWE is curated by DataCurve, a company specializing in data curation for agentic and coding model evaluation.