Benchgen

CursorBench — Results

RankModelScore
1grok-4-669.9
C

CursorBench

1 phaseActive

Cursor's internal agentic coding benchmark testing in-IDE task completion across real developer workflows.

Overview

CursorBench

Category Metric Saturation Created

Quick answer: CursorBench is an internal agentic coding benchmark maintained by Cursor (Anysphere) that measures how well a model completes real in-IDE developer tasks — editing, multi-file refactors, and tool-assisted debugging inside the Cursor editor. Grok 4.6 scores 69.9% as of August 2026.

At a Glance

What it tests: A model's ability to complete real developer tasks inside the Cursor IDE, using Cursor's own agent harness and tool-calling loop.

Why it matters: Because it runs inside the actual product developers use, CursorBench scores correlate closely with how a model will perform for real coding-agent users, rather than an isolated academic task format.

Known limitations: CursorBench is proprietary to Cursor — the task set, harness, and scoring rubric are not publicly published, so scores are only directly comparable across models Cursor itself has evaluated.

What CursorBench Measures

CursorBench is Cursor's (Anysphere's) internal benchmark for evaluating how well large language models perform as the reasoning engine behind its AI coding agent. Rather than testing isolated code generation, it measures end-to-end task completion inside Cursor's own agent harness — multi-file edits, running tests, and iterating on tool output.

Because the benchmark is run using Cursor's production agent harness, scores are frequently cited by model labs (including xAI, Anthropic, and OpenAI) as a proxy for real-world coding-agent usability, even though the underlying task set is not public.

Benchmark Specifications

FieldValue
Task categoryCoding agent
Metric% score
SaturationLow
Created byCursor (Anysphere)
DatasetProprietary — not publicly released
Projectcursor.com

How CursorBench Is Scored

Models are evaluated inside Cursor's agent harness on a suite of real developer tasks and scored as a percentage reflecting successful task completion, verified by Cursor's internal grading pipeline.

State-of-the-Art Results

RankModelScoreSourceDate
1Grok 4.669.9%xAI Grok 4.6 announcement2026-08

Score sourced from xAI's Grok 4.6 announcement, August 2026.

CursorBench on Benchgen

No Benchgen results yet — be the first to run CursorBench.

CursorBench vs Other Benchmarks

BenchmarkWhat it testsSaturation
CursorBenchIn-IDE agentic coding task completionLow
DeepSWEReal-world software engineering agent tasksLow
SWE-Bench VerifiedVerified real-world GitHub issue resolutionMedium

Run CursorBench on Your Model

Benchgen lets you track agentic coding performance across model and harness versions, complementing vendor-reported CursorBench scores with independently repeatable evaluation.

Frequently Asked Questions

What is CursorBench? CursorBench is Cursor's (Anysphere's) internal benchmark measuring how well a model performs as the reasoning engine behind its AI coding agent, testing real in-IDE developer tasks.
What does a good score look like on CursorBench? Grok 4.6 reports 69.9% as of August 2026; scores in this range are considered strong for current frontier coding models.
Who created CursorBench? CursorBench is created and maintained by Cursor (Anysphere Inc.).
Is CursorBench's dataset public? No — the task set and harness are proprietary to Cursor and not publicly released.