| Rank | Model | Score |
|---|---|---|
| 1 | grok-4-6 | 69.9 |
1 phaseActive
Cursor's internal agentic coding benchmark testing in-IDE task completion across real developer workflows.
Quick answer: CursorBench is an internal agentic coding benchmark maintained by Cursor (Anysphere) that measures how well a model completes real in-IDE developer tasks — editing, multi-file refactors, and tool-assisted debugging inside the Cursor editor. Grok 4.6 scores 69.9% as of August 2026.
What it tests: A model's ability to complete real developer tasks inside the Cursor IDE, using Cursor's own agent harness and tool-calling loop.
Why it matters: Because it runs inside the actual product developers use, CursorBench scores correlate closely with how a model will perform for real coding-agent users, rather than an isolated academic task format.
Known limitations: CursorBench is proprietary to Cursor — the task set, harness, and scoring rubric are not publicly published, so scores are only directly comparable across models Cursor itself has evaluated.
CursorBench is Cursor's (Anysphere's) internal benchmark for evaluating how well large language models perform as the reasoning engine behind its AI coding agent. Rather than testing isolated code generation, it measures end-to-end task completion inside Cursor's own agent harness — multi-file edits, running tests, and iterating on tool output.
Because the benchmark is run using Cursor's production agent harness, scores are frequently cited by model labs (including xAI, Anthropic, and OpenAI) as a proxy for real-world coding-agent usability, even though the underlying task set is not public.
| Field | Value |
|---|---|
| Task category | Coding agent |
| Metric | % score |
| Saturation | Low |
| Created by | Cursor (Anysphere) |
| Dataset | Proprietary — not publicly released |
| Project | cursor.com |
Models are evaluated inside Cursor's agent harness on a suite of real developer tasks and scored as a percentage reflecting successful task completion, verified by Cursor's internal grading pipeline.
| Rank | Model | Score | Source | Date |
|---|---|---|---|---|
| 1 | Grok 4.6 | 69.9% | xAI Grok 4.6 announcement | 2026-08 |
Score sourced from xAI's Grok 4.6 announcement, August 2026.
No Benchgen results yet — be the first to run CursorBench.
| Benchmark | What it tests | Saturation |
|---|---|---|
| CursorBench | In-IDE agentic coding task completion | Low |
| DeepSWE | Real-world software engineering agent tasks | Low |
| SWE-Bench Verified | Verified real-world GitHub issue resolution | Medium |
Benchgen lets you track agentic coding performance across model and harness versions, complementing vendor-reported CursorBench scores with independently repeatable evaluation.