Benchgen

SRE-Bench — Results

RankModelScore
1gpt-6-astra88
S

SRE-Bench

1 phaseActive

Reverse-engineering benchmark testing whether models can understand a compiled binary's core logic without source access.

Overview

SRE-Bench

Category Metric Saturation Created

Quick answer: SRE-Bench measures whether AI models can reverse-engineer compiled software binaries — understanding a program's core logic without access to its source code. It first appeared publicly in OpenAI's GPT-6 Astra announcement (September 2026), where Astra solved 88.0% of tasks in a single attempt and 99.2% within four attempts, versus 55.9% and 68.7% respectively for GPT-5.6 Sol.

At a Glance

What it tests: Binary reverse engineering — recovering a compiled program's logic, control flow, and intent purely from its binary form, without source code, a core skill in security research, malware analysis, and vulnerability discovery.

Why it matters: Reverse engineering is a bottleneck skill for both defensive security research (understanding malware, auditing closed-source software) and offensive capability (finding exploitable bugs in binaries). A model's SRE-Bench score is a proxy for how much of this traditionally expert-only skill can now be automated.

Known limitations: Full methodology (task count, binary corpus composition, public availability) has not yet been documented in an independent paper as of this writing — currently known primarily through frontier lab technical reports. Reported scores include both single-attempt and multi-attempt (within-k) results, which should not be conflated.

What SRE-Bench Measures

SRE-Bench presents models with compiled software binaries and asks them to reverse-engineer the underlying logic without access to source code — a task traditionally requiring specialized human expertise in disassembly, decompilation, and program analysis. Results are reported both as a single-attempt solve rate and as a within-four-attempts solve rate, distinguishing models that get it right immediately from those that need iteration to converge on a correct understanding.

Benchmark Specifications

FieldValue
Task categoryAgent / Cybersecurity
Metric% solved, single-attempt and within-k-attempts
SaturationMedium — frontier scores now above 85% single-attempt
Created byOpenAI (introduced via GPT-6 Astra announcement)
Public datasetNot yet publicly released as of Sep 2026

How SRE-Bench Is Scored

Two scores are reported per model: the percentage of tasks solved correctly on the first attempt, and the percentage solved within four attempts total. The gap between these two numbers indicates how much a model benefits from iterative refinement versus getting the analysis right immediately.

State-of-the-Art Results

RankModelSingle-AttemptWithin 4 AttemptsSourceDate
1GPT-6 Astra88.0%99.2%OpenAI: GPT-6 Astra2026-09
2GPT-5.6 Sol55.9%68.7%OpenAI: GPT-6 Astra2026-09
3Claude Opus 512.5%—OpenAI: GPT-6 Astra2026-09

Scores sourced from OpenAI's GPT-6 Astra announcement (September 2026). Leaderboard entries below reflect single-attempt scores for cross-benchmark consistency.

SRE-Bench on Benchgen

No Benchgen results yet — be the first to run SRE-Bench.

SRE-Bench vs Other Benchmarks

BenchmarkWhat it testsSaturation
SRE-BenchBinary reverse engineering without source accessMedium
ExploitBenchOffensive exploit construction (V8 bugs)Low
SEC-Bench ProDefensive security engineeringMedium

SRE-Bench sits between offensive and defensive security work — reverse engineering is a precursor skill useful for both finding new vulnerabilities and understanding existing malware.

Run SRE-Bench on Your Model

Benchgen lets teams track reverse-engineering capability across model versions over time, catching regressions a single vendor-reported snapshot would miss.

Frequently Asked Questions

What is SRE-Bench? A benchmark testing whether AI models can reverse-engineer compiled software binaries without source code access, introduced publicly via OpenAI's GPT-6 Astra announcement in September 2026.
What does a good SRE-Bench score look like? As of September 2026, GPT-6 Astra leads with 88.0% solved on the first attempt and 99.2% within four attempts.
Who created SRE-Bench? SRE-Bench was introduced by OpenAI, first publicly referenced in the GPT-6 Astra model announcement (September 2026).

Benchmark definition based on OpenAI's GPT-6 Astra announcement. State-of-the-art scores sourced from the same and attributed inline. Last updated 2026-09-07.