| Rank | Model | Score |
|---|---|---|
| 1 | gpt-6-astra | 88 |
1 phaseActive
Reverse-engineering benchmark testing whether models can understand a compiled binary's core logic without source access.
Quick answer: SRE-Bench measures whether AI models can reverse-engineer compiled software binaries — understanding a program's core logic without access to its source code. It first appeared publicly in OpenAI's GPT-6 Astra announcement (September 2026), where Astra solved 88.0% of tasks in a single attempt and 99.2% within four attempts, versus 55.9% and 68.7% respectively for GPT-5.6 Sol.
What it tests: Binary reverse engineering — recovering a compiled program's logic, control flow, and intent purely from its binary form, without source code, a core skill in security research, malware analysis, and vulnerability discovery.
Why it matters: Reverse engineering is a bottleneck skill for both defensive security research (understanding malware, auditing closed-source software) and offensive capability (finding exploitable bugs in binaries). A model's SRE-Bench score is a proxy for how much of this traditionally expert-only skill can now be automated.
Known limitations: Full methodology (task count, binary corpus composition, public availability) has not yet been documented in an independent paper as of this writing — currently known primarily through frontier lab technical reports. Reported scores include both single-attempt and multi-attempt (within-k) results, which should not be conflated.
SRE-Bench presents models with compiled software binaries and asks them to reverse-engineer the underlying logic without access to source code — a task traditionally requiring specialized human expertise in disassembly, decompilation, and program analysis. Results are reported both as a single-attempt solve rate and as a within-four-attempts solve rate, distinguishing models that get it right immediately from those that need iteration to converge on a correct understanding.
| Field | Value |
|---|---|
| Task category | Agent / Cybersecurity |
| Metric | % solved, single-attempt and within-k-attempts |
| Saturation | Medium — frontier scores now above 85% single-attempt |
| Created by | OpenAI (introduced via GPT-6 Astra announcement) |
| Public dataset | Not yet publicly released as of Sep 2026 |
Two scores are reported per model: the percentage of tasks solved correctly on the first attempt, and the percentage solved within four attempts total. The gap between these two numbers indicates how much a model benefits from iterative refinement versus getting the analysis right immediately.
| Rank | Model | Single-Attempt | Within 4 Attempts | Source | Date |
|---|---|---|---|---|---|
| 1 | GPT-6 Astra | 88.0% | 99.2% | OpenAI: GPT-6 Astra | 2026-09 |
| 2 | GPT-5.6 Sol | 55.9% | 68.7% | OpenAI: GPT-6 Astra | 2026-09 |
| 3 | Claude Opus 5 | 12.5% | — | OpenAI: GPT-6 Astra | 2026-09 |
Scores sourced from OpenAI's GPT-6 Astra announcement (September 2026). Leaderboard entries below reflect single-attempt scores for cross-benchmark consistency.
No Benchgen results yet — be the first to run SRE-Bench.
| Benchmark | What it tests | Saturation |
|---|---|---|
| SRE-Bench | Binary reverse engineering without source access | Medium |
| ExploitBench | Offensive exploit construction (V8 bugs) | Low |
| SEC-Bench Pro | Defensive security engineering | Medium |
SRE-Bench sits between offensive and defensive security work — reverse engineering is a precursor skill useful for both finding new vulnerabilities and understanding existing malware.
Benchgen lets teams track reverse-engineering capability across model versions over time, catching regressions a single vendor-reported snapshot would miss.
Benchmark definition based on OpenAI's GPT-6 Astra announcement. State-of-the-art scores sourced from the same and attributed inline. Last updated 2026-09-07.