| Rank | Model | Score |
|---|---|---|
| 1 | gpt-6-astra | 85.4 |
| 2 | deepseek-v4-1-flash | 62.8 |
| 3 | deepseek-v4-1-flash | 62.8 |
1 phaseActive
Professional-tier cybersecurity agent benchmark testing secure code review and defensive security engineering tasks.
Quick answer: SEC-Bench Pro evaluates AI agents on professional-grade defensive security engineering tasks — the kind of secure code review and vulnerability-remediation work security teams do day to day, as distinct from offensive exploit-development benchmarks like ExploitBench. It first appeared publicly in OpenAI's GPT-6 Astra announcement (September 2026), where Astra scored 85.4%, ahead of GPT-5.6 Sol's 79.1%.
What it tests: Professional defensive-security workflows — secure code review, vulnerability identification and remediation, and related engineering tasks a security team would handle in production, rather than offensive exploit construction.
Why it matters: Most public cybersecurity benchmarks (ExploitBench, ExploitGym) focus on offensive capability — can a model build a working exploit. SEC-Bench Pro instead measures the defensive counterpart: can a model be trusted to do the security engineering work that keeps systems safe. This distinction matters for responsible deployment, since offensive and defensive capability don't necessarily scale together.
Known limitations: Full methodology (task count, dataset composition, public availability) has not yet been documented in an independent paper as of this writing — currently known primarily through frontier lab technical reports.
SEC-Bench Pro evaluates models on professional-tier security engineering tasks — the practical, defensive work of a security team, such as secure code review and patching, as opposed to constructing novel exploits against hardened targets. This makes it a natural complement to offensive-capability benchmarks like ExploitBench and ExploitGym: a model could plausibly score very differently on offensive versus defensive security tasks, and tracking both gives a fuller picture of a model's net effect on the security landscape.
| Field | Value |
|---|---|
| Task category | Agent / Cybersecurity (defensive) |
| Metric | % task success |
| Saturation | Medium |
| Created by | OpenAI (introduced via GPT-6 Astra announcement) |
| Public dataset | Not yet publicly released as of Sep 2026 |
Scores are reported as a percentage task-success rate across the benchmark's professional security-engineering task set. Exact grading criteria and task composition have not been fully disclosed publicly as of this writing.
| Rank | Model | Score | Source | Date |
|---|---|---|---|---|
| 1 | GPT-6 Astra | 85.4% | OpenAI: GPT-6 Astra | 2026-09 |
| 2 | GPT-5.6 Sol | 79.1% | OpenAI: GPT-6 Astra | 2026-09 |
Scores sourced from OpenAI's GPT-6 Astra announcement (September 2026).
No Benchgen results yet — be the first to run SEC-Bench Pro.
| Benchmark | What it tests | Saturation |
|---|---|---|
| SEC-Bench Pro | Defensive security engineering (code review, remediation) | Medium |
| ExploitBench | Offensive exploit construction (V8 bugs) | Low |
| ExploitGym | Offensive exploit construction (broader gym) | Low |
| SRE-Bench | Reverse engineering / binary analysis | Low |
SEC-Bench Pro fills the defensive-security gap alongside the mostly-offensive existing cybersecurity benchmark set — important for evaluating net security impact, not just raw exploit capability.
Benchgen lets teams track defensive security-engineering performance across model versions, catching regressions a single vendor-reported snapshot would miss.
Benchmark definition based on OpenAI's GPT-6 Astra announcement. State-of-the-art scores sourced from the same and attributed inline. Last updated 2026-09-07.