Benchgen

SEC-Bench Pro — Results

RankModelScore
1gpt-6-astra85.4
2deepseek-v4-1-flash62.8
3deepseek-v4-1-flash62.8
S

SEC-Bench Pro

1 phaseActive

Professional-tier cybersecurity agent benchmark testing secure code review and defensive security engineering tasks.

Overview

SEC-Bench Pro

Category Metric Saturation Created

Quick answer: SEC-Bench Pro evaluates AI agents on professional-grade defensive security engineering tasks — the kind of secure code review and vulnerability-remediation work security teams do day to day, as distinct from offensive exploit-development benchmarks like ExploitBench. It first appeared publicly in OpenAI's GPT-6 Astra announcement (September 2026), where Astra scored 85.4%, ahead of GPT-5.6 Sol's 79.1%.

At a Glance

What it tests: Professional defensive-security workflows — secure code review, vulnerability identification and remediation, and related engineering tasks a security team would handle in production, rather than offensive exploit construction.

Why it matters: Most public cybersecurity benchmarks (ExploitBench, ExploitGym) focus on offensive capability — can a model build a working exploit. SEC-Bench Pro instead measures the defensive counterpart: can a model be trusted to do the security engineering work that keeps systems safe. This distinction matters for responsible deployment, since offensive and defensive capability don't necessarily scale together.

Known limitations: Full methodology (task count, dataset composition, public availability) has not yet been documented in an independent paper as of this writing — currently known primarily through frontier lab technical reports.

What SEC-Bench Pro Measures

SEC-Bench Pro evaluates models on professional-tier security engineering tasks — the practical, defensive work of a security team, such as secure code review and patching, as opposed to constructing novel exploits against hardened targets. This makes it a natural complement to offensive-capability benchmarks like ExploitBench and ExploitGym: a model could plausibly score very differently on offensive versus defensive security tasks, and tracking both gives a fuller picture of a model's net effect on the security landscape.

Benchmark Specifications

FieldValue
Task categoryAgent / Cybersecurity (defensive)
Metric% task success
SaturationMedium
Created byOpenAI (introduced via GPT-6 Astra announcement)
Public datasetNot yet publicly released as of Sep 2026

How SEC-Bench Pro Is Scored

Scores are reported as a percentage task-success rate across the benchmark's professional security-engineering task set. Exact grading criteria and task composition have not been fully disclosed publicly as of this writing.

State-of-the-Art Results

RankModelScoreSourceDate
1GPT-6 Astra85.4%OpenAI: GPT-6 Astra2026-09
2GPT-5.6 Sol79.1%OpenAI: GPT-6 Astra2026-09

Scores sourced from OpenAI's GPT-6 Astra announcement (September 2026).

SEC-Bench Pro on Benchgen

No Benchgen results yet — be the first to run SEC-Bench Pro.

SEC-Bench Pro vs Other Benchmarks

BenchmarkWhat it testsSaturation
SEC-Bench ProDefensive security engineering (code review, remediation)Medium
ExploitBenchOffensive exploit construction (V8 bugs)Low
ExploitGymOffensive exploit construction (broader gym)Low
SRE-BenchReverse engineering / binary analysisLow

SEC-Bench Pro fills the defensive-security gap alongside the mostly-offensive existing cybersecurity benchmark set — important for evaluating net security impact, not just raw exploit capability.

Run SEC-Bench Pro on Your Model

Benchgen lets teams track defensive security-engineering performance across model versions, catching regressions a single vendor-reported snapshot would miss.

Frequently Asked Questions

What is SEC-Bench Pro? A benchmark testing AI agents on professional defensive security engineering tasks — secure code review and vulnerability remediation — introduced publicly via OpenAI's GPT-6 Astra announcement in September 2026.
What does a good SEC-Bench Pro score look like? As of September 2026, frontier models score in the high 70s to mid 80s percent, with GPT-6 Astra leading at 85.4%.
Who created SEC-Bench Pro? SEC-Bench Pro was introduced by OpenAI, first publicly referenced in the GPT-6 Astra model announcement (September 2026).
How does SEC-Bench Pro differ from ExploitBench? ExploitBench measures offensive exploit construction against hardened real-world bugs. SEC-Bench Pro measures the defensive counterpart — secure code review and remediation work.

Benchmark definition based on OpenAI's GPT-6 Astra announcement. State-of-the-art scores sourced from the same and attributed inline. Last updated 2026-09-07.