Benchgen

FORTRESS — Results

RankModelScore
1deepseek-v4-pro98.5
2kimi-k2-598.3
3gpt-5-6-sol98.1
4gemini-3-1-pro98
5kimi-k2-697.2
6claude-fable-596
7inkling95.9
8glm-5-290
9gpt-5-6-sol82.4
10inkling78
11glm-5-271.3
12kimi-k2-665.6
13gemini-3-1-pro65.2
14claude-fable-555.1
15kimi-k2-554.1
16deepseek-v4-pro36

FORTRESS

1 phaseActive

Scale AI's safety benchmark for national security and public safety. 500 adversarial + 500 benign prompts across CBRNE, terrorism, and criminal domains. Two metrics: adversarial refusal rate (higher = safer) and benign pass rate (higher = less over-refusal). CC BY-NC-ND 4.0.

Overview

FORTRESS

Category Prompts Domains License

Paper Dataset

Quick answer: FORTRESS (Frontier Risk Evaluation for National Security and Public Safety) is a Scale AI benchmark designed to measure LLM safeguard robustness on dual-use national security and public safety (NSPS) threats. It contains 500 expert-crafted adversarial prompts across CBRNE, political violence, and criminal/financial domains, each paired with a benign counterpart to detect over-refusals. Two complementary metrics: Adversarial (refusal rate on harmful prompts — higher = safer) and Benign (helpfulness rate on benign look-alikes — higher = less over-refusal). Claude Fable 5 leads on adversarial safety (96.0%) but has significant over-refusal (55.1% benign).

At a Glance

What it tests: Whether LLMs correctly refuse genuinely harmful national security and public safety requests while still being helpful on superficially similar but benign queries. FORTRESS directly surfaces the safety-usefulness trade-off in one benchmark.

Why it matters: Existing safety benchmarks either focus narrowly on general harm (e.g., StrongREJECT) or treat refusal rate as the only success metric. FORTRESS specifically targets high-stakes NSPS domains and provides a parallel benign evaluation, revealing which models over-refuse helpful queries alongside actually harmful ones. The trade-off structure reveals that very safe models (Claude Fable 5: 96% adversarial, 55% benign) and very helpful models (DeepSeek V4 Pro: 36% adversarial, 98.5% benign) are on opposite ends.

Known limitations: A private evaluation set is maintained alongside the public set to preserve leaderboard integrity. The benchmark covers only three NSPS domains (CBRNE, political violence, criminal/financial) using unclassified information. CC BY-NC-ND 4.0 license restricts commercial use and derivatives.

What FORTRESS Measures

FORTRESS contains 500 adversarial prompts paired with 500 benign counterparts across 3 high-stakes domains with 10 subcategories:

DomainSubcategories
CBRNE (Chemical, Biological, Radiological, Nuclear, Explosive)Chemical synthesis, biological agents, explosives, etc.
Political Violence & TerrorismExtremist content, recruitment, operational planning
Criminal & Financial Illicit ActivitiesFraud, money laundering, illegal markets

Each prompt has a 4–7 question rubric enabling automated claim-level evaluation.

Two complementary metrics:

  • Adversarial refusal rate (% harmful prompts correctly refused): Higher = safer
  • Benign pass rate (% benign look-alike prompts correctly answered): Higher = more useful / fewer false refusals

Benchmark Specifications

FieldValue
Task categorySafety / NSPS evaluation
Adversarial metric% refusal (higher = safer)
Benign metric% pass rate (higher = more useful)
Number of tasks500 adversarial + 500 benign
Domains3 (CBRNE, Political Violence, Criminal/Financial)
Subcategories10
ScoringInstance-based rubric (4–7 binary questions per prompt)
LicenseCC BY-NC-ND 4.0
Created byChristina Q. Knight, Kaustubh Deshpande, Ved Sirdeshmukh, Meher Mankikar, Scale Red Team, SEAL Research Team, Julian Michael
AffiliationScale AI
PaperFORTRESS: Frontier Risk Evaluation for National Security and Public Safety (arXiv 2506.14922)
DatasetScaleAI/fortress_public on HuggingFace

State-of-the-Art Results

Scores from Inkling model card (Thinking Machines Lab, July 2026). All 9 comparison models reported.

Adversarial Refusal Rate (% harmful prompts correctly refused, higher = safer)

RankModelAdversarialWeights
1Claude Fable 596.0%Closed
2GPT-5.6 Sol82.4%Closed
3Inkling78.0%Open
4Nemotron 3 Ultra 550B77.6%Open
5GLM 5.271.3%Open
6Gemini 3.1 Pro65.2%Closed
7Kimi K2.665.6%Open
8Kimi K2.554.1%Open
9DeepSeek V4 Pro36.0%Open

Benign Pass Rate (% benign look-alike prompts correctly answered, higher = more useful)

RankModelBenignWeights
1DeepSeek V4 Pro98.5%Open
2Kimi K2.598.3%Open
3Kimi K2.697.2%Open
4Gemini 3.1 Pro98.0%Closed
5GPT-5.6 Sol98.1%Closed
6Inkling95.9%Open
7Nemotron 3 Ultra 550B90.5%Open
8GLM 5.290.0%Open
9Claude Fable 555.1%Closed

Last updated 2026-07-16.