| Rank | Model | Score |
|---|---|---|
| 1 | deepseek-v4-pro | 98.5 |
| 2 | kimi-k2-5 | 98.3 |
| 3 | gpt-5-6-sol | 98.1 |
| 4 | gemini-3-1-pro | 98 |
| 5 | kimi-k2-6 | 97.2 |
| 6 | claude-fable-5 | 96 |
| 7 | inkling | 95.9 |
| 8 | glm-5-2 | 90 |
| 9 | gpt-5-6-sol | 82.4 |
| 10 | inkling | 78 |
| 11 | glm-5-2 | 71.3 |
| 12 | kimi-k2-6 | 65.6 |
| 13 | gemini-3-1-pro | 65.2 |
| 14 | claude-fable-5 | 55.1 |
| 15 | kimi-k2-5 | 54.1 |
| 16 | deepseek-v4-pro | 36 |
1 phaseActive
Scale AI's safety benchmark for national security and public safety. 500 adversarial + 500 benign prompts across CBRNE, terrorism, and criminal domains. Two metrics: adversarial refusal rate (higher = safer) and benign pass rate (higher = less over-refusal). CC BY-NC-ND 4.0.
Quick answer: FORTRESS (Frontier Risk Evaluation for National Security and Public Safety) is a Scale AI benchmark designed to measure LLM safeguard robustness on dual-use national security and public safety (NSPS) threats. It contains 500 expert-crafted adversarial prompts across CBRNE, political violence, and criminal/financial domains, each paired with a benign counterpart to detect over-refusals. Two complementary metrics: Adversarial (refusal rate on harmful prompts — higher = safer) and Benign (helpfulness rate on benign look-alikes — higher = less over-refusal). Claude Fable 5 leads on adversarial safety (96.0%) but has significant over-refusal (55.1% benign).
What it tests: Whether LLMs correctly refuse genuinely harmful national security and public safety requests while still being helpful on superficially similar but benign queries. FORTRESS directly surfaces the safety-usefulness trade-off in one benchmark.
Why it matters: Existing safety benchmarks either focus narrowly on general harm (e.g., StrongREJECT) or treat refusal rate as the only success metric. FORTRESS specifically targets high-stakes NSPS domains and provides a parallel benign evaluation, revealing which models over-refuse helpful queries alongside actually harmful ones. The trade-off structure reveals that very safe models (Claude Fable 5: 96% adversarial, 55% benign) and very helpful models (DeepSeek V4 Pro: 36% adversarial, 98.5% benign) are on opposite ends.
Known limitations: A private evaluation set is maintained alongside the public set to preserve leaderboard integrity. The benchmark covers only three NSPS domains (CBRNE, political violence, criminal/financial) using unclassified information. CC BY-NC-ND 4.0 license restricts commercial use and derivatives.
FORTRESS contains 500 adversarial prompts paired with 500 benign counterparts across 3 high-stakes domains with 10 subcategories:
| Domain | Subcategories |
|---|---|
| CBRNE (Chemical, Biological, Radiological, Nuclear, Explosive) | Chemical synthesis, biological agents, explosives, etc. |
| Political Violence & Terrorism | Extremist content, recruitment, operational planning |
| Criminal & Financial Illicit Activities | Fraud, money laundering, illegal markets |
Each prompt has a 4–7 question rubric enabling automated claim-level evaluation.
Two complementary metrics:
| Field | Value |
|---|---|
| Task category | Safety / NSPS evaluation |
| Adversarial metric | % refusal (higher = safer) |
| Benign metric | % pass rate (higher = more useful) |
| Number of tasks | 500 adversarial + 500 benign |
| Domains | 3 (CBRNE, Political Violence, Criminal/Financial) |
| Subcategories | 10 |
| Scoring | Instance-based rubric (4–7 binary questions per prompt) |
| License | CC BY-NC-ND 4.0 |
| Created by | Christina Q. Knight, Kaustubh Deshpande, Ved Sirdeshmukh, Meher Mankikar, Scale Red Team, SEAL Research Team, Julian Michael |
| Affiliation | Scale AI |
| Paper | FORTRESS: Frontier Risk Evaluation for National Security and Public Safety (arXiv 2506.14922) |
| Dataset | ScaleAI/fortress_public on HuggingFace |
Scores from Inkling model card (Thinking Machines Lab, July 2026). All 9 comparison models reported.
| Rank | Model | Adversarial | Weights |
|---|---|---|---|
| 1 | Claude Fable 5 | 96.0% | Closed |
| 2 | GPT-5.6 Sol | 82.4% | Closed |
| 3 | Inkling | 78.0% | Open |
| 4 | Nemotron 3 Ultra 550B | 77.6% | Open |
| 5 | GLM 5.2 | 71.3% | Open |
| 6 | Gemini 3.1 Pro | 65.2% | Closed |
| 7 | Kimi K2.6 | 65.6% | Open |
| 8 | Kimi K2.5 | 54.1% | Open |
| 9 | DeepSeek V4 Pro | 36.0% | Open |
| Rank | Model | Benign | Weights |
|---|---|---|---|
| 1 | DeepSeek V4 Pro | 98.5% | Open |
| 2 | Kimi K2.5 | 98.3% | Open |
| 3 | Kimi K2.6 | 97.2% | Open |
| 4 | Gemini 3.1 Pro | 98.0% | Closed |
| 5 | GPT-5.6 Sol | 98.1% | Closed |
| 6 | Inkling | 95.9% | Open |
| 7 | Nemotron 3 Ultra 550B | 90.5% | Open |
| 8 | GLM 5.2 | 90.0% | Open |
| 9 | Claude Fable 5 | 55.1% | Closed |
Last updated 2026-07-16.