Benchgen

Aegis v2 (Prompt) — Results

RankModelScore
1nemotron-3-5-content-safety-4b86.3
2shieldstral-1-086.2
3qwen3guard-8b84.6
4gpt-oss-safeguard-20b84.4
5llamaguard-4-12b71.5
6shieldgemma-9b65.8
A

Aegis v2 (Prompt)

1 phaseActive

NVIDIA's Nemotron Content Safety Dataset V2 (formerly Aegis 2.0), prompt-classification task. 1,964-item test set, 12-category taxonomy. Ghosh et al., 2025.

Overview

Aegis v2 (Prompt)

Category Metric Saturation

Paper

Quick answer: Aegis v2 (Prompt) measures prompt-harm classification against NVIDIA's 12-category (plus 9 fine-grained subcategory) safety taxonomy, using a 1,964-item held-out test set from the Nemotron Content Safety Dataset V2 (formerly Aegis AI Content Safety Dataset 2.0).

At a Glance

What it tests: Whether a classifier correctly flags harmful user prompts across a broad, fine-grained hazard taxonomy spanning hate speech, self-harm, weapons, criminal planning, and more. Why it matters: Aegis v2's taxonomy is one of the more comprehensive in guard-model literature (12 core + 9 fine-grained categories), and its responses were generated by Mistral-7B-v0.1 specifically because that model has low built-in refusal rates — producing genuinely harmful completions to label rather than safety-filtered ones. Known limitations: Primarily English; response data generated by a single base model (Mistral-7B-v0.1) rather than sourced from diverse production traffic.

What Aegis v2 (Prompt) Measures

The Nemotron Content Safety Dataset V2 contains 33,416 annotated human–LLM interactions (30,007 train / 1,445 validation / 1,964 test), hybrid-labeled via human annotation with LLM-jury augmentation for response labels. Aegis v2 (Prompt) scores the prompt-harm label specifically — classifying whether the initial user request violates any of the taxonomy's hazard categories, independent of the response.

Benchmark Specifications

FieldValue
Task categorySafety / content moderation
MetricF1 score (%)
Test set size1,964 items
Taxonomy12 core + 9 fine-grained hazard categories
Created byShaona Ghosh, Prasoon Varshney, Makesh Narsimhan Sreedhar, Aishwarya Padmakumar, Traian Rebedea, Jibin Rajan Varghese, Christopher Parisien (NVIDIA)
PaperAEGIS2.0 (NAACL 2025)
Datasetnvidia/Aegis-AI-Content-Safety-Dataset-2.0
LicenseCC-BY-4.0

How Aegis v2 (Prompt) Is Scored

Each of the 1,964 test items' human-annotated prompt-harm label is compared against a classifier's prediction, scored as F1.

State-of-the-Art Results

Scores from the Shieldstral model card (Mistral AI, August 2026). F1 (%).

Scores sourced from Mistral AI's published model card, shown for context. Not Benchgen measurements.

Aegis v2 (Prompt) on Benchgen

No Benchgen results yet — be the first to run Aegis v2 (Prompt).

Aegis v2 (Prompt) vs Other Benchmarks

BenchmarkWhat it testsTest sizeSaturation
Aegis v2 (Prompt)Fine-grained taxonomy prompt classification1,964Medium
WildGuardTest (Prompt)Prompt-harm classification1,725Medium
HarmBench (Prompt)Adversarial jailbreak prompt classification400Medium

Run Aegis v2 (Prompt) on Your Model

Benchgen tracks version-controlled, regression-tested guard-model performance — submit your classifier's Aegis v2 scores to compare against Shieldstral and other guard models.

Frequently Asked Questions

What is Aegis v2? Aegis v2 (now called the Nemotron Content Safety Dataset V2) is NVIDIA's safety classification dataset spanning a 12-category, 9-subcategory hazard taxonomy, with prompt and response classification tasks scored separately.
Why was it renamed from "Aegis" to "Nemotron Content Safety Dataset"? NVIDIA consolidated its safety datasets under the Nemotron branding; the underlying dataset and taxonomy are unchanged.
Who created Aegis v2? NVIDIA's NeMo Guardrails team (Ghosh et al.), published at NAACL 2025.

Last updated 2026-08-12.