Benchgen

Aegis v2 (Response) — Results

RankModelScore
1shieldstral-1-087.2
2qwen3guard-8b86.2
3nemotron-3-5-content-safety-4b84.9
4gpt-oss-safeguard-20b75.2
5llamaguard-4-12b64.7
6shieldgemma-9b59.7
A

Aegis v2 (Response)

1 phaseActive

NVIDIA's Nemotron Content Safety Dataset V2, response-classification task. 1,964-item test set, 12-category taxonomy. Ghosh et al., 2025.

Overview

Aegis v2 (Response)

Category Metric Saturation

Paper

Quick answer: Aegis v2 (Response) measures response-harm classification against NVIDIA's 12-category safety taxonomy, using the same 1,964-item held-out test set as Aegis v2 (Prompt) but scoring the LLM's reply rather than the user's request.

At a Glance

What it tests: Whether a classifier correctly flags harmful model responses across NVIDIA's hazard taxonomy, given the prompt-response pair. Why it matters: Response labels in Aegis v2 are hybrid — human-annotated where possible, augmented via an LLM-jury (Mixtral-8x22B, Mistral-NeMo-12B-Instruct, Gemma-2-27B-it) for scale — making this benchmark a useful test of whether a classifier's judgments align with a multi-model consensus process as well as human raters. Known limitations: Responses were generated by a single base model (Mistral-7B-v0.1), so response style/harm patterns may not generalize to responses from more safety-aligned or differently-trained models.

What Aegis v2 (Response) Measures

This task scores the response_label field on the 1,964-item Aegis v2 test set — judging whether the LLM's reply to a (potentially harmful) prompt itself violates the taxonomy, including cases augmented with synthetic refusal data to test whether classifiers correctly recognize safe refusals as non-harmful.

Benchmark Specifications

FieldValue
Task categorySafety / content moderation
MetricF1 score (%)
Test set size1,964 items
Taxonomy12 core + 9 fine-grained hazard categories
Created byShaona Ghosh, Prasoon Varshney, Makesh Narsimhan Sreedhar, Aishwarya Padmakumar, Traian Rebedea, Jibin Rajan Varghese, Christopher Parisien (NVIDIA)
PaperAEGIS2.0 (NAACL 2025)
Datasetnvidia/Aegis-AI-Content-Safety-Dataset-2.0
LicenseCC-BY-4.0

How Aegis v2 (Response) Is Scored

Each of the 1,964 test items' ground-truth response-harm label is compared against a classifier's prediction, scored as F1.

State-of-the-Art Results

Scores from the Shieldstral model card (Mistral AI, August 2026). F1 (%).

Scores sourced from Mistral AI's published model card, shown for context. Not Benchgen measurements.

Aegis v2 (Response) on Benchgen

No Benchgen results yet — be the first to run Aegis v2 (Response).

Aegis v2 (Response) vs Other Benchmarks

BenchmarkWhat it testsTest sizeSaturation
Aegis v2 (Response)Fine-grained taxonomy response classification1,964Medium
WildGuardTest (Response)Response-harm classification1,725Medium
HarmBench (Response)Adversarial jailbreak response classification400Medium
BeaverTailsQA-pair harm classification~700Medium

Run Aegis v2 (Response) on Your Model

Benchgen tracks version-controlled, regression-tested guard-model performance — submit your classifier's Aegis v2 response scores to compare against Shieldstral and other guard models.

Frequently Asked Questions

What is Aegis v2 (Response)? It's the response-harm classification task scored on NVIDIA's 1,964-item Aegis v2 (Nemotron Content Safety Dataset V2) test set.
How were response labels generated? Through a hybrid process: human annotation for the base labels, augmented by an LLM-jury (Mixtral-8x22B, Mistral-NeMo-12B-Instruct, Gemma-2-27B-it) for additional coverage, plus synthetic refusal data.
Who created Aegis v2? NVIDIA's NeMo Guardrails team (Ghosh et al.), published at NAACL 2025.

Last updated 2026-08-12.