Benchgen

PolyGuard Response — Results

RankModelScore
1gpt-oss-safeguard-20b80
2shieldstral-1-078.3
3qwen3guard-8b78.1
4nemotron-3-5-content-safety-4b75.3
5llamaguard-4-12b54.6
6shieldgemma-9b31.8
P

PolyGuard Response

1 phaseActive

A 29,325-item multilingual safety benchmark spanning 17 languages, response-classification task. Kumar et al., 2025.

Overview

PolyGuard Response

Category Metric Saturation

Paper

Quick answer: PolyGuard Response measures multilingual response-harm classification across 17 languages, using the same 29,325-item PolyGuardPrompts set as PolyGuard Prompt but scoring the model's reply rather than the user's request.

At a Glance

What it tests: Whether a classifier correctly flags harmful model responses across 17 languages, given the prompt-response pair. Why it matters: Response-harm judgments must account for both the response text and its language-specific cultural/linguistic context — a harder generalization test than prompt-only classification, especially across low-resource languages. Known limitations: Some languages in the 17-language set have less training/eval data than others, so per-language reliability may vary even though the aggregate F1 is reported as a single number.

What PolyGuard Response Measures

This task scores the response_harm_label field on the 29,325-item PolyGuardPrompts set, testing whether a classifier's judgment of response harmfulness generalizes across the same 17 languages used for prompt classification.

Benchmark Specifications

FieldValue
Task categorySafety / multilingual content moderation
MetricF1 score (%)
Test set size29,325 items
Languages17
Created byPriyanshu Kumar, Devansh Jain, Akhila Yerukola, Liwei Jiang, Himanshu Beniwal, Thomas Hartvigsen, Maarten Sap
PaperPolyGuard (arXiv 2504.04377)
DatasetToxicityPrompts/PolyGuardPrompts
LicenseCC-BY-4.0

How PolyGuard Response Is Scored

Each item's ground-truth response-harm label is compared against a classifier's prediction and scored as F1, aggregated across all 17 languages.

State-of-the-Art Results

Scores from the Shieldstral model card (Mistral AI, August 2026). F1 (%).

Scores sourced from Mistral AI's published model card, shown for context. Not Benchgen measurements.

PolyGuard Response on Benchgen

No Benchgen results yet — be the first to run PolyGuard Response.

PolyGuard Response vs Other Benchmarks

BenchmarkWhat it testsLanguagesSaturation
PolyGuard ResponseMultilingual response-harm classification17Medium
RTP-LX CompletionMultilingual completion toxicity28Medium
WildGuardTest (Response)English-only response-harm classification1Medium

Run PolyGuard Response on Your Model

Benchgen tracks version-controlled, regression-tested guard-model performance — submit your classifier's PolyGuard response scores to compare against Shieldstral and other guard models.

Frequently Asked Questions

What is PolyGuard Response? It's the response-harm classification task on the 29,325-item, 17-language PolyGuardPrompts benchmark.
Why does PolyGuard-Qwen-7B lead this benchmark? PolyGuard-Qwen-7B is fine-tuned directly on PolyGuard training data, giving it a natural advantage on this specific evaluation relative to models not trained on it.
Who created PolyGuard? Priyanshu Kumar, Devansh Jain, Akhila Yerukola, and colleagues (Kumar et al., 2025).

Last updated 2026-08-12.