Benchgen

PolyGuard (Refusal) — Results

RankModelScore
1gpt-oss-safeguard-20b92.3
2wildguard-7b89.6
3shieldstral-1-089.5
4qwen3guard-8b89.3
5polyguard-qwen-7b83.8
P

PolyGuard (Refusal)

1 phaseActive

A 29,325-item multilingual safety benchmark spanning 17 languages, refusal-detection task. Kumar et al., 2025.

Overview

PolyGuard (Refusal)

Category Metric Saturation

Paper

Quick answer: PolyGuard (Refusal) measures whether a classifier can detect refusal vs. compliance across 17 languages, using the same 29,325-item PolyGuardPrompts set as the prompt and response tasks, but scoring the response_refusal_label field.

At a Glance

What it tests: Whether a classifier correctly labels a model's multilingual response as a refusal or a compliance, independent of harm. Why it matters: Refusal detection at scale across 17 languages helps teams monitor over-refusal and under-refusal patterns in multilingual deployments, where refusal-triggering phrases and cultural norms around sensitive topics vary by language. Known limitations: Refusal phrasing conventions differ across languages and models, which can introduce more label ambiguity than in English-only refusal benchmarks like XSTest.

What PolyGuard (Refusal) Measures

This task scores the response_refusal_label field on the 29,325-item PolyGuardPrompts set — testing whether a classifier's refusal/compliance judgment generalizes across the same 17-language distribution used for prompt- and response-harm classification.

Benchmark Specifications

FieldValue
Task categorySafety / multilingual refusal detection
MetricF1 score (%)
Test set size29,325 items
Languages17
Created byPriyanshu Kumar, Devansh Jain, Akhila Yerukola, Liwei Jiang, Himanshu Beniwal, Thomas Hartvigsen, Maarten Sap
PaperPolyGuard (arXiv 2504.04377)
DatasetToxicityPrompts/PolyGuardPrompts
LicenseCC-BY-4.0

How PolyGuard (Refusal) Is Scored

Each item's ground-truth refusal label is compared against a classifier's prediction and scored as F1, aggregated across all 17 languages.

State-of-the-Art Results

Scores from the Shieldstral model card (Mistral AI, August 2026). F1 (%).

Scores sourced from Mistral AI's published model card, shown for context. Not Benchgen measurements. LlamaGuard-4, ShieldGemma, and Nemotron-3.5-Content-Safety variants do not report refusal-detection scores in this comparison.

PolyGuard (Refusal) on Benchgen

No Benchgen results yet — be the first to run PolyGuard (Refusal).

PolyGuard (Refusal) vs Other Benchmarks

BenchmarkWhat it testsLanguagesSaturation
PolyGuard (Refusal)Multilingual refusal detection17Medium
WildGuardTest (Refusal)English-only refusal detection1Medium
XSTest (Refusal)Over-refusal detection1High

Run PolyGuard (Refusal) on Your Model

Benchgen tracks version-controlled, regression-tested guard-model performance — submit your classifier's PolyGuard refusal scores to compare against Shieldstral and other guard models.

Frequently Asked Questions

What is PolyGuard (Refusal)? It's the refusal-detection task on the 29,325-item, 17-language PolyGuardPrompts benchmark.
How does this differ from XSTest (Refusal)? XSTest is English-only and focused on exaggerated-safety trap prompts; PolyGuard (Refusal) tests general refusal detection across 17 languages on naturally occurring and translated data.
Who created PolyGuard? Priyanshu Kumar, Devansh Jain, Akhila Yerukola, and colleagues (Kumar et al., 2025).

Last updated 2026-08-12.