Benchgen

PolyGuard Prompt — Results

RankModelScore
1shieldstral-1-084.6
2qwen3guard-8b84.3
3gpt-oss-safeguard-20b83
4nemotron-3-5-content-safety-4b80.5
5llamaguard-4-12b62.1
6shieldgemma-9b33.8
P

PolyGuard Prompt

1 phaseActive

A 29,325-item multilingual safety benchmark spanning 17 languages, prompt-classification task. Kumar et al., 2025.

Overview

PolyGuard Prompt

Category Metric Saturation

Paper

Quick answer: PolyGuard Prompt measures prompt-harm classification across 17 languages using PolyGuardPrompts, a 29,325-item evaluation benchmark combining naturally occurring multilingual human-LLM interactions with human-verified machine translations of WildGuardMix. It's one of the largest multilingual safety test sets available.

At a Glance

What it tests: Whether a classifier correctly flags harmful prompts across 17 languages including Arabic, Chinese, Czech, German, Hindi, Japanese, Korean, Russian, Spanish, and Thai, among others. Why it matters: Most safety classifiers are trained and evaluated primarily in English; PolyGuard was built to close that gap, and its authors report that PolyGuard-trained models outperform existing open-weight and commercial safety classifiers by 5.5% on average across languages. Known limitations: Language coverage (17 languages) is broad but not exhaustive; translation-derived items may carry some translation-artifact noise even with human verification.

What PolyGuard Prompt Measures

PolyGuardPrompts contains 29,325 items across 17 languages, built by combining real multilingual human–LLM interactions with human-verified machine translations of the English-only WildGuardMix dataset. This task scores the prompt_harm_label field — testing whether a classifier generalizes its harm judgments beyond English to a genuinely multilingual distribution of prompts.

Benchmark Specifications

FieldValue
Task categorySafety / multilingual content moderation
MetricF1 score (%)
Test set size29,325 items
Languages17
Created byPriyanshu Kumar, Devansh Jain, Akhila Yerukola, Liwei Jiang, Himanshu Beniwal, Thomas Hartvigsen, Maarten Sap
PaperPolyGuard (arXiv 2504.04377)
DatasetToxicityPrompts/PolyGuardPrompts
LicenseCC-BY-4.0

How PolyGuard Prompt Is Scored

Each of the 29,325 items' ground-truth prompt-harm label is compared against a classifier's prediction and scored as F1, aggregated across all 17 languages.

State-of-the-Art Results

Scores from the Shieldstral model card (Mistral AI, August 2026). F1 (%).

Scores sourced from Mistral AI's published model card, shown for context. Not Benchgen measurements.

PolyGuard Prompt on Benchgen

No Benchgen results yet — be the first to run PolyGuard Prompt.

PolyGuard Prompt vs Other Benchmarks

BenchmarkWhat it testsLanguagesSaturation
PolyGuard PromptMultilingual prompt-harm classification17Medium
RTP-LX PromptMultilingual toxicity classification28Medium
WildGuardTest (Prompt)English-only prompt-harm classification1Medium

Run PolyGuard Prompt on Your Model

Benchgen tracks version-controlled, regression-tested guard-model performance — submit your classifier's PolyGuard scores to compare against Shieldstral and other guard models.

Frequently Asked Questions

What is PolyGuard? PolyGuard is a multilingual safety moderation model and dataset family covering 17 languages. PolyGuardPrompts, the evaluation set used here, contains 29,325 human-verified prompts.
Which languages does PolyGuard cover? Arabic, Chinese, Czech, English, German, Hindi, Italian, Japanese, Korean, Dutch, Polish, Portuguese, Russian, Swedish, Thai, and others — 17 in total.
Who created PolyGuard? Priyanshu Kumar, Devansh Jain, Akhila Yerukola, and colleagues (Kumar et al., 2025).

Last updated 2026-08-12.