Benchgen

RTP-LX Prompt — Results

RankModelScore
1nemotron-3-5-content-safety-4b86.1
2gpt-oss-safeguard-20b83.9
3shieldstral-1-070.3
4qwen3guard-8b67.3
5llamaguard-4-12b43.9
6shieldgemma-9b36.7
R

RTP-LX Prompt

1 phaseActive

Microsoft's multilingual extension of RealToxicityPrompts, spanning 28 languages, prompt-classification task. de Wynter et al., 2024.

Overview

RTP-LX Prompt

Category Metric Saturation

Paper GitHub

Quick answer: RTP-LX Prompt measures multilingual toxicity classification across roughly 28 languages, using Microsoft's human-transcreated (not just translated) extension of the original RealToxicityPrompts benchmark. It's a relative weak spot for Shieldstral, where a larger competing model leads by double digits.

At a Glance

What it tests: Whether a classifier correctly flags toxic prompts across a broad set of languages, including many outside the typical high-resource set covered by other multilingual safety benchmarks. Why it matters: RTP-LX uses human "transcreation" (culturally adapted translation, not literal machine translation) so toxicity patterns reflect real linguistic and cultural nuance rather than translation artifacts. Known limitations: Per-language subset sizes vary; some included languages have comparatively little training data available for classifiers to learn from, which can depress aggregate scores.

What RTP-LX Prompt Measures

RTP-LX extends the RealToxicityPrompts methodology to roughly 28 languages via human transcreation and annotation, rather than relying purely on machine translation. This task scores whether a classifier's prompt-toxicity judgment holds up across that broader, culturally-adapted multilingual distribution.

Benchmark Specifications

FieldValue
Task categorySafety / multilingual toxicity
MetricF1 score (%)
Languages~28
Created byAdrian de Wynter, Ishaan Watts, Nektaria Potha, and colleagues (Microsoft)
PaperRTP-LX (arXiv 2404.14397)
GitHubmicrosoft/RTP-LX
LicenseResearch use (see repository terms)

How RTP-LX Prompt Is Scored

Each prompt's ground-truth toxicity label is compared against a classifier's prediction and scored as F1, aggregated across the covered languages.

State-of-the-Art Results

Scores from the Shieldstral model card (Mistral AI, August 2026). F1 (%).

Scores sourced from Mistral AI's published model card, shown for context. Not Benchgen measurements. This is one of the benchmarks where Shieldstral's model card shows a notable gap to the leading model.

RTP-LX Prompt on Benchgen

No Benchgen results yet — be the first to run RTP-LX Prompt.

RTP-LX Prompt vs Other Benchmarks

BenchmarkWhat it testsLanguagesSaturation
RTP-LX PromptMultilingual toxicity classification~28Medium
PolyGuard PromptMultilingual prompt-harm classification17Medium

Run RTP-LX Prompt on Your Model

Benchgen tracks version-controlled, regression-tested guard-model performance — submit your classifier's RTP-LX scores to compare against Shieldstral and other guard models.

Frequently Asked Questions

What is RTP-LX? RTP-LX is Microsoft's multilingual extension of the RealToxicityPrompts benchmark, covering roughly 28 languages via human transcreation rather than literal machine translation.
What does "transcreation" mean? Transcreation adapts content to be culturally and linguistically appropriate in the target language rather than translating it literally — intended to preserve the intent and severity of toxic content across languages.
Who created RTP-LX? Researchers at Microsoft (de Wynter et al., 2024).

Last updated 2026-08-12.