Quick answer: PolyGuard-Qwen-7B is a 7B open-weights multilingual safety classifier fine-tuned on Qwen and trained on PolyGuardMix, a 1.91M-sample multilingual safety corpus spanning 17 languages. It's the strongest performer on PolyGuard's own multilingual benchmarks in Mistral's Shieldstral comparison — an expected result given it's purpose-trained on that exact data distribution.
Where it leads: PolyGuard Prompt and PolyGuard Response — the top score among comparison models on both, reflecting direct fine-tuning on this benchmark's training split. Where it lags: No published scores in Mistral's comparison for English-only benchmarks like WildGuardTest or HarmBench, or for multimodal tasks. Best for: Teams whose primary need is multilingual (17-language) safety classification.
PolyGuard-Qwen-7B is a Qwen-based fine-tune trained on PolyGuardMix, combining naturally occurring multilingual human-LLM interactions with human-verified machine translations of WildGuardMix. Its authors report it outperforms existing open-weight and commercial safety classifiers by an average of 5.5% across the 17 languages it covers.
| Field | Value |
|---|---|
| Organization | ToxicityPrompts (multi-institution research collaboration) |
| Parameters | 7B |
| Base model | Qwen |
| Training data | PolyGuardMix (1.91M samples, 17 languages) |
| License | Apache 2.0 |
| Modality | Text only |
Open weights, free to download.
Scores from Mistral AI's Shieldstral model card, shown for context. Not Benchgen measurements.
| Benchmark | Score |
|---|---|
| PolyGuard (Refusal) | 83.8% |
PolyGuard-Qwen-7B is not among the models reporting prompt/response classification scores on Shieldstral's public comparison table — only its refusal-detection score is directly comparable.
Scores sourced from Mistral AI's Shieldstral model card, shown for context. Last updated 2026-08-12.
This model isn’t on any benchmark leaderboard yet.