| Rank | Model | Score |
|---|---|---|
| 1 | gpt-oss-safeguard-20b | 85 |
| 2 | qwen3guard-8b | 84.2 |
| 3 | shieldstral-1-0 | 82.9 |
| 4 | nemotron-3-5-content-safety-4b | 80 |
| 5 | llamaguard-4-12b | 60.6 |
| 6 | shieldgemma-9b | 38.7 |
1 phaseActive
Alibaba's held-out evaluation set from the Qwen3Guard technical report, used to score prompt/response safety classification. Zhao et al., 2025.
Quick answer: Qwen3GuardTest is Alibaba's held-out safety evaluation benchmark released alongside Qwen3Guard, used to score prompt and response safety classification. It's referenced by competing guard-model releases (including Shieldstral) as a comparison point outside of each vendor's own internal test sets.
What it tests: Safety classification accuracy against Alibaba's own evaluation methodology for the Qwen3Guard family of guard models. Why it matters: As a benchmark released alongside a specific guard model family, Qwen3GuardTest gives an independent-vendor cross-check — seeing how a competing classifier performs on Alibaba's own test set complements evaluation on more neutral, third-party benchmarks like WildGuardTest. Known limitations: Full test set composition and exact size are not broken out in public third-party comparisons; treat scores here as directionally useful rather than precisely reproducible without access to Alibaba's original evaluation harness.
Qwen3Guard is Alibaba's guard-model family (available in Gen and Stream variants across multiple sizes) built for real-time and generation-time safety classification. Qwen3GuardTest is the evaluation methodology and held-out data referenced in the Qwen3Guard technical report, used both to validate Qwen3Guard itself and, increasingly, as a cross-vendor comparison benchmark for other safety classifiers such as Shieldstral.
| Field | Value |
|---|---|
| Task category | Safety / content moderation |
| Metric | F1 score (%) |
| Created by | Qwen Team (Alibaba) |
| Paper | Qwen3Guard Technical Report (arXiv 2510.14276) |
| GitHub | QwenLM/Qwen3Guard |
| License | Apache 2.0 |
Classifiers are scored as F1 against Alibaba's held-out labeled evaluation set, following the methodology described in the Qwen3Guard technical report.
Scores from the Shieldstral model card (Mistral AI, August 2026). F1 (%). Qwen3Guard-8B figures reflect an average over strict/loose label mappings per the model card's methodology notes.
| Rank | Model | Score | Source | Date |
|---|---|---|---|---|
| 1 | GPT-OSS-Safeguard-20B | 85% | Mistral AI model card | 2026-08 |
| 2 | Qwen3Guard-8B | 84.2% | Mistral AI model card | 2026-08 |
| 3 | Shieldstral 1.0 | 82.9% | Mistral AI model card | 2026-08 |
| 4 | Nemotron-3.5-Content-Safety-4B | 80% | Mistral AI model card | 2026-08 |
| 5 | LlamaGuard-4-12B | 60.6% | Mistral AI model card | 2026-08 |
| 6 | ShieldGemma-9B | 38.7% | Mistral AI model card | 2026-08 |
Scores sourced from Mistral AI's published model card, shown for context. Not Benchgen measurements. Qwen3Guard-8B scores near the top here, as expected on its own vendor-released evaluation set — but GPT-OSS-Safeguard-20B edges narrowly ahead.
No Benchgen results yet — be the first to run Qwen3GuardTest.
| Benchmark | What it tests | Saturation |
|---|---|---|
| Qwen3GuardTest | Vendor-released safety classification eval | Medium |
| WildGuardTest (Prompt) | Neutral third-party prompt-harm classification | Medium |
| Aegis v2 (Prompt) | Neutral third-party fine-grained taxonomy classification | Medium |
Benchgen tracks version-controlled, regression-tested guard-model performance — submit your classifier's Qwen3GuardTest scores to compare against Shieldstral and other guard models.
Last updated 2026-08-12.