Benchgen

Qwen3GuardTest — Results

RankModelScore
1gpt-oss-safeguard-20b85
2qwen3guard-8b84.2
3shieldstral-1-082.9
4nemotron-3-5-content-safety-4b80
5llamaguard-4-12b60.6
6shieldgemma-9b38.7
Q

Qwen3GuardTest

1 phaseActive

Alibaba's held-out evaluation set from the Qwen3Guard technical report, used to score prompt/response safety classification. Zhao et al., 2025.

Overview

Qwen3GuardTest

Category Metric Saturation

Paper GitHub

Quick answer: Qwen3GuardTest is Alibaba's held-out safety evaluation benchmark released alongside Qwen3Guard, used to score prompt and response safety classification. It's referenced by competing guard-model releases (including Shieldstral) as a comparison point outside of each vendor's own internal test sets.

At a Glance

What it tests: Safety classification accuracy against Alibaba's own evaluation methodology for the Qwen3Guard family of guard models. Why it matters: As a benchmark released alongside a specific guard model family, Qwen3GuardTest gives an independent-vendor cross-check — seeing how a competing classifier performs on Alibaba's own test set complements evaluation on more neutral, third-party benchmarks like WildGuardTest. Known limitations: Full test set composition and exact size are not broken out in public third-party comparisons; treat scores here as directionally useful rather than precisely reproducible without access to Alibaba's original evaluation harness.

What Qwen3GuardTest Measures

Qwen3Guard is Alibaba's guard-model family (available in Gen and Stream variants across multiple sizes) built for real-time and generation-time safety classification. Qwen3GuardTest is the evaluation methodology and held-out data referenced in the Qwen3Guard technical report, used both to validate Qwen3Guard itself and, increasingly, as a cross-vendor comparison benchmark for other safety classifiers such as Shieldstral.

Benchmark Specifications

FieldValue
Task categorySafety / content moderation
MetricF1 score (%)
Created byQwen Team (Alibaba)
PaperQwen3Guard Technical Report (arXiv 2510.14276)
GitHubQwenLM/Qwen3Guard
LicenseApache 2.0

How Qwen3GuardTest Is Scored

Classifiers are scored as F1 against Alibaba's held-out labeled evaluation set, following the methodology described in the Qwen3Guard technical report.

State-of-the-Art Results

Scores from the Shieldstral model card (Mistral AI, August 2026). F1 (%). Qwen3Guard-8B figures reflect an average over strict/loose label mappings per the model card's methodology notes.

Scores sourced from Mistral AI's published model card, shown for context. Not Benchgen measurements. Qwen3Guard-8B scores near the top here, as expected on its own vendor-released evaluation set — but GPT-OSS-Safeguard-20B edges narrowly ahead.

Qwen3GuardTest on Benchgen

No Benchgen results yet — be the first to run Qwen3GuardTest.

Qwen3GuardTest vs Other Benchmarks

BenchmarkWhat it testsSaturation
Qwen3GuardTestVendor-released safety classification evalMedium
WildGuardTest (Prompt)Neutral third-party prompt-harm classificationMedium
Aegis v2 (Prompt)Neutral third-party fine-grained taxonomy classificationMedium

Run Qwen3GuardTest on Your Model

Benchgen tracks version-controlled, regression-tested guard-model performance — submit your classifier's Qwen3GuardTest scores to compare against Shieldstral and other guard models.

Frequently Asked Questions

What is Qwen3GuardTest? Qwen3GuardTest is Alibaba's held-out safety evaluation benchmark from the Qwen3Guard technical report, used to validate the Qwen3Guard model family and increasingly cited by competing guard models for comparison.
Why does Qwen3Guard-8B score highest here? This is Alibaba's own vendor-released evaluation set for the Qwen3Guard family, so some home-court advantage is expected relative to neutral third-party benchmarks.
Who created Qwen3GuardTest? The Qwen Team at Alibaba, published in the Qwen3Guard technical report (2025).

Last updated 2026-08-12.