Benchgen
Models/toxicityprompts/

PolyGuard-Qwen-7B

DraftPublic

Model Details

PolyGuard-Qwen-7B

Organization Pricing License Modality

Quick answer: PolyGuard-Qwen-7B is a 7B open-weights multilingual safety classifier fine-tuned on Qwen and trained on PolyGuardMix, a 1.91M-sample multilingual safety corpus spanning 17 languages. It's the strongest performer on PolyGuard's own multilingual benchmarks in Mistral's Shieldstral comparison — an expected result given it's purpose-trained on that exact data distribution.

At a Glance

Where it leads: PolyGuard Prompt and PolyGuard Response — the top score among comparison models on both, reflecting direct fine-tuning on this benchmark's training split. Where it lags: No published scores in Mistral's comparison for English-only benchmarks like WildGuardTest or HarmBench, or for multimodal tasks. Best for: Teams whose primary need is multilingual (17-language) safety classification.

What PolyGuard-Qwen-7B Is

PolyGuard-Qwen-7B is a Qwen-based fine-tune trained on PolyGuardMix, combining naturally occurring multilingual human-LLM interactions with human-verified machine translations of WildGuardMix. Its authors report it outperforms existing open-weight and commercial safety classifiers by an average of 5.5% across the 17 languages it covers.

Specifications

FieldValue
OrganizationToxicityPrompts (multi-institution research collaboration)
Parameters7B
Base modelQwen
Training dataPolyGuardMix (1.91M samples, 17 languages)
LicenseApache 2.0
ModalityText only

Pricing

Open weights, free to download.

Public Benchmark Scores

Scores from Mistral AI's Shieldstral model card, shown for context. Not Benchgen measurements.

BenchmarkScore
PolyGuard (Refusal)83.8%

PolyGuard-Qwen-7B is not among the models reporting prompt/response classification scores on Shieldstral's public comparison table — only its refusal-detection score is directly comparable.

Frequently Asked Questions

What is PolyGuard-Qwen-7B?A 7B open-weights multilingual safety classifier fine-tuned on Qwen, trained on the PolyGuardMix dataset spanning 17 languages.
Is it multimodal?No, it's text-only.
Is it open source?Yes, released under Apache 2.0.

Scores sourced from Mistral AI's Shieldstral model card, shown for context. Last updated 2026-08-12.

Benchmark Leaderboards

This model isn’t on any benchmark leaderboard yet.