Quick answer: ShieldGemma-9B is Google's 9B open-weights text safety classifier, built on the Gemma architecture and released alongside the original Gemma model family. It trails the newer, smaller Shieldstral, Qwen3Guard, and GPT-OSS-Safeguard models across most benchmarks in Mistral's comparison, reflecting its earlier (2024) release relative to those 2025-2026 classifiers.
Where it leads: Not the top performer on any benchmark in Shieldstral's comparison set. Where it lags: Notably behind on WildGuardTest (Prompt) and HarmBench (Prompt), where newer guard models have made substantial gains. Best for: Teams already using Gemma-family models who want a same-family, open-weights baseline classifier.
ShieldGemma-9B is one of Google's earliest dedicated open-weights safety classifiers, built on the Gemma architecture to classify prompts and responses against a defined harm taxonomy (including categories like sexually explicit content, hate speech, harassment, and dangerous content).
| Field | Value |
|---|---|
| Organization | |
| Parameters | 9B |
| Base model | Gemma |
| License | Gemma Terms of Use |
| Modality | Text only |
Open weights, free to download under Gemma's Terms of Use.
Scores from Mistral AI's Shieldstral model card, shown for context. Not Benchgen measurements.
| Benchmark | Score |
|---|---|
| WildGuardTest (Prompt) | 46.0% |
| HarmBench (Prompt) | 50.2% |
Scores sourced from Mistral AI's Shieldstral model card, shown for context. Last updated 2026-08-12.
This model isn’t on any benchmark leaderboard yet.