Quick answer: ShieldGemma 2 (4B) is Google's 4B multimodal safety classifier built on Gemma 3, focused primarily on image content moderation (unsafe image categories like dangerous content, sexually explicit material, and violence). It's Google's successor to the text-only ShieldGemma-9B, adding native vision support in a smaller footprint.
Where it leads: Not a top scorer in Shieldstral's published comparison set (limited public score overlap). Where it lags: No text-only prompt/response classification scores reported in Mistral's comparison — ShieldGemma 2 focuses on image classification specifically. Best for: Teams needing a lightweight, Gemma-family image safety classifier alongside a separate text moderation stack.
ShieldGemma 2 is built on Gemma 3's 4B vision-language model, specialized for classifying images as safe or unsafe across a defined harm taxonomy. It represents Google's shift toward native multimodal safety classification, succeeding the text-only ShieldGemma-9B.
| Field | Value |
|---|---|
| Organization | |
| Parameters | 4B |
| Base model | Gemma 3 |
| License | Gemma Terms of Use |
| Modality | Multimodal (primarily image) |
Open weights, free to download under Gemma's Terms of Use.
Scores from Mistral AI's Shieldstral model card, shown for context. Not Benchgen measurements.
| Benchmark | Score |
|---|---|
| VLGuard | 61.3% |
| UnsafeBench | 54.9% |
| LlavaGuard | 56.2% |
Last updated 2026-08-12.
This model isn’t on any benchmark leaderboard yet.