Quick answer: LlamaGuard-4-12B is Meta's 12B multimodal safety classifier, the successor to the LlamaGuard line, supporting both text and image moderation. It's the largest model in Shieldstral's comparison set that reports scores across all task categories including multimodal, but generally trails Shieldstral, GPT-OSS-Safeguard, and Qwen3Guard on text tasks.
Where it leads: Not the top performer on any individual benchmark in Shieldstral's comparison, but the only large (12B) multimodal model that reports scores across the full range of text and image tasks. Where it lags: WildGuardTest (Prompt/Response), OpenAI Moderation, and BeaverTails, where smaller Shieldstral and Qwen3Guard-8B outperform it. Best for: Teams already in the Llama ecosystem needing a single multimodal guard model with broad task coverage.
LlamaGuard-4-12B builds on Meta's Llama Guard series (LlamaGuard, LlamaGuard 2, LlamaGuard 3), extending native multimodal support to the 12B parameter class. Like its predecessors, it classifies content against a defined taxonomy of unsafe categories for both prompts and responses, now extended to images.
| Field | Value |
|---|---|
| Organization | Meta |
| Parameters | 12B |
| License | Llama 4 Community License |
| Modality | Multimodal (text + image) |
Open weights, free to download under the Llama 4 Community License (usage restrictions apply for very large deployments).
Scores from Mistral AI's Shieldstral model card, shown for context. Not Benchgen measurements.
| Benchmark | Score |
|---|---|
| WildGuardTest (Prompt) | 74.3% |
| HarmBench (Prompt) | 97.9% |
| VLGuard | 59.9% |
| UnsafeBench | 68.1% |
| LlavaGuard | 74.5% |
Scores sourced from Mistral AI's Shieldstral model card, shown for context. Last updated 2026-08-12.
This model isn’t on any benchmark leaderboard yet.