Quick answer: LlavaGuard-7B is TU Darmstadt and hessian.AI's 7B open-weights vision-language safety classifier, judging image content against a detailed, structured safety policy rather than a fixed label set — the model behind the LlavaGuard benchmark. It's the strongest performer on its own LlavaGuard benchmark and competitive on VLGuard and UnsafeBench in Mistral's Shieldstral comparison.
Where it leads: LlavaGuard — naturally the top scorer on its own namesake benchmark and evaluation methodology. Where it lags: VLGuard and UnsafeBench, where Shieldstral's smaller, more recent multimodal design scores higher. Best for: Teams needing an academic-grade, policy-conditioned image moderation classifier with a research paper trail.
LlavaGuard-7B is built on the LLaVA vision-language architecture, fine-tuned to judge image content against detailed written safety policies (category, rationale, and safety rating) rather than a fixed category label — an approach conceptually similar to Shieldstral's policy-adaptive design but purpose-built for image moderation specifically.
| Field | Value |
|---|---|
| Organization | TU Darmstadt / hessian.AI |
| Parameters | 7B |
| Base architecture | LLaVA |
| License | Apache 2.0 |
| Modality | Multimodal (image + policy text) |
Open weights, free to download.
Scores from Mistral AI's Shieldstral model card, shown for context. Not Benchgen measurements.
| Benchmark | Score |
|---|---|
| LlavaGuard | 81.4% |
| VLGuard | 69.5% |
| UnsafeBench | 63.9% |
Scores sourced from Mistral AI's Shieldstral model card, shown for context. Last updated 2026-08-12.
This model isn’t on any benchmark leaderboard yet.