Benchgen

LlavaGuard — Results

RankModelScore
1llavaguard-7b81.4
2shieldstral-1-072
3omniguard-7b71.7
4nemotron-3-5-content-safety-4b70
5shieldgemma-2-4b56.2
6llamaguard-4-12b21.9
L

LlavaGuard

1 phaseActive

A vision-language safety benchmark with fine-grained, policy-based safety taxonomy for image content. Helff et al., 2024.

Overview

LlavaGuard

Category Metric Saturation

Paper GitHub

Quick answer: LlavaGuard measures fine-grained, policy-based multimodal safety classification — going beyond a binary safe/unsafe label to judge image content against a detailed, customizable safety policy, similar in spirit to Shieldstral's own policy-adaptive approach but purpose-built for vision-language moderation.

At a Glance

What it tests: Whether a classifier can apply a detailed, structured safety policy (rather than a fixed category label) to image content, judging compliance the way a human moderator using written guidelines would. Why it matters: LlavaGuard's policy-driven format anticipated the kind of flexible, instruction-conditioned moderation that Shieldstral itself is built around — making this benchmark a natural point of comparison for policy-adaptive safety classifiers on images. Known limitations: Smaller-scale academic release relative to the large industrial guard datasets (e.g., Aegis v2, WildGuardMix); exact held-out test set size is not broken out precisely in public third-party comparisons.

What LlavaGuard Measures

LlavaGuard evaluates vision-language models against a structured safety taxonomy, judging image content according to a detailed written policy (categories, rationale, and safety rating) rather than a simple binary label — testing whether a classifier's judgment aligns with policy-grounded human annotation on visual content.

Benchmark Specifications

FieldValue
Task categorySafety / multimodal policy-based moderation
MetricF1 score (%)
ModalityImage + policy text
Created byLukas Helff, Felix Friedrich, Manuel Brack, Kristian Kersting, Patrick Schramowski
AffiliationTU Darmstadt, hessian.AI
PaperLlavaGuard (arXiv 2406.05113)
GitHubml-research/LlavaGuard
DatasetAIML-TUDA/LlavaGuard

How LlavaGuard Is Scored

Each image-policy pair's ground-truth safety judgment is compared against a classifier's prediction and scored as F1.

State-of-the-Art Results

Scores from the Shieldstral model card (Mistral AI, August 2026). F1 (%). Only models with native multimodal support report scores here.

Scores sourced from Mistral AI's published model card, shown for context. Not Benchgen measurements. LlavaGuard-7B naturally leads on this vendor-adjacent evaluation, being purpose-trained on this dataset's methodology; text-only guard models do not report scores here.

LlavaGuard on Benchgen

No Benchgen results yet — be the first to run LlavaGuard.

LlavaGuard vs Other Benchmarks

BenchmarkWhat it testsModalitySaturation
LlavaGuardPolicy-based multimodal safety judgmentImage + policy textMedium
VLGuardImage-instruction pair safetyImage + textHigh
UnsafeBenchStandalone unsafe image classificationImage onlyMedium

Run LlavaGuard on Your Model

Benchgen tracks version-controlled, regression-tested guard-model performance — submit your classifier's LlavaGuard scores to compare against Shieldstral and other multimodal guard models.

Frequently Asked Questions

What is LlavaGuard? LlavaGuard is a vision-language safety benchmark that judges image content against a detailed, structured safety policy rather than a simple binary label.
How does this compare to Shieldstral's own approach? Both use policy-conditioned judgments rather than fixed categories, though LlavaGuard was purpose-built for image moderation specifically, while Shieldstral applies the same idea across text, image, and combined moderation.
Who created LlavaGuard? Researchers at TU Darmstadt and hessian.AI (Helff et al., 2024).

Last updated 2026-08-12.