Benchgen
Models/allen-institute-for-ai/

WildGuard-7B

DraftPublic

Model Details

WildGuard-7B

Organization Pricing License Modality

Quick answer: WildGuard-7B is the Allen Institute for AI's 7B open-weights safety classifier, the original model trained on WildGuardMix to jointly handle prompt-harm classification, response-harm classification, and refusal detection in a single model — the same test set (WildGuardTest) that many later guard models, including Shieldstral, are benchmarked against.

At a Glance

Where it leads: Historically the reference model for the WildGuardTest benchmark it introduced; naturally a strong baseline on that specific test. Where it lags: Superseded on raw scores by newer, larger, and policy-adaptive models like Shieldstral and Qwen3Guard on several tasks. Best for: Teams wanting the original open-source reference implementation for the WildGuard moderation methodology.

What WildGuard-7B Is

WildGuard-7B was introduced alongside the WildGuardMix dataset (86,759 training examples, 1,725 held-out test items) as AI2's "one-stop" moderation tool — a single model handling prompt harm, response harm, and refusal classification, evaluated against the WildGuardTest benchmark that has since become a common reference point across the guard-model field.

Specifications

FieldValue
OrganizationAllen Institute for AI (AI2)
Parameters7B
Training dataWildGuardMix (86,759 examples)
LicenseApache 2.0
ModalityText only

Pricing

Open weights, free to download.

Public Benchmark Scores

Scores from Mistral AI's Shieldstral model card, shown for context. Not Benchgen measurements. WildGuard-7B is also the original reference model for the WildGuardTest benchmark suite.

Frequently Asked Questions

What is WildGuard-7B?The Allen Institute for AI's 7B open-weights safety classifier, jointly handling prompt-harm, response-harm, and refusal classification.
Is it multimodal?No, WildGuard-7B is text-only.
Is it open source?Yes, released under Apache 2.0.

Last updated 2026-08-12.

Benchmark Leaderboards

This model isn’t on any benchmark leaderboard yet.