Benchgen
Models/nvidia/

Nemotron-3.5-Content-Safety-4B

DraftPublic

Model Details

Nemotron-3.5-Content-Safety-4B

Organization Pricing License Modality

Quick answer: Nemotron-3.5-Content-Safety-4B is NVIDIA's 4B open-weights safety classifier, trained on the Nemotron Content Safety Dataset V2 (formerly Aegis v2). It's the standout performer on multilingual toxicity (RTP-LX Prompt) in Mistral's Shieldstral comparison, leading all other models by a wide margin there.

At a Glance

Where it leads: RTP-LX Prompt (multilingual toxicity classification) — the clear leader among all comparison models. Where it lags: Refusal-detection and multilingual PolyGuard tasks relative to purpose-built classifiers for those specific tasks. Best for: Teams prioritizing multilingual toxicity detection in a small (4B), self-hostable footprint.

What Nemotron-3.5-Content-Safety-4B Is

Nemotron-3.5-Content-Safety-4B is NVIDIA's compact guard model trained on the Nemotron Content Safety Dataset V2 (Aegis v2), NVIDIA's 12-category, 9-subcategory hazard taxonomy dataset. It's part of NVIDIA's broader NeMo Guardrails ecosystem for content moderation, topic control, and jailbreak detection.

Specifications

FieldValue
OrganizationNVIDIA
Parameters4B
Training dataNemotron Content Safety Dataset V2 (Aegis v2)
LicenseNVIDIA Open Model License
ModalityText only

Pricing

Open weights, free to download.

Public Benchmark Scores

Scores from Mistral AI's Shieldstral model card, shown for context. Not Benchgen measurements.

Frequently Asked Questions

What is Nemotron-3.5-Content-Safety-4B?NVIDIA's 4B open-weights safety classifier trained on the Nemotron Content Safety Dataset V2 (Aegis v2), part of the NeMo Guardrails ecosystem.
Is it multimodal?No, it's text-only.
Is it open source?Yes, released under the NVIDIA Open Model License.

Scores sourced from Mistral AI's Shieldstral model card, shown for context. Last updated 2026-08-12.

Benchmark Leaderboards

This model isn’t on any benchmark leaderboard yet.