Quick answer: Nemotron-3.5-Content-Safety-4B is NVIDIA's 4B open-weights safety classifier, trained on the Nemotron Content Safety Dataset V2 (formerly Aegis v2). It's the standout performer on multilingual toxicity (RTP-LX Prompt) in Mistral's Shieldstral comparison, leading all other models by a wide margin there.
Where it leads: RTP-LX Prompt (multilingual toxicity classification) — the clear leader among all comparison models. Where it lags: Refusal-detection and multilingual PolyGuard tasks relative to purpose-built classifiers for those specific tasks. Best for: Teams prioritizing multilingual toxicity detection in a small (4B), self-hostable footprint.
Nemotron-3.5-Content-Safety-4B is NVIDIA's compact guard model trained on the Nemotron Content Safety Dataset V2 (Aegis v2), NVIDIA's 12-category, 9-subcategory hazard taxonomy dataset. It's part of NVIDIA's broader NeMo Guardrails ecosystem for content moderation, topic control, and jailbreak detection.
| Field | Value |
|---|---|
| Organization | NVIDIA |
| Parameters | 4B |
| Training data | Nemotron Content Safety Dataset V2 (Aegis v2) |
| License | NVIDIA Open Model License |
| Modality | Text only |
Open weights, free to download.
Scores from Mistral AI's Shieldstral model card, shown for context. Not Benchgen measurements.
| Benchmark | Score |
|---|---|
| RTP-LX Prompt | 86.1% |
| WildGuardTest (Prompt) | 84.4% |
| Aegis v2 (Response) | 80.9% |
| RTP-LX Completion | 89.7% |
| HarmBench (Prompt) | 91.7% |
Scores sourced from Mistral AI's Shieldstral model card, shown for context. Last updated 2026-08-12.
This model isn’t on any benchmark leaderboard yet.