Quick answer: Llama 3.1 Nemotron 70B Instruct is NVIDIA's October 2024 instruction-aligned fine-tune of Llama 3.1 70B, scoring 91.4% on GSM8K, 85.6% on MMLU, and 9.0/10 on MT-Bench. The 9.0 MT-Bench score was notable — among the highest for an open-weight model at launch — demonstrating NVIDIA's RLHF alignment expertise.
Where Nemotron 70B leads
Where it lags
Best for: High-quality instruction following at 70B scale; teams prioritising alignment and response quality; NVIDIA GPU deployments.
Llama 3.1 Nemotron 70B Instruct is NVIDIA's fine-tuned variant of Meta's Llama 3.1 70B, applying NVIDIA's RLHF (Reinforcement Learning from Human Feedback) and alignment techniques. Released October 15, 2024, it was designed to demonstrate how post-training alignment can improve instruction following over the base model.
The 9.0 MT-Bench score represented a significant achievement — surpassing GPT-4-level MT-Bench scores while remaining open-weight. The RLHF process improves response quality, instruction adherence, and helpfulness compared to the base Llama 3.1 70B (8.70 MT-Bench).
NVIDIA has since released Nemotron Super 49B (April 2025) with even higher MT-Bench (9.17) at fewer parameters, and Nemotron Ultra 253B for maximum performance. The Nemotron 70B remains relevant for teams that have deployed it in production.
| Field | Value |
|---|---|
| Organization | NVIDIA |
| Parameters | 70B (dense) |
| Context window | 128,000 tokens |
| License | Llama 3.1 Community License |
| HuggingFace | nvidia/Llama-3.1-Nemotron-70B-Instruct-HF |
| Base model | Meta Llama 3.1 70B |
| Release date | October 15, 2024 |
| Knowledge cutoff | December 2023 |
| Modality | Text only |
Open weights under Llama 3.1 Community License. Available via NVIDIA API Catalog (build.nvidia.com).
| Benchmark | Score | Source | Date |
|---|---|---|---|
| GSM8K | 91.4% | Benchgen evaluation | 2025-07 |
| MMLU | 85.6% | Benchgen evaluation | 2025-07 |
| HellaSwag | 85.6% | Benchgen evaluation | 2025-07 |
| MT-Bench | 9.0 / 10 | Benchgen evaluation | 2025-07 |
| Model | MT-Bench | MMLU | GSM8K | License |
|---|---|---|---|---|
| Llama 3.1 Nemotron 70B | 9.0 | 85.6% | 91.4% | Llama 3.1 |
| Llama 3.1 70B Instruct | 8.70 | 86.0% | — | Llama 3.1 |
| Llama 3.3 70B Instruct | — | 86.0% | — | Llama 3.3 |
| Nemotron Super 49B | 9.17 | — | — | Llama 3.3 |
Nemotron 70B vs Llama 3.1 70B: +0.3 MT-Bench (9.0 vs 8.70) from NVIDIA's alignment. Nemotron Super 49B achieves 9.17 MT-Bench at smaller parameter count. For new deployments, Nemotron Super 49B or Llama 3.3 70B are preferred.
Specs from NVIDIA's official Nemotron 70B release (October 2024) and Benchgen evaluations. Last updated 2026-07-24.
This model isn’t on any benchmark leaderboard yet.