Benchgen
Models/nvidia/

Llama 3.1 Nemotron Ultra 253B v1

DraftPublic

Model Details

Llama 3.1 Nemotron Ultra 253B

Organization License Released Params

Quick answer: Llama 3.1 Nemotron Ultra 253B is NVIDIA's April 2025 post-trained model built on Meta's Llama 3.1 405B base, scoring 74.1% BFCL v2 and 66.3% LiveCodeBench. It is NVIDIA's largest Nemotron model — above Nemotron Super 49B (73.7% BFCL v2) in function calling.

At a Glance

Where Nemotron Ultra 253B leads

  • 74.1% BFCL v2 — strong function calling, above Nemotron Super 49B (73.7%)
  • 66.3% LiveCodeBench — competitive coding performance at this size
  • Built on Llama 3.1 405B base — full Meta ecosystem compatibility
  • 128K context window (via Llama 3.1)

Where it lags

  • 253B parameters: requires significant GPU memory (~500GB+ for inference)
  • Llama 3.1 license restrictions (not fully permissive for all commercial uses)
  • Less benchmark data than OpenAI/Google frontier models
  • April 2025 release — pre-dates some newer SOTA

Best for: Enterprise deployments requiring large-scale reasoning with NVIDIA optimisation; tool-use and function-calling agent pipelines; teams using NVIDIA infrastructure.

What Llama 3.1 Nemotron Ultra 253B Is

Llama 3.1 Nemotron Ultra 253B is NVIDIA's April 2025 post-training of Meta's Llama 3.1 405B. NVIDIA's Nemotron series involves RLHF and preference optimisation on top of Meta's base models, producing variants that excel on tool-use, instruction-following, and agentic tasks.

"Ultra" indicates the top tier in NVIDIA's Nemotron lineup for this generation, with "Super" (49B) being the mid-tier variant. At 253B parameters on Llama 3.1 base, it inherits the 128K context window while benefiting from NVIDIA's alignment training.

Specifications

FieldValue
OrganizationNVIDIA
LicenseLlama 3.1 License
HuggingFacenvidia/Llama-3.1-Nemotron-Ultra-253B-v1
Release dateApril 14, 2025
Base modelLlama 3.1 405B
Parameters253B
ModalityText only
Context window128K tokens

Pricing

Open weights (Llama 3.1 License) — self-host via NVIDIA NIM or HuggingFace. Available via NVIDIA API Catalog.

Public Benchmark Scores

BenchmarkScoreSourceDate
BFCL v274.1%Benchgen evaluation2025-04
LiveCodeBench66.3%Benchgen evaluation2025-04

Nemotron Ultra 253B vs Alternatives

ModelBFCL v2LiveCodeBenchParamsLicense
Nemotron Ultra 253B74.1%66.3%253BLlama 3.1
Nemotron Super 49B73.7%49BLlama 3.3
Llama 3.1 405B Instruct405BLlama 3.1
QwQ-32B66.4%32BApache 2.0

Nemotron Ultra 253B vs Super 49B: marginal BFCL gain (74.1% vs 73.7%) at ~5x the parameter cost. Super 49B is more efficient for most deployments.

Frequently Asked Questions

What is Llama 3.1 Nemotron Ultra 253B? NVIDIA's April 2025 post-training of Llama 3.1 405B, scoring 74.1% BFCL v2 and 66.3% LiveCodeBench. NVIDIA's largest Nemotron model optimised for function-calling and agentic tasks.
Is Nemotron Ultra 253B better than Nemotron Super 49B? Marginally on BFCL (74.1% vs 73.7%) at ~5x parameter cost. For most deployments, Nemotron Super 49B offers better efficiency. Choose Ultra 253B only if you need maximum BFCL-class function calling and have the hardware.

Specs from NVIDIA's Llama 3.1 Nemotron Ultra 253B release (April 2025) and Benchgen evaluations. Last updated 2026-07-24.

Benchmark Leaderboards

This model isn’t on any benchmark leaderboard yet.