Benchgen
Models/meta/

Llama 3.3 70B Instruct

DraftPublic

Model Details

Llama 3.3 70B Instruct

Organization Parameters Context License Weights Released

Quick answer: Llama 3.3 70B Instruct is Meta's December 2024 improved 70B model, scoring 96.0% on AI2 Reasoning Challenge, 86.0% on MMLU, 77.0% on MATH, 88.4% on HumanEval, and 68.9% on MMLU-Pro. It significantly improves on Llama 3.1 70B — notably +27pp MATH (77% vs 50%) — while maintaining the same 128K context window.

At a Glance

Where Llama 3.3 70B leads

  • 77.0% MATH — strong competition mathematics, +27pp vs Llama 3.1 70B
  • 88.4% HumanEval — excellent code completion
  • 68.9% MMLU-Pro — improved academic knowledge
  • 96.0% AI2 Reasoning Challenge
  • 128K context window
  • Open weights with strong ecosystem support

Where it lags

  • Llama 3.3 License (not Apache 2.0) — commercial restrictions
  • Text-only: no vision capability
  • Superseded by Llama 4 Maverick for multimodal and 1M context needs
  • August 2024 knowledge cutoff

Best for: Open-weight 70B deployments requiring strong math/coding; teams on Llama 3.1 70B seeking improvements without switching to Llama 4; complex STEM tasks.

What Llama 3.3 70B Instruct Is

Llama 3.3 70B Instruct (released December 6, 2024) is Meta's iterative improvement to the Llama 3.1 70B architecture, trained with enhanced data and improved instruction tuning. The most notable improvement is MATH: 77.0% vs approximately 50% for Llama 3.1 70B — a 27-percentage-point gain on competition mathematics.

The model achieves 88.4% HumanEval (vs ~72% for 3.1 70B) and 68.9% MMLU-Pro, placing it solidly in the frontier open-weight tier at the 70B scale. It maintains the same 128K context window and architecture, making it a drop-in upgrade for most Llama 3.1 70B deployments.

For teams choosing between Llama 3.1 70B and 3.3 70B, Llama 3.3 is the clear choice: substantially better math, coding, and knowledge with minimal migration overhead.

Specifications

FieldValue
OrganizationMeta
Parameters70B (dense)
Context window128,000 tokens
LicenseLlama 3.3 License
HuggingFacemeta-llama/Llama-3.3-70B-Instruct
Release dateDecember 6, 2024
Knowledge cutoffAugust 2024
ModalityText only

Pricing

Open weights under Llama 3.3 License on Hugging Face. Available via hosted providers (Together AI, Groq, etc.) at market rates.

Public Benchmark Scores

BenchmarkScoreSourceDate
AI2 Reasoning Challenge96.0%Benchgen evaluation2025-07
MMLU86.0%Benchgen evaluation2025-07
MATH77.0%Benchgen evaluation2025-07
HumanEval88.4%Benchgen evaluation2025-07
MMLU-Pro68.9%Benchgen evaluation2025-07
BigCodeBench28.4%Benchgen evaluation2025-07

Llama 3.3 70B vs Alternatives

ModelMATHHumanEvalMMLU-ProLicense
Llama 3.3 70B Instruct77.0%88.4%68.9%Llama 3.3
Llama 3.1 70B Instruct~50%~72%Llama 3.1
Llama 4 Maverick80.5%Llama 4
DeepSeek-V390.2%MIT

Llama 3.3 70B vs 3.1 70B: +27pp MATH, +16pp HumanEval — a clear upgrade with minimal migration overhead. vs Llama 4 Maverick: 4 Maverick wins on MMLU-Pro (80.5% vs 68.9%) and context (1M vs 128K) but requires a MoE-aware inference stack.

Frequently Asked Questions

What is Llama 3.3 70B Instruct? Llama 3.3 70B Instruct is Meta's December 2024 improved 70B model, scoring 96.0% AI2 RC, 86.0% MMLU, 77.0% MATH, and 88.4% HumanEval with 128K context. It significantly outperforms Llama 3.1 70B on mathematics and coding.
Should I use Llama 3.3 70B or Llama 3.1 70B? Llama 3.3 70B for new deployments: +27pp MATH (77.0% vs ~50%), +16pp HumanEval — a significant upgrade with minimal migration overhead as it uses the same architecture.
Should I use Llama 3.3 70B or Llama 4 Maverick? Llama 4 Maverick achieves better MMLU-Pro (80.5% vs 68.9%), 1M context, and multimodal capability. Llama 3.3 70B is preferred when a dense 70B architecture is specifically required.

Specs from Meta's official Llama 3.3 announcement (December 2024) and Benchgen evaluations. Last updated 2026-07-24.

Benchmark Leaderboards

This model isn’t on any benchmark leaderboard yet.