Benchgen
Models/meta/

Llama 3.1 405B Instruct

DraftPublic

Model Details

Llama 3.1 405B Instruct

Organization Parameters Context License Weights Released

Quick answer: Llama 3.1 405B Instruct is Meta's July 2024 largest open-weight model, scoring 96.9% on AI2 Reasoning Challenge, 88.6% on MMLU, 73.3% on MMLU-Pro, and 26.4% on BigCodeBench. With a 128K-token context window under the Llama 3.1 License, it was the first open-weight model to match GPT-4-class performance at the 400B+ parameter scale.

At a Glance

Where Llama 3.1 405B leads

  • 96.9% AI2 Reasoning Challenge — near-perfect scientific reasoning
  • 88.6% MMLU — matches GPT-4 Turbo (86.4%) at open-weight pricing
  • 73.3% MMLU-Pro — strong academic knowledge
  • Largest open-weight dense model at launch
  • 128K context window
  • Establishes frontier capability ceiling for open-weight models at launch

Where it lags

  • Dense 405B: massive infrastructure requirement for self-hosting
  • Superseded by Llama 4 Maverick for most tasks at far lower cost
  • Text-only: no vision
  • March 2024 knowledge cutoff
  • 26.4% BigCodeBench — moderate coding performance

Best for: Research requiring maximum open-weight capability; self-hosted frontier AI where data sovereignty is paramount; fine-tuning at maximum scale.

What Llama 3.1 405B Instruct Is

Llama 3.1 405B Instruct was the largest open-weight model available when released July 23, 2024 — and the first to directly compete with GPT-4 on MMLU. Its release was a landmark event: for the first time, teams could access GPT-4-class performance through self-hosted open weights.

At 88.6% MMLU, the 405B surpassed the then-state-of-the-art GPT-4 Turbo (86.4%) on the primary academic knowledge benchmark. This performance came at substantial infrastructure cost: a dense 405B model requires multiple high-end GPUs for inference. For most production use cases, Llama 4 Maverick (17B active MoE, 1M context, better overall performance) is the recommended replacement.

Llama 3.1 405B remains relevant for teams with specific requirements: maximum open-weight capability, dense architecture for quantisation research, or use cases where the Llama 3.1 instruction style has been validated in production.

Specifications

FieldValue
OrganizationMeta
Parameters405B (dense)
Context window128,000 tokens
LicenseLlama 3.1 License
HuggingFacemeta-llama/Meta-Llama-3.1-405B-Instruct
Release dateJuly 23, 2024
Knowledge cutoffMarch 2024
ModalityText only

Pricing

Llama 3.1 405B is available as open weights under the Llama 3.1 License — free to self-host. Hosted API access via Together AI, Fireworks, and others at market rates.

Public Benchmark Scores

BenchmarkScoreSourceDate
AI2 Reasoning Challenge96.9%Benchgen evaluation2024-07
MMLU88.6%Benchgen evaluation2024-07
MMLU-Pro73.3%Benchgen evaluation2024-07
BigCodeBench26.4%Benchgen evaluation2024-07

Llama 3.1 405B vs Alternatives

ModelMMLUMMLU-ProContextWeights
Llama 3.1 405B Instruct88.6%73.3%128KOpen (dense 405B)
Llama 3.1 70B Instruct86.0%128KOpen (dense 70B)
Llama 4 Maverick80.5%1MOpen (17B active MoE)
GPT-4 Turbo86.4%128KClosed

Llama 4 Maverick (17B active MoE) achieves 80.5% MMLU-Pro vs Llama 3.1 405B's 73.3%, with a 1M context window, multimodal capability, and dramatically lower inference cost. For new deployments, Llama 4 Maverick is the recommended open-weight choice.

Frequently Asked Questions

What is Llama 3.1 405B? Llama 3.1 405B Instruct is Meta's July 2024 largest open-weight model, scoring 96.9% AI2 RC, 88.6% MMLU, and 73.3% MMLU-Pro with 128K context. It was the first open-weight model to match GPT-4-class MMLU performance.
Is Llama 3.1 405B open source? The weights are open on Hugging Face under the Llama 3.1 License. Commercial use is permitted with some restrictions.
Should I use Llama 3.1 405B or Llama 4 Maverick? For new deployments, Llama 4 Maverick is strongly preferred: better benchmark scores, 1M context, multimodal capability, and dramatically lower inference cost (17B active MoE vs 405B dense). Llama 3.1 405B is appropriate for existing integrations or when a dense 405B architecture is specifically required.

Specs from Meta's official Llama 3.1 announcement (July 2024) and Benchgen evaluations. Last updated 2026-07-24.

Benchmark Leaderboards

This model isn’t on any benchmark leaderboard yet.