Benchgen
Models/qwen/

Qwen2.5 72B Instruct

DraftPublic

Model Details

Qwen2.5 72B Instruct

Organization Parameters Context License Weights Released

Quick answer: Qwen2.5 72B Instruct is Alibaba's September 2024 improved 72B model, scoring 95.8% on GSM8K, 86.6% on HumanEval, 81.2% on Arena Hard, and 55.5% on LiveCodeBench. Significant improvements over Qwen2 72B — particularly in coding (+22pp LiveCodeBench) and math (+6pp GSM8K).

At a Glance

Where Qwen2.5 72B leads

  • 95.8% GSM8K — near-ceiling math reasoning at 72B scale
  • 86.6% HumanEval — strong code completion
  • 81.2% Arena Hard — competitive instruction following
  • 55.5% LiveCodeBench — significant improvement over Qwen2 72B
  • 128K context window
  • Strong Chinese-English bilingual capability

Where it lags

  • Qianwen License — not Apache 2.0; commercial restrictions
  • 25.4% BigCodeBench — moderate complex coding
  • Knowledge cutoff July 2024
  • Superseded by Qwen3 family for new deployments

Best for: Existing Qwen2.5 72B deployments; strong math/coding tasks in Chinese-English bilingual contexts; teams in Alibaba Cloud ecosystem.

What Qwen2.5 72B Instruct Is

Qwen2.5 72B Instruct (released September 19, 2024) represents Alibaba's Q3 2024 flagship open-weight model — the successor to Qwen2 72B with improved training data and instruction tuning.

Key improvements over Qwen2 72B: GSM8K +6.3pp (95.8% vs 89.5%), HumanEval +0.6pp (86.6% vs 86.0%), and LiveCodeBench significantly improved (+22pp, representing a major coding capability leap). Arena Hard jumped from ~75% to 81.2%.

The 128K context window matches other flagship open models. Qwen2.5 72B was among the strongest open-weight 72B models at launch, competing closely with Llama 3.3 70B (released December 2024) on most benchmarks.

Specifications

FieldValue
OrganizationAlibaba / Qwen
Parameters72B (dense)
Context window128,000 tokens
LicenseQianwen License
HuggingFaceQwen/Qwen2.5-72B-Instruct
Release dateSeptember 19, 2024
Knowledge cutoffJuly 2024
ModalityText only

Pricing

Open weights under Qianwen License on Hugging Face. Available via Alibaba Cloud DashScope at market rates.

Public Benchmark Scores

BenchmarkScoreSourceDate
GSM8K95.8%Benchgen evaluation2025-07
HumanEval86.6%Benchgen evaluation2025-07
Arena Hard81.2%Benchgen evaluation2025-07
LiveCodeBench55.5%Benchgen evaluation2025-07
BigCodeBench25.4%Benchgen evaluation2025-07

Qwen2.5 72B vs Alternatives

ModelGSM8KHumanEvalLiveCodeBenchLicense
Qwen2.5 72B Instruct95.8%86.6%55.5%Qianwen
Qwen2 72B Instruct89.5%86.0%Qianwen
Llama 3.3 70B Instruct88.4%Llama 3.3
DeepSeek-V327.2%MIT

Qwen2.5 72B vs Qwen2 72B: +6.3pp GSM8K, significantly higher LiveCodeBench — a meaningful upgrade. For new open-weight 72B deployments, Qwen3 32B (Apache 2.0, better performance at smaller size) is the recommended choice.

Frequently Asked Questions

What is Qwen2.5 72B Instruct? Qwen2.5 72B Instruct is Alibaba's September 2024 flagship 72B model scoring 95.8% GSM8K, 86.6% HumanEval, 81.2% Arena Hard, and 55.5% LiveCodeBench with 128K context.
Is Qwen2.5 72B open source? The weights are open under the Qianwen License — permits commercial use with conditions. Not Apache 2.0. For fully permissive commercial use, consider Qwen3 32B (Apache 2.0).
Should I use Qwen2.5 72B or Qwen2 72B? Qwen2.5 72B is a clear upgrade: +6.3pp GSM8K, significantly better LiveCodeBench. Prefer Qwen2.5 72B for new deployments within the Qwen2.x family.

Specs from Alibaba's official Qwen2.5 release (September 2024) and Benchgen evaluations. Last updated 2026-07-24.

Benchmark Leaderboards

This model isn’t on any benchmark leaderboard yet.