Quick answer: Qwen2.5 72B Instruct is Alibaba's September 2024 improved 72B model, scoring 95.8% on GSM8K, 86.6% on HumanEval, 81.2% on Arena Hard, and 55.5% on LiveCodeBench. Significant improvements over Qwen2 72B — particularly in coding (+22pp LiveCodeBench) and math (+6pp GSM8K).
Where Qwen2.5 72B leads
Where it lags
Best for: Existing Qwen2.5 72B deployments; strong math/coding tasks in Chinese-English bilingual contexts; teams in Alibaba Cloud ecosystem.
Qwen2.5 72B Instruct (released September 19, 2024) represents Alibaba's Q3 2024 flagship open-weight model — the successor to Qwen2 72B with improved training data and instruction tuning.
Key improvements over Qwen2 72B: GSM8K +6.3pp (95.8% vs 89.5%), HumanEval +0.6pp (86.6% vs 86.0%), and LiveCodeBench significantly improved (+22pp, representing a major coding capability leap). Arena Hard jumped from ~75% to 81.2%.
The 128K context window matches other flagship open models. Qwen2.5 72B was among the strongest open-weight 72B models at launch, competing closely with Llama 3.3 70B (released December 2024) on most benchmarks.
| Field | Value |
|---|---|
| Organization | Alibaba / Qwen |
| Parameters | 72B (dense) |
| Context window | 128,000 tokens |
| License | Qianwen License |
| HuggingFace | Qwen/Qwen2.5-72B-Instruct |
| Release date | September 19, 2024 |
| Knowledge cutoff | July 2024 |
| Modality | Text only |
Open weights under Qianwen License on Hugging Face. Available via Alibaba Cloud DashScope at market rates.
| Benchmark | Score | Source | Date |
|---|---|---|---|
| GSM8K | 95.8% | Benchgen evaluation | 2025-07 |
| HumanEval | 86.6% | Benchgen evaluation | 2025-07 |
| Arena Hard | 81.2% | Benchgen evaluation | 2025-07 |
| LiveCodeBench | 55.5% | Benchgen evaluation | 2025-07 |
| BigCodeBench | 25.4% | Benchgen evaluation | 2025-07 |
| Model | GSM8K | HumanEval | LiveCodeBench | License |
|---|---|---|---|---|
| Qwen2.5 72B Instruct | 95.8% | 86.6% | 55.5% | Qianwen |
| Qwen2 72B Instruct | 89.5% | 86.0% | — | Qianwen |
| Llama 3.3 70B Instruct | — | 88.4% | — | Llama 3.3 |
| DeepSeek-V3 | — | — | 27.2% | MIT |
Qwen2.5 72B vs Qwen2 72B: +6.3pp GSM8K, significantly higher LiveCodeBench — a meaningful upgrade. For new open-weight 72B deployments, Qwen3 32B (Apache 2.0, better performance at smaller size) is the recommended choice.
Specs from Alibaba's official Qwen2.5 release (September 2024) and Benchgen evaluations. Last updated 2026-07-24.
This model isn’t on any benchmark leaderboard yet.