Quick answer: Qwen2 72B Instruct is Alibaba's June 2024 flagship open-weight model, scoring 89.5% on GSM8K, 82.3% on MMLU, 64.4% on MMLU-Pro, and 86.0% on HumanEval. With a 128K context window and Qianwen License, it was among the strongest open-weight 72B models at launch and has since been succeeded by Qwen2.5 and Qwen3.
Where Qwen2 72B leads
Where it lags
Best for: Existing Qwen2 72B deployments; strong Chinese-English bilingual tasks; foundation for Qwen2-based fine-tunes.
Qwen2 72B Instruct (released June 7, 2024) was Alibaba's state-of-the-art open-weight model at launch, competing directly with Llama 3 70B and GPT-4o-mini on standard benchmarks. It demonstrated particularly strong performance in mathematics (89.5% GSM8K) and code (86.0% HumanEval) relative to its parameter count.
The Qwen2 family was designed with bilingual Chinese-English capability as a first-class requirement — outperforming Llama 3 on Chinese benchmarks while remaining competitive on English tasks. This made it a leading choice for multilingual applications in 2024.
For new deployments in 2025, Qwen3 32B offers better performance across all benchmarks under Apache 2.0 with hybrid thinking capability. Qwen2 72B Instruct is maintained for backwards compatibility and existing integrations.
| Field | Value |
|---|---|
| Organization | Alibaba / Qwen |
| Parameters | 72B (dense) |
| Context window | 128,000 tokens |
| License | Qianwen License |
| HuggingFace | Qwen/Qwen2-72B-Instruct |
| Release date | June 7, 2024 |
| Knowledge cutoff | March 2024 |
| Modality | Text only |
Qwen2 72B Instruct is available as open weights under the Qianwen License on Hugging Face. Available via Alibaba Cloud DashScope and third-party providers.
| Benchmark | Score | Source | Date |
|---|---|---|---|
| GSM8K | 89.5% | Benchgen evaluation | 2025-07 |
| MMLU | 82.3% | Benchgen evaluation | 2025-07 |
| MMLU-Pro | 64.4% | Benchgen evaluation | 2025-07 |
| HumanEval | 86.0% | Benchgen evaluation | 2025-07 |
| MATH | 59.7% | Benchgen evaluation | 2025-07 |
| BigCodeBench | 20.6% | Benchgen evaluation | 2025-07 |
| Model | MMLU | MMLU-Pro | HumanEval | License |
|---|---|---|---|---|
| Qwen2 72B Instruct | 82.3% | 64.4% | 86.0% | Qianwen |
| Llama 3.1 70B Instruct | 86.0% | — | — | Llama 3.1 |
| Qwen3 32B | — | — | — | Apache 2.0 |
| DeepSeek-V3 | 88.5% | — | — | MIT |
Qwen3 32B (April 2025, Apache 2.0) is preferred for new Qwen deployments: better performance at lower parameter count, hybrid thinking, fully permissive license. Qwen2 72B is appropriate for existing workflows.
Specs from Alibaba's official Qwen2 release (June 2024) and Benchgen evaluations. Last updated 2026-07-24.
This model isn’t on any benchmark leaderboard yet.