Quick answer: Qwen2.5 7B Instruct is Alibaba's September 2024 compact instruction-tuned model, scoring 91.6% GSM8K, 84.8% HumanEval, 75.5% MATH, and 52% Arena Hard. Apache 2.0 with a 128K token context window.
Where Qwen2.5 7B Instruct leads
Where it lags
Best for: Compact Apache 2.0 deployments requiring good math and coding; edge/local inference; cost-efficient API hosting.
Qwen2.5 7B Instruct is the instruction-tuned variant of Alibaba's Qwen2.5 7B base model, released September 2024. The Qwen2.5 series offers models from 0.5B to 72B, with the 7B being the primary compact option for most deployment scenarios.
With 91.6% GSM8K and 84.8% HumanEval, the 7B Instruct model punches above its weight on math and coding benchmarks relative to models from 2023-era 7B class. The 128K context window is unusually large for this size tier.
| Field | Value |
|---|---|
| Organization | Alibaba |
| License | Apache 2.0 |
| HuggingFace | Qwen/Qwen2.5-7B-Instruct |
| Release date | September 19, 2024 |
| Parameters | 7B |
| Modality | Text only |
| Context window | 128K tokens |
Open weights under Apache 2.0 — self-host at no cost. Available via Alibaba Cloud (DashScope) and major providers (Together AI, Fireworks, Groq, etc.).
| Benchmark | Score | Source | Date |
|---|---|---|---|
| GSM8K | 91.6% | Benchgen evaluation | 2024-09 |
| HumanEval | 84.8% | Benchgen evaluation | 2024-09 |
| MATH | 75.5% | Benchgen evaluation | 2024-09 |
| Arena Hard | 52% | Benchgen evaluation | 2024-09 |
| ACEBench | 57.8% | Benchgen evaluation | 2024-09 |
| BigCodeBench | 14.2% | Benchgen evaluation | 2024-09 |
| Model | GSM8K | HumanEval | Params | License |
|---|---|---|---|---|
| Qwen2.5 7B Instruct | 91.6% | 84.8% | 7B | Apache 2.0 |
| Qwen2.5 Coder 7B Instruct | 83.9% | 88.4% | 7B | Apache 2.0 |
| Llama 3.1 8B Instruct | 84.5% | — | 8B | Llama 3.1 |
| Gemma 3 12B | — | 85.4% | 12B | Gemma ToU |
Qwen2.5 7B Instruct vs Coder 7B: better GSM8K (91.6% vs 83.9%), lower HumanEval (84.8% vs 88.4%). For coding-first use: Coder 7B. For math-first: 7B Instruct.
Specs from Alibaba's Qwen2.5 7B Instruct release (September 2024) and Benchgen evaluations. Last updated 2026-07-24.
This model isn’t on any benchmark leaderboard yet.