Quick answer: Qwen2.5 14B Instruct is Alibaba's September 2024 mid-tier 14B model scoring 83.5% HumanEval, 94.8% GSM8K, 80.0% MATH, and 64.0% ACEBench. Apache 2.0 — sits between 7B and 32B in the Qwen2.5 family.
Where Qwen2.5 14B Instruct leads
Where it lags
Best for: Mid-tier Apache 2.0 deployments balancing math and coding; applications where 7B isn't enough but 32B is too large.
Qwen2.5 14B Instruct is the 14B instruction-tuned model in the Qwen2.5 series. It fills the gap between the compact 7B (simpler but lower quality) and the larger 32B/72B (better but more compute). With 94.8% GSM8K and 80.0% MATH, the 14B offers competitive math for its size.
| Field | Value |
|---|---|
| Organization | Alibaba |
| License | Apache 2.0 |
| HuggingFace | Qwen/Qwen2.5-14B-Instruct |
| Release date | September 19, 2024 |
| Parameters | 14B |
| Modality | Text only |
| Context window | 128K tokens |
Open weights under Apache 2.0 — self-host at no cost.
| Benchmark | Score | Source | Date |
|---|---|---|---|
| GSM8K | 94.8% | Benchgen evaluation | 2024-09 |
| HumanEval | 83.5% | Benchgen evaluation | 2024-09 |
| MATH | 80.0% | Benchgen evaluation | 2024-09 |
| ACEBench | 64.0% | Benchgen evaluation | 2024-09 |
| BigCodeBench | 20.9% | Benchgen evaluation | 2024-09 |
| Model | GSM8K | HumanEval | MATH | License |
|---|---|---|---|---|
| Qwen2.5 14B Instruct | 94.8% | 83.5% | 80.0% | Apache 2.0 |
| Qwen2.5 7B Instruct | 91.6% | 84.8% | 75.5% | Apache 2.0 |
| Qwen2.5 32B Instruct | 95.9% | 88.4% | — | Apache 2.0 |
Qwen2.5 14B vs 7B: higher GSM8K (94.8% vs 91.6%), higher MATH (80% vs 75.5%), slightly lower HumanEval (83.5% vs 84.8%). Use 14B for math-heavy; 7B for coding-primary with minimal size.
Specs from Alibaba's Qwen2.5 14B Instruct release (September 2024) and Benchgen evaluations. Last updated 2026-07-24.
This model isn’t on any benchmark leaderboard yet.