Benchgen
Models/alibaba/

Qwen2.5 14B Instruct

DraftPublic

Model Details

Qwen2.5 14B Instruct

Organization License Released Params

Quick answer: Qwen2.5 14B Instruct is Alibaba's September 2024 mid-tier 14B model scoring 83.5% HumanEval, 94.8% GSM8K, 80.0% MATH, and 64.0% ACEBench. Apache 2.0 — sits between 7B and 32B in the Qwen2.5 family.

At a Glance

Where Qwen2.5 14B Instruct leads

  • 94.8% GSM8K — excellent math at 14B
  • 80.0% MATH — solid competition math
  • 83.5% HumanEval — good coding
  • Apache 2.0 — fully open
  • Good balance between 7B (limited) and 32B (larger deployment cost)

Where it lags

  • 20.9% BigCodeBench — complex code struggles
  • 64.0% ACEBench — moderate tool use
  • September 2024: older than Qwen3 generation

Best for: Mid-tier Apache 2.0 deployments balancing math and coding; applications where 7B isn't enough but 32B is too large.

What Qwen2.5 14B Instruct Is

Qwen2.5 14B Instruct is the 14B instruction-tuned model in the Qwen2.5 series. It fills the gap between the compact 7B (simpler but lower quality) and the larger 32B/72B (better but more compute). With 94.8% GSM8K and 80.0% MATH, the 14B offers competitive math for its size.

Specifications

FieldValue
OrganizationAlibaba
LicenseApache 2.0
HuggingFaceQwen/Qwen2.5-14B-Instruct
Release dateSeptember 19, 2024
Parameters14B
ModalityText only
Context window128K tokens

Pricing

Open weights under Apache 2.0 — self-host at no cost.

Public Benchmark Scores

BenchmarkScoreSourceDate
GSM8K94.8%Benchgen evaluation2024-09
HumanEval83.5%Benchgen evaluation2024-09
MATH80.0%Benchgen evaluation2024-09
ACEBench64.0%Benchgen evaluation2024-09
BigCodeBench20.9%Benchgen evaluation2024-09

Qwen2.5 14B Instruct vs Alternatives

ModelGSM8KHumanEvalMATHLicense
Qwen2.5 14B Instruct94.8%83.5%80.0%Apache 2.0
Qwen2.5 7B Instruct91.6%84.8%75.5%Apache 2.0
Qwen2.5 32B Instruct95.9%88.4%Apache 2.0

Qwen2.5 14B vs 7B: higher GSM8K (94.8% vs 91.6%), higher MATH (80% vs 75.5%), slightly lower HumanEval (83.5% vs 84.8%). Use 14B for math-heavy; 7B for coding-primary with minimal size.

Frequently Asked Questions

What is Qwen2.5 14B Instruct? Alibaba's September 2024 14B instruction model scoring 94.8% GSM8K, 83.5% HumanEval, 80.0% MATH. Apache 2.0. Mid-tier option in the Qwen2.5 family.

Specs from Alibaba's Qwen2.5 14B Instruct release (September 2024) and Benchgen evaluations. Last updated 2026-07-24.

Benchmark Leaderboards

This model isn’t on any benchmark leaderboard yet.