Benchgen
Models/alibaba/

Qwen2.5 32B Instruct

DraftPublic

Model Details

Qwen2.5 32B Instruct

Organization License Released Params

Quick answer: Qwen2.5 32B Instruct is Alibaba's September 2024 general-purpose 32B model scoring 88.4% HumanEval, 95.9% GSM8K, 79.9% ACEBench, and 69% MMLU-Pro. Apache 2.0 with 128K context.

At a Glance

Where Qwen2.5 32B Instruct leads

  • 95.9% GSM8K — excellent math
  • 88.4% HumanEval — strong coding
  • 79.9% ACEBench — good agent/tool use
  • Apache 2.0 — fully open
  • 32B: middle-tier efficient deployment

Where it lags

  • 69% MMLU-Pro — moderate academic knowledge
  • 24.6% BigCodeBench — limited complex code generation
  • September 2024: superseded by Qwen2.5 Coder 32B and Qwen3

Best for: General instruction-following at 32B scale; balanced math+coding+agent deployments; Apache 2.0 mid-tier deployment.

What Qwen2.5 32B Instruct Is

Qwen2.5 32B Instruct is the general-purpose instruction-tuned 32B model in the Qwen2.5 series, released September 2024. It occupies the middle tier between Qwen2.5 7B and Qwen2.5 72B.

With 95.9% GSM8K, 88.4% HumanEval, and 79.9% ACEBench, this model provides a balanced profile for math, coding, and tool-use — without the code-specialisation trade-offs of Qwen2.5 Coder 32B. The 128K context window is standard for this generation.

Specifications

FieldValue
OrganizationAlibaba
LicenseApache 2.0
HuggingFaceQwen/Qwen2.5-32B-Instruct
Release dateSeptember 19, 2024
Parameters32B
ModalityText only
Context window128K tokens

Pricing

Open weights under Apache 2.0 — self-host. Available via Alibaba Cloud and major providers.

Public Benchmark Scores

BenchmarkScoreSourceDate
GSM8K95.9%Benchgen evaluation2024-09
HumanEval88.4%Benchgen evaluation2024-09
ACEBench79.9%Benchgen evaluation2024-09
MMLU-Pro69%Benchgen evaluation2024-09
BigCodeBench24.6%Benchgen evaluation2024-09

Qwen2.5 32B Instruct vs Alternatives

ModelGSM8KHumanEvalACEBenchLicense
Qwen2.5 32B Instruct95.9%88.4%79.9%Apache 2.0
Qwen2.5 Coder 32B Instruct91.1%92.7%85.3%Apache 2.0
Qwen2.5 72B Instruct95.8%86.6%Apache 2.0
Qwen2.5 7B Instruct91.6%84.8%57.8%Apache 2.0

Qwen2.5 32B vs Coder 32B: higher GSM8K (95.9% vs 91.1%), lower HumanEval (88.4% vs 92.7%), lower ACEBench (79.9% vs 85.3%). For coding agents: Coder 32B. For general balance: 32B Instruct.

Frequently Asked Questions

What is Qwen2.5 32B Instruct? Alibaba's September 2024 general-purpose 32B model scoring 95.9% GSM8K, 88.4% HumanEval, 79.9% ACEBench. Apache 2.0 with 128K context.
Should I use Qwen2.5 32B or Qwen2.5 Coder 32B? Qwen2.5 Coder 32B for coding agents and SWE tasks (higher HumanEval: 92.7% vs 88.4%, higher ACEBench: 85.3% vs 79.9%). Qwen2.5 32B Instruct for general tasks with balanced math and moderate coding.

Specs from Alibaba's Qwen2.5 32B Instruct release (September 2024) and Benchgen evaluations. Last updated 2026-07-24.

Benchmark Leaderboards

This model isn’t on any benchmark leaderboard yet.