Benchgen
Models/alibaba/

Qwen2.5-Coder 32B Instruct

DraftPublic

Model Details

Qwen2.5 Coder 32B Instruct

Organization License Released Focus Params

Quick answer: Qwen2.5 Coder 32B Instruct is Alibaba's October 2024 code-specialised 32B model, scoring 92.7% HumanEval, 85.3% ACEBench, 91.1% GSM8K, and 30.8% BigCodeBench. Apache 2.0 — the largest Qwen2.5 Coder variant and strongest open-weight coder for its era.

At a Glance

Where Qwen2.5 Coder 32B Instruct leads

  • 92.7% HumanEval — excellent code generation
  • 85.3% ACEBench — strong agent/tool calling
  • 91.1% GSM8K — solid math
  • 83.0% HellaSwag — good language understanding
  • Apache 2.0 — fully open
  • 32B: more capacity than 7B Coder while staying efficient

Where it lags

  • 30.8% BigCodeBench — complex code generation is limited
  • 27% BigCodeBench-Hard — difficult code problems are challenging
  • Released October 2024 — pre-dates Qwen3 series

Best for: Coding assistants, code review, and code generation at 32B scale; open-source IDE integration; agent frameworks requiring strong tool use.

What Qwen2.5 Coder 32B Instruct Is

Qwen2.5 Coder 32B Instruct is the flagship model in the Qwen2.5 Coder series — the 32B parameter code-specialised variant released October 2024. It was the strongest open-weight coding model at its release, surpassing models like DeepSeek-Coder and CodeLlama on HumanEval.

The 85.3% ACEBench score reflects solid agentic and tool-calling performance, making it suitable for coding agent frameworks. The 30.8% BigCodeBench indicates limitations on very complex code generation tasks — an area where larger reasoning models (QwQ-32B, DeepSeek-R1) have since improved.

Specifications

FieldValue
OrganizationAlibaba
LicenseApache 2.0
HuggingFaceQwen/Qwen2.5-Coder-32B-Instruct
Release dateOctober 7, 2024
Parameters32B
ModalityText / Code
Context window128K tokens

Pricing

Open weights under Apache 2.0 — self-host at no cost. Available via major inference providers.

Public Benchmark Scores

BenchmarkScoreSourceDate
HumanEval92.7%Benchgen evaluation2024-10
ACEBench85.3%Benchgen evaluation2024-10
GSM8K91.1%Benchgen evaluation2024-10
HellaSwag83.0%Benchgen evaluation2024-10
BigCodeBench30.8%Benchgen evaluation2024-10
BigCodeBench Hard27%Benchgen evaluation2024-10

Qwen2.5 Coder 32B vs Alternatives

ModelHumanEvalACEBenchParamsLicense
Qwen2.5 Coder 32B Instruct92.7%85.3%32BApache 2.0
Qwen2.5 Coder 7B Instruct88.4%49.6%7BApache 2.0
Qwen2.5 72B Instruct86.6%72BQwen
Kimi K2 Instruct 090576.5%MoEApache 2.0

Qwen2.5 Coder 32B vs 7B Coder: higher HumanEval (92.7% vs 88.4%) and much higher ACEBench (85.3% vs 49.6%) at ~4.6x the parameter count. Use 7B for lighter deployments, 32B for coding agents.

Frequently Asked Questions

What is Qwen2.5 Coder 32B Instruct? Alibaba's October 2024 code-specialised 32B model scoring 92.7% HumanEval, 85.3% ACEBench, and 91.1% GSM8K. Apache 2.0 — the flagship Qwen2.5 Coder model.
Is Qwen2.5 Coder 32B better than Qwen2.5 Coder 7B? Yes on most metrics: 92.7% vs 88.4% HumanEval, 85.3% vs 49.6% ACEBench, 91.1% vs 83.9% GSM8K. Use 7B for resource-constrained deployments; 32B for production coding agents.

Specs from Alibaba's Qwen2.5 Coder 32B Instruct release (October 2024) and Benchgen evaluations. Last updated 2026-07-24.

Benchmark Leaderboards

This model isn’t on any benchmark leaderboard yet.