Quick answer: Qwen2.5 Coder 32B Instruct is Alibaba's October 2024 code-specialised 32B model, scoring 92.7% HumanEval, 85.3% ACEBench, 91.1% GSM8K, and 30.8% BigCodeBench. Apache 2.0 — the largest Qwen2.5 Coder variant and strongest open-weight coder for its era.
Where Qwen2.5 Coder 32B Instruct leads
Where it lags
Best for: Coding assistants, code review, and code generation at 32B scale; open-source IDE integration; agent frameworks requiring strong tool use.
Qwen2.5 Coder 32B Instruct is the flagship model in the Qwen2.5 Coder series — the 32B parameter code-specialised variant released October 2024. It was the strongest open-weight coding model at its release, surpassing models like DeepSeek-Coder and CodeLlama on HumanEval.
The 85.3% ACEBench score reflects solid agentic and tool-calling performance, making it suitable for coding agent frameworks. The 30.8% BigCodeBench indicates limitations on very complex code generation tasks — an area where larger reasoning models (QwQ-32B, DeepSeek-R1) have since improved.
| Field | Value |
|---|---|
| Organization | Alibaba |
| License | Apache 2.0 |
| HuggingFace | Qwen/Qwen2.5-Coder-32B-Instruct |
| Release date | October 7, 2024 |
| Parameters | 32B |
| Modality | Text / Code |
| Context window | 128K tokens |
Open weights under Apache 2.0 — self-host at no cost. Available via major inference providers.
| Benchmark | Score | Source | Date |
|---|---|---|---|
| HumanEval | 92.7% | Benchgen evaluation | 2024-10 |
| ACEBench | 85.3% | Benchgen evaluation | 2024-10 |
| GSM8K | 91.1% | Benchgen evaluation | 2024-10 |
| HellaSwag | 83.0% | Benchgen evaluation | 2024-10 |
| BigCodeBench | 30.8% | Benchgen evaluation | 2024-10 |
| BigCodeBench Hard | 27% | Benchgen evaluation | 2024-10 |
| Model | HumanEval | ACEBench | Params | License |
|---|---|---|---|---|
| Qwen2.5 Coder 32B Instruct | 92.7% | 85.3% | 32B | Apache 2.0 |
| Qwen2.5 Coder 7B Instruct | 88.4% | 49.6% | 7B | Apache 2.0 |
| Qwen2.5 72B Instruct | 86.6% | — | 72B | Qwen |
| Kimi K2 Instruct 0905 | — | 76.5% | MoE | Apache 2.0 |
Qwen2.5 Coder 32B vs 7B Coder: higher HumanEval (92.7% vs 88.4%) and much higher ACEBench (85.3% vs 49.6%) at ~4.6x the parameter count. Use 7B for lighter deployments, 32B for coding agents.
Specs from Alibaba's Qwen2.5 Coder 32B Instruct release (October 2024) and Benchgen evaluations. Last updated 2026-07-24.
This model isn’t on any benchmark leaderboard yet.