Benchgen
Models/alibaba/

Qwen2.5-Coder 7B Instruct

DraftPublic

Model Details

Qwen2.5 Coder 7B Instruct

Organization License Released Focus

Quick answer: Qwen2.5 Coder 7B Instruct is Alibaba's October 2024 code-specialised 7B model, scoring 88.4% HumanEval, 83.9% GSM8K, 49.6% ACEBench, and 20.3% BigCodeBench. Apache 2.0 — the compact Qwen coding model.

At a Glance

Where Qwen2.5 Coder 7B Instruct leads

  • 88.4% HumanEval — strong coding at 7B size
  • 76.8% HellaSwag — good language understanding
  • Apache 2.0 — fully open
  • Code-first training on code datasets

Where it lags

  • 20.3% BigCodeBench — struggles with complex code generation
  • 49.6% ACEBench — moderate agent/tool calling
  • 83.9% GSM8K — slightly below Qwen2.5 7B Instruct (91.6%)
  • 7B parameters: limited vs larger coder models

Best for: Compact code assistance and code completion; local/edge IDE plugins; cost-efficient coding agents; lightweight code review tools.

What Qwen2.5 Coder 7B Instruct Is

Qwen2.5 Coder 7B Instruct is the code-specialised variant in the Qwen2.5 Coder series, released October 2024. Unlike Qwen2.5 7B Instruct (general instruction), the Coder variant is trained on an extended code corpus with specific code instruction fine-tuning.

The code-specific training shows in HumanEval (88.4% vs 84.8% for 7B Instruct) but trades off some math performance (83.9% GSM8K vs 91.6%). This is the standard tradeoff for code-specialised models.

Specifications

FieldValue
OrganizationAlibaba
LicenseApache 2.0
HuggingFaceQwen/Qwen2.5-Coder-7B-Instruct
Release dateOctober 7, 2024
Parameters7B
ModalityText / Code
Context window128K tokens

Pricing

Open weights under Apache 2.0 — self-host at no cost. Available via major inference providers.

Public Benchmark Scores

BenchmarkScoreSourceDate
HumanEval88.4%Benchgen evaluation2024-10
GSM8K83.9%Benchgen evaluation2024-10
HellaSwag76.8%Benchgen evaluation2024-10
ACEBench49.6%Benchgen evaluation2024-10
BigCodeBench20.3%Benchgen evaluation2024-10

Qwen2.5 Coder 7B vs Alternatives

ModelHumanEvalGSM8KLicenseFocus
Qwen2.5 Coder 7B Instruct88.4%83.9%Apache 2.0Code
Qwen2.5 7B Instruct84.8%91.6%Apache 2.0General
Granite 3.3 8B Instruct89.7%80.9%Apache 2.0General
Llama 3.1 8B Instruct84.5%Llama 3.1General

Qwen2.5 Coder 7B vs Qwen2.5 7B Instruct: higher HumanEval (88.4% vs 84.8%) but lower GSM8K (83.9% vs 91.6%). If coding is primary: Coder 7B. If balanced math + code: 7B Instruct.

Frequently Asked Questions

What is Qwen2.5 Coder 7B Instruct? Alibaba's October 2024 code-specialised 7B model scoring 88.4% HumanEval, 83.9% GSM8K, and 20.3% BigCodeBench. Apache 2.0.
What is the difference between Qwen2.5 7B Instruct and Qwen2.5 Coder 7B Instruct? Both are 7B Apache 2.0 models. Coder 7B is trained on extended code data with code-specific fine-tuning (higher HumanEval: 88.4% vs 84.8%). 7B Instruct focuses on general instruction-following with better math (91.6% vs 83.9% GSM8K).

Specs from Alibaba's Qwen2.5 Coder 7B Instruct release (October 2024) and Benchgen evaluations. Last updated 2026-07-24.

Benchmark Leaderboards

This model isn’t on any benchmark leaderboard yet.