Benchgen
Models/alibaba/

Qwen2.5 7B Instruct

DraftPublic

Model Details

Qwen2.5 7B Instruct

Organization License Released Params

Quick answer: Qwen2.5 7B Instruct is Alibaba's September 2024 compact instruction-tuned model, scoring 91.6% GSM8K, 84.8% HumanEval, 75.5% MATH, and 52% Arena Hard. Apache 2.0 with a 128K token context window.

At a Glance

Where Qwen2.5 7B Instruct leads

  • 91.6% GSM8K — excellent math for a 7B model
  • 84.8% HumanEval — strong coding
  • 75.5% MATH — competitive competition math
  • Apache 2.0 — fully open
  • 128K context window

Where it lags

  • 57.8% ACEBench — moderate tool/agent performance
  • 14.2% BigCodeBench — complex code generation struggles
  • 52% Arena Hard — modest instruction quality

Best for: Compact Apache 2.0 deployments requiring good math and coding; edge/local inference; cost-efficient API hosting.

What Qwen2.5 7B Instruct Is

Qwen2.5 7B Instruct is the instruction-tuned variant of Alibaba's Qwen2.5 7B base model, released September 2024. The Qwen2.5 series offers models from 0.5B to 72B, with the 7B being the primary compact option for most deployment scenarios.

With 91.6% GSM8K and 84.8% HumanEval, the 7B Instruct model punches above its weight on math and coding benchmarks relative to models from 2023-era 7B class. The 128K context window is unusually large for this size tier.

Specifications

FieldValue
OrganizationAlibaba
LicenseApache 2.0
HuggingFaceQwen/Qwen2.5-7B-Instruct
Release dateSeptember 19, 2024
Parameters7B
ModalityText only
Context window128K tokens

Pricing

Open weights under Apache 2.0 — self-host at no cost. Available via Alibaba Cloud (DashScope) and major providers (Together AI, Fireworks, Groq, etc.).

Public Benchmark Scores

BenchmarkScoreSourceDate
GSM8K91.6%Benchgen evaluation2024-09
HumanEval84.8%Benchgen evaluation2024-09
MATH75.5%Benchgen evaluation2024-09
Arena Hard52%Benchgen evaluation2024-09
ACEBench57.8%Benchgen evaluation2024-09
BigCodeBench14.2%Benchgen evaluation2024-09

Qwen2.5 7B Instruct vs Alternatives

ModelGSM8KHumanEvalParamsLicense
Qwen2.5 7B Instruct91.6%84.8%7BApache 2.0
Qwen2.5 Coder 7B Instruct83.9%88.4%7BApache 2.0
Llama 3.1 8B Instruct84.5%8BLlama 3.1
Gemma 3 12B85.4%12BGemma ToU

Qwen2.5 7B Instruct vs Coder 7B: better GSM8K (91.6% vs 83.9%), lower HumanEval (84.8% vs 88.4%). For coding-first use: Coder 7B. For math-first: 7B Instruct.

Frequently Asked Questions

What is Qwen2.5 7B Instruct? Alibaba's September 2024 compact 7B instruction model scoring 91.6% GSM8K, 84.8% HumanEval, and 75.5% MATH. Apache 2.0 with 128K context.
Should I use Qwen2.5 7B Instruct or Qwen2.5 72B Instruct? Qwen2.5 72B Instruct for maximum performance (95.8% GSM8K, 86.6% HumanEval). Qwen2.5 7B Instruct for deployments where lower compute cost or local inference is required.

Specs from Alibaba's Qwen2.5 7B Instruct release (September 2024) and Benchgen evaluations. Last updated 2026-07-24.

Benchmark Leaderboards

This model isn’t on any benchmark leaderboard yet.