Quick answer: Qwen2 7B Instruct is Alibaba's June 2024 compact model scoring 82.3% GSM8K, 79.9% HumanEval, 70.6% MMLU, 44.1% MMLU-Pro, and 8.41 MT-Bench. Apache 2.0 — the 7B instruction-tuned model in the Qwen2 family.
Where Qwen2 7B Instruct leads
Where it lags
Best for: Legacy Qwen2 ecosystem; compact coding at 7B; simple multi-turn chatbot deployments; baselines and comparisons.
Qwen2 7B Instruct is the June 2024 7B instruction model from Alibaba's Qwen2 generation. It was competitive at launch but is now superseded by Qwen2.5 7B (Sep 2024) and Qwen3 variants — both of which improve substantially on all metrics.
For new deployments, Qwen2.5 7B Instruct or Qwen3 7B are recommended over Qwen2 7B.
| Field | Value |
|---|---|
| Organization | Alibaba |
| License | Apache 2.0 |
| HuggingFace | Qwen/Qwen2-7B-Instruct |
| Release date | June 7, 2024 |
| Parameters | 7B |
| Modality | Text only |
| Context window | 128K tokens |
Open weights under Apache 2.0 — self-host at no cost.
| Benchmark | Score | Source | Date |
|---|---|---|---|
| GSM8K | 82.3% | Benchgen evaluation | 2024-06 |
| HumanEval | 79.9% | Benchgen evaluation | 2024-06 |
| MMLU | 70.6% | Benchgen evaluation | 2024-06 |
| MMLU-Pro | 44.1% | Benchgen evaluation | 2024-06 |
| MT-Bench | 8.41 | Benchgen evaluation | 2024-06 |
| Model | GSM8K | HumanEval | MMLU-Pro | License |
|---|---|---|---|---|
| Qwen2 7B Instruct | 82.3% | 79.9% | 44.1% | Apache 2.0 |
| Qwen2.5 7B Instruct | 91.6% | 84.8% | 56.3% | Apache 2.0 |
| Qwen2.5 14B Instruct | 94.8% | 83.5% | 64.0% | Apache 2.0 |
Qwen2.5 7B (Sep 2024) uniformly outperforms Qwen2 7B. For all new 7B deployments: use Qwen2.5 7B or Qwen3 7B.
Specs from Alibaba's Qwen2 7B Instruct release (June 2024) and Benchgen evaluations. Last updated 2026-07-24.
This model isn’t on any benchmark leaderboard yet.