Quick answer: Qwen3 Next 80B A3B Thinking is Alibaba's April 2026 thinking MoE model scoring 82.7% MMLU-Pro, 72.0% BFCL-v3, and 62.3% Arena Hard v2. Apache 2.0 — 80B/3B active.
Where Qwen3 Next 80B A3B Thinking leads
Where it lags
Best for: MMLU-Pro knowledge + BFCL function calling at 3B active cost; reasoning-heavy tasks.
The thinking (chain-of-thought reasoning) variant of Qwen3 Next 80B A3B. Thinking raises MMLU-Pro to 82.7% (vs 80.6% instruct) and BFCL to 72.0% (vs 70.3%) but reduces Arena Hard significantly (62.3% vs 82.7%).
| Field | Value |
|---|---|
| Organization | Alibaba |
| License | Apache 2.0 |
| HuggingFace | Qwen/Qwen3-Next-80B-A3B-Thinking |
| Release date | April 2026 |
| Parameters | 80B total / 3B active (MoE) |
| Modality | Text only |
Open weights under Apache 2.0 — self-host at no cost.
| Benchmark | Score | Source | Date |
|---|---|---|---|
| MMLU-Pro | 82.7% | Benchgen evaluation | 2026-04 |
| BFCL-v3 | 72.0% | Benchgen evaluation | 2026-04 |
| Arena Hard v2 | 62.3% | Benchgen evaluation | 2026-04 |
| Model | MMLU-Pro | BFCL-v3 | Arena Hard | Type |
|---|---|---|---|---|
| Qwen3 Next 80B A3B Thinking | 82.7% | 72.0% | 62.3% | Thinking MoE |
| Qwen3 Next 80B A3B Instruct | 80.6% | 70.3% | 82.7% | Instruct MoE |
Thinking vs Instruct tradeoff: Thinking +2.1% MMLU-Pro, +1.7% BFCL, but -20.4% Arena Hard. Use Instruct for instruction following; Thinking for knowledge/reasoning.
Specs from Alibaba's Qwen3 Next 80B A3B Thinking release (April 2026) and Benchgen evaluations. Last updated 2026-07-24.
This model isn’t on any benchmark leaderboard yet.