Benchgen
Models/alibaba/

Qwen3-Next-80B-A3B-Thinking

DraftPublic

Model Details

Qwen3 Next 80B A3B Thinking

Organization License Released Thinking

Quick answer: Qwen3 Next 80B A3B Thinking is Alibaba's April 2026 thinking MoE model scoring 82.7% MMLU-Pro, 72.0% BFCL-v3, and 62.3% Arena Hard v2. Apache 2.0 — 80B/3B active.

At a Glance

Where Qwen3 Next 80B A3B Thinking leads

  • 82.7% MMLU-Pro — best MMLU-Pro in Qwen3 Next 80B family
  • 72.0% BFCL-v3 — strong function calling
  • Apache 2.0 — fully open
  • 3B active: low inference cost

Where it lags

  • 62.3% Arena Hard — significantly lower than Instruct (82.7%)

Best for: MMLU-Pro knowledge + BFCL function calling at 3B active cost; reasoning-heavy tasks.

What Qwen3 Next 80B A3B Thinking Is

The thinking (chain-of-thought reasoning) variant of Qwen3 Next 80B A3B. Thinking raises MMLU-Pro to 82.7% (vs 80.6% instruct) and BFCL to 72.0% (vs 70.3%) but reduces Arena Hard significantly (62.3% vs 82.7%).

Specifications

FieldValue
OrganizationAlibaba
LicenseApache 2.0
HuggingFaceQwen/Qwen3-Next-80B-A3B-Thinking
Release dateApril 2026
Parameters80B total / 3B active (MoE)
ModalityText only

Pricing

Open weights under Apache 2.0 — self-host at no cost.

Public Benchmark Scores

BenchmarkScoreSourceDate
MMLU-Pro82.7%Benchgen evaluation2026-04
BFCL-v372.0%Benchgen evaluation2026-04
Arena Hard v262.3%Benchgen evaluation2026-04

Qwen3 Next 80B Thinking vs Alternatives

ModelMMLU-ProBFCL-v3Arena HardType
Qwen3 Next 80B A3B Thinking82.7%72.0%62.3%Thinking MoE
Qwen3 Next 80B A3B Instruct80.6%70.3%82.7%Instruct MoE

Thinking vs Instruct tradeoff: Thinking +2.1% MMLU-Pro, +1.7% BFCL, but -20.4% Arena Hard. Use Instruct for instruction following; Thinking for knowledge/reasoning.

Frequently Asked Questions

What is Qwen3 Next 80B A3B Thinking? Alibaba's April 2026 thinking MoE model scoring 82.7% MMLU-Pro, 72.0% BFCL-v3. Apache 2.0 — 80B/3B active.

Specs from Alibaba's Qwen3 Next 80B A3B Thinking release (April 2026) and Benchgen evaluations. Last updated 2026-07-24.

Benchmark Leaderboards

This model isn’t on any benchmark leaderboard yet.