Quick answer: Qwen3 VL 30B A3B Thinking is Alibaba's February 2026 compact MoE vision-language thinking model scoring 80.5% MMLU-Pro, 68.6% BFCL-v3, and 56.7% Arena Hard v2. Apache 2.0 — 3B active params.
Where Qwen3 VL 30B A3B Thinking leads
Where it lags
Best for: Cost-efficient VL thinking at 3B active params; MMLU-Pro heavy vision tasks.
A 30B MoE (3B active params) thinking variant for vision-language tasks. Near-equal MMLU-Pro to the 32B dense thinking at a fraction of the inference cost.
| Field | Value |
|---|---|
| Organization | Alibaba |
| License | Apache 2.0 |
| HuggingFace | Qwen/Qwen3-VL-30B-A3B-Thinking |
| Release date | February 2026 |
| Parameters | 30B total / 3B active (MoE) |
| Modality | Text and vision |
Open weights under Apache 2.0 — self-host at no cost.
| Benchmark | Score | Source | Date |
|---|---|---|---|
| MMLU-Pro | 80.5% | Benchgen evaluation | 2026-02 |
| BFCL-v3 | 68.6% | Benchgen evaluation | 2026-02 |
| Arena Hard v2 | 56.7% | Benchgen evaluation | 2026-02 |
| Model | MMLU-Pro | BFCL-v3 | Arena Hard | Active Params |
|---|---|---|---|---|
| Qwen3 VL 30B A3B Thinking | 80.5% | 68.6% | 56.7% | 3B |
| Qwen3 VL 32B Thinking | 82.1% | 71.7% | 60.5% | 32B |
| Qwen3 VL 30B A3B Instruct | 77.8% | 66.3% | 58.5% | 3B |
30B A3B Thinking vs 32B Thinking: nearly same MMLU-Pro (80.5% vs 82.1%) at 1/10 the active params. Best cost/quality in Qwen3 VL thinking family.
Specs from Alibaba's Qwen3 VL 30B A3B Thinking release (February 2026) and Benchgen evaluations. Last updated 2026-07-24.
This model isn’t on any benchmark leaderboard yet.