Quick answer: Qwen3 VL 32B Thinking is Alibaba's February 2026 vision-language reasoning model scoring 82.1% MMLU-Pro, 71.7% BFCL-v3, and 60.5% Arena Hard v2. Apache 2.0.
Where Qwen3 VL 32B Thinking leads
Where it lags
Best for: Open-source vision-language reasoning at 32B; BFCL tool-calling with multimodal input; Qwen3 VL pipeline deployments.
Qwen3 VL 32B Thinking is the 32B vision-language model with thinking (chain-of-thought reasoning) from Alibaba's Qwen3 VL generation, released February 2026. It is the top tier in the Qwen3 VL thinking family (32B, 8B, 4B).
| Field | Value |
|---|---|
| Organization | Alibaba |
| License | Apache 2.0 |
| HuggingFace | Qwen/Qwen3-VL-32B-Thinking |
| Release date | February 2026 |
| Parameters | 32B |
| Modality | Text and vision |
Open weights under Apache 2.0 — self-host at no cost.
| Benchmark | Score | Source | Date |
|---|---|---|---|
| MMLU-Pro | 82.1% | Benchgen evaluation | 2026-02 |
| BFCL-v3 | 71.7% | Benchgen evaluation | 2026-02 |
| Arena Hard v2 | 60.5% | Benchgen evaluation | 2026-02 |
| Model | MMLU-Pro | BFCL-v3 | Arena Hard | Size |
|---|---|---|---|---|
| Qwen3 VL 32B Thinking | 82.1% | 71.7% | 60.5% | 32B |
| Qwen3 VL 8B Thinking | 77.3% | 63.0% | 51.1% | 8B |
| Qwen3 VL 4B Thinking | 73.6% | 67.3% | 36.8% | 4B |
Qwen3 VL 32B leads on MMLU-Pro (82.1%) and Arena Hard (60.5%) vs smaller siblings. For cost-efficient multimodal: 8B or 4B variants.
Specs from Alibaba's Qwen3 VL 32B Thinking release (February 2026) and Benchgen evaluations. Last updated 2026-07-24.
This model isn’t on any benchmark leaderboard yet.