Quick answer: Qwen3 VL 32B Instruct is Alibaba's February 2026 32B vision-language instruct model scoring 78.6% MMLU-Pro, 70.2% BFCL-v3, and 64.7% Arena Hard v2. Apache 2.0.
Where Qwen3 VL 32B Instruct leads
Where it lags
Best for: Vision-language function calling with 32B quality; non-thinking instruct deployments in Qwen3 VL.
Qwen3 VL 32B Instruct is the standard instruct (non-thinking) 32B vision-language model in the Qwen3 VL family. Compared to the 32B Thinking variant, it scores lower on MMLU-Pro (78.6% vs 82.1%) but is faster to run (no long chain-of-thought).
| Field | Value |
|---|---|
| Organization | Alibaba |
| License | Apache 2.0 |
| HuggingFace | Qwen/Qwen3-VL-32B-Instruct |
| Release date | February 2026 |
| Parameters | 32B (dense) |
| Modality | Text and vision |
Open weights under Apache 2.0 — self-host at no cost.
| Benchmark | Score | Source | Date |
|---|---|---|---|
| MMLU-Pro | 78.6% | Benchgen evaluation | 2026-02 |
| BFCL-v3 | 70.2% | Benchgen evaluation | 2026-02 |
| Arena Hard v2 | 64.7% | Benchgen evaluation | 2026-02 |
| Model | MMLU-Pro | BFCL-v3 | Arena Hard | Type |
|---|---|---|---|---|
| Qwen3 VL 32B Instruct | 78.6% | 70.2% | 64.7% | Instruct |
| Qwen3 VL 32B Thinking | 82.1% | 71.7% | 60.5% | Thinking |
| Qwen3 VL 30B A3B Instruct | 77.8% | 66.3% | 58.5% | MoE Instruct |
32B Instruct vs 32B Thinking: Instruct has higher Arena Hard (64.7% vs 60.5%) but lower MMLU-Pro (78.6% vs 82.1%). Use Thinking for knowledge/reasoning; Instruct for instruction following speed.
Specs from Alibaba's Qwen3 VL 32B Instruct release (February 2026) and Benchgen evaluations. Last updated 2026-07-24.
This model isn’t on any benchmark leaderboard yet.