Quick answer: Qwen3 VL 235B A22B Instruct is Alibaba's February 2026 flagship vision-language MoE model scoring 81.8% MMLU-Pro, 77.4% Arena Hard v2, and 67.7% BFCL-v3. Apache 2.0.
Where Qwen3 VL 235B leads
Where it lags
Best for: Highest-quality Qwen3 VL multimodal deployments; Arena Hard instruction quality + vision; enterprise vision-language applications.
Qwen3 VL 235B A22B Instruct is the flagship model in Alibaba's Qwen3 VL family — a 235B MoE vision-language model with 22B active parameters. It leads the Qwen3 VL family on Arena Hard (77.4%) and MMLU-Pro (81.8%).
| Field | Value |
|---|---|
| Organization | Alibaba |
| License | Apache 2.0 |
| HuggingFace | Qwen/Qwen3-VL-235B-A22B-Instruct |
| Release date | February 2026 |
| Parameters | 235B total / 22B active (MoE) |
| Modality | Text and vision |
Open weights under Apache 2.0 — self-host at no cost.
| Benchmark | Score | Source | Date |
|---|---|---|---|
| MMLU-Pro | 81.8% | Benchgen evaluation | 2026-02 |
| Arena Hard v2 | 77.4% | Benchgen evaluation | 2026-02 |
| BFCL-v3 | 67.7% | Benchgen evaluation | 2026-02 |
| Model | MMLU-Pro | Arena Hard | BFCL-v3 | Size |
|---|---|---|---|---|
| Qwen3 VL 235B A22B | 81.8% | 77.4% | 67.7% | 235B/22B |
| Qwen3 VL 32B Thinking | 82.1% | 60.5% | 71.7% | 32B |
| Qwen3 VL 30B A3B Instruct | 77.8% | 58.5% | 66.3% | 30B/3B |
Qwen3 VL 235B leads on Arena Hard (77.4%) among VL variants. For MMLU-Pro: 32B Thinking slightly edges ahead (82.1% vs 81.8%).
Specs from Alibaba's Qwen3 VL 235B A22B Instruct release (February 2026) and Benchgen evaluations. Last updated 2026-07-24.
This model isn’t on any benchmark leaderboard yet.