Quick answer: Qwen3 VL 8B Thinking is Alibaba's February 2026 compact vision-language reasoning model scoring 77.3% MMLU-Pro, 63.0% BFCL-v3, and 51.1% Arena Hard v2. Apache 2.0 — 8B params.
Where Qwen3 VL 8B Thinking leads
Where it lags
Best for: Compact open-source multimodal pipelines; BFCL function calling at 8B inference cost.
Qwen3 VL 8B Thinking is the mid-tier 8B model in Alibaba's Qwen3 VL thinking family (32B, 8B, 4B). Compact deployment at 8B with vision-language + thinking capabilities.
| Field | Value |
|---|---|
| Organization | Alibaba |
| License | Apache 2.0 |
| HuggingFace | Qwen/Qwen3-VL-8B-Thinking |
| Release date | February 2026 |
| Parameters | 8B |
| Modality | Text and vision |
Open weights under Apache 2.0 — self-host at no cost.
| Benchmark | Score | Source | Date |
|---|---|---|---|
| MMLU-Pro | 77.3% | Benchgen evaluation | 2026-02 |
| BFCL-v3 | 63.0% | Benchgen evaluation | 2026-02 |
| Arena Hard v2 | 51.1% | Benchgen evaluation | 2026-02 |
| Model | MMLU-Pro | BFCL-v3 | Arena Hard | Size |
|---|---|---|---|---|
| Qwen3 VL 8B Thinking | 77.3% | 63.0% | 51.1% | 8B |
| Qwen3 VL 32B Thinking | 82.1% | 71.7% | 60.5% | 32B |
| Qwen3 VL 4B Thinking | 73.6% | 67.3% | 36.8% | 4B |
For cost/quality tradeoff in Qwen3 VL: 8B is the recommended mid-tier. 4B is compact but lower Arena Hard (36.8%).
Specs from Alibaba's Qwen3 VL 8B Thinking release (February 2026) and Benchgen evaluations. Last updated 2026-07-24.
This model isn’t on any benchmark leaderboard yet.