Quick answer: Qwen3 VL 4B Thinking is Alibaba's February 2026 ultra-compact vision-language model scoring 73.6% MMLU-Pro, 67.3% BFCL-v3, and 36.8% Arena Hard v2. Apache 2.0 — 4B params.
Where Qwen3 VL 4B Thinking leads
Where it lags
Best for: Ultra-compact multimodal deployments; BFCL function calling at minimal inference cost; edge/embedded vision-language tasks.
Qwen3 VL 4B Thinking is the smallest model in Alibaba's Qwen3 VL thinking family. Notably, it has the highest BFCL-v3 (67.3%) among the three Qwen3 VL variants, suggesting specialised function-calling training effectiveness at 4B scale.
| Field | Value |
|---|---|
| Organization | Alibaba |
| License | Apache 2.0 |
| HuggingFace | Qwen/Qwen3-VL-4B-Thinking |
| Release date | February 2026 |
| Parameters | 4B |
| Modality | Text and vision |
Open weights under Apache 2.0 — self-host at no cost.
| Benchmark | Score | Source | Date |
|---|---|---|---|
| MMLU-Pro | 73.6% | Benchgen evaluation | 2026-02 |
| BFCL-v3 | 67.3% | Benchgen evaluation | 2026-02 |
| Arena Hard v2 | 36.8% | Benchgen evaluation | 2026-02 |
| Model | MMLU-Pro | BFCL-v3 | Arena Hard | Size |
|---|---|---|---|---|
| Qwen3 VL 4B Thinking | 73.6% | 67.3% | 36.8% | 4B |
| Qwen3 VL 8B Thinking | 77.3% | 63.0% | 51.1% | 8B |
| Qwen3 VL 32B Thinking | 82.1% | 71.7% | 60.5% | 32B |
Qwen3 VL 4B has highest BFCL-v3 (67.3%) vs the 8B (63.0%), despite being smaller — good choice for function-calling at minimum compute.
Specs from Alibaba's Qwen3 VL 4B Thinking release (February 2026) and Benchgen evaluations. Last updated 2026-07-24.
This model isn’t on any benchmark leaderboard yet.