Benchgen
Models/alibaba/

Qwen3 VL 4B Thinking

DraftPublic

Model Details

Qwen3 VL 4B Thinking

Organization License Released VL

Quick answer: Qwen3 VL 4B Thinking is Alibaba's February 2026 ultra-compact vision-language model scoring 73.6% MMLU-Pro, 67.3% BFCL-v3, and 36.8% Arena Hard v2. Apache 2.0 — 4B params.

At a Glance

Where Qwen3 VL 4B Thinking leads

  • 73.6% MMLU-Pro — excellent for 4B multimodal
  • 67.3% BFCL-v3 — best BFCL-v3 in the Qwen3 VL family
  • Apache 2.0 — fully open
  • 4B: minimal inference cost

Where it lags

  • 36.8% Arena Hard — weakest instruction quality in Qwen3 VL family
  • Smallest tier: limited complex reasoning

Best for: Ultra-compact multimodal deployments; BFCL function calling at minimal inference cost; edge/embedded vision-language tasks.

What Qwen3 VL 4B Thinking Is

Qwen3 VL 4B Thinking is the smallest model in Alibaba's Qwen3 VL thinking family. Notably, it has the highest BFCL-v3 (67.3%) among the three Qwen3 VL variants, suggesting specialised function-calling training effectiveness at 4B scale.

Specifications

FieldValue
OrganizationAlibaba
LicenseApache 2.0
HuggingFaceQwen/Qwen3-VL-4B-Thinking
Release dateFebruary 2026
Parameters4B
ModalityText and vision

Pricing

Open weights under Apache 2.0 — self-host at no cost.

Public Benchmark Scores

BenchmarkScoreSourceDate
MMLU-Pro73.6%Benchgen evaluation2026-02
BFCL-v367.3%Benchgen evaluation2026-02
Arena Hard v236.8%Benchgen evaluation2026-02

Qwen3 VL 4B vs Alternatives

ModelMMLU-ProBFCL-v3Arena HardSize
Qwen3 VL 4B Thinking73.6%67.3%36.8%4B
Qwen3 VL 8B Thinking77.3%63.0%51.1%8B
Qwen3 VL 32B Thinking82.1%71.7%60.5%32B

Qwen3 VL 4B has highest BFCL-v3 (67.3%) vs the 8B (63.0%), despite being smaller — good choice for function-calling at minimum compute.

Frequently Asked Questions

What is Qwen3 VL 4B Thinking? Alibaba's February 2026 4B vision-language model scoring 73.6% MMLU-Pro, 67.3% BFCL-v3. Apache 2.0 — smallest Qwen3 VL, best BFCL-v3 in family.

Specs from Alibaba's Qwen3 VL 4B Thinking release (February 2026) and Benchgen evaluations. Last updated 2026-07-24.

Benchmark Leaderboards

This model isn’t on any benchmark leaderboard yet.