Benchgen
Models/alibaba/

Qwen3 VL 8B Thinking

DraftPublic

Model Details

Qwen3 VL 8B Thinking

Organization License Released VL

Quick answer: Qwen3 VL 8B Thinking is Alibaba's February 2026 compact vision-language reasoning model scoring 77.3% MMLU-Pro, 63.0% BFCL-v3, and 51.1% Arena Hard v2. Apache 2.0 — 8B params.

At a Glance

Where Qwen3 VL 8B Thinking leads

  • 77.3% MMLU-Pro — excellent for 8B multimodal
  • 63.0% BFCL-v3 — good function calling
  • Apache 2.0 — fully open
  • 8B: low deployment cost

Where it lags

  • Below 32B variant on all metrics
  • 51.1% Arena Hard — moderate instruction

Best for: Compact open-source multimodal pipelines; BFCL function calling at 8B inference cost.

What Qwen3 VL 8B Thinking Is

Qwen3 VL 8B Thinking is the mid-tier 8B model in Alibaba's Qwen3 VL thinking family (32B, 8B, 4B). Compact deployment at 8B with vision-language + thinking capabilities.

Specifications

FieldValue
OrganizationAlibaba
LicenseApache 2.0
HuggingFaceQwen/Qwen3-VL-8B-Thinking
Release dateFebruary 2026
Parameters8B
ModalityText and vision

Pricing

Open weights under Apache 2.0 — self-host at no cost.

Public Benchmark Scores

BenchmarkScoreSourceDate
MMLU-Pro77.3%Benchgen evaluation2026-02
BFCL-v363.0%Benchgen evaluation2026-02
Arena Hard v251.1%Benchgen evaluation2026-02

Qwen3 VL 8B vs Alternatives

ModelMMLU-ProBFCL-v3Arena HardSize
Qwen3 VL 8B Thinking77.3%63.0%51.1%8B
Qwen3 VL 32B Thinking82.1%71.7%60.5%32B
Qwen3 VL 4B Thinking73.6%67.3%36.8%4B

For cost/quality tradeoff in Qwen3 VL: 8B is the recommended mid-tier. 4B is compact but lower Arena Hard (36.8%).

Frequently Asked Questions

What is Qwen3 VL 8B Thinking? Alibaba's February 2026 8B vision-language model scoring 77.3% MMLU-Pro, 63.0% BFCL-v3. Apache 2.0 — compact multimodal reasoning.

Specs from Alibaba's Qwen3 VL 8B Thinking release (February 2026) and Benchgen evaluations. Last updated 2026-07-24.

Benchmark Leaderboards

This model isn’t on any benchmark leaderboard yet.