Benchgen
Models/alibaba/

Qwen3 VL 32B Thinking

DraftPublic

Model Details

Qwen3 VL 32B Thinking

Organization License Released VL

Quick answer: Qwen3 VL 32B Thinking is Alibaba's February 2026 vision-language reasoning model scoring 82.1% MMLU-Pro, 71.7% BFCL-v3, and 60.5% Arena Hard v2. Apache 2.0.

At a Glance

Where Qwen3 VL 32B Thinking leads

  • 82.1% MMLU-Pro — strong academic knowledge in multimodal context
  • 71.7% BFCL-v3 — solid function calling
  • Vision-language: text + image reasoning
  • Apache 2.0 — fully open
  • 32B: largest Qwen3 VL thinking variant

Where it lags

  • 60.5% Arena Hard v2 — moderate instruction quality vs text-only models

Best for: Open-source vision-language reasoning at 32B; BFCL tool-calling with multimodal input; Qwen3 VL pipeline deployments.

What Qwen3 VL 32B Thinking Is

Qwen3 VL 32B Thinking is the 32B vision-language model with thinking (chain-of-thought reasoning) from Alibaba's Qwen3 VL generation, released February 2026. It is the top tier in the Qwen3 VL thinking family (32B, 8B, 4B).

Specifications

FieldValue
OrganizationAlibaba
LicenseApache 2.0
HuggingFaceQwen/Qwen3-VL-32B-Thinking
Release dateFebruary 2026
Parameters32B
ModalityText and vision

Pricing

Open weights under Apache 2.0 — self-host at no cost.

Public Benchmark Scores

BenchmarkScoreSourceDate
MMLU-Pro82.1%Benchgen evaluation2026-02
BFCL-v371.7%Benchgen evaluation2026-02
Arena Hard v260.5%Benchgen evaluation2026-02

Qwen3 VL 32B vs Alternatives

ModelMMLU-ProBFCL-v3Arena HardSize
Qwen3 VL 32B Thinking82.1%71.7%60.5%32B
Qwen3 VL 8B Thinking77.3%63.0%51.1%8B
Qwen3 VL 4B Thinking73.6%67.3%36.8%4B

Qwen3 VL 32B leads on MMLU-Pro (82.1%) and Arena Hard (60.5%) vs smaller siblings. For cost-efficient multimodal: 8B or 4B variants.

Frequently Asked Questions

What is Qwen3 VL 32B Thinking? Alibaba's February 2026 vision-language reasoning model scoring 82.1% MMLU-Pro, 71.7% BFCL-v3. Apache 2.0 — top tier in Qwen3 VL family.

Specs from Alibaba's Qwen3 VL 32B Thinking release (February 2026) and Benchgen evaluations. Last updated 2026-07-24.

Benchmark Leaderboards

This model isn’t on any benchmark leaderboard yet.