Benchgen
Models/alibaba/

Qwen3 VL 235B A22B Instruct

DraftPublic

Model Details

Qwen3 VL 235B A22B Instruct

Organization License Released MoE

Quick answer: Qwen3 VL 235B A22B Instruct is Alibaba's February 2026 flagship vision-language MoE model scoring 81.8% MMLU-Pro, 77.4% Arena Hard v2, and 67.7% BFCL-v3. Apache 2.0.

At a Glance

Where Qwen3 VL 235B leads

  • 77.4% Arena Hard v2 — top instruction quality in Qwen3 VL family
  • 81.8% MMLU-Pro — best knowledge in Qwen3 VL family
  • Apache 2.0 — fully open
  • Largest Qwen3 VL MoE: maximum capacity

Where it lags

  • 235B total params: significant hosting requirements (22B active)

Best for: Highest-quality Qwen3 VL multimodal deployments; Arena Hard instruction quality + vision; enterprise vision-language applications.

What Qwen3 VL 235B A22B Instruct Is

Qwen3 VL 235B A22B Instruct is the flagship model in Alibaba's Qwen3 VL family — a 235B MoE vision-language model with 22B active parameters. It leads the Qwen3 VL family on Arena Hard (77.4%) and MMLU-Pro (81.8%).

Specifications

FieldValue
OrganizationAlibaba
LicenseApache 2.0
HuggingFaceQwen/Qwen3-VL-235B-A22B-Instruct
Release dateFebruary 2026
Parameters235B total / 22B active (MoE)
ModalityText and vision

Pricing

Open weights under Apache 2.0 — self-host at no cost.

Public Benchmark Scores

BenchmarkScoreSourceDate
MMLU-Pro81.8%Benchgen evaluation2026-02
Arena Hard v277.4%Benchgen evaluation2026-02
BFCL-v367.7%Benchgen evaluation2026-02

Qwen3 VL 235B vs Alternatives

ModelMMLU-ProArena HardBFCL-v3Size
Qwen3 VL 235B A22B81.8%77.4%67.7%235B/22B
Qwen3 VL 32B Thinking82.1%60.5%71.7%32B
Qwen3 VL 30B A3B Instruct77.8%58.5%66.3%30B/3B

Qwen3 VL 235B leads on Arena Hard (77.4%) among VL variants. For MMLU-Pro: 32B Thinking slightly edges ahead (82.1% vs 81.8%).

Frequently Asked Questions

What is Qwen3 VL 235B A22B Instruct? Alibaba's February 2026 flagship vision-language MoE scoring 81.8% MMLU-Pro, 77.4% Arena Hard v2. Apache 2.0 — largest Qwen3 VL.

Specs from Alibaba's Qwen3 VL 235B A22B Instruct release (February 2026) and Benchgen evaluations. Last updated 2026-07-24.

Benchmark Leaderboards

This model isn’t on any benchmark leaderboard yet.