Benchgen
Models/alibaba/

Qwen2.5 VL 32B Instruct

DraftPublic

Model Details

Qwen2.5 VL 32B Instruct

Organization License Released VL

Quick answer: Qwen2.5 VL 32B Instruct is Alibaba's March 2025 vision-language model scoring 91.5% HumanEval, 82.2% MATH, and 68.8% MMLU-Pro. Apache 2.0 — 32B multimodal.

At a Glance

Where Qwen2.5 VL 32B Instruct leads

  • 91.5% HumanEval — excellent coding with vision input
  • 82.2% MATH — strong math at 32B VL scale
  • Apache 2.0 — fully open
  • Multimodal: text + vision understanding

Where it lags

  • 68.8% MMLU-Pro — moderate academic breadth
  • March 2025: superseded by Qwen3 VL series

Best for: Open-source VL coding and math pipelines; Apache 2.0 vision-language 32B; image-understanding + math tasks.

What Qwen2.5 VL 32B Instruct Is

Qwen2.5 VL 32B Instruct is the 32B vision-language model from Alibaba's Qwen2.5 VL generation (March 2025). It extends Qwen2.5 32B with visual understanding, while maintaining strong HumanEval (91.5%) and MATH (82.2%) performance.

Specifications

FieldValue
OrganizationAlibaba
LicenseApache 2.0
HuggingFaceQwen/Qwen2.5-VL-32B-Instruct
Release dateMarch 6, 2025
Parameters32B
ModalityText and vision

Pricing

Open weights under Apache 2.0 — self-host at no cost.

Public Benchmark Scores

BenchmarkScoreSourceDate
HumanEval91.5%Benchgen evaluation2025-03
MATH82.2%Benchgen evaluation2025-03
MMLU-Pro68.8%Benchgen evaluation2025-03

Qwen2.5 VL 32B vs Alternatives

ModelHumanEvalMATHMMLU-ProMultimodal
Qwen2.5 VL 32B Instruct91.5%82.2%68.8%Yes
Qwen2.5 32B Instruct88.4%No
Qwen3 VL 32B Instruct78.6%Yes

Qwen2.5 VL 32B leads on HumanEval (91.5%) among 32B VL models. For higher MMLU-Pro: Qwen3 VL 32B Instruct (78.6%).

Frequently Asked Questions

What is Qwen2.5 VL 32B Instruct? Alibaba's March 2025 vision-language 32B model scoring 91.5% HumanEval, 82.2% MATH. Apache 2.0 — strong coding + VL.

Specs from Alibaba's Qwen2.5 VL 32B Instruct release (March 2025) and Benchgen evaluations. Last updated 2026-07-24.

Benchmark Leaderboards

This model isn’t on any benchmark leaderboard yet.