Benchgen
Models/alibaba/

Qwen2-VL 72B

DraftPublic

Model Details

Qwen2-VL 72B

Organization Modality License Released

Quick answer: Qwen2-VL 72B is Alibaba's 72-billion-parameter open-weight multimodal model, notable for outperforming GPT-4o and achieving the best overall performance on the MTVQA multilingual scene-text comprehension benchmark at release.

At a Glance

Where Qwen2-VL 72B leads

  • Strong multilingual OCR and scene-text comprehension, outperforming GPT-4o on MTVQA at release
  • Fully open weights, enabling self-hosted deployment and fine-tuning

Where it lags

  • Requires substantial GPU infrastructure to self-host at 72B parameters

Best for: teams needing strong open-weight multilingual visual document/OCR understanding with full deployment control.

What Qwen2-VL 72B Is

Qwen2-VL 72B is part of Alibaba's Qwen2-VL family of open-weight multimodal models. At release, it achieved the best overall performance on the MTVQA multilingual text-centric visual QA benchmark, outperforming GPT-4o and other leading proprietary and open-source MLLMs on multilingual scene-text comprehension tasks.

Specifications

FieldValue
OrganizationAlibaba
ModalityMultimodal (text + vision)
LicenseOpen weights (Tongyi Qianwen license)
Release dateAugust 2024