Quick answer: Qwen2-VL 72B is Alibaba's 72-billion-parameter open-weight multimodal model, notable for outperforming GPT-4o and achieving the best overall performance on the MTVQA multilingual scene-text comprehension benchmark at release.
Where Qwen2-VL 72B leads
Where it lags
Best for: teams needing strong open-weight multilingual visual document/OCR understanding with full deployment control.
Qwen2-VL 72B is part of Alibaba's Qwen2-VL family of open-weight multimodal models. At release, it achieved the best overall performance on the MTVQA multilingual text-centric visual QA benchmark, outperforming GPT-4o and other leading proprietary and open-source MLLMs on multilingual scene-text comprehension tasks.
| Field | Value |
|---|---|
| Organization | Alibaba |
| Modality | Multimodal (text + vision) |
| License | Open weights (Tongyi Qianwen license) |
| Release date | August 2024 |