Quick answer: Qwen2.5 Omni 7B is Alibaba's March 2025 end-to-end multimodal 7B model scoring 88.7% GSM8K, 78.7% HumanEval, 71.5% MATH, and 47% MMLU-Pro. Apache 2.0 — text, image, audio, and video input.
Where Qwen2.5 Omni 7B leads
Where it lags
Best for: Open-source multimodal 7B pipelines; audio/video processing at minimal compute; Apache 2.0 omni-modal tasks.
Qwen2.5 Omni 7B is Alibaba's March 2025 omni-modal model — a 7B model accepting text, image, audio, and video, with text and audio output. It is part of the Qwen2.5 Omni family, designed to be a compact general-purpose multimodal model.
| Field | Value |
|---|---|
| Organization | Alibaba |
| License | Apache 2.0 |
| HuggingFace | Qwen/Qwen2.5-Omni-7B |
| Release date | March 19, 2025 |
| Parameters | 7B |
| Modality | Text, image, audio, video → text + audio |
| Context window | 128K tokens |
Open weights under Apache 2.0 — self-host at no cost.
| Benchmark | Score | Source | Date |
|---|---|---|---|
| GSM8K | 88.7% | Benchgen evaluation | 2025-03 |
| HumanEval | 78.7% | Benchgen evaluation | 2025-03 |
| MATH | 71.5% | Benchgen evaluation | 2025-03 |
| MMLU-Pro | 47% | Benchgen evaluation | 2025-03 |
| Model | GSM8K | HumanEval | MATH | Multimodal |
|---|---|---|---|---|
| Qwen2.5 Omni 7B | 88.7% | 78.7% | 71.5% | Full omni |
| Qwen2.5 7B Instruct | 91.6% | 84.8% | 75.5% | Text only |
Qwen2.5 Omni 7B vs Qwen2.5 7B Instruct: Omni adds multimodal at small benchmark cost (88.7% vs 91.6% GSM8K). Use Omni for any audio/video needs; Instruct for pure text performance.
Specs from Alibaba's Qwen2.5 Omni 7B release (March 2025) and Benchgen evaluations. Last updated 2026-07-24.
This model isn’t on any benchmark leaderboard yet.