| Rank | Model | Score |
|---|---|---|
| 1 | kimi-k3 | 82.1 |
1 phaseActive
Expert-level, multi-discipline video understanding benchmark — 3,000 questions across 27 subjects in science, healthcare, humanities, and engineering. Metric: % accuracy.
Quick answer: MMVU (Zhao et al., Yale NLP, 2025) is a comprehensive expert-level, multi-discipline benchmark for video understanding, with 3,000 expert-annotated questions spanning 27 subjects across Science, Healthcare, Humanities & Social Sciences, and Engineering. Each example includes expert-annotated reasoning rationales and relevant domain knowledge. Kimi K3 scores 82.1% as of July 2026.
What it tests: Domain-specific expert knowledge combined with video understanding — models must analyze specialized-domain videos (e.g., a surgical procedure or an engineering demonstration) and apply expert reasoning, not just basic visual perception.
Why it matters: Most video benchmarks test surface-level perception. MMVU requires applying real domain expertise to video content, and every example is annotated by human experts from scratch with strict quality controls.
Known limitations: As with most expert-level benchmarks, coverage is limited to the 27 selected subjects; performance may vary significantly across disciplines not represented in the dataset.
MMVU evaluates foundation models on expert-level, multi-discipline video understanding, requiring both domain-specific knowledge and expert-level reasoning to analyze specialized videos. Its four core disciplines — Science, Healthcare, Humanities & Social Sciences, and Engineering — span 27 subjects in total. Each of the 3,000 questions is annotated from scratch by human experts and enriched with reasoning rationales and relevant domain knowledge, enabling in-depth error analysis beyond simple accuracy scores.
| Field | Value |
|---|---|
| Task category | Video / expert-domain reasoning |
| Metric | % accuracy |
| Number of questions | 3,000 |
| Disciplines | 4 core (27 subjects) |
| Saturation | Low |
| Created by | Zhao et al. (Yale NLP) |
| Source paper | MMVU (arXiv 2501.12380) |
Models answer expert-level questions about specialized-domain videos and are scored on % accuracy against expert-annotated ground truth. The dataset's reasoning rationales allow deeper analysis of why a model's answer is correct or incorrect, beyond the headline accuracy number.
| Rank | Model | Score | Source | Date |
|---|---|---|---|---|
| 1 | Kimi K3 | 82.1% | Kimi K3 technical report | 2026-07 |
Score sourced from Moonshot AI's Kimi K3 announcement, July 2026.
No Benchgen results yet — be the first to run MMVU.
| Benchmark | What it tests | Tasks | Saturation |
|---|---|---|---|
| MMVU | Expert-level, multi-discipline video understanding | 3,000 | Low |
| Video-MME | General multimodal video understanding | 2,700 | Low |
| MMMU-Pro | Multi-discipline multimodal image understanding | — | Low |
| SciCode | Scientific coding | — | Low |
Benchgen lets you run MMVU against your own multimodal model, tracking expert-domain video reasoning accuracy across disciplines to identify where fine-tuning improves real-world domain understanding.