Benchgen

MMVU — Results

RankModelScore
1kimi-k382.1
M

MMVU

1 phaseActive

Expert-level, multi-discipline video understanding benchmark — 3,000 questions across 27 subjects in science, healthcare, humanities, and engineering. Metric: % accuracy.

Overview

MMVU

Category Metric Questions Saturation Created

Paper

Quick answer: MMVU (Zhao et al., Yale NLP, 2025) is a comprehensive expert-level, multi-discipline benchmark for video understanding, with 3,000 expert-annotated questions spanning 27 subjects across Science, Healthcare, Humanities & Social Sciences, and Engineering. Each example includes expert-annotated reasoning rationales and relevant domain knowledge. Kimi K3 scores 82.1% as of July 2026.

At a Glance

What it tests: Domain-specific expert knowledge combined with video understanding — models must analyze specialized-domain videos (e.g., a surgical procedure or an engineering demonstration) and apply expert reasoning, not just basic visual perception.

Why it matters: Most video benchmarks test surface-level perception. MMVU requires applying real domain expertise to video content, and every example is annotated by human experts from scratch with strict quality controls.

Known limitations: As with most expert-level benchmarks, coverage is limited to the 27 selected subjects; performance may vary significantly across disciplines not represented in the dataset.

What MMVU Measures

MMVU evaluates foundation models on expert-level, multi-discipline video understanding, requiring both domain-specific knowledge and expert-level reasoning to analyze specialized videos. Its four core disciplines — Science, Healthcare, Humanities & Social Sciences, and Engineering — span 27 subjects in total. Each of the 3,000 questions is annotated from scratch by human experts and enriched with reasoning rationales and relevant domain knowledge, enabling in-depth error analysis beyond simple accuracy scores.

Benchmark Specifications

FieldValue
Task categoryVideo / expert-domain reasoning
Metric% accuracy
Number of questions3,000
Disciplines4 core (27 subjects)
SaturationLow
Created byZhao et al. (Yale NLP)
Source paperMMVU (arXiv 2501.12380)

How MMVU Is Scored

Models answer expert-level questions about specialized-domain videos and are scored on % accuracy against expert-annotated ground truth. The dataset's reasoning rationales allow deeper analysis of why a model's answer is correct or incorrect, beyond the headline accuracy number.

State-of-the-Art Results

RankModelScoreSourceDate
1Kimi K382.1%Kimi K3 technical report2026-07

Score sourced from Moonshot AI's Kimi K3 announcement, July 2026.

MMVU on Benchgen

No Benchgen results yet — be the first to run MMVU.

MMVU vs Other Benchmarks

BenchmarkWhat it testsTasksSaturation
MMVUExpert-level, multi-discipline video understanding3,000Low
Video-MMEGeneral multimodal video understanding2,700Low
MMMU-ProMulti-discipline multimodal image understandingLow
SciCodeScientific codingLow

Run MMVU on Your Model

Benchgen lets you run MMVU against your own multimodal model, tracking expert-domain video reasoning accuracy across disciplines to identify where fine-tuning improves real-world domain understanding.

Frequently Asked Questions

What is MMVU? MMVU is an expert-level, multi-discipline video understanding benchmark with 3,000 questions across 27 subjects in Science, Healthcare, Humanities & Social Sciences, and Engineering.
What does a good score look like on MMVU? At release, System-2-capable models like o1 and Gemini 2.0 Flash Thinking led the pack but still fell short of human expertise. By mid-2026, frontier models such as Kimi K3 report scores around 82%.
Who created MMVU? MMVU was created by Yilun Zhao and collaborators at Yale NLP (arXiv:2501.12380, January 2025).