Benchgen

MMMU — Results

RankModelScore
1lfm2-5-vl-3b48.4
2north-micro-vision-instruct32.9
3qwen3-6-plus0.86
4gpt-5-1-instant0.854
5gpt-5-1-thinking0.854
6gpt-5-10.854
7gpt-50.842
8qwen3-5-122b-a10b0.839
9o30.829
10qwen3-6-27b0.829
11qwen3-5-27b0.823
12qwen3-6-35b-a3b0.817
13o4-mini0.816
14qwen3-5-35b-a3b0.814
15gemini-2-5-flash0.797
16gemini-2-5-pro0.796
17o10.776
18gpt-4-50.752
19gpt-4-10.748
20claude-sonnet-40.744
21llama-4-maverick0.734
22gpt-4-1-mini0.727
23gpt-4o0.722
24gemini-2-0-flash0.707
25kimi-k1-50.7

MMMU

1 phaseActive

Massive Multi-discipline Multimodal Understanding and Reasoning benchmark — 11.5K college-level exam, quiz, and textbook questions across 30 subjects. Metric: accuracy.

Overview

MMMU

Category Metric Saturation Tasks

Dataset

Quick answer: MMMU (Massive Multi-discipline Multimodal Understanding and Reasoning) is a benchmark of 11.5K college-level exam, quiz, and textbook questions spanning six core disciplines and 183 subfields, designed to test multimodal models on expert-level knowledge and deliberate reasoning. Qwen3.6 Plus leads with 86.0% across 63 evaluated models.


What Does MMMU Test?

MMMU evaluates multimodal AI models on college-level subject knowledge that requires interpreting diagrams, charts, medical images, and other visual content alongside text. Questions span Art & Design, Business, Science, Health & Medicine, Humanities & Social Science, and Tech & Engineering — demanding expert-level reasoning rather than simple visual recognition.

Focus areaExamples
Science & engineeringDiagram interpretation, technical schematics, physics/chemistry problems
Health & medicineMedical imaging, anatomical diagrams, clinical reasoning
Business & humanitiesCharts, tables, historical images, document analysis
Cross-modal reasoningCombining visual and textual information for expert-level answers

How Is MMMU Scored?

Questions are drawn from real college exams, quizzes, and textbooks, formatted as multiple-choice or open-ended with verifiable answers. Accuracy is the fraction of questions answered correctly, normalized to 0–1, across 30 subjects and 183 subfields.


Key Facts

PropertyValue
ReleasedNovember 2023
Questions11,500
Disciplines6 core, 30 subjects, 183 subfields
MetricAccuracy
Score range0–1
Top modelQwen3.6 Plus (0.860)
Models evaluated63

FAQ

What is MMMU? MMMU is a benchmark evaluating multimodal AI models on college-level subject knowledge and deliberate reasoning, using 11.5K questions from real exams, quizzes, and textbooks across six core disciplines.

Who created MMMU? MMMU was created by Xiang Yue, Yuansheng Ni, Kai Zhang, Tianyu Zheng, and 18 other collaborators, published in November 2023.

What subjects does MMMU cover? MMMU spans Art & Design, Business, Science, Health & Medicine, Humanities & Social Science, and Tech & Engineering, across 30 subjects and 183 subfields.

What score does the best model achieve on MMMU? Qwen3.6 Plus leads with 0.860 (86.0%), followed closely by OpenAI's GPT-5.1 variants at 0.854 and GPT-5 at 0.842.

Is MMMU saturated? Getting there at the frontier — top models now exceed 0.85 — but the average across all 63 evaluated models is around 0.7, with a long tail of older and smaller models scoring well below 0.6, so meaningful headroom remains for weaker systems.