| Rank | Model | Score |
|---|---|---|
| 1 | lfm2-5-vl-3b | 48.4 |
| 2 | north-micro-vision-instruct | 32.9 |
| 3 | qwen3-6-plus | 0.86 |
| 4 | gpt-5-1-instant | 0.854 |
| 5 | gpt-5-1-thinking | 0.854 |
| 6 | gpt-5-1 | 0.854 |
| 7 | gpt-5 | 0.842 |
| 8 | qwen3-5-122b-a10b | 0.839 |
| 9 | o3 | 0.829 |
| 10 | qwen3-6-27b | 0.829 |
| 11 | qwen3-5-27b | 0.823 |
| 12 | qwen3-6-35b-a3b | 0.817 |
| 13 | o4-mini | 0.816 |
| 14 | qwen3-5-35b-a3b | 0.814 |
| 15 | gemini-2-5-flash | 0.797 |
| 16 | gemini-2-5-pro | 0.796 |
| 17 | o1 | 0.776 |
| 18 | gpt-4-5 | 0.752 |
| 19 | gpt-4-1 | 0.748 |
| 20 | claude-sonnet-4 | 0.744 |
| 21 | llama-4-maverick | 0.734 |
| 22 | gpt-4-1-mini | 0.727 |
| 23 | gpt-4o | 0.722 |
| 24 | gemini-2-0-flash | 0.707 |
| 25 | kimi-k1-5 | 0.7 |
1 phaseActive
Massive Multi-discipline Multimodal Understanding and Reasoning benchmark — 11.5K college-level exam, quiz, and textbook questions across 30 subjects. Metric: accuracy.
Quick answer: MMMU (Massive Multi-discipline Multimodal Understanding and Reasoning) is a benchmark of 11.5K college-level exam, quiz, and textbook questions spanning six core disciplines and 183 subfields, designed to test multimodal models on expert-level knowledge and deliberate reasoning. Qwen3.6 Plus leads with 86.0% across 63 evaluated models.
MMMU evaluates multimodal AI models on college-level subject knowledge that requires interpreting diagrams, charts, medical images, and other visual content alongside text. Questions span Art & Design, Business, Science, Health & Medicine, Humanities & Social Science, and Tech & Engineering — demanding expert-level reasoning rather than simple visual recognition.
| Focus area | Examples |
|---|---|
| Science & engineering | Diagram interpretation, technical schematics, physics/chemistry problems |
| Health & medicine | Medical imaging, anatomical diagrams, clinical reasoning |
| Business & humanities | Charts, tables, historical images, document analysis |
| Cross-modal reasoning | Combining visual and textual information for expert-level answers |
Questions are drawn from real college exams, quizzes, and textbooks, formatted as multiple-choice or open-ended with verifiable answers. Accuracy is the fraction of questions answered correctly, normalized to 0–1, across 30 subjects and 183 subfields.
| Property | Value |
|---|---|
| Released | November 2023 |
| Questions | 11,500 |
| Disciplines | 6 core, 30 subjects, 183 subfields |
| Metric | Accuracy |
| Score range | 0–1 |
| Top model | Qwen3.6 Plus (0.860) |
| Models evaluated | 63 |
What is MMMU? MMMU is a benchmark evaluating multimodal AI models on college-level subject knowledge and deliberate reasoning, using 11.5K questions from real exams, quizzes, and textbooks across six core disciplines.
Who created MMMU? MMMU was created by Xiang Yue, Yuansheng Ni, Kai Zhang, Tianyu Zheng, and 18 other collaborators, published in November 2023.
What subjects does MMMU cover? MMMU spans Art & Design, Business, Science, Health & Medicine, Humanities & Social Science, and Tech & Engineering, across 30 subjects and 183 subfields.
What score does the best model achieve on MMMU? Qwen3.6 Plus leads with 0.860 (86.0%), followed closely by OpenAI's GPT-5.1 variants at 0.854 and GPT-5 at 0.842.
Is MMMU saturated? Getting there at the frontier — top models now exceed 0.85 — but the average across all 63 evaluated models is around 0.7, with a long tail of older and smaller models scoring well below 0.6, so meaningful headroom remains for weaker systems.