| Rank | Model | Score |
|---|---|---|
| 1 | north-micro-vision-instruct | 77.5 |
| 2 | claude-3-5-sonnet | 0.947 |
| 3 | qwen3-6-plus | 0.944 |
| 4 | gpt-4o | 0.942 |
| 5 | qwen3-5-122b-a10b | 0.933 |
| 6 | mistral-small-3-2-24b-instruct | 0.929 |
| 7 | qwen3-5-27b | 0.929 |
| 8 | qwen3-6-35b-a3b | 0.927 |
| 9 | qwen3-5-35b-a3b | 0.926 |
| 10 | llama-3-2-90b-instruct | 0.923 |
| 11 | qwen3-vl-235b-a22b-instruct | 0.897 |
| 12 | qwen3-vl-32b-instruct | 0.895 |
| 13 | qwen3-vl-235b-a22b-thinking | 0.892 |
| 14 | qwen3-vl-32b-thinking | 0.889 |
| 15 | qwen3-vl-30b-a3b-thinking | 0.869 |
| 16 | qwen3-vl-8b-instruct | 0.857 |
| 17 | qwen3-vl-30b-a3b-instruct | 0.85 |
| 18 | qwen3-vl-4b-thinking | 0.849 |
| 19 | qwen3-vl-8b-thinking | 0.849 |
| 20 | gemma-3-27b | 0.845 |
| 21 | gemma-3-12b | 0.842 |
| 22 | qwen3-vl-4b-instruct | 0.841 |
| 23 | qwen2-5-omni-7b | 0.832 |
| 24 | gemma-3-4b | 0.748 |
1 phaseActive
4,903 grade-school science diagrams with 15,000+ multiple-choice questions — evaluates diagram understanding and visual reasoning about food webs, physiology, and life cycles. Metric: accuracy.
Quick answer: AI2D (AI2 Diagrams) is a visual reasoning benchmark by Kembhavi et al. (2016) from the Allen Institute for AI, with 4,903 grade-school science diagrams and over 15,000 multiple-choice questions about food webs, human physiology, life cycles, and other natural science concepts. Claude 3.5 Sonnet leads with 94.7% across 32 evaluated models.
AI2D tests multimodal AI on diagram understanding — a fundamental skill for interpreting scientific figures in textbooks and papers. The diagrams are from grade-school natural sciences and require models to understand diagrammatic notation, arrows, labels, and structural relationships.
| Science topic | Diagram types |
|---|---|
| Biology | Food webs, life cycles, human anatomy |
| Earth science | Rock cycles, water cycles, weather systems |
| Physics | Force diagrams, circuit diagrams |
| Chemistry | Molecular structures, reaction diagrams |
Models must interpret not just the visual elements but the relational structure encoded in the diagram — arrows showing causation, labels denoting parts, and hierarchies representing classification.
Questions are multiple-choice with a single correct answer. Accuracy is the fraction of questions answered correctly, reported on a 0–1 scale.
| Property | Value |
|---|---|
| Published | March 2016 |
| Diagrams | 4,903 |
| Questions | 15,000+ |
| Science level | Grade school |
| Metric | Accuracy |
| Score range | 0–1 |
| Top model | Claude 3.5 Sonnet (0.947) |
| Models evaluated | 32 |
What is AI2D? AI2D is a dataset of 4,903 illustrative diagrams from grade-school natural sciences with over 15,000 multiple-choice questions. It evaluates diagram understanding and visual reasoning, requiring models to interpret diagrammatic elements, relationships, and structure.
Who created AI2D? AI2D was created by Aniruddha Kembhavi, Mike Salvato, Eric Kolve, Minjoon Seo, Hannaneh Hajishirzi, and Ali Farhadi at the Allen Institute for AI, published in March 2016 (arXiv 1603.07396).
Why are AI2D scores so high across models? AI2D is a relatively saturated benchmark — even mid-tier models score above 80%. It remains useful for comparing vision understanding across model families, but frontier models are clustered at the top.
What score does the best model achieve on AI2D? Claude 3.5 Sonnet leads with 0.947 (94.7%), followed by Qwen3.6 Plus at 0.944 and GPT-4o at 0.942.