Benchgen

AI2D — Results

RankModelScore
1north-micro-vision-instruct77.5
2claude-3-5-sonnet0.947
3qwen3-6-plus0.944
4gpt-4o0.942
5qwen3-5-122b-a10b0.933
6mistral-small-3-2-24b-instruct0.929
7qwen3-5-27b0.929
8qwen3-6-35b-a3b0.927
9qwen3-5-35b-a3b0.926
10llama-3-2-90b-instruct0.923
11qwen3-vl-235b-a22b-instruct0.897
12qwen3-vl-32b-instruct0.895
13qwen3-vl-235b-a22b-thinking0.892
14qwen3-vl-32b-thinking0.889
15qwen3-vl-30b-a3b-thinking0.869
16qwen3-vl-8b-instruct0.857
17qwen3-vl-30b-a3b-instruct0.85
18qwen3-vl-4b-thinking0.849
19qwen3-vl-8b-thinking0.849
20gemma-3-27b0.845
21gemma-3-12b0.842
22qwen3-vl-4b-instruct0.841
23qwen2-5-omni-7b0.832
24gemma-3-4b0.748

AI2D

1 phaseActive

4,903 grade-school science diagrams with 15,000+ multiple-choice questions — evaluates diagram understanding and visual reasoning about food webs, physiology, and life cycles. Metric: accuracy.

Overview

AI2D

Category Metric Saturation Tasks

Paper Dataset

Quick answer: AI2D (AI2 Diagrams) is a visual reasoning benchmark by Kembhavi et al. (2016) from the Allen Institute for AI, with 4,903 grade-school science diagrams and over 15,000 multiple-choice questions about food webs, human physiology, life cycles, and other natural science concepts. Claude 3.5 Sonnet leads with 94.7% across 32 evaluated models.


What Does AI2D Test?

AI2D tests multimodal AI on diagram understanding — a fundamental skill for interpreting scientific figures in textbooks and papers. The diagrams are from grade-school natural sciences and require models to understand diagrammatic notation, arrows, labels, and structural relationships.

Science topicDiagram types
BiologyFood webs, life cycles, human anatomy
Earth scienceRock cycles, water cycles, weather systems
PhysicsForce diagrams, circuit diagrams
ChemistryMolecular structures, reaction diagrams

Models must interpret not just the visual elements but the relational structure encoded in the diagram — arrows showing causation, labels denoting parts, and hierarchies representing classification.


How Is AI2D Scored?

Questions are multiple-choice with a single correct answer. Accuracy is the fraction of questions answered correctly, reported on a 0–1 scale.


Key Facts

PropertyValue
PublishedMarch 2016
Diagrams4,903
Questions15,000+
Science levelGrade school
MetricAccuracy
Score range0–1
Top modelClaude 3.5 Sonnet (0.947)
Models evaluated32

FAQ

What is AI2D? AI2D is a dataset of 4,903 illustrative diagrams from grade-school natural sciences with over 15,000 multiple-choice questions. It evaluates diagram understanding and visual reasoning, requiring models to interpret diagrammatic elements, relationships, and structure.

Who created AI2D? AI2D was created by Aniruddha Kembhavi, Mike Salvato, Eric Kolve, Minjoon Seo, Hannaneh Hajishirzi, and Ali Farhadi at the Allen Institute for AI, published in March 2016 (arXiv 1603.07396).

Why are AI2D scores so high across models? AI2D is a relatively saturated benchmark — even mid-tier models score above 80%. It remains useful for comparing vision understanding across model families, but frontier models are clustered at the top.

What score does the best model achieve on AI2D? Claude 3.5 Sonnet leads with 0.947 (94.7%), followed by Qwen3.6 Plus at 0.944 and GPT-4o at 0.942.