Benchgen

ARC-AGI v2 — Results

RankModelScore
1gpt-5-50.85
2gemini-3-1-pro0.771
3gpt-5-40.733
4claude-opus-4-60.688
5claude-sonnet-4-60.583
6gpt-5-2-pro-2025-12-110.542
7gpt-5-20.529
8muse-spark0.425
9claude-opus-4-50.376
10gemini-3-flash0.336
11gemini-3-pro0.311
12grok-40.159
13claude-opus-40.086
14o30.065
15gemini-2-5-pro0.049
A

ARC-AGI v2

1 phaseActive

Chollet et al.'s harder successor to ARC-AGI: visual grid transformations measuring fluid reasoning. 800 tasks, 0–1 scale. Not saturated — top score 85.0%.

Overview

ARC-AGI v2

Category Metric Tasks Saturation Created

Paper Leaderboard

Quick answer: ARC-AGI v2 is the second-generation Abstraction and Reasoning Corpus benchmark, introduced by François Chollet and collaborators in May 2025 to address the near-saturation of the original ARC-AGI. It uses the same visual grid transformation format — models must infer hidden rules from input/output examples and apply them to a test grid — but with significantly harder task designs that remain challenging even for frontier models. GPT-5.5 leads at 85.0%.

At a Glance

What it tests: Abstract visual reasoning: given a small number of input→output grid examples, infer the underlying transformation rule and apply it to a novel test grid. Requires spatial reasoning, pattern recognition, and compositional generalisation.

Why it matters: ARC-AGI v2 was designed specifically because the original ARC-AGI approached saturation for frontier models. V2 maintains the same evaluation philosophy — tests should be easy for humans but hard for current AI — while being substantially harder, preserving its value as a signal for true generalisation.

Known limitations: Scored on a 0–1 scale; tasks are not representative of all reasoning types (domain is visual/spatial). The benchmark is still early and the leaderboard is small relative to more established benchmarks.

What ARC-AGI v2 Measures

ARC-AGI v2 inherits the core ARC format: each task provides 2–5 input/output grid pairs demonstrating a hidden transformation. The model must identify the rule and produce the correct output for a held-out test input. Grids use up to 10 colours in cells ranging from 1×1 to 30×30.

V2 introduces harder task designs with more complex compositional rules, multi-step transformations, and higher sensitivity to spatial configuration. Where frontier models could approach 95%+ on the original ARC-AGI using test-time compute scaling, V2 creates meaningful differentiation — with a 50-point gap between the leader (85%) and models like Gemini 2.5 Pro (4.9%).

The benchmark is multimodal: models receive grids as visual inputs, making it relevant for evaluating vision-language models' abstract reasoning capabilities alongside text-only reasoning systems.

Benchmark Specifications

FieldValue
Task categoryAbstract reasoning / Vision
MetricAccuracy (fraction of tasks solved, 0–1)
Number of tasks~800 (public evaluation set)
Grid sizeUp to 30×30
Colours10
SaturationMedium
Created byChollet et al. (ARC Prize Foundation)
Source paperChollet et al. 2025
ModalityMultimodal (vision + text)

How ARC-AGI v2 Is Scored

Each task is binary: solved (1) or not solved (0). The score is the fraction of tasks solved. Models are typically allowed multiple attempts per task, and a task is counted correct if any attempt produces the right output. The 0–1 normalised scale means a score of 0.850 = 85% of tasks solved.

A score above 0.7 is considered frontier-tier. Human performance on ARC-AGI tasks approaches 1.0, making it the target ceiling.

ARC-AGI v2 on Benchgen

No Benchgen results yet — be the first to run ARC-AGI v2.

ARC-AGI v2 vs Other Benchmarks

BenchmarkWhat it testsTasksSaturation
ARC-AGI v2Visual grid transformation (harder)~800Medium
ARC-AGIVisual grid transformation (original)400High
GPQA DiamondExpert-level science reasoning198Low
Humanity's Last ExamFrontier academic knowledge3000Low

ARC-AGI v2 is the go-to benchmark when you want to evaluate generalisation and fluid reasoning without saturation concerns that affect the original ARC-AGI.

Run ARC-AGI v2 on Your Model

Use Benchgen to run ARC-AGI v2 against your model continuously — tracking how reasoning capability evolves across fine-tuning runs, prompt changes, or model versions. Unlike one-time leaderboard submissions, Benchgen gives you a regression-tracked history of abstract reasoning performance.

Frequently Asked Questions

What is ARC-AGI v2? ARC-AGI v2 is the second-generation Abstraction and Reasoning Corpus benchmark, introduced by François Chollet and collaborators in May 2025. It uses visual grid transformation tasks to test abstract reasoning and fluid intelligence, with harder task designs than the original ARC-AGI to prevent saturation among frontier models.
What does a good ARC-AGI v2 score look like? Scores above 0.7 (70%) are frontier-tier. The current leader is GPT-5.5 at 0.850 (85%). Human performance approaches 1.0. Unlike the original ARC-AGI, scores below 0.3 remain common even for capable models, making ARC-AGI v2 a strong differentiator.
Who created ARC-AGI v2? ARC-AGI v2 was created by François Chollet, Mike Knoop, Gregory Kamradt, Bryan Landers and collaborators at the ARC Prize Foundation. The benchmark is described in the paper "ARC-AGI-2: A New Challenge for Frontier AI Reasoning Systems" (arXiv 2505.11831, May 2025).
Is ARC-AGI v2 saturated? No. With the top score at 85% and a large spread between top models and the rest of the field, ARC-AGI v2 remains an active and meaningful differentiator. It was specifically designed to avoid the saturation that affected the original ARC-AGI.
How does ARC-AGI v2 differ from the original ARC-AGI? Both use the same visual grid transformation format. ARC-AGI v2 introduces harder compositional rules and more complex multi-step transformations, reducing the effectiveness of brute-force test-time compute scaling that helped models achieve near-perfect scores on the original.

Benchmark definition paraphrased from Chollet et al. 2025. State-of-the-art scores sourced from llm-stats.com and attributed inline. Last updated 2026-07-23.