Benchgen

ARC-AGI-3 — Results

RankModelScore
1gpt-5-6-sol0.078
2gpt-5-6-terra0.008
3gpt-5-6-luna0.002
A

ARC-AGI-3

1 phaseActive

Third-generation ARC-AGI: interactive fluid reasoning evaluation, near-zero scores for all current frontier models. Multimodal, 0–1 scale.

Overview

ARC-AGI-3

Category Metric Saturation Created

Leaderboard

Quick answer: ARC-AGI-3 is the third generation of the Abstraction and Reasoning Corpus benchmark, designed as an interactive reasoning evaluation that measures fluid, novel problem-solving ability in AI systems. It is far from saturated — the current leader, GPT-5.6 Sol, scores only 7.8%, making it one of the hardest publicly tracked benchmarks for frontier models.

At a Glance

What it tests: Fluid, novel problem-solving through abstract visual reasoning tasks — the same grid-transformation format as ARC-AGI and ARC-AGI v2 but at a level of difficulty that current frontier models barely clear.

Why it matters: ARC-AGI-3 serves as the frontier of abstract reasoning evaluation. With top scores below 10%, it provides maximum headroom to distinguish progress in genuine reasoning capability over the coming years, without the saturation that affected earlier ARC generations.

Known limitations: Very few models have been evaluated (3 as of July 2026), all from OpenAI. The benchmark is early-stage and evaluation methodology is still being established.

What ARC-AGI-3 Measures

ARC-AGI-3 is the successor to ARC-AGI v2, further increasing task difficulty to maintain a meaningful challenge as AI reasoning capabilities improve. Like its predecessors, it uses visual grid transformation tasks where models must infer a hidden rule from a small set of input/output examples and apply it to a novel test case.

The "interactive" framing suggests ARC-AGI-3 may involve multi-turn evaluation where models can probe or query aspects of the task — a departure from the static example-based format of previous versions. This makes it a closer proxy for how AI systems reason through genuinely novel problems in practice.

Current results show a dramatic difficulty increase: while GPT-5.5 scores 85% on ARC-AGI v2, the same class of models scores below 8% on ARC-AGI-3, indicating a fundamental leap in required reasoning capability.

Benchmark Specifications

FieldValue
Task categoryAbstract reasoning / Vision
MetricAccuracy (fraction of tasks solved, 0–1)
SaturationLow — top score 7.8%
Created byARC Prize Foundation
ModalityMultimodal
Parent benchmarkARC-AGI v2

How ARC-AGI-3 Is Scored

Tasks are scored on a 0–1 binary scale — either the model produces the correct output (1) or it does not (0). The overall score is the fraction of tasks solved. A score of 0.078 means 7.8% of tasks were correctly solved.

ARC-AGI-3 on Benchgen

No Benchgen results yet — be the first to run ARC-AGI-3.

ARC-AGI-3 vs Other Benchmarks

BenchmarkWhat it testsSaturationTop score
ARC-AGI-3Interactive fluid reasoning (gen 3)Low7.8%
ARC-AGI v2Visual grid transformation (gen 2)Medium85.0%
ARC-AGIVisual grid transformation (gen 1)High~95%
GPQA DiamondExpert science reasoningLow~80%

ARC-AGI-3 is the most challenging publicly tracked abstract reasoning benchmark as of mid-2026, with essentially all current frontier models clustered near zero.

Run ARC-AGI-3 on Your Model

Track ARC-AGI-3 performance on Benchgen to get a regression-tested view of your model's abstract reasoning frontier over time. Given the extremely low current scores, even marginal improvements are significant signals.

Frequently Asked Questions

What is ARC-AGI-3? ARC-AGI-3 is the third-generation Abstraction and Reasoning Corpus benchmark, created by the ARC Prize Foundation. It uses interactive reasoning tasks to evaluate fluid, novel problem-solving ability in AI systems. It is far harder than ARC-AGI v2 — current frontier models score below 8%.
What does a good ARC-AGI-3 score look like? Any score above 0 is notable — the current leader is GPT-5.6 Sol at 7.8%. Human performance is expected to approach 1.0. As of mid-2026, ARC-AGI-3 remains essentially unsolved.
Is ARC-AGI-3 saturated? No. With a top score of only 7.8% and most frontier models near 0%, ARC-AGI-3 is one of the least saturated reasoning benchmarks currently tracked.
How does ARC-AGI-3 differ from ARC-AGI v2? ARC-AGI-3 features harder task designs and an interactive evaluation format. While frontier models score 49–85% on ARC-AGI v2, the same models score below 8% on ARC-AGI-3, indicating a fundamental jump in required reasoning capability.