| Rank | Model | Score |
|---|---|---|
| 1 | gpt-5-6-sol | 0.078 |
| 2 | gpt-5-6-terra | 0.008 |
| 3 | gpt-5-6-luna | 0.002 |
1 phaseActive
Third-generation ARC-AGI: interactive fluid reasoning evaluation, near-zero scores for all current frontier models. Multimodal, 0–1 scale.
Quick answer: ARC-AGI-3 is the third generation of the Abstraction and Reasoning Corpus benchmark, designed as an interactive reasoning evaluation that measures fluid, novel problem-solving ability in AI systems. It is far from saturated — the current leader, GPT-5.6 Sol, scores only 7.8%, making it one of the hardest publicly tracked benchmarks for frontier models.
What it tests: Fluid, novel problem-solving through abstract visual reasoning tasks — the same grid-transformation format as ARC-AGI and ARC-AGI v2 but at a level of difficulty that current frontier models barely clear.
Why it matters: ARC-AGI-3 serves as the frontier of abstract reasoning evaluation. With top scores below 10%, it provides maximum headroom to distinguish progress in genuine reasoning capability over the coming years, without the saturation that affected earlier ARC generations.
Known limitations: Very few models have been evaluated (3 as of July 2026), all from OpenAI. The benchmark is early-stage and evaluation methodology is still being established.
ARC-AGI-3 is the successor to ARC-AGI v2, further increasing task difficulty to maintain a meaningful challenge as AI reasoning capabilities improve. Like its predecessors, it uses visual grid transformation tasks where models must infer a hidden rule from a small set of input/output examples and apply it to a novel test case.
The "interactive" framing suggests ARC-AGI-3 may involve multi-turn evaluation where models can probe or query aspects of the task — a departure from the static example-based format of previous versions. This makes it a closer proxy for how AI systems reason through genuinely novel problems in practice.
Current results show a dramatic difficulty increase: while GPT-5.5 scores 85% on ARC-AGI v2, the same class of models scores below 8% on ARC-AGI-3, indicating a fundamental leap in required reasoning capability.
| Field | Value |
|---|---|
| Task category | Abstract reasoning / Vision |
| Metric | Accuracy (fraction of tasks solved, 0–1) |
| Saturation | Low — top score 7.8% |
| Created by | ARC Prize Foundation |
| Modality | Multimodal |
| Parent benchmark | ARC-AGI v2 |
Tasks are scored on a 0–1 binary scale — either the model produces the correct output (1) or it does not (0). The overall score is the fraction of tasks solved. A score of 0.078 means 7.8% of tasks were correctly solved.
No Benchgen results yet — be the first to run ARC-AGI-3.
| Benchmark | What it tests | Saturation | Top score |
|---|---|---|---|
| ARC-AGI-3 | Interactive fluid reasoning (gen 3) | Low | 7.8% |
| ARC-AGI v2 | Visual grid transformation (gen 2) | Medium | 85.0% |
| ARC-AGI | Visual grid transformation (gen 1) | High | ~95% |
| GPQA Diamond | Expert science reasoning | Low | ~80% |
ARC-AGI-3 is the most challenging publicly tracked abstract reasoning benchmark as of mid-2026, with essentially all current frontier models clustered near zero.
Track ARC-AGI-3 performance on Benchgen to get a regression-tested view of your model's abstract reasoning frontier over time. Given the extremely low current scores, even marginal improvements are significant signals.