Benchgen

Agents' Last Exam — Results

RankModelScore
1gpt-5-6-sol0.527
2gpt-5-6-terra0.504
3gpt-5-6-luna0.503
4claude-fable-50.487
5gpt-5-50.479
6claude-opus-4-80.451
7claude-opus-4-70.418
8seed-2-1-pro0.414
9glm-5-20.406
10gpt-5-40.373
11gemini-3-1-pro0.32
12qwen3-7-max0.311
13deepseek-v4-pro0.276
14qwen3-6-plus0.243
15kimi-k2-60.217
16grok-4-200.201
17minimax-m20.142
A

Agents' Last Exam

1 phaseActive

Berkeley RDI's benchmark of 1K+ economically valuable long-horizon agent tasks across 13 industries, designed to close the gap between benchmark scores and GDP impact.

Overview

Agents' Last Exam

Category Metric Tasks Saturation Created

Paper GitHub Website

Quick answer: Agents' Last Exam (ALE) is a benchmark from UC Berkeley's RDI group (Sun et al., June 2026) that evaluates AI agents on 1,000+ long-horizon, economically valuable real-world tasks spanning 55 sub-fields across 13 industry clusters — built with 250+ industry experts using the O*NET/SOC 2018 occupational taxonomy. Unlike academic benchmarks, ALE is designed as a living benchmark specifically to close the gap between leaderboard performance and GDP-relevant impact.

At a Glance

What it tests: An AI agent's ability to complete sustained, multi-step professional workflows with verifiable outcomes across non-physical industries — from legal research to financial analysis, data science, and beyond.

Why it matters: Most agent benchmarks use synthetic or simplified tasks that don't map to actual work that generates economic value. ALE was co-designed with 250+ industry practitioners to ensure every task represents a genuine, high-value workflow. The hardest tier sees average pass rates below 1%, making it a strong differentiator between frontier models even as easier benchmarks saturate.

Known limitations: As a living benchmark, task difficulty and coverage evolve over time, making direct comparisons across evaluation dates unreliable without version pinning. The full hardest tier is extremely difficult, so aggregate scores primarily reflect performance on easier sub-tiers.

What Agents' Last Exam Measures

ALE organises tasks using the U.S. federal occupational taxonomy (O*NET/SOC 2018) as a structural backbone, covering non-physical industries across 13 clusters and 55 sub-fields. This grounds the benchmark in work that is already defined, scoped, and valued in the real economy — rather than researcher-invented synthetic tasks. Each task has a verifiable outcome, enabling automated evaluation without LLM-as-judge.

The benchmark is structured into difficulty tiers. The full-task ("hardest") tier requires complete end-to-end execution of a professional workflow, where current state-of-the-art models average below 1% pass rate. Easier tiers decompose tasks into sub-steps, producing more informative scores for current models. The leaderboard reports results across harness and backbone configurations, reflecting that agent benchmarks evaluate harness-model pairs, not models in isolation.

ALE is explicitly designed as a living benchmark: its task pool grows continuously as new industries and workflows are onboarded, and the benchmark is intended to remain unsaturated as AI capabilities advance.

Benchmark Specifications

FieldValue
Task categoryAgent / Long-horizon
MetricPass rate (0–1)
Number of tasks1,000+ (growing)
Industry clusters13 clusters, 55 sub-fields
Taxonomy basisO*NET / SOC 2018
SaturationLow (hardest tier <1% avg)
Created bySun et al. (UC Berkeley RDI)
Source paperSun et al. 2026
GitHubrdi-berkeley/agents-last-exam
DatasetNot yet public

How Agents' Last Exam Is Scored

Each task is scored pass/fail based on whether the agent's output satisfies the verifiable outcome criterion. The benchmark reports a pass rate (fraction of tasks passed) per tier and per industry cluster. Because ALE evaluates harness-model pairs, results depend heavily on the scaffolding used — the same backbone model can achieve substantially different scores with different agent frameworks.

State-of-the-Art Results

RankModelScoreSourceDate
1GPT-5.6 Sol0.527llm-stats.com2026-07
2GPT-5.6 Terra0.504llm-stats.com2026-07
3GPT-5.6 Luna0.503llm-stats.com2026-07
4Seed 2.1 Pro0.414llm-stats.com2026-07

Scores sourced from llm-stats.com leaderboard. ALE evaluates harness-model pairs — scores vary significantly by scaffolding. See each source for methodology.

Agents' Last Exam on Benchgen

No Benchgen results yet — be the first to run Agents' Last Exam.

Agents' Last Exam vs Other Benchmarks

BenchmarkWhat it testsTasksSaturation
Agents' Last ExamEconomically valuable professional workflows1,000+Low
Tau3 BankingCustomer service agentic tasks (banking)~200Low
SWE-Bench ProSoftware engineering workflows~500Low
MCP AtlasMCP tool-use agentic workflows~500Low

ALE's unique strength is its breadth across 13 real-economy industries and its grounding in O*NET occupational taxonomy — use it for cross-domain agent capability assessment. Use domain-specific benchmarks when you need detailed coverage of a single vertical.

Run Agents' Last Exam on Your Model

Benchgen lets you run Agents' Last Exam against your own model and harness configurations, version-tracking results across deployments so you can detect regressions and measure the real impact of fine-tuning on economically relevant task performance.

Frequently Asked Questions

What is Agents' Last Exam?

Agents' Last Exam (ALE) is a benchmark for evaluating AI agents on 1,000+ economically valuable, long-horizon professional tasks across 13 industry clusters. Developed by UC Berkeley RDI with 250+ industry experts, it uses O*NET/SOC 2018 as a structural taxonomy and is designed as a living benchmark to remain unsaturated as capabilities advance.

What does a good score look like on Agents' Last Exam?

ALE reports a pass rate (0–1). On the hardest full-task tier, average pass rates are below 1% for current models. The aggregate scores visible on leaderboards (GPT-5.6 Sol at 52.7%) reflect performance across easier sub-tiers. Any score above 50% on aggregate is frontier-class as of mid-2026.

Who created Agents' Last Exam?

Agents' Last Exam was created by Yiyou Sun, Xinyang Han, and 285+ co-authors at UC Berkeley's RDI (Center for Responsible Decentralized Intelligence), in collaboration with 250+ industry experts. The paper was submitted in June 2026 (arXiv:2606.05405).