| Rank | Model | Score |
|---|---|---|
| 1 | gpt-5-6-sol | 0.527 |
| 2 | gpt-5-6-terra | 0.504 |
| 3 | gpt-5-6-luna | 0.503 |
| 4 | claude-fable-5 | 0.487 |
| 5 | gpt-5-5 | 0.479 |
| 6 | claude-opus-4-8 | 0.451 |
| 7 | claude-opus-4-7 | 0.418 |
| 8 | seed-2-1-pro | 0.414 |
| 9 | glm-5-2 | 0.406 |
| 10 | gpt-5-4 | 0.373 |
| 11 | gemini-3-1-pro | 0.32 |
| 12 | qwen3-7-max | 0.311 |
| 13 | deepseek-v4-pro | 0.276 |
| 14 | qwen3-6-plus | 0.243 |
| 15 | kimi-k2-6 | 0.217 |
| 16 | grok-4-20 | 0.201 |
| 17 | minimax-m2 | 0.142 |
1 phaseActive
Berkeley RDI's benchmark of 1K+ economically valuable long-horizon agent tasks across 13 industries, designed to close the gap between benchmark scores and GDP impact.
Quick answer: Agents' Last Exam (ALE) is a benchmark from UC Berkeley's RDI group (Sun et al., June 2026) that evaluates AI agents on 1,000+ long-horizon, economically valuable real-world tasks spanning 55 sub-fields across 13 industry clusters — built with 250+ industry experts using the O*NET/SOC 2018 occupational taxonomy. Unlike academic benchmarks, ALE is designed as a living benchmark specifically to close the gap between leaderboard performance and GDP-relevant impact.
What it tests: An AI agent's ability to complete sustained, multi-step professional workflows with verifiable outcomes across non-physical industries — from legal research to financial analysis, data science, and beyond.
Why it matters: Most agent benchmarks use synthetic or simplified tasks that don't map to actual work that generates economic value. ALE was co-designed with 250+ industry practitioners to ensure every task represents a genuine, high-value workflow. The hardest tier sees average pass rates below 1%, making it a strong differentiator between frontier models even as easier benchmarks saturate.
Known limitations: As a living benchmark, task difficulty and coverage evolve over time, making direct comparisons across evaluation dates unreliable without version pinning. The full hardest tier is extremely difficult, so aggregate scores primarily reflect performance on easier sub-tiers.
ALE organises tasks using the U.S. federal occupational taxonomy (O*NET/SOC 2018) as a structural backbone, covering non-physical industries across 13 clusters and 55 sub-fields. This grounds the benchmark in work that is already defined, scoped, and valued in the real economy — rather than researcher-invented synthetic tasks. Each task has a verifiable outcome, enabling automated evaluation without LLM-as-judge.
The benchmark is structured into difficulty tiers. The full-task ("hardest") tier requires complete end-to-end execution of a professional workflow, where current state-of-the-art models average below 1% pass rate. Easier tiers decompose tasks into sub-steps, producing more informative scores for current models. The leaderboard reports results across harness and backbone configurations, reflecting that agent benchmarks evaluate harness-model pairs, not models in isolation.
ALE is explicitly designed as a living benchmark: its task pool grows continuously as new industries and workflows are onboarded, and the benchmark is intended to remain unsaturated as AI capabilities advance.
| Field | Value |
|---|---|
| Task category | Agent / Long-horizon |
| Metric | Pass rate (0–1) |
| Number of tasks | 1,000+ (growing) |
| Industry clusters | 13 clusters, 55 sub-fields |
| Taxonomy basis | O*NET / SOC 2018 |
| Saturation | Low (hardest tier <1% avg) |
| Created by | Sun et al. (UC Berkeley RDI) |
| Source paper | Sun et al. 2026 |
| GitHub | rdi-berkeley/agents-last-exam |
| Dataset | Not yet public |
Each task is scored pass/fail based on whether the agent's output satisfies the verifiable outcome criterion. The benchmark reports a pass rate (fraction of tasks passed) per tier and per industry cluster. Because ALE evaluates harness-model pairs, results depend heavily on the scaffolding used — the same backbone model can achieve substantially different scores with different agent frameworks.
| Rank | Model | Score | Source | Date |
|---|---|---|---|---|
| 1 | GPT-5.6 Sol | 0.527 | llm-stats.com | 2026-07 |
| 2 | GPT-5.6 Terra | 0.504 | llm-stats.com | 2026-07 |
| 3 | GPT-5.6 Luna | 0.503 | llm-stats.com | 2026-07 |
| 4 | Seed 2.1 Pro | 0.414 | llm-stats.com | 2026-07 |
Scores sourced from llm-stats.com leaderboard. ALE evaluates harness-model pairs — scores vary significantly by scaffolding. See each source for methodology.
No Benchgen results yet — be the first to run Agents' Last Exam.
| Benchmark | What it tests | Tasks | Saturation |
|---|---|---|---|
| Agents' Last Exam | Economically valuable professional workflows | 1,000+ | Low |
| Tau3 Banking | Customer service agentic tasks (banking) | ~200 | Low |
| SWE-Bench Pro | Software engineering workflows | ~500 | Low |
| MCP Atlas | MCP tool-use agentic workflows | ~500 | Low |
ALE's unique strength is its breadth across 13 real-economy industries and its grounding in O*NET occupational taxonomy — use it for cross-domain agent capability assessment. Use domain-specific benchmarks when you need detailed coverage of a single vertical.
Benchgen lets you run Agents' Last Exam against your own model and harness configurations, version-tracking results across deployments so you can detect regressions and measure the real impact of fine-tuning on economically relevant task performance.
Agents' Last Exam (ALE) is a benchmark for evaluating AI agents on 1,000+ economically valuable, long-horizon professional tasks across 13 industry clusters. Developed by UC Berkeley RDI with 250+ industry experts, it uses O*NET/SOC 2018 as a structural taxonomy and is designed as a living benchmark to remain unsaturated as capabilities advance.
ALE reports a pass rate (0–1). On the hardest full-task tier, average pass rates are below 1% for current models. The aggregate scores visible on leaderboards (GPT-5.6 Sol at 52.7%) reflect performance across easier sub-tiers. Any score above 50% on aggregate is frontier-class as of mid-2026.
Agents' Last Exam was created by Yiyou Sun, Xinyang Han, and 285+ co-authors at UC Berkeley's RDI (Center for Responsible Decentralized Intelligence), in collaboration with 250+ industry experts. The paper was submitted in June 2026 (arXiv:2606.05405).