Benchgen

Humanity's Last Exam — Results

RankModelScore
1claude-sonnet-557.4
2kimi-k356
3seed-2-1-pro55.7
4glm-5-254.7
5seed-2-1-turbo54.6
6kimi-k2-thinking-090551
7fugu-ultra50
8qwen3-5-27b48.5
9deepseek-v4-pro-max48.2
10qwen3-5-122b-a10b47.5
11qwen3-5-35b-a3b47.4
12fugu47.2
13gemini-3-1-pro46.44
14inkling46
15deepseek-v4-flash-max45.1
16gpt-5-444.32
17qwen3-8-max43.6
18gemini-3-flash43.5
19claude-sonnet-543.2
20glm-4-742.8
21gpt-5-541.4
22qwen3-7-max41.4
23deepseek-v3-240.8
24muse-spark40.56
25grok-440

Humanity's Last Exam

1 phaseActive

Extreme-difficulty benchmark of ~3,000 expert-level questions from PhD researchers — the hardest publicly available AI knowledge test. Metric: % correct.

Overview

Humanity's Last Exam

Category Metric Tasks Saturation Created

Paper Dataset

Quick answer: Humanity's Last Exam (HLE) is an extreme-difficulty benchmark of approximately 3,000 expert-level questions spanning science, mathematics, and humanities, created by the Center for AI Safety and Scale AI (Phan et al., 2025). Questions were contributed by PhD researchers and are designed to be at or beyond the frontier of current AI capability. When first released, frontier models scored below 10%; by June 2026, Fugu Ultra reaches 50.0%.

At a Glance

What it tests: Broad, expert-level knowledge and reasoning across science (physics, chemistry, biology), mathematics, history, law, economics, and other disciplines — questions that require PhD-level understanding to answer correctly.

Why it matters: As models rapidly saturate MMLU and GPQA, HLE provides a long-lived discrimination signal at the very frontier of AI capability. A model that scores 50% on HLE is genuinely competing with PhD-level human experts across diverse fields.

Known limitations: Questions are extremely heterogeneous in difficulty and domain; a high overall score may mask large domain-specific gaps. Some questions may become contaminated over time as they circulate. Results vary significantly with and without tool use.

What Humanity's Last Exam Measures

HLE was built by inviting domain experts — primarily PhD students, postdocs, and faculty — to submit questions that they believed current frontier AI models would fail. The approximately 3,000 questions span over 100 academic disciplines, from advanced mathematics and cutting-edge physics to obscure historical and legal questions. Each question has a verifiable, unambiguous correct answer, and the benchmark uses exact-match or closed-form scoring.

When HLE launched in January 2025, the best frontier models scored below 10%, making it the most informative new hard-knowledge benchmark since MMLU. By mid-2026, scores have risen dramatically — Fugu Ultra at 50.0% represents a 5× improvement over early 2025 baselines — reflecting both model capability gains and the effect of multi-agent coordination and tool use.

Note: HLE scores vary substantially based on whether tool use (code execution, web search) is permitted. Scores reported without tools are generally lower. Always check the source for the evaluation conditions used.

Benchmark Specifications

FieldValue
Task categoryReasoning / knowledge
Metric% correct (accuracy)
Number of questions~3,000
Domain coverage100+ academic disciplines
SaturationLow
Created byCenter for AI Safety / Scale AI (Phan et al.)
Source paperHumanity's Last Exam (2025)
DatasetHuggingFace — cais/hle

How HLE Is Scored

Each question has a single verified correct answer. A model receives 1 point for each correct response and 0 for incorrect. The final score is the percentage of questions answered correctly. Because questions vary enormously in domain, an overall score provides a broad capability signal but may disguise specific domain weaknesses. Vendor-reported scores often differ from independent evaluations due to differences in prompting, tool access, and retries.

State-of-the-Art Results

RankModelScoreSourceDate
1Fugu Ultra50.0%Sakana Fugu technical report2026-06
2Fable 5 / Mythos Preview (max)49.8%Sakana Fugu technical report2026-06
3Fugu47.2%Sakana Fugu technical report2026-06

Scores sourced from Sakana AI's Fugu technical report, June 2026. Evaluation conditions (tool use, etc.) follow Sakana AI's methodology — see technical report for details.