| Rank | Model | Score |
|---|---|---|
| 1 | claude-sonnet-5 | 57.4 |
| 2 | kimi-k3 | 56 |
| 3 | seed-2-1-pro | 55.7 |
| 4 | glm-5-2 | 54.7 |
| 5 | seed-2-1-turbo | 54.6 |
| 6 | kimi-k2-thinking-0905 | 51 |
| 7 | fugu-ultra | 50 |
| 8 | qwen3-5-27b | 48.5 |
| 9 | deepseek-v4-pro-max | 48.2 |
| 10 | qwen3-5-122b-a10b | 47.5 |
| 11 | qwen3-5-35b-a3b | 47.4 |
| 12 | fugu | 47.2 |
| 13 | gemini-3-1-pro | 46.44 |
| 14 | inkling | 46 |
| 15 | deepseek-v4-flash-max | 45.1 |
| 16 | gpt-5-4 | 44.32 |
| 17 | qwen3-8-max | 43.6 |
| 18 | gemini-3-flash | 43.5 |
| 19 | claude-sonnet-5 | 43.2 |
| 20 | glm-4-7 | 42.8 |
| 21 | gpt-5-5 | 41.4 |
| 22 | qwen3-7-max | 41.4 |
| 23 | deepseek-v3-2 | 40.8 |
| 24 | muse-spark | 40.56 |
| 25 | grok-4 | 40 |
1 phaseActive
Extreme-difficulty benchmark of ~3,000 expert-level questions from PhD researchers — the hardest publicly available AI knowledge test. Metric: % correct.
Quick answer: Humanity's Last Exam (HLE) is an extreme-difficulty benchmark of approximately 3,000 expert-level questions spanning science, mathematics, and humanities, created by the Center for AI Safety and Scale AI (Phan et al., 2025). Questions were contributed by PhD researchers and are designed to be at or beyond the frontier of current AI capability. When first released, frontier models scored below 10%; by June 2026, Fugu Ultra reaches 50.0%.
What it tests: Broad, expert-level knowledge and reasoning across science (physics, chemistry, biology), mathematics, history, law, economics, and other disciplines — questions that require PhD-level understanding to answer correctly.
Why it matters: As models rapidly saturate MMLU and GPQA, HLE provides a long-lived discrimination signal at the very frontier of AI capability. A model that scores 50% on HLE is genuinely competing with PhD-level human experts across diverse fields.
Known limitations: Questions are extremely heterogeneous in difficulty and domain; a high overall score may mask large domain-specific gaps. Some questions may become contaminated over time as they circulate. Results vary significantly with and without tool use.
HLE was built by inviting domain experts — primarily PhD students, postdocs, and faculty — to submit questions that they believed current frontier AI models would fail. The approximately 3,000 questions span over 100 academic disciplines, from advanced mathematics and cutting-edge physics to obscure historical and legal questions. Each question has a verifiable, unambiguous correct answer, and the benchmark uses exact-match or closed-form scoring.
When HLE launched in January 2025, the best frontier models scored below 10%, making it the most informative new hard-knowledge benchmark since MMLU. By mid-2026, scores have risen dramatically — Fugu Ultra at 50.0% represents a 5× improvement over early 2025 baselines — reflecting both model capability gains and the effect of multi-agent coordination and tool use.
Note: HLE scores vary substantially based on whether tool use (code execution, web search) is permitted. Scores reported without tools are generally lower. Always check the source for the evaluation conditions used.
| Field | Value |
|---|---|
| Task category | Reasoning / knowledge |
| Metric | % correct (accuracy) |
| Number of questions | ~3,000 |
| Domain coverage | 100+ academic disciplines |
| Saturation | Low |
| Created by | Center for AI Safety / Scale AI (Phan et al.) |
| Source paper | Humanity's Last Exam (2025) |
| Dataset | HuggingFace — cais/hle |
Each question has a single verified correct answer. A model receives 1 point for each correct response and 0 for incorrect. The final score is the percentage of questions answered correctly. Because questions vary enormously in domain, an overall score provides a broad capability signal but may disguise specific domain weaknesses. Vendor-reported scores often differ from independent evaluations due to differences in prompting, tool access, and retries.
| Rank | Model | Score | Source | Date |
|---|---|---|---|---|
| 1 | Fugu Ultra | 50.0% | Sakana Fugu technical report | 2026-06 |
| 2 | Fable 5 / Mythos Preview (max) | 49.8% | Sakana Fugu technical report | 2026-06 |
| 3 | Fugu | 47.2% | Sakana Fugu technical report | 2026-06 |
Scores sourced from Sakana AI's Fugu technical report, June 2026. Evaluation conditions (tool use, etc.) follow Sakana AI's methodology — see technical report for details.