| Rank | Model | Score |
|---|---|---|
| 1 | gpt-5-6-sol | 0.89 |
| 2 | gpt-5-6-terra | 0.849 |
| 3 | gpt-5-6-luna | 0.786 |
| 4 | gpt-5-4 | 0.476 |
| 5 | gpt-5-2 | 0.403 |
| 6 | gpt-5-5 | 0.354 |
| 7 | gpt-5-1-instant | 0.267 |
| 8 | gpt-5-1-thinking | 0.267 |
| 9 | gpt-5-1 | 0.267 |
| 10 | gpt-5 | 0.263 |
| 11 | gpt-5-mini | 0.221 |
| 12 | o3 | 0.158 |
| 13 | o3-mini | 0.092 |
| 14 | o1 | 0.055 |
1 phaseActive
Hundreds of original, exceptionally challenging mathematics problems crafted and vetted by expert mathematicians, spanning modern math from number theory to algebraic geometry. Metric: accuracy.
Quick answer: FrontierMath is a benchmark of hundreds of original, exceptionally challenging mathematics problems crafted and vetted by expert mathematicians, covering most major branches of modern mathematics — from number theory and real analysis to algebraic geometry and category theory. GPT-5.6 Sol leads with 89.0% across 17 evaluated models.
FrontierMath evaluates advanced mathematical reasoning at the level of research mathematicians. Unlike saturated benchmarks such as GSM8K or MATH, FrontierMath problems are designed to remain unsolved by current frontier models for years, requiring deep domain expertise, creative problem-solving, and often hours of effort even for expert humans.
| Focus area | Examples |
|---|---|
| Number theory | Original problems requiring deep number-theoretic insight |
| Real & complex analysis | Advanced analytic reasoning problems |
| Algebraic geometry | Graduate/research-level geometric reasoning |
| Category theory | Abstract structural mathematics problems |
Each problem has a verifiable, unambiguous final answer that can be automatically checked, avoiding the grading ambiguity common in open-ended math benchmarks. Accuracy is the fraction of problems solved correctly, normalized to 0–1. Problems are held out and unpublished to prevent contamination of model training data.
| Property | Value |
|---|---|
| Released | November 2024 |
| Created by | Epoch AI |
| Metric | Accuracy |
| Score range | 0–1 |
| Top model | GPT-5.6 Sol (0.890) |
| Models evaluated | 17 |
What is FrontierMath? FrontierMath is a benchmark of hundreds of original, exceptionally challenging mathematics problems crafted and vetted by expert mathematicians, spanning most major branches of modern mathematics.
Who created FrontierMath? FrontierMath was created by Epoch AI — Elliot Glazer, Ege Erdil, Tamay Besiroglu, Diego Chicharro, and 20 other collaborators — published in November 2024.
Why is FrontierMath harder than other math benchmarks? Problems are original and unpublished, vetted by expert mathematicians, and designed to require genuine mathematical insight rather than pattern matching — most problems take expert mathematicians substantial time to solve, and the questions are held out to prevent training data contamination.
What score does the best model achieve on FrontierMath? GPT-5.6 Sol leads with 0.890 (89.0%), followed by GPT-5.6 Terra at 0.849 and GPT-5.6 Luna at 0.786 — a rapid jump from earlier frontier models like GPT-5 which scored 0.263.
Is FrontierMath saturated? No. Despite recent gains at the frontier, average scores across all evaluated models remain around 0.3, and many older frontier models score below 0.1, indicating substantial remaining headroom.