Benchgen

FrontierMath — Results

RankModelScore
1gpt-5-6-sol0.89
2gpt-5-6-terra0.849
3gpt-5-6-luna0.786
4gpt-5-40.476
5gpt-5-20.403
6gpt-5-50.354
7gpt-5-1-instant0.267
8gpt-5-1-thinking0.267
9gpt-5-10.267
10gpt-50.263
11gpt-5-mini0.221
12o30.158
13o3-mini0.092
14o10.055

FrontierMath

1 phaseActive

Hundreds of original, exceptionally challenging mathematics problems crafted and vetted by expert mathematicians, spanning modern math from number theory to algebraic geometry. Metric: accuracy.

Overview

FrontierMath

Category Metric Saturation

Quick answer: FrontierMath is a benchmark of hundreds of original, exceptionally challenging mathematics problems crafted and vetted by expert mathematicians, covering most major branches of modern mathematics — from number theory and real analysis to algebraic geometry and category theory. GPT-5.6 Sol leads with 89.0% across 17 evaluated models.


What Does FrontierMath Test?

FrontierMath evaluates advanced mathematical reasoning at the level of research mathematicians. Unlike saturated benchmarks such as GSM8K or MATH, FrontierMath problems are designed to remain unsolved by current frontier models for years, requiring deep domain expertise, creative problem-solving, and often hours of effort even for expert humans.

Focus areaExamples
Number theoryOriginal problems requiring deep number-theoretic insight
Real & complex analysisAdvanced analytic reasoning problems
Algebraic geometryGraduate/research-level geometric reasoning
Category theoryAbstract structural mathematics problems

How Is FrontierMath Scored?

Each problem has a verifiable, unambiguous final answer that can be automatically checked, avoiding the grading ambiguity common in open-ended math benchmarks. Accuracy is the fraction of problems solved correctly, normalized to 0–1. Problems are held out and unpublished to prevent contamination of model training data.


Key Facts

PropertyValue
ReleasedNovember 2024
Created byEpoch AI
MetricAccuracy
Score range0–1
Top modelGPT-5.6 Sol (0.890)
Models evaluated17

FAQ

What is FrontierMath? FrontierMath is a benchmark of hundreds of original, exceptionally challenging mathematics problems crafted and vetted by expert mathematicians, spanning most major branches of modern mathematics.

Who created FrontierMath? FrontierMath was created by Epoch AI — Elliot Glazer, Ege Erdil, Tamay Besiroglu, Diego Chicharro, and 20 other collaborators — published in November 2024.

Why is FrontierMath harder than other math benchmarks? Problems are original and unpublished, vetted by expert mathematicians, and designed to require genuine mathematical insight rather than pattern matching — most problems take expert mathematicians substantial time to solve, and the questions are held out to prevent training data contamination.

What score does the best model achieve on FrontierMath? GPT-5.6 Sol leads with 0.890 (89.0%), followed by GPT-5.6 Terra at 0.849 and GPT-5.6 Luna at 0.786 — a rapid jump from earlier frontier models like GPT-5 which scored 0.263.

Is FrontierMath saturated? No. Despite recent gains at the frontier, average scores across all evaluated models remain around 0.3, and many older frontier models score below 0.1, indicating substantial remaining headroom.