Benchgen

AIME 2026 — Results

RankModelScore
1gpt-5-6-sol99.9
2glm-5-299.2
3gemini-3-1-pro98.3
4inkling97.1
5deepseek-v4-pro96.7
6namazu96.67
7kimi-k2-696.4
8kimi-k2-595.8
9qwen3-6-plus95.3
10solar-pro-495.3
11muse-glimmer94.7
12nemotron-3-ultra-550b-a55b94.2
13qwen3-6-27b94.1
14qwen3-6-35b-a3b92.7
15qwen3-5-397b-a17b91.3
16maple-preview87.5

AIME 2026

1 phaseActive

Competition math benchmark — 30 problems from AIME I & II 2026, integer answers 0–999. Tests frontier model mathematical reasoning. Metric: % correct (pass@1).

Overview

AIME 2026

Category Metric Tasks Saturation Administered

Problems

Quick answer: AIME 2026 is an AI evaluation benchmark built from the 30 problems of the 2026 American Invitational Mathematics Examination (15 problems from AIME I, administered February 5 2026, and 15 from AIME II). It tests frontier model mathematical reasoning at competition level — problems require algebra, combinatorics, number theory, and geometry, with integer answers in the range 000–999. GLM 5.2 leads the Inkling comparison set at 99.2%.

At a Glance

What it tests: Multi-step mathematical reasoning across algebra, combinatorics, number theory, geometry, and probability at the high-school olympiad level. Problems require deriving a specific integer answer (000–999), preventing partial credit or lucky guessing.

Why it matters: AIME sits above AMC 10/12 but below USAMO in difficulty — it is the most widely-used intermediate competition math benchmark for AI evaluation. Every problem requires a non-trivial derivation chain, making it a reliable signal for genuine mathematical reasoning rather than pattern matching to memorised solutions.

Known limitations: 30 problems is a small evaluation set, making scores high-variance (each problem is worth ~3.3%). Since AIME 2026 problems are publicly available, results for models trained after February 2026 carry contamination risk — scores may reflect memorised solutions rather than reasoning. Models from labs that explicitly test on the full problem set (not a held-out split) should be compared cautiously.

What AIME 2026 Measures

The American Invitational Mathematics Examination is an annual U.S. high-school mathematics competition administered by the Mathematical Association of America (MAA). Students who qualify via the AMC 10 or AMC 12 sit a 3-hour, 15-problem exam; answers are integers from 000 to 999, with no multiple-choice options. AIME 2026 ran in two versions: AIME I (February 5, 2026) and AIME II, each with 15 problems of escalating difficulty.

For AI evaluation, the standard protocol is to run each model on all 30 problems (both AIME I and AIME II) and report the pass@1 accuracy — the percentage of problems where the model's first attempt produces the correct integer answer. Some labs also report pass@k (majority vote over multiple samples). Scores in the Inkling table are pass@1 at effort=0.99.

Problem categories broadly cover: algebra and polynomials, combinatorics and counting, number theory and modular arithmetic, Euclidean and coordinate geometry, and probability. Problems are self-contained and do not require specialised knowledge beyond pre-calculus.

Benchmark Specifications

FieldValue
Task categoryMathematical reasoning
Metric% correct (pass@1, integer answer matching)
Number of problems30 (15 × AIME I + 15 × AIME II)
Answer rangeInteger 000–999
AdministeredFebruary 2026
DifficultyHigh-school olympiad (post-AMC, pre-USAMO)
SaturationLow
Created byMathematical Association of America (MAA)
SourceArt of Problem Solving AIME archive

How AIME 2026 Is Scored

A model is prompted with each problem statement and asked to produce an integer answer. The answer is graded as correct if it exactly matches the official answer key (000–999). Scores are reported as % of 30 problems correct (pass@1). Chain-of-thought reasoning is standard; temperature 0 or majority-vote sampling is common. Some evaluations allow tool use (e.g. Python code execution for computation); the Inkling results are at effort=0.99 with no external tools specified.

State-of-the-Art Results

Scores from Inkling model card (Thinking Machines Lab, July 2026), evaluated at effort=0.99. Claude Fable 5 score not reported in source table.

RankModelScoreWeights
1GLM 5.299.2%Open
2GPT-5.6 Sol99.9%Closed
3Gemini 3.1 Pro98.3%Closed
4Inkling97.1%Open
5DeepSeek V4 Pro96.7%Open
6Kimi K2.696.4%Open
7Kimi K2.595.8%Open
8Nemotron 3 Ultra94.2%Open

Frontier model scores above 94% reflect near-saturation — the benchmark has limited discriminative power at the very top of the leaderboard as of mid-2026.

BenchmarkProblemsFormatDifficultySaturation
AIME 202630Integer answerHigh-school olympiadLow
AMC 10/12305-choice MCQHigh-school competitionHigh
Humanity's Last Exam~3,000MixedPhD+ levelLow
GPQA Diamond1984-choice MCQPhD scienceLow
MATH-500500Free-formCompetition mathMedium

Last updated 2026-07-16.