Benchgen

AIME 2025 — Results

RankModelScore
1granite-4-2-30b89.17
2granite-4-2-8b86.67
3granite-4-2-3b78.33
4lfm2-5-2-6b51.87
5k2-horizon-0-9b41.7
6gemini-3-pro1
7gpt-5-2-pro-2025-12-111
8gpt-5-21
9kimi-k2-thinking-09051
10claude-opus-4-60.998
11gemini-3-flash0.997
12seed-2-0-pro0.983
13kimi-k2-50.961
14deepseek-v3-2-speciale0.96
15glm-4-70.957
16gpt-50.946
17gpt-5-1-instant0.94
18gpt-5-1-thinking0.94
19gpt-5-10.94
20deepseek-v3-2-thinking0.931
21deepseek-v3-20.931
22seed-2-0-lite0.93
23o4-mini0.927
24gpt-oss-120b-high0.925
25qwen3-235b-a22b-thinking-25070.923

AIME 2025

1 phaseActive

Competition math benchmark — 30 problems from AIME I & II 2025, integer answers 0–999. Tests frontier model mathematical reasoning. Metric: % correct (pass@1).

Overview

AIME 2025

Category Metric Tasks Saturation Administered

Problems

Quick answer: AIME 2025 is an AI evaluation benchmark built from the 30 problems of the 2025 American Invitational Mathematics Examination (15 problems from AIME I, administered February 2025, and 15 from AIME II). It tests frontier model mathematical reasoning at competition level — problems require algebra, combinatorics, number theory, and geometry, with integer answers in the range 000–999. With 114+ models evaluated, it is one of the most widely tracked math benchmarks for frontier AI.

At a Glance

What it tests: Multi-step mathematical reasoning across algebra, combinatorics, number theory, geometry, and probability at the high-school olympiad level. Problems require deriving a specific integer answer (000–999), preventing partial credit or lucky guessing.

Why it matters: AIME 2025 became a critical benchmark for tracking reasoning model progress through 2025–2026, with o3, Gemini 3 Pro, and GPT-5 families all scoring above 90%. It sits above AMC but below USAMO, making it ideal for discriminating among top-tier models.

Known limitations: 30 problems is a small evaluation set, making scores high-variance (each problem is worth ~3.3%). Models trained after February 2025 carry contamination risk.

What AIME 2025 Measures

The American Invitational Mathematics Examination is an annual U.S. high-school mathematics competition administered by the Mathematical Association of America (MAA). For AI evaluation, the standard protocol is to run each model on all 30 problems (both AIME I and AIME II) and report the pass@1 accuracy — the percentage of problems where the model's first attempt produces the correct integer answer.

Problem categories broadly cover: algebra and polynomials, combinatorics and counting, number theory and modular arithmetic, Euclidean and coordinate geometry, and probability.

Benchmark Specifications

FieldValue
Problems30 (15 AIME I + 15 AIME II)
Answer formatInteger 000–999
Primary metricPass@1 accuracy (% correct)
Year2025
Administered byMathematical Association of America
SaturationLow — frontier models approach 90%+

Frequently Asked Questions

What score does a top AI model get on AIME 2025? As of mid-2026, leading reasoning models achieve 80–95% on AIME 2025. o3, Gemini 3 Pro, and GPT-5 families score at the high end.

Is AIME 2025 contaminated? Models with training cutoffs after February 2025 may have seen the problems. Pass@1 scores should be interpreted cautiously for post-cutoff models; some labs report separate held-out results.

How does AIME 2025 compare to AIME 2024? The problem sets are independently generated each year. AIME 2025 is generally considered comparable in difficulty to 2024, though difficulty can vary by year and problem category.