Benchgen

AIME 2025 — Results

RankModelScore
1lfm2-5-2-6b51.87
2gemini-3-pro1
3gpt-5-2-pro-2025-12-111
4gpt-5-21
5kimi-k2-thinking-09051
6claude-opus-4-60.998
7gemini-3-flash0.997
8seed-2-0-pro0.983
9kimi-k2-50.961
10deepseek-v3-2-speciale0.96
11glm-4-70.957
12gpt-50.946
13gpt-5-1-instant0.94
14gpt-5-1-thinking0.94
15gpt-5-10.94
16deepseek-v3-2-thinking0.931
17deepseek-v3-20.931
18seed-2-0-lite0.93
19o4-mini0.927
20gpt-oss-120b-high0.925
21qwen3-235b-a22b-thinking-25070.923
22grok-40.917
23gpt-5-mini0.911
24grok-3-mini0.908
25qwen3-vl-235b-a22b-thinking0.897

AIME 2025

1 phaseActive

Competition math benchmark — 30 problems from AIME I & II 2025, integer answers 0–999. Tests frontier model mathematical reasoning. Metric: % correct (pass@1).

Overview

AIME 2025

Category Metric Tasks Saturation Administered

Problems

Quick answer: AIME 2025 is an AI evaluation benchmark built from the 30 problems of the 2025 American Invitational Mathematics Examination (15 problems from AIME I, administered February 2025, and 15 from AIME II). It tests frontier model mathematical reasoning at competition level — problems require algebra, combinatorics, number theory, and geometry, with integer answers in the range 000–999. With 114+ models evaluated, it is one of the most widely tracked math benchmarks for frontier AI.

At a Glance

What it tests: Multi-step mathematical reasoning across algebra, combinatorics, number theory, geometry, and probability at the high-school olympiad level. Problems require deriving a specific integer answer (000–999), preventing partial credit or lucky guessing.

Why it matters: AIME 2025 became a critical benchmark for tracking reasoning model progress through 2025–2026, with o3, Gemini 3 Pro, and GPT-5 families all scoring above 90%. It sits above AMC but below USAMO, making it ideal for discriminating among top-tier models.

Known limitations: 30 problems is a small evaluation set, making scores high-variance (each problem is worth ~3.3%). Models trained after February 2025 carry contamination risk.

What AIME 2025 Measures

The American Invitational Mathematics Examination is an annual U.S. high-school mathematics competition administered by the Mathematical Association of America (MAA). For AI evaluation, the standard protocol is to run each model on all 30 problems (both AIME I and AIME II) and report the pass@1 accuracy — the percentage of problems where the model's first attempt produces the correct integer answer.

Problem categories broadly cover: algebra and polynomials, combinatorics and counting, number theory and modular arithmetic, Euclidean and coordinate geometry, and probability.

Benchmark Specifications

FieldValue
Problems30 (15 AIME I + 15 AIME II)
Answer formatInteger 000–999
Primary metricPass@1 accuracy (% correct)
Year2025
Administered byMathematical Association of America
SaturationLow — frontier models approach 90%+

Frequently Asked Questions

What score does a top AI model get on AIME 2025? As of mid-2026, leading reasoning models achieve 80–95% on AIME 2025. o3, Gemini 3 Pro, and GPT-5 families score at the high end.

Is AIME 2025 contaminated? Models with training cutoffs after February 2025 may have seen the problems. Pass@1 scores should be interpreted cautiously for post-cutoff models; some labs report separate held-out results.

How does AIME 2025 compare to AIME 2024? The problem sets are independently generated each year. AIME 2025 is generally considered comparable in difficulty to 2024, though difficulty can vary by year and problem category.