Benchgen

AIME 2024 — Results

RankModelScore
1grok-3-mini0.958
2o4-mini0.934
3gemini-2-5-pro0.92
4o30.916
5deepseek-r1-05280.914
6glm-4-50.91
7glm-4-5-air0.894
8gemini-2-5-flash0.88
9o3-mini0.873
10qwen3-235b-a22b0.857
11qwen3-32b0.814
12phi-4-reasoning-plus0.813
13qwen3-30b-a3b0.804
14qwq-32b0.795
15kimi-k1-50.775
16phi-4-reasoning0.753
17o10.743
18kimi-k2-09050.72
19kimi-k2-instruct-09050.696
20kimi-k2-instruct0.696
21deepseek-v3-10.663
22deepseek-v3-03240.594
23gpt-4-1-mini0.496
24gpt-4-10.481
25deepseek-v30.392

AIME 2024

1 phaseActive

Competition math benchmark — 30 problems from AIME I & II 2024, integer answers 0–999. Tests frontier model mathematical reasoning. Metric: % correct (pass@1).

Overview

AIME 2024

Category Metric Tasks Saturation Administered

Problems

Quick answer: AIME 2024 is an AI evaluation benchmark built from the 30 problems of the 2024 American Invitational Mathematics Examination (15 from AIME I and 15 from AIME II). It became a key early benchmark for tracking the emergence of strong mathematical reasoning in large language models, with 53+ models evaluated. Problems require algebra, combinatorics, number theory, and geometry, with integer answers in the range 000–999.

At a Glance

What it tests: Multi-step mathematical reasoning across algebra, combinatorics, number theory, geometry, and probability. Each problem requires deriving a specific integer answer (000–999).

Why it matters: AIME 2024 was one of the first widely used math competition benchmarks to show a clear separation between standard LLMs (10–30% pass@1) and reasoning-capable models (50–90%+). It established the baseline for comparing frontier math reasoning through 2024–2025.

Known limitations: 30 problems is a small evaluation set. Models trained after February 2024 may have seen the problem set, particularly as solutions were widely published.

What AIME 2024 Measures

The American Invitational Mathematics Examination is an annual U.S. high-school mathematics competition. For AI evaluation, models are run on all 30 problems and scored on pass@1 accuracy. Problem categories include algebra, combinatorics, number theory, Euclidean geometry, and probability.

Benchmark Specifications

FieldValue
Problems30 (15 AIME I + 15 AIME II)
Answer formatInteger 000–999
Primary metricPass@1 accuracy (% correct)
Year2024
Administered byMathematical Association of America
SaturationMedium — frontier models reach 80–90%

Frequently Asked Questions

What score does a top AI model get on AIME 2024? Leading reasoning models as of 2025–2026 achieve 70–90%+ on AIME 2024. GPT-4o and Claude 3.5 Sonnet scored in the 20–40% range; o1, o3, and subsequent reasoning models pushed well above 70%.

How does AIME 2024 compare to later AIME years? AIME 2024 was among the first AIME years to be extensively evaluated for AI. Later editions (2025, 2026) are generally treated as less contaminated given the later training cutoffs of newer models.