| Rank | Model | Score |
|---|---|---|
| 1 | grok-3-mini | 0.958 |
| 2 | o4-mini | 0.934 |
| 3 | gemini-2-5-pro | 0.92 |
| 4 | o3 | 0.916 |
| 5 | deepseek-r1-0528 | 0.914 |
| 6 | glm-4-5 | 0.91 |
| 7 | glm-4-5-air | 0.894 |
| 8 | gemini-2-5-flash | 0.88 |
| 9 | o3-mini | 0.873 |
| 10 | qwen3-235b-a22b | 0.857 |
| 11 | qwen3-32b | 0.814 |
| 12 | phi-4-reasoning-plus | 0.813 |
| 13 | qwen3-30b-a3b | 0.804 |
| 14 | qwq-32b | 0.795 |
| 15 | kimi-k1-5 | 0.775 |
| 16 | phi-4-reasoning | 0.753 |
| 17 | o1 | 0.743 |
| 18 | kimi-k2-0905 | 0.72 |
| 19 | kimi-k2-instruct-0905 | 0.696 |
| 20 | kimi-k2-instruct | 0.696 |
| 21 | deepseek-v3-1 | 0.663 |
| 22 | deepseek-v3-0324 | 0.594 |
| 23 | gpt-4-1-mini | 0.496 |
| 24 | gpt-4-1 | 0.481 |
| 25 | deepseek-v3 | 0.392 |
1 phaseActive
Competition math benchmark — 30 problems from AIME I & II 2024, integer answers 0–999. Tests frontier model mathematical reasoning. Metric: % correct (pass@1).
Quick answer: AIME 2024 is an AI evaluation benchmark built from the 30 problems of the 2024 American Invitational Mathematics Examination (15 from AIME I and 15 from AIME II). It became a key early benchmark for tracking the emergence of strong mathematical reasoning in large language models, with 53+ models evaluated. Problems require algebra, combinatorics, number theory, and geometry, with integer answers in the range 000–999.
What it tests: Multi-step mathematical reasoning across algebra, combinatorics, number theory, geometry, and probability. Each problem requires deriving a specific integer answer (000–999).
Why it matters: AIME 2024 was one of the first widely used math competition benchmarks to show a clear separation between standard LLMs (10–30% pass@1) and reasoning-capable models (50–90%+). It established the baseline for comparing frontier math reasoning through 2024–2025.
Known limitations: 30 problems is a small evaluation set. Models trained after February 2024 may have seen the problem set, particularly as solutions were widely published.
The American Invitational Mathematics Examination is an annual U.S. high-school mathematics competition. For AI evaluation, models are run on all 30 problems and scored on pass@1 accuracy. Problem categories include algebra, combinatorics, number theory, Euclidean geometry, and probability.
| Field | Value |
|---|---|
| Problems | 30 (15 AIME I + 15 AIME II) |
| Answer format | Integer 000–999 |
| Primary metric | Pass@1 accuracy (% correct) |
| Year | 2024 |
| Administered by | Mathematical Association of America |
| Saturation | Medium — frontier models reach 80–90% |
What score does a top AI model get on AIME 2024? Leading reasoning models as of 2025–2026 achieve 70–90%+ on AIME 2024. GPT-4o and Claude 3.5 Sonnet scored in the 20–40% range; o1, o3, and subsequent reasoning models pushed well above 70%.
How does AIME 2024 compare to later AIME years? AIME 2024 was among the first AIME years to be extensively evaluated for AI. Later editions (2025, 2026) are generally treated as less contaminated given the later training cutoffs of newer models.