| Rank | Model | Score |
|---|---|---|
| 1 | lfm2-5-2-6b | 51.87 |
| 2 | gemini-3-pro | 1 |
| 3 | gpt-5-2-pro-2025-12-11 | 1 |
| 4 | gpt-5-2 | 1 |
| 5 | kimi-k2-thinking-0905 | 1 |
| 6 | claude-opus-4-6 | 0.998 |
| 7 | gemini-3-flash | 0.997 |
| 8 | seed-2-0-pro | 0.983 |
| 9 | kimi-k2-5 | 0.961 |
| 10 | deepseek-v3-2-speciale | 0.96 |
| 11 | glm-4-7 | 0.957 |
| 12 | gpt-5 | 0.946 |
| 13 | gpt-5-1-instant | 0.94 |
| 14 | gpt-5-1-thinking | 0.94 |
| 15 | gpt-5-1 | 0.94 |
| 16 | deepseek-v3-2-thinking | 0.931 |
| 17 | deepseek-v3-2 | 0.931 |
| 18 | seed-2-0-lite | 0.93 |
| 19 | o4-mini | 0.927 |
| 20 | gpt-oss-120b-high | 0.925 |
| 21 | qwen3-235b-a22b-thinking-2507 | 0.923 |
| 22 | grok-4 | 0.917 |
| 23 | gpt-5-mini | 0.911 |
| 24 | grok-3-mini | 0.908 |
| 25 | qwen3-vl-235b-a22b-thinking | 0.897 |
1 phaseActive
Competition math benchmark — 30 problems from AIME I & II 2025, integer answers 0–999. Tests frontier model mathematical reasoning. Metric: % correct (pass@1).
Quick answer: AIME 2025 is an AI evaluation benchmark built from the 30 problems of the 2025 American Invitational Mathematics Examination (15 problems from AIME I, administered February 2025, and 15 from AIME II). It tests frontier model mathematical reasoning at competition level — problems require algebra, combinatorics, number theory, and geometry, with integer answers in the range 000–999. With 114+ models evaluated, it is one of the most widely tracked math benchmarks for frontier AI.
What it tests: Multi-step mathematical reasoning across algebra, combinatorics, number theory, geometry, and probability at the high-school olympiad level. Problems require deriving a specific integer answer (000–999), preventing partial credit or lucky guessing.
Why it matters: AIME 2025 became a critical benchmark for tracking reasoning model progress through 2025–2026, with o3, Gemini 3 Pro, and GPT-5 families all scoring above 90%. It sits above AMC but below USAMO, making it ideal for discriminating among top-tier models.
Known limitations: 30 problems is a small evaluation set, making scores high-variance (each problem is worth ~3.3%). Models trained after February 2025 carry contamination risk.
The American Invitational Mathematics Examination is an annual U.S. high-school mathematics competition administered by the Mathematical Association of America (MAA). For AI evaluation, the standard protocol is to run each model on all 30 problems (both AIME I and AIME II) and report the pass@1 accuracy — the percentage of problems where the model's first attempt produces the correct integer answer.
Problem categories broadly cover: algebra and polynomials, combinatorics and counting, number theory and modular arithmetic, Euclidean and coordinate geometry, and probability.
| Field | Value |
|---|---|
| Problems | 30 (15 AIME I + 15 AIME II) |
| Answer format | Integer 000–999 |
| Primary metric | Pass@1 accuracy (% correct) |
| Year | 2025 |
| Administered by | Mathematical Association of America |
| Saturation | Low — frontier models approach 90%+ |
What score does a top AI model get on AIME 2025? As of mid-2026, leading reasoning models achieve 80–95% on AIME 2025. o3, Gemini 3 Pro, and GPT-5 families score at the high end.
Is AIME 2025 contaminated? Models with training cutoffs after February 2025 may have seen the problems. Pass@1 scores should be interpreted cautiously for post-cutoff models; some labs report separate held-out results.
How does AIME 2025 compare to AIME 2024? The problem sets are independently generated each year. AIME 2025 is generally considered comparable in difficulty to 2024, though difficulty can vary by year and problem category.