| Rank | Model | Score |
|---|---|---|
| 1 | gpt-5-2-pro-2025-12-11 | 1 |
| 2 | gpt-5-2 | 0.994 |
| 3 | deepseek-v3-2-speciale | 0.992 |
| 4 | kimi-k2-thinking-0905 | 0.975 |
| 5 | qwen3-6-plus | 0.967 |
| 6 | kimi-k2-5 | 0.954 |
| 7 | qwen3-5-397b-a17b | 0.948 |
| 8 | nemotron-3-super-120b-a12b | 0.947 |
| 9 | glm-5-2 | 0.944 |
| 10 | glm-5-1 | 0.94 |
| 11 | qwen3-6-27b | 0.938 |
| 12 | gpt-5 | 0.933 |
| 13 | grok-4-20 | 0.933 |
| 14 | qwen3-5-27b | 0.92 |
| 15 | qwen3-5-122b-a10b | 0.914 |
| 16 | qwen3-6-35b-a3b | 0.907 |
| 17 | deepseek-v3-2-thinking | 0.902 |
| 18 | deepseek-v3-2 | 0.902 |
| 19 | qwen3-5-35b-a3b | 0.89 |
| 20 | gpt-5-mini | 0.878 |
| 21 | sarvam-105b | 0.858 |
| 22 | mimo-v2-flash | 0.844 |
| 23 | deepseek-v3-2-exp | 0.836 |
| 24 | qwen3-5-9b | 0.832 |
| 25 | deepseek-r1-0528 | 0.794 |
1 phaseActive
AI evaluation on Harvard-MIT Mathematics Tournament 2025 problems — competition-grade math covering algebra, geometry, combinatorics, and guts rounds. Metric: accuracy.
Quick answer: HMMT 2025 measures AI performance on the Harvard-MIT Mathematics Tournament — one of the most prestigious high-school math competitions in the United States. Problems span algebra, geometry, combinatorics, number theory, and multi-round team formats. GPT-5.2 Pro achieves a perfect 100% score as of August 2026.
The Harvard-MIT Mathematics Tournament (HMMT) is a student-organized competition featuring two annual tournaments: one at MIT in November 2025 and one at Harvard in February 2026. The benchmark evaluates AI models on actual HMMT problem sets, which include:
| Round | Format |
|---|---|
| Individual rounds | Subject-specific tests (Algebra, Geometry, Combinatorics, etc.) |
| Team round | Collaborative multi-step problems |
| Guts round | Rapid-fire sequential problems with partial scoring |
Problems are significantly harder than standard olympiad benchmarks like AMC/AIME, making HMMT one of the most challenging mathematics evaluations for AI models.
Models are evaluated on accuracy across the problem set. Scores range from 0 to 1 (0–100%). Given the extreme difficulty of competition problems, even top models scored below 80% on earlier math benchmarks — HMMT 2025 provides meaningful separation at the frontier.
| Benchmark | Difficulty | Focus |
|---|---|---|
| GSM8K | Easy | Grade-school arithmetic |
| MATH-500 | Medium | High-school math |
| AIME 2025 | Hard | AMC competition (30 problems) |
| HMMT 2025 | Very Hard | Elite team competition |
| FrontierMath | Extreme | Research-level mathematics |
| Property | Value |
|---|---|
| Tournament | Harvard-MIT Mathematics Tournament 2025 |
| Metric | Accuracy |
| Score range | 0–1 |
| Top model | GPT-5.2 Pro (1.000) |
| Models evaluated | 33 |
What is HMMT 2025? HMMT 2025 is a benchmark that evaluates AI models on problems from the Harvard-MIT Mathematics Tournament, a prestigious student-organized math competition with two events in 2025–2026.
How hard is HMMT 2025? HMMT is significantly harder than AIME and closer in difficulty to Putnam-level competition math. Top high-school teams solve only a fraction of problems correctly, making it a demanding frontier for AI evaluation.
Who runs HMMT? HMMT is organized by Harvard and MIT students. It has no affiliation with a specific AI research paper; the benchmark uses published competition problems.
What score does the best model achieve on HMMT 2025? GPT-5.2 Pro achieves a perfect score of 1.000 (100%), with GPT-5.2 close behind at 0.994. DeepSeek-V3.2-Speciale ranks third at 0.992.