Benchgen

HMMT 2025 — Results

RankModelScore
1granite-4-2-30b89.17
2granite-4-2-8b78.33
3granite-4-2-3b66.67
4gpt-5-2-pro-2025-12-111
5gpt-5-20.994
6deepseek-v3-2-speciale0.992
7kimi-k2-thinking-09050.975
8qwen3-6-plus0.967
9kimi-k2-50.954
10qwen3-5-397b-a17b0.948
11nemotron-3-super-120b-a12b0.947
12glm-5-20.944
13glm-5-10.94
14qwen3-6-27b0.938
15gpt-50.933
16grok-4-200.933
17qwen3-5-27b0.92
18qwen3-5-122b-a10b0.914
19qwen3-6-35b-a3b0.907
20deepseek-v3-2-thinking0.902
21deepseek-v3-20.902
22qwen3-5-35b-a3b0.89
23gpt-5-mini0.878
24sarvam-105b0.858
25mimo-v2-flash0.844
H

HMMT 2025

1 phaseActive

AI evaluation on Harvard-MIT Mathematics Tournament 2025 problems — competition-grade math covering algebra, geometry, combinatorics, and guts rounds. Metric: accuracy.

Overview

HMMT 2025

Category Metric Saturation Level

Quick answer: HMMT 2025 measures AI performance on the Harvard-MIT Mathematics Tournament — one of the most prestigious high-school math competitions in the United States. Problems span algebra, geometry, combinatorics, number theory, and multi-round team formats. GPT-5.2 Pro achieves a perfect 100% score as of August 2026.


What Does HMMT 2025 Test?

The Harvard-MIT Mathematics Tournament (HMMT) is a student-organized competition featuring two annual tournaments: one at MIT in November 2025 and one at Harvard in February 2026. The benchmark evaluates AI models on actual HMMT problem sets, which include:

RoundFormat
Individual roundsSubject-specific tests (Algebra, Geometry, Combinatorics, etc.)
Team roundCollaborative multi-step problems
Guts roundRapid-fire sequential problems with partial scoring

Problems are significantly harder than standard olympiad benchmarks like AMC/AIME, making HMMT one of the most challenging mathematics evaluations for AI models.


How Is HMMT 2025 Scored?

Models are evaluated on accuracy across the problem set. Scores range from 0 to 1 (0–100%). Given the extreme difficulty of competition problems, even top models scored below 80% on earlier math benchmarks — HMMT 2025 provides meaningful separation at the frontier.


HMMT 2025 vs. Other Math Benchmarks

BenchmarkDifficultyFocus
GSM8KEasyGrade-school arithmetic
MATH-500MediumHigh-school math
AIME 2025HardAMC competition (30 problems)
HMMT 2025Very HardElite team competition
FrontierMathExtremeResearch-level mathematics

Key Facts

PropertyValue
TournamentHarvard-MIT Mathematics Tournament 2025
MetricAccuracy
Score range0–1
Top modelGPT-5.2 Pro (1.000)
Models evaluated33

FAQ

What is HMMT 2025? HMMT 2025 is a benchmark that evaluates AI models on problems from the Harvard-MIT Mathematics Tournament, a prestigious student-organized math competition with two events in 2025–2026.

How hard is HMMT 2025? HMMT is significantly harder than AIME and closer in difficulty to Putnam-level competition math. Top high-school teams solve only a fraction of problems correctly, making it a demanding frontier for AI evaluation.

Who runs HMMT? HMMT is organized by Harvard and MIT students. It has no affiliation with a specific AI research paper; the benchmark uses published competition problems.

What score does the best model achieve on HMMT 2025? GPT-5.2 Pro achieves a perfect score of 1.000 (100%), with GPT-5.2 close behind at 0.994. DeepSeek-V3.2-Speciale ranks third at 0.992.