Benchgen

HMMT 2025 — Results

RankModelScore
1gpt-5-2-pro-2025-12-111
2gpt-5-20.994
3deepseek-v3-2-speciale0.992
4kimi-k2-thinking-09050.975
5qwen3-6-plus0.967
6kimi-k2-50.954
7qwen3-5-397b-a17b0.948
8nemotron-3-super-120b-a12b0.947
9glm-5-20.944
10glm-5-10.94
11qwen3-6-27b0.938
12gpt-50.933
13grok-4-200.933
14qwen3-5-27b0.92
15qwen3-5-122b-a10b0.914
16qwen3-6-35b-a3b0.907
17deepseek-v3-2-thinking0.902
18deepseek-v3-20.902
19qwen3-5-35b-a3b0.89
20gpt-5-mini0.878
21sarvam-105b0.858
22mimo-v2-flash0.844
23deepseek-v3-2-exp0.836
24qwen3-5-9b0.832
25deepseek-r1-05280.794
H

HMMT 2025

1 phaseActive

AI evaluation on Harvard-MIT Mathematics Tournament 2025 problems — competition-grade math covering algebra, geometry, combinatorics, and guts rounds. Metric: accuracy.

Overview

HMMT 2025

Category Metric Saturation Level

Quick answer: HMMT 2025 measures AI performance on the Harvard-MIT Mathematics Tournament — one of the most prestigious high-school math competitions in the United States. Problems span algebra, geometry, combinatorics, number theory, and multi-round team formats. GPT-5.2 Pro achieves a perfect 100% score as of August 2026.


What Does HMMT 2025 Test?

The Harvard-MIT Mathematics Tournament (HMMT) is a student-organized competition featuring two annual tournaments: one at MIT in November 2025 and one at Harvard in February 2026. The benchmark evaluates AI models on actual HMMT problem sets, which include:

RoundFormat
Individual roundsSubject-specific tests (Algebra, Geometry, Combinatorics, etc.)
Team roundCollaborative multi-step problems
Guts roundRapid-fire sequential problems with partial scoring

Problems are significantly harder than standard olympiad benchmarks like AMC/AIME, making HMMT one of the most challenging mathematics evaluations for AI models.


How Is HMMT 2025 Scored?

Models are evaluated on accuracy across the problem set. Scores range from 0 to 1 (0–100%). Given the extreme difficulty of competition problems, even top models scored below 80% on earlier math benchmarks — HMMT 2025 provides meaningful separation at the frontier.


HMMT 2025 vs. Other Math Benchmarks

BenchmarkDifficultyFocus
GSM8KEasyGrade-school arithmetic
MATH-500MediumHigh-school math
AIME 2025HardAMC competition (30 problems)
HMMT 2025Very HardElite team competition
FrontierMathExtremeResearch-level mathematics

Key Facts

PropertyValue
TournamentHarvard-MIT Mathematics Tournament 2025
MetricAccuracy
Score range0–1
Top modelGPT-5.2 Pro (1.000)
Models evaluated33

FAQ

What is HMMT 2025? HMMT 2025 is a benchmark that evaluates AI models on problems from the Harvard-MIT Mathematics Tournament, a prestigious student-organized math competition with two events in 2025–2026.

How hard is HMMT 2025? HMMT is significantly harder than AIME and closer in difficulty to Putnam-level competition math. Top high-school teams solve only a fraction of problems correctly, making it a demanding frontier for AI evaluation.

Who runs HMMT? HMMT is organized by Harvard and MIT students. It has no affiliation with a specific AI research paper; the benchmark uses published competition problems.

What score does the best model achieve on HMMT 2025? GPT-5.2 Pro achieves a perfect score of 1.000 (100%), with GPT-5.2 close behind at 0.994. DeepSeek-V3.2-Speciale ranks third at 0.992.