| Rank | Model | Score |
|---|---|---|
| 1 | maple-preview | 78.8 |
1 phaseActive
AI evaluation on Harvard-MIT Mathematics Tournament 2026 problems — the newest elite competition math set, covering algebra, geometry, combinatorics, and guts rounds.
Quick answer: HMMT 2026 evaluates AI models on the newest problem set from the Harvard-MIT Mathematics Tournament, one of the most prestigious student-organized math competitions in the United States. As the freshest HMMT cycle, it carries the lowest contamination risk of the series, making it a sharper stress test of genuine mathematical reasoning than older HMMT years.
What it tests: Accuracy on real Harvard-MIT Mathematics Tournament 2026 problems, spanning individual subject rounds (algebra, geometry, combinatorics, number theory), a collaborative team round, and a rapid-fire guts round.
Why it matters: HMMT problems are consistently harder than AMC/AIME-tier competition math, and each year's fresh problem set gives a contamination-resistant read on frontier models' genuine mathematical reasoning rather than memorized training data.
Known limitations: As the most recent HMMT cycle, fewer models have been evaluated against it than against earlier years (HMMT 2025), and scores may shift as more labs report results.
HMMT 2026 uses the actual problem sets from the 2026 Harvard-MIT Mathematics Tournament cycle, evaluating models the same way human competitors are scored: across individual subject-specific rounds, a collaborative team round, and a rapid-fire guts round with partial scoring. Because HMMT problems are written fresh each year by student organizers and are significantly harder than standard olympiad-style benchmarks like AMC or AIME, the tournament provides meaningful separation even among frontier reasoning models.
Using the newest available HMMT year rather than an older one reduces the risk that a model's training data included the problems (and their solutions) ahead of evaluation — an increasingly important consideration as competition math benchmarks saturate and models are trained on ever-larger web crawls that may include past tournament archives.
| Field | Value |
|---|---|
| Task category | Math |
| Metric | Accuracy (0–1, i.e. 0–100%) |
| Tournament | Harvard-MIT Mathematics Tournament 2026 |
| Saturation | Low |
| Created by | Harvard-MIT Mathematics Tournament (student-organized) |
Models are evaluated on accuracy across the HMMT 2026 problem set, scored from 0 to 1 (equivalently, 0–100%). Given the extreme difficulty of HMMT problems relative to standard competition math benchmarks, scores in the 70–90% range represent frontier-tier performance, with only the strongest reasoning models approaching saturation.
No Benchgen results yet — be the first to run HMMT 2026.
| Benchmark | Difficulty | Focus |
|---|---|---|
| AIME 2026 | Hard | AMC-tier competition math (30 problems) |
| HMMT 2026 | Very Hard | Elite team competition, newest cycle |
| HMMT 2025 | Very Hard | Elite team competition, prior cycle |
HMMT 2026 is the right choice when contamination-resistance matters most — it's the newest HMMT cycle available. HMMT 2025 remains useful for tracking a model's performance trend across tournament years, and AIME 2026 is a slightly less extreme competition-math benchmark for broader comparison.
What is HMMT 2026? HMMT 2026 is a benchmark evaluating AI models on problems from the 2026 Harvard-MIT Mathematics Tournament, a prestigious student-organized math competition.
How hard is HMMT 2026? HMMT is significantly harder than AIME-tier competition math; only the strongest frontier reasoning models score above 70-80% on a given HMMT year's problem set.
Who created HMMT 2026? HMMT is organized by Harvard and MIT students; it isn't tied to a specific AI research paper, but uses freshly published competition problems each year.
Is HMMT 2026 saturated? No — as the newest tournament cycle, it has the lowest contamination risk of any HMMT year and meaningful headroom remains for most models.
How does HMMT 2026 differ from HMMT 2025? Both test the same tournament format and difficulty level; HMMT 2026 uses the newest problem set, giving a fresher, less contamination-prone read on genuine reasoning ability.
Benchmark definition based on the Harvard-MIT Mathematics Tournament 2026 problem sets. Last updated 2026-08-05.