Benchgen

ArXivMath — Results

RankModelScore
1hy4-preview66.6
A

ArXivMath

1 phaseActive

MathArena's rolling, contamination-resistant track testing final-answer math questions freshly derived from new arXiv papers. Metric: automatically graded accuracy.

Overview

ArXivMath

Category Metric Saturation Created

Paper GitHub Leaderboard

Quick answer: ArXivMath is a track within ETH SRI's MathArena platform that derives final-answer mathematical questions from newly published arXiv papers on a rolling monthly basis, curated to be nontrivial and robustly gradable — designed specifically to resist training-data contamination.

At a Glance

What it tests: Final-answer mathematical reasoning on genuinely novel problems, sourced from mathematics papers published after any given model's training cutoff.

Why it matters: Static math benchmarks eventually leak into training data; ArXivMath's rolling, freshly-sourced design gives a much cleaner read on real reasoning ability rather than memorization.

Known limitations: As a rolling monthly collection, the exact task set (and difficulty) changes over time, making strict month-to-month score comparison imprecise; no fixed aggregate task count is publicly documented.

What ArXivMath Measures

ArXivMath is part of MathArena, a continuously updated platform built by researchers at ETH Zurich's SRI lab (and INSAIT) specifically to combat contamination in math benchmarks. Rather than drawing from a fixed, static problem set, ArXivMath curates final-answer questions from mathematics papers newly posted to arXiv each month — questions are selected and adapted to be both nontrivial (not solvable by pattern-matching) and automatically gradable (verifiable final answers).

A high ArXivMath score is a stronger signal of genuine mathematical reasoning than scores on older, static benchmarks, precisely because the questions can't have been seen during pretraining. MathArena reports scores as an average over multiple independent runs (typically four) per problem to reduce variance from sampling.

Benchmark Specifications

FieldValue
Task categoryMath
MetricAutomatically graded final-answer accuracy (avg. of 4 runs)
Number of tasksRolling monthly collection; no fixed aggregate count published
SaturationLow
Created byDekoninck, Jovanović, Gehrunger, Rögnvaldsson, Petrov, Sun, Vechev (MathArena / ETH SRI)
Source paperDekoninck et al. 2026
GitHubeth-sri/matharena
Leaderboardmatharena.ai

How ArXivMath Is Scored

Each question has an automatically-checkable final answer. Models are typically run four times per problem, and the reported score is the average accuracy across those runs — reducing the chance a single lucky (or unlucky) sample distorts the result.

State-of-the-Art Results

RankModelScoreSourceDate
1Hy4 Preview66.6%Tencent Hunyuan model card2026-08

Scores sourced from published technical reports and model cards. Results depend on harness, prompt format, and effort settings — see each source for methodology.

ArXivMath on Benchgen

No Benchgen results yet — be the first to run ArXivMath.

ArXivMath vs Other Benchmarks

BenchmarkWhat it testsTasksSaturation
ArXivMathFresh, uncontaminated final-answer math from new arXiv papersRollingLow
BrokenArXivIdentifying flawed reasoning/proofs in new arXiv papersRollingLow
MathArena Apex 2025Competition mathematics (Apex 2025 set)Low
HMMT 2026Static competition mathMedium

Use ArXivMath specifically when contamination resistance matters more than a stable, fixed task set — for a stable year-over-year comparison, a static competition benchmark like HMMT is more appropriate.

Run ArXivMath on Your Model

Benchgen lets teams run ArXivMath against their own model versions, compare results across runs, and catch regressions in fresh, uncontaminated math reasoning — rather than relying on a single vendor-reported number.

Frequently Asked Questions

What is ArXivMath? ArXivMath is a MathArena track that derives final-answer math questions from newly published arXiv papers each month, designed to resist training-data contamination.
What does a good ArXivMath score look like? As of August 2026, frontier models score in the 43–80% range; scores above 66% represent strong performance on genuinely novel math reasoning.
Who created ArXivMath? ArXivMath is part of MathArena, created by Dekoninck, Jovanović, Gehrunger, Rögnvaldsson, Petrov, Sun, and Vechev at ETH Zurich's SRI lab and INSAIT.