| Rank | Model | Score |
|---|---|---|
| 1 | hy4-preview | 66.6 |
1 phaseActive
MathArena's rolling, contamination-resistant track testing final-answer math questions freshly derived from new arXiv papers. Metric: automatically graded accuracy.
Quick answer: ArXivMath is a track within ETH SRI's MathArena platform that derives final-answer mathematical questions from newly published arXiv papers on a rolling monthly basis, curated to be nontrivial and robustly gradable — designed specifically to resist training-data contamination.
What it tests: Final-answer mathematical reasoning on genuinely novel problems, sourced from mathematics papers published after any given model's training cutoff.
Why it matters: Static math benchmarks eventually leak into training data; ArXivMath's rolling, freshly-sourced design gives a much cleaner read on real reasoning ability rather than memorization.
Known limitations: As a rolling monthly collection, the exact task set (and difficulty) changes over time, making strict month-to-month score comparison imprecise; no fixed aggregate task count is publicly documented.
ArXivMath is part of MathArena, a continuously updated platform built by researchers at ETH Zurich's SRI lab (and INSAIT) specifically to combat contamination in math benchmarks. Rather than drawing from a fixed, static problem set, ArXivMath curates final-answer questions from mathematics papers newly posted to arXiv each month — questions are selected and adapted to be both nontrivial (not solvable by pattern-matching) and automatically gradable (verifiable final answers).
A high ArXivMath score is a stronger signal of genuine mathematical reasoning than scores on older, static benchmarks, precisely because the questions can't have been seen during pretraining. MathArena reports scores as an average over multiple independent runs (typically four) per problem to reduce variance from sampling.
| Field | Value |
|---|---|
| Task category | Math |
| Metric | Automatically graded final-answer accuracy (avg. of 4 runs) |
| Number of tasks | Rolling monthly collection; no fixed aggregate count published |
| Saturation | Low |
| Created by | Dekoninck, Jovanović, Gehrunger, Rögnvaldsson, Petrov, Sun, Vechev (MathArena / ETH SRI) |
| Source paper | Dekoninck et al. 2026 |
| GitHub | eth-sri/matharena |
| Leaderboard | matharena.ai |
Each question has an automatically-checkable final answer. Models are typically run four times per problem, and the reported score is the average accuracy across those runs — reducing the chance a single lucky (or unlucky) sample distorts the result.
| Rank | Model | Score | Source | Date |
|---|---|---|---|---|
| 1 | Hy4 Preview | 66.6% | Tencent Hunyuan model card | 2026-08 |
Scores sourced from published technical reports and model cards. Results depend on harness, prompt format, and effort settings — see each source for methodology.
No Benchgen results yet — be the first to run ArXivMath.
| Benchmark | What it tests | Tasks | Saturation |
|---|---|---|---|
| ArXivMath | Fresh, uncontaminated final-answer math from new arXiv papers | Rolling | Low |
| BrokenArXiv | Identifying flawed reasoning/proofs in new arXiv papers | Rolling | Low |
| MathArena Apex 2025 | Competition mathematics (Apex 2025 set) | — | Low |
| HMMT 2026 | Static competition math | — | Medium |
Use ArXivMath specifically when contamination resistance matters more than a stable, fixed task set — for a stable year-over-year comparison, a static competition benchmark like HMMT is more appropriate.
Benchgen lets teams run ArXivMath against their own model versions, compare results across runs, and catch regressions in fresh, uncontaminated math reasoning — rather than relying on a single vendor-reported number.