| Rank | Model | Score |
|---|---|---|
| 1 | hy4-preview | 74.2 |
1 phaseActive
MathArena's Apex 2025 competition-math configuration — final-answer scoring on ETH SRI's continuously updated, anti-contamination math platform.
Quick answer: MathArena Apex 2025 is a competition-mathematics configuration within ETH SRI's MathArena platform — final-answer problems from the "Apex" 2025 competition set, evaluated using MathArena's standard anti-contamination protocol of averaging results over four independent runs.
What it tests: Final-answer competition mathematics, drawn from a curated 2025 competition ("Apex") set within the broader MathArena platform.
Why it matters: MathArena's whole premise is evaluating models on math competitions that are unlikely to already be in training data, giving a cleaner read on reasoning ability than older, widely-leaked competition benchmarks.
Known limitations: No authoritative public item count for this specific configuration has been published; treat it as one dated configuration within MathArena's continuously evolving set, not a fixed standalone benchmark.
MathArena is a continuously updated platform built by researchers at ETH Zurich's SRI lab and INSAIT to evaluate LLMs on math competitions that resist training-data contamination. The "Apex 2025" configuration is one specific competition set tracked within that platform — final-answer competition mathematics questions, graded automatically and averaged over repeated runs (MathArena's standard protocol uses four runs per problem) to reduce single-sample variance.
A high score indicates strong competition-level mathematical problem-solving on a set specifically chosen to minimize the chance the model has memorized the answers from pretraining data.
| Field | Value |
|---|---|
| Task category | Math |
| Metric | Final-answer accuracy (avg. of 4 runs) |
| Number of tasks | Not publicly documented as a fixed count for this configuration |
| Saturation | Medium |
| Created by | MathArena / ETH SRI |
| Source paper | MathArena team 2025/2026 |
| GitHub | eth-sri/matharena |
| Leaderboard | matharena.ai |
Each problem has an automatically-checkable final answer. MathArena's standard protocol runs each model four times per problem and reports the average accuracy, reducing variance from any single sampling run.
| Rank | Model | Score | Source | Date |
|---|---|---|---|---|
| 1 | Hy4 Preview | 74.2% | Tencent Hunyuan model card | 2026-08 |
Scores sourced from published technical reports and model cards. Results depend on harness, prompt format, and effort settings — see each source for methodology.
No Benchgen results yet — be the first to run MathArena Apex 2025.
| Benchmark | What it tests | Tasks | Saturation |
|---|---|---|---|
| MathArena Apex 2025 | Apex 2025 competition mathematics | — | Medium |
| ArXivMath | Fresh math from new arXiv papers | Rolling | Low |
| HMMT 2026 | Static competition math | — | Medium |
| AIME 2026 | Static competition math (AIME) | — | High |
Use MathArena Apex 2025 alongside ArXivMath for a fuller picture of MathArena's anti-contamination approach to math evaluation — competition-style problems here, freshly-sourced paper-derived problems there.
Benchgen lets teams run MathArena Apex 2025 against their own model versions, compare results across runs, and catch regressions in competition-math reasoning — rather than relying on a single vendor-reported number.