Benchgen

MathArena Apex 2025 — Results

RankModelScore
1hy4-preview74.2
M

MathArena Apex 2025

1 phaseActive

MathArena's Apex 2025 competition-math configuration — final-answer scoring on ETH SRI's continuously updated, anti-contamination math platform.

Overview

MathArena Apex 2025

Category Metric Saturation Created

Paper GitHub Leaderboard

Quick answer: MathArena Apex 2025 is a competition-mathematics configuration within ETH SRI's MathArena platform — final-answer problems from the "Apex" 2025 competition set, evaluated using MathArena's standard anti-contamination protocol of averaging results over four independent runs.

At a Glance

What it tests: Final-answer competition mathematics, drawn from a curated 2025 competition ("Apex") set within the broader MathArena platform.

Why it matters: MathArena's whole premise is evaluating models on math competitions that are unlikely to already be in training data, giving a cleaner read on reasoning ability than older, widely-leaked competition benchmarks.

Known limitations: No authoritative public item count for this specific configuration has been published; treat it as one dated configuration within MathArena's continuously evolving set, not a fixed standalone benchmark.

What MathArena Apex 2025 Measures

MathArena is a continuously updated platform built by researchers at ETH Zurich's SRI lab and INSAIT to evaluate LLMs on math competitions that resist training-data contamination. The "Apex 2025" configuration is one specific competition set tracked within that platform — final-answer competition mathematics questions, graded automatically and averaged over repeated runs (MathArena's standard protocol uses four runs per problem) to reduce single-sample variance.

A high score indicates strong competition-level mathematical problem-solving on a set specifically chosen to minimize the chance the model has memorized the answers from pretraining data.

Benchmark Specifications

FieldValue
Task categoryMath
MetricFinal-answer accuracy (avg. of 4 runs)
Number of tasksNot publicly documented as a fixed count for this configuration
SaturationMedium
Created byMathArena / ETH SRI
Source paperMathArena team 2025/2026
GitHubeth-sri/matharena
Leaderboardmatharena.ai

How MathArena Apex 2025 Is Scored

Each problem has an automatically-checkable final answer. MathArena's standard protocol runs each model four times per problem and reports the average accuracy, reducing variance from any single sampling run.

State-of-the-Art Results

RankModelScoreSourceDate
1Hy4 Preview74.2%Tencent Hunyuan model card2026-08

Scores sourced from published technical reports and model cards. Results depend on harness, prompt format, and effort settings — see each source for methodology.

MathArena Apex 2025 on Benchgen

No Benchgen results yet — be the first to run MathArena Apex 2025.

MathArena Apex 2025 vs Other Benchmarks

BenchmarkWhat it testsTasksSaturation
MathArena Apex 2025Apex 2025 competition mathematicsMedium
ArXivMathFresh math from new arXiv papersRollingLow
HMMT 2026Static competition mathMedium
AIME 2026Static competition math (AIME)High

Use MathArena Apex 2025 alongside ArXivMath for a fuller picture of MathArena's anti-contamination approach to math evaluation — competition-style problems here, freshly-sourced paper-derived problems there.

Run MathArena Apex 2025 on Your Model

Benchgen lets teams run MathArena Apex 2025 against their own model versions, compare results across runs, and catch regressions in competition-math reasoning — rather than relying on a single vendor-reported number.

Frequently Asked Questions

What is MathArena Apex 2025? MathArena Apex 2025 is a competition-mathematics configuration within ETH SRI's MathArena platform, evaluating final-answer problems from the Apex 2025 competition set.
What does a good MathArena Apex 2025 score look like? As of August 2026, frontier models score in the 66–91% range; scores above 74% represent strong competition-math performance.
Who created MathArena? MathArena was created by researchers at ETH Zurich's SRI lab and INSAIT, including Dekoninck, Jovanović, Gehrunger, Rögnvaldsson, Petrov, Sun, and Vechev.