Benchgen

BrokenArXiv — Results

RankModelScore
1hy4-preview54.6
B

BrokenArXiv

1 phaseActive

MathArena's companion track to ArXivMath — tests whether a model can identify invalid or broken mathematical reasoning in freshly published arXiv papers.

Overview

BrokenArXiv

Category Metric Saturation Created

Paper GitHub Leaderboard

Quick answer: BrokenArXiv is a companion track to MathArena's ArXivMath, built by ETH SRI. Instead of solving a final-answer question, it tests whether a model can correctly identify invalid, flawed, or broken mathematical reasoning and proofs in freshly published arXiv papers, scored by a judge model.

At a Glance

What it tests: Critical mathematical review — can a model spot where a piece of published mathematical reasoning is actually wrong, rather than just producing its own answer?

Why it matters: Verifying a proof or derivation is a distinct (and arguably harder) skill than generating one; this measures a model's ability to critically audit mathematical content, relevant for any agentic use case involving reviewing or fact-checking technical work.

Known limitations: Uses freshly-sourced papers on a rolling basis (no fixed task count published) and relies on judge-based scoring rather than a fully automatic check, so evaluation consistency depends on the judge's reliability.

What BrokenArXiv Measures

BrokenArXiv is part of the MathArena platform (the same ETH SRI project behind ArXivMath), built specifically to test critical mathematical reasoning rather than generative problem-solving. Where ArXivMath asks a model to solve a fresh, automatically-gradable final-answer question derived from a new arXiv paper, BrokenArXiv instead presents reasoning or proof steps from freshly published papers — some intentionally containing flaws — and asks the model to correctly identify whether the reasoning is valid or broken.

Because this requires a judgment call rather than a single checkable numeric answer, BrokenArXiv is scored using a judge model rather than fully automatic grading. A high score indicates a model can reliably distinguish sound mathematical reasoning from subtly flawed reasoning — a skill distinct from, and arguably harder than, generating correct proofs from scratch.

Benchmark Specifications

FieldValue
Task categoryMath
MetricJudge-based scoring / accuracy
Number of tasksRolling monthly collection; no fixed aggregate count published
SaturationLow
Created byMathArena / ETH SRI
Source paperMathArena team 2026
GitHubeth-sri/matharena
Leaderboardmatharena.ai

How BrokenArXiv Is Scored

A judge model evaluates whether the model under test correctly identified valid vs. flawed reasoning in each sampled excerpt from a freshly published arXiv math paper. Scores are reported as an accuracy percentage.

State-of-the-Art Results

RankModelScoreSourceDate
1Hy4 Preview54.6%Tencent Hunyuan model card2026-08

Scores sourced from published technical reports and model cards. Results depend on harness, prompt format, and effort settings — see each source for methodology.

BrokenArXiv on Benchgen

No Benchgen results yet — be the first to run BrokenArXiv.

BrokenArXiv vs Other Benchmarks

BenchmarkWhat it testsTasksSaturation
BrokenArXivIdentifying flawed reasoning/proofs in new arXiv papersRollingLow
ArXivMathSolving fresh final-answer math from new arXiv papersRollingLow
MathArena Apex 2025Competition mathematicsMedium

Use BrokenArXiv specifically to evaluate a model's critical-review capability on mathematics, distinct from its generative problem-solving ability measured by ArXivMath.

Run BrokenArXiv on Your Model

Benchgen lets teams run BrokenArXiv against their own model versions, compare results across runs, and catch regressions in critical mathematical reasoning — rather than relying on a single vendor-reported number.

Frequently Asked Questions

What is BrokenArXiv? BrokenArXiv is a MathArena track that tests whether a model can identify invalid or flawed mathematical reasoning and proofs in freshly published arXiv papers, scored by a judge model.
What does a good BrokenArXiv score look like? As of August 2026, frontier models score in the 42–78% range; scores above 54% represent solid critical mathematical review capability.
Who created BrokenArXiv? BrokenArXiv is part of MathArena, created by researchers at ETH Zurich's SRI lab and INSAIT as a companion track to ArXivMath.