| Rank | Model | Score |
|---|---|---|
| 1 | hy4-preview | 54.6 |
1 phaseActive
MathArena's companion track to ArXivMath — tests whether a model can identify invalid or broken mathematical reasoning in freshly published arXiv papers.
Quick answer: BrokenArXiv is a companion track to MathArena's ArXivMath, built by ETH SRI. Instead of solving a final-answer question, it tests whether a model can correctly identify invalid, flawed, or broken mathematical reasoning and proofs in freshly published arXiv papers, scored by a judge model.
What it tests: Critical mathematical review — can a model spot where a piece of published mathematical reasoning is actually wrong, rather than just producing its own answer?
Why it matters: Verifying a proof or derivation is a distinct (and arguably harder) skill than generating one; this measures a model's ability to critically audit mathematical content, relevant for any agentic use case involving reviewing or fact-checking technical work.
Known limitations: Uses freshly-sourced papers on a rolling basis (no fixed task count published) and relies on judge-based scoring rather than a fully automatic check, so evaluation consistency depends on the judge's reliability.
BrokenArXiv is part of the MathArena platform (the same ETH SRI project behind ArXivMath), built specifically to test critical mathematical reasoning rather than generative problem-solving. Where ArXivMath asks a model to solve a fresh, automatically-gradable final-answer question derived from a new arXiv paper, BrokenArXiv instead presents reasoning or proof steps from freshly published papers — some intentionally containing flaws — and asks the model to correctly identify whether the reasoning is valid or broken.
Because this requires a judgment call rather than a single checkable numeric answer, BrokenArXiv is scored using a judge model rather than fully automatic grading. A high score indicates a model can reliably distinguish sound mathematical reasoning from subtly flawed reasoning — a skill distinct from, and arguably harder than, generating correct proofs from scratch.
| Field | Value |
|---|---|
| Task category | Math |
| Metric | Judge-based scoring / accuracy |
| Number of tasks | Rolling monthly collection; no fixed aggregate count published |
| Saturation | Low |
| Created by | MathArena / ETH SRI |
| Source paper | MathArena team 2026 |
| GitHub | eth-sri/matharena |
| Leaderboard | matharena.ai |
A judge model evaluates whether the model under test correctly identified valid vs. flawed reasoning in each sampled excerpt from a freshly published arXiv math paper. Scores are reported as an accuracy percentage.
| Rank | Model | Score | Source | Date |
|---|---|---|---|---|
| 1 | Hy4 Preview | 54.6% | Tencent Hunyuan model card | 2026-08 |
Scores sourced from published technical reports and model cards. Results depend on harness, prompt format, and effort settings — see each source for methodology.
No Benchgen results yet — be the first to run BrokenArXiv.
| Benchmark | What it tests | Tasks | Saturation |
|---|---|---|---|
| BrokenArXiv | Identifying flawed reasoning/proofs in new arXiv papers | Rolling | Low |
| ArXivMath | Solving fresh final-answer math from new arXiv papers | Rolling | Low |
| MathArena Apex 2025 | Competition mathematics | — | Medium |
Use BrokenArXiv specifically to evaluate a model's critical-review capability on mathematics, distinct from its generative problem-solving ability measured by ArXivMath.
Benchgen lets teams run BrokenArXiv against their own model versions, compare results across runs, and catch regressions in critical mathematical reasoning — rather than relying on a single vendor-reported number.