Benchgen

HorizonMath — Results

RankModelScore
1hy4-preview8.8
H

HorizonMath

1 phaseActive

Oxford/Harvard/Princeton benchmark of 136 predominantly unsolved computational and applied-math problems across 8 domains. Metric: pass@4.

Overview

HorizonMath

Category Metric Tasks Saturation Created

Paper GitHub Dataset

Quick answer: HorizonMath is a benchmark of 136 predominantly unsolved computational and applied-mathematics problems across eight domains, built by researchers from Oxford, Harvard, Princeton, and the Ellison Institute of Technology. Scored pass@4 — success in at least one of four independent attempts.

At a Glance

What it tests: Whether a model can make genuine progress on real, currently-unsolved computational and applied math problems — not curated exercises with known answers.

Why it matters: Most math benchmarks test known problems with known solutions; HorizonMath specifically targets the frontier of open problems, where correctness must be automatically verified against numerical, constructive, or best-known-baseline criteria rather than a pre-known answer key.

Known limitations: Extremely hard by design (low scores across the board are expected); 136 tasks is a modest sample; pass@4 rewards any success across 4 tries rather than single-run reliability.

What HorizonMath Measures

HorizonMath was built by Erik Y. Wang and collaborators to test models against genuinely unsolved (or only partially solved) problems in computational and applied mathematics, spanning eight domains. Each problem is constructed so a candidate solution — a numerical answer, an explicit construction, or an improvement over a published best-known result — can be checked automatically without a human-authored answer key, since for many of the problems no complete answer key exists.

Because most of the problems have no fully known solution, scores across all models are typically very low; a small absolute improvement here represents genuine progress at the edge of what's currently computationally achievable, rather than incremental gains against a saturating benchmark.

Benchmark Specifications

FieldValue
Task categoryMath
MetricPass@4 (success in ≥1 of 4 attempts)
Number of tasks136
SaturationLow
Created byErik Y. Wang et al. (Oxford, Harvard, Princeton, Ellison Institute of Technology)
Source paperWang et al. 2026
GitHubewang26/HorizonMath
DatasetHugging Face

How HorizonMath Is Scored

Validation is automatic: numerical answers are checked for exact/tolerance-based correctness, constructive answers for validity, and improvement-type answers against the best previously published baseline. The pass@4 metric counts a problem as solved if the model succeeds in at least one of four independent attempts.

State-of-the-Art Results

RankModelScoreSourceDate
1Hy4 Preview8.8%Tencent Hunyuan model card2026-08

Scores sourced from published technical reports and model cards. Results depend on harness, prompt format, and effort settings — see each source for methodology.

HorizonMath on Benchgen

No Benchgen results yet — be the first to run HorizonMath.

HorizonMath vs Other Benchmarks

BenchmarkWhat it testsTasksSaturation
HorizonMathUnsolved computational/applied math136Low
FrontierMathExtremely hard research-level mathLow
MathArena Apex 2025Competition mathematicsLow

Use HorizonMath specifically to probe the very frontier of open mathematical problems — expect single-digit-to-low-double-digit scores from even the strongest models, unlike competition-math benchmarks where top models can exceed 70-90%.

Run HorizonMath on Your Model

Benchgen lets teams run HorizonMath against their own model versions, compare results across runs, and track incremental progress on genuinely unsolved problems — rather than relying on a single vendor-reported number.

Frequently Asked Questions

What is HorizonMath? HorizonMath is a benchmark of 136 predominantly unsolved computational and applied-math problems across eight domains, scored pass@4, created by Erik Y. Wang et al.
What does a good HorizonMath score look like? As of August 2026, frontier models score in the low single digits to low teens (roughly 3.5–10.6%) — this benchmark is intentionally unsaturated, so any measurable progress is notable.
Who created HorizonMath? HorizonMath was created by Erik Y. Wang and collaborators from the University of Oxford, Harvard, Princeton, and the Ellison Institute of Technology, published in 2026.