| Rank | Model | Score |
|---|---|---|
| 1 | hy4-preview | 8.8 |
1 phaseActive
Oxford/Harvard/Princeton benchmark of 136 predominantly unsolved computational and applied-math problems across 8 domains. Metric: pass@4.
Quick answer: HorizonMath is a benchmark of 136 predominantly unsolved computational and applied-mathematics problems across eight domains, built by researchers from Oxford, Harvard, Princeton, and the Ellison Institute of Technology. Scored pass@4 — success in at least one of four independent attempts.
What it tests: Whether a model can make genuine progress on real, currently-unsolved computational and applied math problems — not curated exercises with known answers.
Why it matters: Most math benchmarks test known problems with known solutions; HorizonMath specifically targets the frontier of open problems, where correctness must be automatically verified against numerical, constructive, or best-known-baseline criteria rather than a pre-known answer key.
Known limitations: Extremely hard by design (low scores across the board are expected); 136 tasks is a modest sample; pass@4 rewards any success across 4 tries rather than single-run reliability.
HorizonMath was built by Erik Y. Wang and collaborators to test models against genuinely unsolved (or only partially solved) problems in computational and applied mathematics, spanning eight domains. Each problem is constructed so a candidate solution — a numerical answer, an explicit construction, or an improvement over a published best-known result — can be checked automatically without a human-authored answer key, since for many of the problems no complete answer key exists.
Because most of the problems have no fully known solution, scores across all models are typically very low; a small absolute improvement here represents genuine progress at the edge of what's currently computationally achievable, rather than incremental gains against a saturating benchmark.
| Field | Value |
|---|---|
| Task category | Math |
| Metric | Pass@4 (success in ≥1 of 4 attempts) |
| Number of tasks | 136 |
| Saturation | Low |
| Created by | Erik Y. Wang et al. (Oxford, Harvard, Princeton, Ellison Institute of Technology) |
| Source paper | Wang et al. 2026 |
| GitHub | ewang26/HorizonMath |
| Dataset | Hugging Face |
Validation is automatic: numerical answers are checked for exact/tolerance-based correctness, constructive answers for validity, and improvement-type answers against the best previously published baseline. The pass@4 metric counts a problem as solved if the model succeeds in at least one of four independent attempts.
| Rank | Model | Score | Source | Date |
|---|---|---|---|---|
| 1 | Hy4 Preview | 8.8% | Tencent Hunyuan model card | 2026-08 |
Scores sourced from published technical reports and model cards. Results depend on harness, prompt format, and effort settings — see each source for methodology.
No Benchgen results yet — be the first to run HorizonMath.
| Benchmark | What it tests | Tasks | Saturation |
|---|---|---|---|
| HorizonMath | Unsolved computational/applied math | 136 | Low |
| FrontierMath | Extremely hard research-level math | — | Low |
| MathArena Apex 2025 | Competition mathematics | — | Low |
Use HorizonMath specifically to probe the very frontier of open mathematical problems — expect single-digit-to-low-double-digit scores from even the strongest models, unlike competition-math benchmarks where top models can exceed 70-90%.
Benchgen lets teams run HorizonMath against their own model versions, compare results across runs, and track incremental progress on genuinely unsolved problems — rather than relying on a single vendor-reported number.