Measuring Progress in Reasoning Toward Mathematical Discovery with Automatic Verification
cs.LG
Submitted: 2026-03-16
Updated: 2026-09-10
Comments: ICML AI4Math Best Paper Award
Code: https://github.com/ewang26/HorizonMath
License: http://creativecommons.org/licenses/by/4.0/
The gist: Can AI make progress on important, unsolved mathematical problems? Large language models are now capable of sophisticated mathematical and scientific reasoning, but whether they can perform novel
Terminology
Abstract
Can AI make progress on important, unsolved mathematical problems? Large language models are now capable of sophisticated mathematical and scientific reasoning, but whether they can perform novel research is still widely debated and underexplored. We introduce HorizonMath, a benchmark of 113 predominantly unsolved problems spanning eight domains in mathematics and the mathematical sciences, paired with an open-source evaluation framework for automated verification. Our benchmark targets the generator-verifier gap: problems where discovery is hard and requires meaningful mathematical insight, but verification is computationally straightforward. This contrasts with most existing research-level benchmarks, which instead rely on formal proof verification or manual review, both of which are expensive to scale. Because these solutions are unknown, HorizonMath is resistant to data contamination, and most state-of-the-art models score under 10%. Using this framework, we identify six novel solutions to research problems that either resolve previously open questions or improve on the best-known published results, with GPT-5.4 Pro and GPT-5.6 Sol each discovering three of these solutions. Across seven frontier model families, reasoning efficiency and behavior also vary substantially. We release HorizonMath as an open challenge and a growing community resource, where each verified solution is a candidate contribution to the mathematical literature.
Sources
- First Proof
- Training Verifiers to Solve Math Word Problems
- Semi-Autonomous Mathematics Discovery with Gemini: A Case Study on the Erd\H{o}s Problems
- Mathematical exploration and discovery at scale
- FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
- Single-minus gluon tree amplitudes are nonzero
- FrontierCS: Evolving Challenges for Evolving Intelligence
- AlphaEvolve: A coding agent for scientific and algorithmic discovery
- IMProofBench: Benchmarking AI on Research-Level Mathematical Proof Generation
- Resolution of Erd\H{o}s Problem #728: a writeup of Aristotle's Lean proof
- Accelerating Scientific Research with Gemini: Case Studies and Common Techniques
- Learning to Discover at Test Time
- Optimizing the CGMS upper bound on Ramsey numbers
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks