The Endless Exam: Mathematical Constructions from Today's Models toward Superintelligence
cs.AI, math.HO
Submitted: 2026-09-21
Updated: 2026-09-28
Comments: 51 pages, 13 figures, 25 tables
Code: https://github.com/ricolalemon/endless-exam
License: http://creativecommons.org/licenses/by/4.0/
The gist: We introduce the Endless Exam, a benchmark for measuring mathematical progress from today's models toward artificial superintelligence through fourteen parameterised construction families.
Terminology
Abstract
We introduce the Endless Exam, a benchmark for measuring mathematical progress from today's models toward artificial superintelligence through fourteen parameterised construction families. Each submitted object is checked automatically for validity and assigned a relative quality score against a published frontier or construction baseline, without capping improvements at 1. The families draw on open mathematical problems for long-term targets and generate new instances at larger parameters, where compact certificates keep large constructions verifiable. Across eight models evaluated on 69 distinct instances, continuous quality scores distinguish performance even though no evaluated system surpasses a published frontier. Size-quality curves show how construction quality changes as problem size increases. We release the generators, verifiers, references, model responses and analysis to support continued measurement before and beyond human frontiers.
Sources
- First Proof
- When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
- Shifted S-templates and improved lower bounds for Schur numbers
- Blocking sets, minimal codes and trifferent codes
- The generalized trifference problem
- On the Measure of Intelligence
- ARC-AGI-2: A New Challenge for Frontier AI Reasoning Systems
- Caps and progression-free sets in $\mathbb{Z}_m^n$
- Mathematical exploration and discovery at scale
- FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
- New constructions for covering designs
- Towards Robust Mathematical Reasoning
- AlphaEvolve: A coding agent for scientific and algorithmic discovery
- MathConstraint: Automated Generation of Verified Combinatorial Reasoning Instances for LLMs
- Humanity's Last Exam
- New lower bound on the Shannon capacity of C7 from circular graphs
- An Improved Lower Bound for $S(7)$ and Some Interesting Templates
- IMProofBench: Benchmarking AI on Research-Level Mathematical Proof Generation
- Strengthening Recursive Constructions for Zero-Error Shannon Capacity
- Measuring Progress in Reasoning Toward Mathematical Discovery with Automatic Verification
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection