FrontierMath Erd s
cs.CL, cs.AI
Submitted: 2026-09-06
Updated: 2026-09-06
Code: https://github.com/leanprover/comparator
License: http://creativecommons.org/licenses/by/4.0/
The gist: We introduce FrontierMath Erdős (FME), a benchmark of 68 Erdős problems that are open as of August 2026.
Terminology
Abstract
We introduce FrontierMath Erdős (FME), a benchmark of 68 Erdős problems that are open as of August 2026. To solve a task in FME, AI systems must resolve (prove or disprove) one of the 68 conjectures in the proof assistant Lean. Our 68 problems were selected by the second author among 652 open problems on erdosproblems.com for their mathematical interest and difficulty. AIs have recently resolved several open problems in mathematics, but these demonstrations fall short of a systematic study of AI capabilities. FME evaluates every AI model on the same fixed problems, autonomously and under the same budget. We evaluated five AIs with a budget of 300 per problem. One (GPT-6 Astra) scored 3%, and all others scored 0%.
Sources
- Measuring Progress in Reasoning Toward Mathematical Discovery with Automatic Verification
- Advancing Mathematics Research with AI-Driven Formal Proof Search
- OEIS Open: How many conjectures can language models turn into theorems?
- First Proof
- First Proof Second Batch
- Formal Conjectures: An Open and Evolving Benchmark for Verified Discovery in Mathematics
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering