MathAdv: What Theorem Provers Know, Reason, Formalize, and Generalize
cs.CL, cs.AI, cs.LO
Submitted: 2026-08-26
Updated: 2026-08-28
Code: https://github.com/margotyjx/MathAdv
Terminology
Sources
- ProofNet: Autoformalizing and Formally Proving Undergraduate-Level Mathematics
- Llemma: An Open Language Model For Mathematics
- IndiMathBench: Autoformalizing Mathematical Reasoning Problems with a Human Touch
- DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models
- Herald: A Natural Language Annotated Lean 4 Dataset
- Solving Quantitative Reasoning Problems with Language Models
- Lean-STaR: Learning to Interleave Thinking and Proving
- Goedel-Prover: A Frontier Model for Open-Source Automated Theorem Proving
- FIMO: A Challenge Formal Dataset for Automated Theorem Proving
- Numina-Lean-Agent: An Open and General Agentic Reasoning System for Formal Mathematics
- WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-Instruct
- Generative Language Modeling for Automated Theorem Proving
- AI scientists produce results without reasoning scientifically
- REAL-Prover: Retrieval Augmented Lean Prover for Mathematical Reasoning
- OpenAI GPT-5 System Card
- PutnamBench: Evaluating Neural Theorem-Provers on the Putnam Mathematical Competition
- MA-LoT: Model-Collaboration Lean-based Long Chain-of-Thought Reasoning enhances Formal Theorem Proving
- InternLM2.5-StepProver: Advancing Automated Theorem Proving via Critic-Guided Search
- DeepSeek-Prover-V1.5: Harnessing Proof Assistant Feedback for Reinforcement Learning and Monte-Carlo Tree Search
- FormalMATH: Benchmarking Formal Mathematical Reasoning of Large Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering