From Atomic to Agentic: Towards Interpretable Evaluation of LLMs' Agentic Mathematical Capabilities
cs.AI, cs.CL
Submitted: 2026-08-27
Updated: 2026-08-27
Comments: EMNLP 2026
Code: https://github.com/Eternity-gaga/Agentic-Math-Bench
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- GPT-4 Technical Report
- AMO-Bench: Large Language Models Still Struggle in High School Math Competitions
- Toward Native Multimodal Modeling: A Roadmap
- ProofNet: Autoformalizing and Formally Proving Undergraduate-Level Mathematics
- A Survey on Mathematical Reasoning and Optimization with Large Language Models
- Atomic Thinking of LLMs: Decoupling and Exploring Mathematical Reasoning Abilities
- Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models
- ToRA: A Tool-Integrated Reasoning Agent for Mathematical Problem Solving
- EvoConfig: Self-Evolving Multi-Agent Systems for Efficient Autonomous Environment Configuration
- OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems
- DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
- Refine Knowledge of Large Language Models via Adaptive Contrastive Learning
- One Example Shown, Many Concepts Known! Counterexample-Driven Conceptual Reasoning in Mathematical LLMs
- Cognitive Mismatch in Multimodal Large Language Models for Discrete Symbol Understanding
- DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models
- FIMO: A Challenge Formal Dataset for Automated Theorem Proving
- MM-Agent: LLM as Agents for Real-world Mathematical Modeling Problem
- Process-Driven Autoformalization in Lean 4
- Youtu-LLM: Unlocking the Native Agentic Potential for Lightweight Large Language Models
- CriticLean: Critic-Guided Reinforcement Learning for Mathematical Formalization
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection