A Survey of Agentic Reasoning for Large Language Models: Towards Recursively Self-Improving and Collective Agents
cs.AI, cs.CL
Submitted: 2026-01-18
Updated: 2026-09-20
Comments: Accepted by TMLR. Project: https://github.com/weitianxin/Awesome-Agentic-Reasoning
Code: https://github.com/weitianxin/Awesome-Agentic-Reasoning
License: http://creativecommons.org/licenses/by/4.0/
The gist: Reasoning is a fundamental cognitive process underlying inference, problem-solving, and decision-making.
Terminology
Abstract
Reasoning is a fundamental cognitive process underlying inference, problem-solving, and decision-making. While large language models (LLMs) demonstrate strong reasoning capabilities in closed-world settings, they struggle in open-ended and dynamic environments. Agentic reasoning marks a paradigm shift by reframing LLMs as autonomous agents that plan, act, and learn through continual interaction. In this survey, we organize agentic reasoning along three complementary dimensions. First, we characterize environmental dynamics through three layers: foundational agentic reasoning, which establishes core single-agent capabilities including planning, tool use, and search in stable environments; self-evolving agentic reasoning, which studies how agents refine these capabilities through feedback, memory, and adaptation; and collective multi-agent reasoning, which extends intelligence to collaborative settings involving coordination, knowledge sharing, and shared goals. Across these layers, we distinguish in-context reasoning, which scales test-time interaction through structured orchestration, from post-training reasoning, which optimizes behaviors via reinforcement learning and supervised fine-tuning. We further review representative agentic reasoning frameworks across real-world applications and benchmarks, including science, robotics, healthcare, autonomous research, and mathematics. This survey synthesizes agentic reasoning methods into a unified roadmap bridging thought and action, and outlines open challenges and future directions, including personalization, long-horizon interaction, world modeling, scalable multi-agent training, and governance for real-world deployment.
Sources
- Least-to-Most Prompting Enables Complex Reasoning in Large Language Models
- Agentic Retrieval-Augmented Generation: A Survey on Agentic RAG
- A Survey on Retrieval-Augmented Text Generation for Large Language Models
- OpenHands: An Open Platform for AI Software Developers as Generalist Agents
- Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory
- MemOS: An Operating System for Memory-Augmented Generation (MAG) in Large Language Models
- Memory-R1: Enhancing Large Language Model Agents to Manage and Utilize Memories via Reinforcement Learning
- AutoAgents: A Framework for Automatic Agent Generation
- BattleAgentBench: A Benchmark for Evaluating Cooperation and Competition Capabilities of Language Models in Multi-Agent Systems
- AgentBench: Evaluating LLMs as Agents
- MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents
- Search-o1: Agentic Search-Enhanced Large Reasoning Models
- A-MEM: Agentic Memory for LLM Agents
- Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory
- Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
- WebAgent-R1: Training Web Agents via End-to-End Multi-Turn Reinforcement Learning
- ChemCrow: Augmenting large-language models with chemistry tools
- Physical AI Agents: Integrating Cognitive Intelligence with Real-World Action
- MatExpert: Decomposing Materials Discovery by Mimicking Human Experts
- Voyager: An Open-Ended Embodied Agent with Large Language Models
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection