The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents
Mingguang Chen, Licheng Wang, Bo Qu
cs.CL
Submitted: 2026-08-07
Updated: 2026-08-10
Comments: 39 pages, 6 figures
License: http://creativecommons.org/licenses/by/4.0/
The gist: Frontier language models solve reasoning problems in a single forward pass that would have been research contributions years ago, yet fail at multi-hour tasks: losing track of earlier decisions,
Terminology
Abstract
Frontier language models solve reasoning problems in a single forward pass that would have been research contributions years ago, yet fail at multi-hour tasks: losing track of earlier decisions, declaring half-finished work done, or drifting from goals. We call this the horizon gap and survey 1,547 arXiv papers (2024-2026) collected via systematic seed harvest with a disclosed 26.8% bleed filter, extended by targeted supplementation. We disambiguate three routinely conflated properties: long-horizon (task property: required steps), long-context (model property: token capacity), and long-term memory (system property: persistence across steps/sessions). We organize the corpus into six categories tracking a long-horizon task's lifecycle -- planning, memory, execution, training, evaluation, and foundations/safety -- crossed with an axis capturing where horizons are carried (within-context, within-task-beyond-context, or cross-task-persistent). Across all categories, we find the same pattern: outcome-only signals grow uninformative as horizons lengthen, and the field's response -- whether process reward models, credit assignment, or trajectory-level diagnostics -- manufactures denser step-level signals. We treat critical and diagnostic literature as first-class threads throughout, arguing that segregating critique from method would routinely split single papers across chapters. We close by naming open measurement problems: decomposing model versus harness capability, managing correlated bias in process-level signals used for both training and evaluation, and whether long-horizon reliability admits general predictive theory.
Sources
- Measuring AI Ability to Complete Long Software Tasks
- Large Language Models Cannot Self-Correct Reasoning Yet
- Are "Solved Issues" in SWE-bench Really Solved Correctly? An Empirical Study
- Establishing Best Practices for Building Rigorous Agentic Benchmarks
- Understanding the planning of LLM agents: A survey
- LLMs as Planning Formalizers: A Survey for Leveraging Large Language Models to Construct Automated Planning Models
- A Survey on the Memory Mechanism of Large Language Model based Agents
- A Survey of Agent Memory in the Second Half: Towards Self-Evolving and Long-Horizon Agents
- Large Language Model-Brained GUI Agents: A Survey
- GUI Agents: A Survey
- When Can LLMs Actually Correct Their Own Mistakes? A Critical Survey of Self-Correction of LLMs
- LLM+P: Empowering Large Language Models with Optimal Planning Proficiency
- Do As I Can, Not As I Say: Grounding Language in Robotic Affordances
- Tree-of-Code: A Hybrid Approach for Robust Complex Task Planning and Execution
- Plan-and-Act: Improving Planning of Agents for Long-Horizon Tasks
- Harnesses for Inference-Time Alignment over Execution Trajectories
- PRompt Optimization in Multi-Step Tasks (PROMST): Integrating Human Feedback and Heuristic-based Sampling
- ReAgent: Reversible Multi-Agent Reasoning for Knowledge-Enhanced Multi-Hop QA
- Do Agents Need to Plan Step-by-Step? Rethinking Planning Horizon in Data-Centric Tool Calling
- The Cognitive Bandwidth Bottleneck: Shifting Long-Horizon Agent from Planning with Actions to Planning with Schemas
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering