An Architecture for Long-Horizon Agents: Levels, Ticks and Cascaded Intelligence
cs.AI, cs.LG
Submitted: 2026-09-17
Updated: 2026-09-17
License: http://creativecommons.org/licenses/by/4.0/
The gist: Language-model agents are increasingly asked to carry out work spanning days or weeks, such as an operations remediation or a research programme.
Terminology
Abstract
Language-model agents are increasingly asked to carry out work spanning days or weeks, such as an operations remediation or a research programme. Such a task outlives any context window, any process and any interval at which a person can attend. In this paper, we argue that a long-horizon agent must run continually without forgetting before it can learn continually. This ability lies in the harness around the model rather than in the model itself. We derive seven bottlenecks from the long-horizon setting and answer them with a hierarchical architecture of three parts: (i) levels indexed by time scale, each keeping a bounded file summarising the level below; (ii) a clocked tick as the unit of autonomous action; and (iii) cascaded intelligence, where work is escalated to a more capable model only after failing review. We report on a ten-day campaign in which an agent built on this architecture reproduced a published reinforcement-learning result with a human attending once a day, and show (1) the agent kept the thread across every context reset and session boundary of the campaign, (2) operating knowledge written early changed later behaviour with no change to model weights, and (3) where learned components would enter such a system. Overall, our experience suggests continual learning for these agents needs a substrate outliving every context and process, and the checks the harness already runs are where a learner belongs.
Sources
- Technical Report: Evaluating Goal Drift in Language Model Agents
- Is Escalation Worth It? On the Depth of LLM Cascades
- Why Do Multi-Agent LLM Systems Fail?
- Governance Decay: How Context Compaction Silently Erases Safety Constraints in Long-Horizon LLM Agents
- Cascaded Language Models for Cost-effective Human-AI Decision-Making
- MLReplicate: Benchmarking Autonomous Research Systems for Machine Learning Reproducibility
- Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning
- Memory in the Age of AI Agents
- HiMAC: Hierarchical Macro-Micro Learning for Long-Horizon LLM Agents
- Do Proactive Agents Need an LLM to Decide When to Act?
- The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery
- MemGPT: Towards LLMs as Operating Systems
- Model or Harness? An Interaction-Centric Taxonomy for Localizing Agent Failures
- Context Compaction Theory
- The Long-Horizon Task Mirage? Diagnosing Where and Why Agentic Systems Break
- RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts
- ReAct: Synergizing Reasoning and Acting in Language Models
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection