How Fast Do Agents Rot? An Empirical Study of Long-Horizon Degradation in LLM Agents for Production Decision-Making
physics.soc-ph, cs.AI
Submitted: 2026-08-31
Updated: 2026-08-31
Code: https://github.com/shubmittal/agent-horizon-degradation
Terminology
Sources
- ReAct: Synergizing Reasoning and Acting in Language Models
- Reflexion: Language Agents with Verbal Reinforcement Learning
- Faith and Fate: Limits of Transformers on Compositionality
- GAIA: a benchmark for General AI Assistants
- SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
- $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
- WebArena: A Realistic Web Environment for Building Autonomous Agents
- Measuring AI Ability to Complete Long Software Tasks
- AI Agents That Matter
- LLMs Get Lost In Multi-Turn Conversation
- Mind2Web: Towards a Generalist Agent for the Web
- RULER: What's the Real Context Size of Your Long-Context Language Models?
- A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions
- Same Task, More Tokens: the Impact of Input Length on the Reasoning Performance of Large Language Models
- AgentBench: Evaluating LLMs as Agents
- AppWorld: A Controllable World of Apps and People for Benchmarking Interactive Coding Agents
- OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
- TheAgentCompany: Benchmarking LLM Agents on Consequential Real World Tasks
Related papers
- URBAN-SPIN: A street-level bikeability index to inform design implementations in historical city centres
- Civilizational Metamaterials: Engineering Coordination Under Capability Gradients and Structural Turbulence
- Collective Behavior of AI Agents: the Case of Moltbook
- A Distinct Communication Strategies Model of the Double Empathy Problem
- Measuring Primitive Accumulation: An Information-Theoretic Approach to Capitalist Enclosure in PIK2, Indonesia
- Diffusion-induced instabilities promote cooperation in eco-evolutionary networks