Dude, Where's My State? Execution Information Requirements for Stateful Agents
cs.AI
Submitted: 2026-09-26
Updated: 2026-09-26
Terminology
Sources
- Training Deep Nets with Sublinear Memory Cost
- What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents
- MemoryArena: Benchmarking Agent Memory in Interdependent Multi-Session Agentic Tasks
- SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
- ACON: Optimizing Context Compression for Long-horizon LLM Agents
- Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
- MemGPT: Towards LLMs as Operating Systems
- The Illusion of Procedural Reasoning: Measuring Long-Horizon FSM Execution in LLMs
- Mem2ActBench: A Benchmark for Evaluating Long-Term Memory Utilization in Task-Oriented Autonomous Agents
- World-Model Collapse as a Phase Transition
- Scaling Long-Horizon LLM Agent via Context-Folding
- Toward a Theory of Hierarchical Memory for Language Agents
- Context Compaction Theory
- Michelangelo: Long Context Evaluations Beyond Haystacks via Latent Structure Queries
- Remember When It Matters: Proactive Memory Agent for Long-Horizon Agents
- $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
- Your Agents Are Aging Too: Agent Lifespan Engineering for Deployed Systems
- Remember the Decision, Not the Description: A Rate-Distortion Framework for Agent Memory
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection