Toward Reliable Context Compression for Long-Horizon Agents: An Empirical Study of Execution Instability
Guanghui Min, Liang Wu, Mayank Darbari, Chen Chen, Liangjie Hong
cs.LG
Submitted: 2026-08-06
Updated: 2026-08-10
Comments: 31 pages, 6 figures
Code: https://github.com/nokia-applied-research/Trace
License: http://creativecommons.org/licenses/by/4.0/
The gist: Recurrent context compression controls context growth in long-horizon agents, but its behavioral effects remain poorly understood.
Terminology
Abstract
Recurrent context compression controls context growth in long-horizon agents, but its behavioral effects remain poorly understood. In this preliminary empirical study, we show that compression can weaken the influence of recent interactions, increasing blocked actions, repeated exploration, and instability across runs. Motivated by these observations, we introduce TRACE, a verifier-guided framework that evaluates individual compaction events through paired closed-loop continuations from the same environment state and uses summary preferences to optimize a natural-language compression prompt while keeping all models frozen. Initial results on AppWorld show improvements over existing compression baselines in task performance, multi-run reliability, and context--execution efficiency. These findings provide early evidence for boundary-local evaluation as a promising direction for reliable agent context compression.
Sources
- $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment
- Scaling LLM Multi-turn RL with End-to-end Summarization-based Context Management
- OfficeBench: Benchmarking Language Agents across Multiple Applications for Office Automation
- ReSum: Unlocking Long-Horizon Search Intelligence via Context Summarization
- Fix the Structural Bottleneck: Context Compression via Explicit Information Transmission
- WebArena: A Realistic Web Environment for Building Autonomous Agents
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks