Beyond Token Savings: A Systematic Study of Context Compression in LLM Agents
cs.LG, cs.AI
Submitted: 2026-09-26
Updated: 2026-09-26
Code: https://github.com/google-gemini/gemini-cli
Terminology
Sources
- CoACT: Action-Preserving Observation Compression for Coding Agents
- Context Length Alone Hurts LLM Performance Despite Perfect Retrieval
- ACON: Optimizing Context Compression for Long-horizon LLM Agents
- Characterization of Multi-Model Agentic AI Systems on General Tasks via Trace-Driven Simulation
- The Cost of Dynamic Reasoning: Demystifying AI Agents and Test-Time Scaling from an AI Infrastructure Perspective
- Rethinking AI Cloud Infrastructure for Agentic Serving Systems with the Aries Experimentation Framework
- The Complexity Trap: Simple Observation Masking Is as Efficient as LLM Summarization for Agent Context Management
- Agentic Coding in the Wild: Characterizing GitHub Copilot Traces at Production Scale
- LMCache: An Efficient KV Cache Layer for Enterprise-Scale LLM Inference
- Toward Reliable Context Compression for Long-Horizon Agents: An Empirical Study of Execution Instability
- Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
- Towards Understanding, Analyzing, and Optimizing Agentic AI Execution: A CPU-Centric Perspective
- Preble: Efficient Distributed Prompt Scheduling for LLM Serving
- How Coding Agents Fail Their Users: A Large-Scale Analysis of Developer-Agent Misalignment in 20,574 Real-World Sessions
- TokenPilot: Cache-Efficient Context Management for LLM Agents
- Architectural Implications of Agentic AI Workflows
- KAIROS: Stateful, Context-Aware Power-Efficient Agentic Inference Serving
- Agentic AI Workload Characteristics
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks