CliffCompaction: Cost-Efficient Compaction for Long-Horizon Coding Agents
cs.AI, cs.LG, cs.SE
Submitted: 2026-09-22
Updated: 2026-09-22
Code: https://github.com/nguyenvuthientrang/cliffcompaction
License: http://creativecommons.org/licenses/by/4.0/
The gist: Agents often work on complex problems that require millions of tokens of context, which necessitates compacting across sessions due to limited context windows.
Terminology
Abstract
Agents often work on complex problems that require millions of tokens of context, which necessitates compacting across sessions due to limited context windows. We develop CliffCompaction, an autocompaction technique that reduces cost by up to 50% under a bounded context while maintaining or improving performance on Terminal-Bench and achieving new levels of efficiency for test-time scaling and state-of-the-art results on KernelBench. The per-rollout savings of CliffCompaction make the performance--cost trade-off of test-time scaling more efficient, adding over 10 percentage points on Terminal-Bench for less than the cost of two full-context runs. Under parallel test-time scaling, CliffCompaction lets Kimi K2.6 match Opus 4.7, and exceed Opus 4.6 and GPT-5.3 Codex at lower cost. The key to CliffCompaction's effectiveness is that it keeps compacted information faithful by only truncating or dropping content, never rephrasing or rewriting it. We never compact a compaction---each pass operates only on original content, and prior compacted output is discarded, preventing context drift from accumulating. These properties sustain continual learning over sessions exceeding a million tokens: on KernelBench, CliffCompaction reaches CUDA kernel speedups of 2.23 times after 200 steps and 3.58 times after 400 steps, surpassing specialized search algorithms and trained agents despite being a general-purpose compaction technique. We open-source a scaffold-agnostic API-proxy implementation of CliffCompaction usable with Claude Code, Codex and other harnesses.
Sources
- CUDA Agent: Large-Scale Agentic RL for High-Performance CUDA Kernel Generation
- DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models
- SWE-Replay: Efficient Test-Time Scaling for Software Engineering Agents
- AdaExplore: Failure-Driven Adaptation and Diversity-Preserving Search for Efficient Kernel Generation
- CodeMonkeys: Scaling Test-Time Compute for Software Engineering
- Recursive Agent Optimization
- Trae Agent: An LLM-based Agent for Software Engineering with Test-time Scaling
- GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models
- The Larger the Better? Improved LLM Code-Generation via Budget Reallocation
- PABU: Progress-Aware Belief Update for Efficient LLM Agents
- ACON: Optimizing Context Compression for Long-horizon LLM Agents
- Scaling Test-Time Compute for Agentic Coding
- LLM-as-a-Verifier: A General-Purpose Verification Framework
- S*: Test Time Scaling for Code Generation
- Self-Compacting Language Model Agents
- Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
- AlphaEvolve: A coding agent for scientific and algorithmic discovery
- Beyond Compaction: Structured Context Eviction for Long-Horizon Agents
- Kimi K2.5: Visual Agentic Intelligence
- Active Context Compression: Autonomous Memory Management in LLM Agents
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection