Persistent Context Graphs for Efficient Memory Compaction in LLM Agents
cs.CL
Submitted: 2026-09-30
Updated: 2026-09-30
Code: https://github.com/UCSB-NLP-Chang/ReCAP
Terminology
Sources
- ACON: Optimizing Context Compression for Long-horizon LLM Agents
- IntentKV: Cross-Turn Intent-Aware KV Cache Pruning for Agent Inference
- The Complexity Trap: Simple Observation Masking Is as Efficient as LLM Summarization for Agent Context Management
- Practical Online KV Cache Compaction for LLM Agents: An Empirical Study
- Mix-Quant: Quantized Prefilling, Precise Decoding for Agentic LLMs
- gpt-oss-120b & gpt-oss-20b Model Card
- Scaling Long-Horizon LLM Agent via Context-Folding
- Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference
- ContextWeaver: Selective and Dependency-Structured Memory Construction for LLM Agents
- SWE-Together: Evaluating Coding Agents in Interactive User Sessions
- Qwen3 Technical Report
- AttnCompress: Dynamic Attention-Guided Trajectory Compression for Software Engineering Agents
- TraceLab: Characterizing Coding Agent Workloads for LLM Serving
- Fast KV Compaction via Attention Matching
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering