Compact-Memory LLM Agents via Online Max-Member Clustering and Atom-Aware Packing
cs.AI
Submitted: 2026-09-04
Updated: 2026-09-04
License: http://creativecommons.org/licenses/by/4.0/
The gist: Many long-horizon LLM deployments face tight prompt budgets: latency, cost, and context limits make full-context prompting impractical as interaction length grows.
Terminology
Abstract
Many long-horizon LLM deployments face tight prompt budgets: latency, cost, and context limits make full-context prompting impractical as interaction length grows. The key question is then not raw recall alone, but which memory design gives the best quality--token trade-off in the compact-memory regime. We present RSM-full, an online clustered-memory pipeline designed for a strong quality--token Pareto point. RSM-full combines two design choices: a cosine-gated max-member merge write rule and an atom-aware grouped context packer. On AMA-Bench, our primary compact-memory benchmark, it reaches 83% of Full-Context quality at 32% of the token cost at a 4 k budget; under four-seed averaging it beats the closest streaming-clustered baseline (Online K-Means) by +3.5 -- 6.0,pp (p <.001) across the whole about 2.6 k-- about 5 k regime. Three-seed ablations show most of this gain comes from the merge rule (+5.7,pp over Online K-Means and matched- τ DP-means) and the grouped packer (+5.0,pp over flat concatenation). The pattern reproduces on RealMem, an independent long-horizon persona-memory benchmark: RSM-full improves on Budget-RAG (+0.69,pp, p =.006), is on par with BM25-RAG (paired Δ = + 0.27,pp, p =.47; we do not claim BM25 equivalence in the equivalence-test sense), and significantly outperforms Streaming-Proto (+2.97,pp) and the closest reproduced 2025 agentic-memory baseline A-MEM (+1.65,pp, p <.001). Across benchmarks the message is consistent: under tight budgets, compact-memory performance is driven mainly by how streaming memories are merged and how retrieved content is assembled. Overall, RSM-full is most useful when answeroughly 2k -- 5k prompt tokens, where itdefines a strong compact-memory Pareto point; higher-token baselines remain stronger outside this regime.
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection