Correct Now, Insufficient Later: Auditing Update Sufficiency in Context Compression
cs.LG, cs.AI
Submitted: 2026-09-17
Updated: 2026-09-17
Comments: 20 pages, 9 tables, 2 figures. Code and reproducibility materials to be released separately
License: http://creativecommons.org/licenses/by/4.0/
The gist: A memory can answer a current query correctly while discarding distinctions required by a later update.
Terminology
Abstract
A memory can answer a current query correctly while discarding distinctions required by a later update. We investigate this failure with a paired-history audit: two histories have the same current answer, receive a shared future update, and require different subsequent answers. A pilot evaluates 24 history pairs across six synthetic mechanisms, 12 memory conditions, two repeats, and two model backends. A deterministic frontier selector obtains strict reveal accuracy of 96/96 on DeepSeek and 82/96 on GLM; a structured writer obtains 62 successes with one unresolved outcome and 56/96. The configured four-outcome joint contrast has finite-sample identification intervals of [0.521, 0.542] and [0.292, 0.313], not confidence intervals. A record-level audit distinguishes retained-state adequacy, response delivery, and answer-schema compliance without changing those original scores. It finds 26 and 25 well-formed but semantically wrong structured reveal memories, while all 14 GLM frontier reveal failures contain correct values in the wrong wrapper. Tombstone removal produces 16/16 exact replay failures in the targeted mechanism. Identifier renaming then exposes a separate flaw: original frontier late-reference adequacy falls from 8/8 to 94/320 transformed instances. We provide and test a label-equivariant repair, but it preserves only 2/8 original late-reference answers: eliminating a naming shortcut does not solve unknown future relevance. These results support a scoped evaluation methodology and reproducible failure analysis, not general superiority of the repaired algorithm. Paid pilot evidence, retrospective diagnostics, and new offline tests are reported separately; no independent held-out or natural-task validation is claimed.
Sources
- MEMAUDIT: An Exact Package-Oracle Evaluation Protocol for Budgeted Long-Term LLM Memory Writing
- Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory
- Reasoning or Not? A Comprehensive Evaluation of Reasoning LLMs for Dialogue Summarization
- ACON: Optimizing Context Compression for Long-horizon LLM Agents
- The Complexity Trap: Simple Observation Masking Is as Efficient as LLM Summarization for Agent Context Management
- Toward Reliable Context Compression for Long-Horizon Agents: An Empirical Study of Execution Instability
- Zep: A Temporal Knowledge Graph Architecture for Agent Memory
- AppWorld: A Controllable World of Apps and People for Benchmarking Interactive Coding Agents
- LongMemEval-V2: Evaluating Long-Term Agent Memory Toward Experienced Colleagues
- LongMemEval: Benchmarking Chat Assistants on Long-Term Interactive Memory
- A-MEM: Agentic Memory for LLM Agents
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks