The Sleeping Agent: What Gist-Based Context Compression Loses and Why
Nicholas E. Kyrkewood
cs.AI, cs.CL
Submitted: 2026-08-12
Updated: 2026-08-13
Comments: 7 pages, 5 tables, appendices. Code and results at https://github.com/kyrkewood/sleeping-agent
Code: https://github.com/kyrkewood/sleeping-agent
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 75/100
The gist: The paper introduces Salience-Weighted Consolidation (SWC), a biologically-inspired compression framework motivated by sleep-based memory consolidation, used as a diagnostic probe to study when
Terminology
Summary
The paper introduces Salience-Weighted Consolidation (SWC), a biologically-inspired compression framework motivated by sleep-based memory consolidation, used as a diagnostic probe to study when gist-based context compression helps and when it hurts. SWC scores conversation history chunks by estimated salience, partitions them into priority tiers, and applies structured gist abstraction to mid-priority content. The framework draws three principles from sleep neuroscience: proactive scheduling, salience-weighted retention, and staged compression.
Evaluating four conditions—truncation, sliding window, SWC-Full, and SWC-Temporal—on all ten LoCoMo conversations (1,935 matched text-only questions total, 1,501 used in the primary aggregate after excluding Category 5 adversarial questions) at temperature 0, the paper finds a consistent task-type interaction: "gist compression substantially outperforms truncation on multi-hop reasoning and single-hop factual questions, but temporal questions remain substantially harder under compression, with compressed conditions scoring well below the full-context reference on the conversations where both are evaluated."
The key results are: SWC-Temporal achieves 0.468 aggregate judge accuracy on categories 1–4, compared to SWC-Full at 0.379, sliding window at 0.238, and truncation at 0.171. On category 2 (temporal) questions specifically, SWC-Full fails dramatically at 0.156 accuracy, while SWC-Temporal reaches 0.470. On multi-hop questions, both SWC conditions score 0.407–0.429 against baselines of 0.157–0.271; on single-hop questions, SWC conditions score 0.459–0.486 against 0.170–0.238.
The paper traces the temporal failure to a specific mechanism: the gist abstraction prompt preserves relational and event structure while discarding dates and times.
A preservation analysis across all ten conversations confirms this: "an approximately 20-fold increase in temporal expression preservation (3.05%→62.39%) with a one-sentence prompt modification, while named entity and event preservation rates barely change (×1.02 and ×1.11), demonstrating that the fix is a precision instrument."
The prompt modification—adding one sentence to the gist abstraction prompt: You MUST preserve verbatim: all specific dates, times, durations, ages, and temporal expressions (e.g. '7 May', 'last Tuesday', 'at 9pm')
—recovers +0.314 [0.254, 0.375] judge accuracy on category-2 (temporal) questions in the matched set,
with the direction positive in all ten conversations (per-conversation deltas: +0.433, +0.160, +0.259, +0.379, +0.385, +0.291, +0.177, +0.341, +0.364, +0.226; unweighted mean +0.301).
The paper argues the failure mode is a correctable design omission
rather than a fundamental limitation. The operative distinction is temporal versus non-temporal, not episodic versus semantic
—SWC-Full scores 0.459 on single-hop factual questions (episodic content in the conventional sense), so it does not fail on episodic content generally; it fails specifically on temporal information because temporal anchors fall outside all protected categories by default.
The paper also reports that category 3 (open-domain) confidence intervals overlap completely across all conditions, so no reliable conclusion can be drawn for commonsense inference questions. Category 5 (adversarial) questions are excluded because their binary scoring rule rewards content removal and conflates memory quality with response calibration.
The conclusion states: Generic gist abstraction preserves relational, event, and factual content while selectively losing temporal anchors, because they are absent from standard prompt specifications.
The paper recommends that any memory system using gist abstraction should explicitly protect temporal anchors
and that temporal anchors should be treated as a protected category in any gist compression component.
It also emphasizes that category-level evaluation paired with preservation rate analysis provides a more complete picture of what a memory strategy preserves—and what it leaves behind,
since systems that achieve high aggregate scores may mask systematic failures on specific question types.
Improvements for AI systems
Improvements to AI Systems:
- Add explicit temporal-anchor protection to all gist-compression prompts.
- The improved system will preserve verbatim dates, times, durations, ages, and temporal expressions (e.g.,
7 May
,last Tuesday
,at 9pm
) during context summarization, preventing the silent loss of time-critical information that current generic abstraction causes.
- Implement salience-weighted tiered memory consolidation with proactive scheduling.
- The system will score conversation chunks by estimated salience, partition them into priority tiers, and apply staged compression only to mid-priority content—while keeping high-priority content verbatim and low-priority content aggressively abstracted. This mimics sleep-based consolidation, improving long-context reasoning without sacrificing critical details.
- Add a temporal-preservation audit layer to memory systems.
- After any compression step, the system will run a preservation-rate check for temporal expressions (targeting >60% preservation, up from the observed 3.05% baseline). If preservation falls below threshold, it will automatically re-abstract with the protected-category prompt. This prevents silent temporal degradation in deployed systems.
- Introduce category-aware evaluation and adaptive compression selection.
- The system will classify incoming questions into task types (temporal, multi-hop, single-hop factual, open-domain) and dynamically choose between full-context, temporal-protected gist, or sliding-window strategies based on predicted category. For temporal questions, it will default to full-context or SWC-Temporal; for multi-hop, it will use gist compression to reduce noise.
- Treat temporal anchors as a first-class protected category in all memory schemas.
- The system will maintain a structured index of temporal facts separate from gist summaries, allowing fast retrieval of exact dates/times even when the main context is compressed. This enables accurate temporal reasoning without requiring full-context retention.
- Add per-category confidence reporting to prevent masked failures.
- The improved system will output accuracy metrics broken down by question type (temporal, multi-hop, single-hop, open-domain) rather than a single aggregate score. This exposes systematic weaknesses (e.g., temporal accuracy dropping to 0.156) that aggregate numbers hide, enabling targeted retraining or fallback strategies.
What the improved AI system can do:
-
Answer temporal questions (e.g.,
When did X happen relative to Y?
) with high accuracy even after long conversations, because dates and times are never lost during compression. -
Maintain strong multi-hop reasoning performance (0.4+ accuracy) while using 70–80% less context memory, by applying gist abstraction only to non-temporal, mid-priority content.
-
Automatically detect when a query is temporal and switch to a full-context or temporal-protected mode, avoiding the 0.156 accuracy collapse seen in naive compression.
-
Provide transparent, per-category performance reports so users and developers know exactly which question types the system handles well and which it fails on—preventing overconfidence from aggregate scores.
-
Scale to arbitrarily long conversations without degrading temporal fidelity, because temporal anchors are stored in a dedicated, lossless index rather than being subject to generic summarization.
Abstract
Gist-based context compression---summarising older conversation history into compact representations---is a common approach in long-horizon language model agents, yet its effect on different types of memory retrieval is poorly understood. We use Salience-Weighted Consolidation (SWC), a biologically-inspired compression framework motivated by sleep-based memory consolidation, as a diagnostic probe to study when gist compression helps and when it hurts. SWC scores conversation history by salience, partitions it into priority tiers, and applies structured gist abstraction to mid-priority content. Evaluating four conditions on all ten LoCoMo conversations---1,935 matched text-only questions in total, 1,501 used in the primary aggregate after excluding Category 5 (adversarial) questions---at temperature 0, we find a consistent task-type interaction: gist compression substantially outperforms truncation on multi-hop reasoning and single-hop factual questions, but temporal questions remain substantially harder under compression, with compressed conditions scoring well below the full-context reference on the conversations where both are evaluated. We trace this failure to a specific mechanism: the gist abstraction prompt preserves relational and event structure while discarding dates and times. A preservation analysis across all ten conversations confirms the mechanism: an approximately 20-fold increase in temporal expression preservation (3.05% to 62.39%) with a one-sentence prompt modification, while named entity and event preservation rates barely change (x1.02 and x1.11), demonstrating that the fix is a precision instrument. The prompt modification recovers +0.314 [0.254, 0.375] judge accuracy on category-2 (temporal) questions in the matched set. Code and results: https://github.com/kyrkewood/sleeping-agent.
Sources
- MemMachine: A Ground-Truth-Preserving Memory System for Personalized AI Agents
- Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory
- Learning to Forget: Sleep-Inspired Memory Consolidation for Resolving Proactive Interference in Large Language Models
- Adaptive Focus Memory for Language Models
- LightMem: Lightweight and Efficient Memory-Augmented Generation
- ACON: Optimizing Context Compression for Long-horizon LLM Agents
- ENGRAM: Effective, Lightweight Memory Orchestration for Conversational Agents
- Unable to Forget: Proactive Interference Reveals Working Memory Limits in LLMs Beyond Context Length
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection