The Sleeping Agent: What Gist-Based Context Compression Loses and Why

arXiv:2608.11775 · cs.AI, cs.CL · Submitted 2026-08-12 · Read on arXiv

Nicholas E. Kyrkewood

cs.AI, cs.CL

Submitted: 2026-08-12

Updated: 2026-08-13

Comments: 7 pages, 5 tables, appendices. Code and results at https://github.com/kyrkewood/sleeping-agent

Code: https://github.com/kyrkewood/sleeping-agent

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 75/100

The gist: The paper introduces Salience-Weighted Consolidation (SWC), a biologically-inspired compression framework motivated by sleep-based memory consolidation, used as a diagnostic probe to study when

Terminology

Summary

The paper introduces Salience-Weighted Consolidation (SWC), a biologically-inspired compression framework motivated by sleep-based memory consolidation, used as a diagnostic probe to study when gist-based context compression helps and when it hurts. SWC scores conversation history chunks by estimated salience, partitions them into priority tiers, and applies structured gist abstraction to mid-priority content. The framework draws three principles from sleep neuroscience: proactive scheduling, salience-weighted retention, and staged compression.

Evaluating four conditions—truncation, sliding window, SWC-Full, and SWC-Temporal—on all ten LoCoMo conversations (1,935 matched text-only questions total, 1,501 used in the primary aggregate after excluding Category 5 adversarial questions) at temperature 0, the paper finds a consistent task-type interaction: "gist compression substantially outperforms truncation on multi-hop reasoning and single-hop factual questions, but temporal questions remain substantially harder under compression, with compressed conditions scoring well below the full-context reference on the conversations where both are evaluated."

The key results are: SWC-Temporal achieves 0.468 aggregate judge accuracy on categories 1–4, compared to SWC-Full at 0.379, sliding window at 0.238, and truncation at 0.171. On category 2 (temporal) questions specifically, SWC-Full fails dramatically at 0.156 accuracy, while SWC-Temporal reaches 0.470. On multi-hop questions, both SWC conditions score 0.407–0.429 against baselines of 0.157–0.271; on single-hop questions, SWC conditions score 0.459–0.486 against 0.170–0.238.

The paper traces the temporal failure to a specific mechanism: the gist abstraction prompt preserves relational and event structure while discarding dates and times. A preservation analysis across all ten conversations confirms this: "an approximately 20-fold increase in temporal expression preservation (3.05%→62.39%) with a one-sentence prompt modification, while named entity and event preservation rates barely change (×1.02 and ×1.11), demonstrating that the fix is a precision instrument."

The prompt modification—adding one sentence to the gist abstraction prompt: You MUST preserve verbatim: all specific dates, times, durations, ages, and temporal expressions (e.g. '7 May', 'last Tuesday', 'at 9pm')—recovers +0.314 [0.254, 0.375] judge accuracy on category-2 (temporal) questions in the matched set, with the direction positive in all ten conversations (per-conversation deltas: +0.433, +0.160, +0.259, +0.379, +0.385, +0.291, +0.177, +0.341, +0.364, +0.226; unweighted mean +0.301).

The paper argues the failure mode is a correctable design omission rather than a fundamental limitation. The operative distinction is temporal versus non-temporal, not episodic versus semantic—SWC-Full scores 0.459 on single-hop factual questions (episodic content in the conventional sense), so it does not fail on episodic content generally; it fails specifically on temporal information because temporal anchors fall outside all protected categories by default.

The paper also reports that category 3 (open-domain) confidence intervals overlap completely across all conditions, so no reliable conclusion can be drawn for commonsense inference questions. Category 5 (adversarial) questions are excluded because their binary scoring rule rewards content removal and conflates memory quality with response calibration.

The conclusion states: Generic gist abstraction preserves relational, event, and factual content while selectively losing temporal anchors, because they are absent from standard prompt specifications. The paper recommends that any memory system using gist abstraction should explicitly protect temporal anchors and that temporal anchors should be treated as a protected category in any gist compression component. It also emphasizes that category-level evaluation paired with preservation rate analysis provides a more complete picture of what a memory strategy preserves—and what it leaves behind, since systems that achieve high aggregate scores may mask systematic failures on specific question types.

Improvements for AI systems

Improvements to AI Systems:

  1. Add explicit temporal-anchor protection to all gist-compression prompts.
  • The improved system will preserve verbatim dates, times, durations, ages, and temporal expressions (e.g., 7 May, last Tuesday, at 9pm) during context summarization, preventing the silent loss of time-critical information that current generic abstraction causes.
  1. Implement salience-weighted tiered memory consolidation with proactive scheduling.
  • The system will score conversation chunks by estimated salience, partition them into priority tiers, and apply staged compression only to mid-priority content—while keeping high-priority content verbatim and low-priority content aggressively abstracted. This mimics sleep-based consolidation, improving long-context reasoning without sacrificing critical details.
  1. Add a temporal-preservation audit layer to memory systems.
  • After any compression step, the system will run a preservation-rate check for temporal expressions (targeting >60% preservation, up from the observed 3.05% baseline). If preservation falls below threshold, it will automatically re-abstract with the protected-category prompt. This prevents silent temporal degradation in deployed systems.
  1. Introduce category-aware evaluation and adaptive compression selection.
  • The system will classify incoming questions into task types (temporal, multi-hop, single-hop factual, open-domain) and dynamically choose between full-context, temporal-protected gist, or sliding-window strategies based on predicted category. For temporal questions, it will default to full-context or SWC-Temporal; for multi-hop, it will use gist compression to reduce noise.
  1. Treat temporal anchors as a first-class protected category in all memory schemas.
  • The system will maintain a structured index of temporal facts separate from gist summaries, allowing fast retrieval of exact dates/times even when the main context is compressed. This enables accurate temporal reasoning without requiring full-context retention.
  1. Add per-category confidence reporting to prevent masked failures.
  • The improved system will output accuracy metrics broken down by question type (temporal, multi-hop, single-hop, open-domain) rather than a single aggregate score. This exposes systematic weaknesses (e.g., temporal accuracy dropping to 0.156) that aggregate numbers hide, enabling targeted retraining or fallback strategies.

What the improved AI system can do:

  • Answer temporal questions (e.g., When did X happen relative to Y?) with high accuracy even after long conversations, because dates and times are never lost during compression.

  • Maintain strong multi-hop reasoning performance (0.4+ accuracy) while using 70–80% less context memory, by applying gist abstraction only to non-temporal, mid-priority content.

  • Automatically detect when a query is temporal and switch to a full-context or temporal-protected mode, avoiding the 0.156 accuracy collapse seen in naive compression.

  • Provide transparent, per-category performance reports so users and developers know exactly which question types the system handles well and which it fails on—preventing overconfidence from aggregate scores.

  • Scale to arbitrarily long conversations without degrading temporal fidelity, because temporal anchors are stored in a dedicated, lossless index rather than being subject to generic summarization.

Abstract

Gist-based context compression---summarising older conversation history into compact representations---is a common approach in long-horizon language model agents, yet its effect on different types of memory retrieval is poorly understood. We use Salience-Weighted Consolidation (SWC), a biologically-inspired compression framework motivated by sleep-based memory consolidation, as a diagnostic probe to study when gist compression helps and when it hurts. SWC scores conversation history by salience, partitions it into priority tiers, and applies structured gist abstraction to mid-priority content. Evaluating four conditions on all ten LoCoMo conversations---1,935 matched text-only questions in total, 1,501 used in the primary aggregate after excluding Category 5 (adversarial) questions---at temperature 0, we find a consistent task-type interaction: gist compression substantially outperforms truncation on multi-hop reasoning and single-hop factual questions, but temporal questions remain substantially harder under compression, with compressed conditions scoring well below the full-context reference on the conversations where both are evaluated. We trace this failure to a specific mechanism: the gist abstraction prompt preserves relational and event structure while discarding dates and times. A preservation analysis across all ten conversations confirms the mechanism: an approximately 20-fold increase in temporal expression preservation (3.05% to 62.39%) with a one-sentence prompt modification, while named entity and event preservation rates barely change (x1.02 and x1.11), demonstrating that the fix is a precision instrument. The prompt modification recovers +0.314 [0.254, 0.375] judge accuracy on category-2 (temporal) questions in the matched set. Code and results: https://github.com/kyrkewood/sleeping-agent.

Sources

Related papers