AI Guardrail Survival under Single-Cycle Agentic Self-Summarization

arXiv:2608.11392 · cs.CR, cs.AI · Submitted 2026-08-13 · Read on arXiv

Ted Kwartler, Alan Aqrawi, Arian Abbasi

Harvard University · Accenture

cs.CR, cs.AI

Submitted: 2026-08-13

Updated: 2026-08-14

Code: https://github.com/alanaqrawi/guardrail-survivalunder-compaction

License: http://creativecommons.org/licenses/by-nc-sa/4.0/

Importance score: 75/100

The gist: This paper investigates how a standing safety rule is lost during a single-cycle agentic self-summarization (context compaction), where an agent's interaction history is replaced with a

Terminology

Summary

This paper investigates how a standing safety rule is lost during a single-cycle agentic self-summarization (context compaction), where an agent's interaction history is replaced with a model-generated summary. The study is motivated by a widely-circulated 2026 incident in which an autonomous email agent (OpenClaw) deleted over 200 emails after its history was compacted mid-task, losing an explicit do not act until I confirm instruction, and by concurrent work (Governance Decay; Chen, 2026) showing that dropped constraints drive behavioral violations across seven model families.

The paper asks four research questions: RQ1 (rule survival morphology), RQ2 (rule-versus-fact retention), RQ3 (whether degraded textual residues still protect behavior), and RQ4 (evaluation reliability of LLM judges).

Central finding: a presence check is not a safety check. When compaction does not drop a rule outright, it often leaves something that looks like a rule but does not act like one. On behavioral replay, a degraded residue leads the model to perform the prohibited action far more often than an intact welded rule does: all-case gaps of +34 and +57 points under two replay models, both positive (Qwen and Llama respectively), and +50 and +56 among the cases with replay headroom. Category-level survival behaves like a degraded residue, and even textually intact rules sometimes fail to fire, so an audit that checks only textual presence gives false assurance.

Rule-versus-fact retention (RQ2): Rule-form items are retained substantially more often than prominence-matched facts: β ≈ 2.5 in the canonical Qwen run (unmarked condition: 27% vs 4% survival), which reproduces on a second summarizer from a different provider and model family (Claude under a hard output cap, β ≈ 2.29). The paper notes this is exactly why presence-based checking feels adequate even though survival is not protection. However, the authors are careful to report this as descriptive: we do not claim that this advantage is independent of salience: the test that would have separated normativity from prominence did not yield a stable result (the marked condition was inconclusive).

Morphology of single-cycle loss (RQ1): The textual form of loss is regime-dependent (weld-or-drop with a single rule; degraded predicate-loss residues under a tighter budget). The hypothesized textual severing mode—a rule staying present while its referent is silently generalized off-target—was not observed: across the between-items runs the referent was generalized ten times and none was non-covering (8 clearly covering, 2 unresolved under presence-not-inference). The stress probe had G = 0 (no referent generalizations at all), offering no opportunity for this mode. The paper reports this as not observed rather than impossible. Additionally, one open model (Llama) sometimes declines the summarization task outright, returning a refusal or meta-description instead of a summary.

Model-dependence: the summarizers that lost rules to input volume were the tested open models (Qwen and Llama), while Claude resisted until forced by a hard output cap. Claude's robustness comes from selecting far less aggressively, retaining most content even though its output is still much shorter than the input. Under a hard 150-token output cap, Claude does drop a minority of rules (38%) but still favors them over facts.

Evaluation pitfalls (RQ4): Two documented cases where LLM-judge labels alone would have reversed a conclusion. Most consequentially, an over-permissive rubric clause inflated marked epistemic survival from 3 to 12 items, which would have made the marked condition appear comfortably analyzable. Author adjudication (blinded to judge labels) and behavioral replay caught both errors. The paper reports this as a concrete caution for the growing practice of judge-only safety evaluation.

Practical implication: Textual loss is silent at runtime and detectable only by comparison with retained external ground truth (such as a constraint registry), which reveals textual absence but not whether a surviving rule still fires. The paper proposes a post-compaction audit against an external constraint registry as a motivated hypothesis, not a validated recommendation, noting that a registry alone guarantees none of retrieval, application, or enforcement.

Scope and limitations: All results concern a single compaction cycle (matching the motivating incident). The primary rule-fact estimate comes from Qwen; the Claude cell is corroborating but pressured differently (output cap vs input volume). The paper flags several limitations: the normativity-versus-salience distinction remains unresolved; the synthetic setting uses one constraint family (standing prohibitions); author adjudication concentrated on decisive cases rather than a preregistered random sample; the salience rater and judge share a model family with a summarizer; and each summary/replay is a single stochastic realization.

Improvements for AI systems

Improvement 1: Post-Compaction Constraint Audit Module

The improved AI system includes a mandatory audit step after any context compaction or summarization event. It compares the generated summary against an external, immutable constraint registry (e.g., a list of standing prohibitions like do not act until I confirm). The system flags any rule that is absent, textually degraded, or present but behaviorally non-firing. It then either (a) refuses to proceed with the task until a human re-confirms the rule, or (b) automatically re-injects the original rule text from the registry into the active context. This prevents silent loss of safety constraints during mid-task history compression.

Improvement 2: Behavioral Replay Validation for Surviving Rules

The system does not trust textual presence of a rule after compaction. Instead, it runs a lightweight behavioral probe: it replays a short, simulated action sequence that would violate the rule (e.g., send email to all contacts) and checks whether the model actually withholds the action. If the probe shows the rule does not fire (i.e., the model proceeds with the prohibited action), the system treats the rule as lost and triggers the audit module from Improvement 1. This converts presence checks into safety checks, closing the +34 to +57 point behavioral gap observed in the paper.

Improvement 3: Rule-Priority Compaction with Hard Guarantees

The summarizer is modified to treat rule-form items (standing prohibitions, safety constraints) as non-compressible tokens. During summarization, it reserves a fixed output budget for rules before compressing any factual content. If the output cap is too tight to include all rules, the system either (a) refuses to compact and instead truncates less critical facts, or (b) escalates to a human for manual rule selection. This mimics Claude's observed robustness (retaining rules under pressure) but makes it a hard architectural guarantee rather than a model-dependent behavior.

Improvement 4: Judge-Rubric Sanity Check with Behavioral Ground Truth

When using LLM judges to evaluate rule survival post-compaction, the system automatically cross-validates judge labels against behavioral replay results on a small, stratified sample (e.g., 10% of cases). If judge labels disagree with behavioral outcomes (e.g., judge says rule present but replay shows violation), the system flags the rubric as over-permissive or under-permissive and recalibrates the judge prompt. This prevents the documented failure where an over-permissive rubric inflated survival rates from 3 to 12 items, which would have reversed conclusions.

Improved AI System Capabilities:

  • Self-auditing after every compaction: The system can detect and recover lost or degraded safety rules without human intervention, using an external registry and behavioral probes.

  • Guaranteed rule retention under memory pressure: The system will never drop a standing prohibition due to token limits; it will sacrifice factual details or halt compaction instead.

  • Reliable safety evaluation: The system's own evaluation of rule survival is behaviorally validated, so it will not falsely report a rule as intact when it does not actually constrain actions.

  • Resilience to model-dependent summarization weaknesses: Even if the underlying summarizer is an open model prone to rule loss (like Qwen or Llama), the audit and replay layers compensate, making the overall system robust across model families.

Abstract

Long-running agents periodically compact their context, replacing the transcript with a model-generated summary. Recent work shows that dropping a standing safety constraint during compaction drives behavioral violations across many models (Governance Decay; Chen, 2026). We ask a finer question: under a single compaction cycle, how is a safety rule lost, and what does that imply for detection and evaluation? Our central finding is that a presence check is not a safety check: when compaction does not drop a rule outright, it often leaves something that looks like a rule but does not act like one. On behavioral replay, a degraded residue leads the model to perform the prohibited action far more often than an intact welded rule does (all-case gaps of +34 and +57 points under two replay models, both positive), category-level survival behaves like a residue, and even intact rules sometimes fail to fire, so an audit that checks only textual presence gives false assurance. Sharpening this, rule-form items are retained substantially more often than prominence-matched facts, which is exactly why presence-based checking feels adequate even though survival is not protection. Textual loss is regime-dependent (weld-or-drop with a single rule; degraded predicate-loss residues under a tighter budget), and we did not observe the hypothesized textual severing mode. Such loss is silent at runtime and detectable only by comparison with retained external ground truth (such as a constraint registry), which reveals textual absence but not whether a surviving rule still fires. We also document evaluation pitfalls where LLM-judge labels alone would have reversed a conclusion. All results concern a single compaction cycle.

Sources

Related papers