From Faulty Memories to Corrected Actions: Dependency-Guided Rollback Repair for Memory-Augmented Agents
Caili Yu, Yiqi Wang, Jiaqi Zhang, Yiqun Duan, Mingkai Zheng, Zhangkai Wu, Kaize Shi, Taotao Cai
cs.AI
Submitted: 2026-08-11
Updated: 2026-08-12
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 95/100
The gist: Persistent memory in language-model agents enables cross-session reuse of preferences, observations, and experience, but it also makes errors durable: "a poisoned, stale, or misattributed record can
Terminology
Summary
Persistent memory in language-model agents enables cross-session reuse of preferences, observations, and experience, but it also makes errors durable: a poisoned, stale, or misattributed record can alter reasoning, tool use, answers, and subsequent memory writes.
Existing defenses mainly detect or delete suspicious memories, or revise the current response,
but deleting the source leaves already propagated claims, actions, and derived memories active, whereas resetting the store or replaying the full trace destroys benign state and repeats unnecessary computation.
The paper therefore formulates post-failure memory recovery: given a failed execution and diagnosed faulty memories, recover both the answer and persistent state while retaining unaffected work.
The authors propose dependency-guided rollback repair, which proceeds in five stages:
-
Dependency graph construction: Builds a typed memory-to-action graph from runtime provenance, with nodes for user inputs, execution steps, and memory records, and edges encoding initiate, cite, support, produce, delete, update, consolidate, supersede, and derive relations.
-
Fault provenance and affected subgraph tracing: From each diagnosed faulty memory, traces explicit downstream dependencies through propagation edges to form a raw affected candidate set, over-approximating truly unsupported state.
-
Independent-support checking:
Because reachability alone would over-invalidate, it preserves candidates whose complete content or decision has independent trusted support.
A candidate is preserved only if it has a sufficient evidentiary predecessor outside the faulty set and affected set, with admissible sources beingan explicit current-turn instruction, an active memory not diagnosed as faulty, a successful tool observation, or an unaffected validated claim.
-
Rule-guided rollback planning: A deterministic planner
deletes diagnosed faults, quarantines unsupported derived state, invalidates unsupported trace outputs, and selects the answer-relevant affected computation.
It identifies answer-relevant affected nodes as those that can still contribute to the regenerated final answer, then computes an execution closure for replay. -
Selective replay: The executor
reuses safe context and replays selected steps in trace order under the repaired store,
regenerating the final answer and producing a repaired memory store and execution trace.
The method is evaluated on a 150-case controlled benchmark spanning shopping, travel, and customer support domains, covering four memory fault types: poisoned, stale, wrong-user, and summary-drift. It is also evaluated on a 50-case trajectory-derived stress test adapted from LongMemEval-V2, which is predominantly multi-fault (45 of 50 cases contain 2–4 faults). All methods use GPT-4o and share the same tool environment, task inputs, faulty memory store, failed execution trace, and diagnosed faulty memory identifiers.
On the controlled benchmark, the method achieves 85.3% recovery versus 77.3% for the best competing recovery method, removes all diagnosed faulty memories, preserves all benign memories, and requires only selective replay with modest LLM-call cost.
Specifically, compared to LLM-judge repair, it achieves higher recovery while reducing replay ratio by 43.3% and LLM calls by 41.8%. Compared to AgentTrace-style, it improves recovery by 24.6 percentage points with better faulty memory removal.
On the adapted LongMemEval-V2 subset, it reaches 68.0% recovery versus 54.0% for the next best method, while also achieving the highest claim invalidation F1, 0.669 versus 0.603.
-
Rollback planner is most responsible for answer recovery: removing it reduces recovery from 85.3% to 71.3% and increases recurrence from 26.6% to 43.0%.
-
Support checking improves preservation and selectivity: removing it increases recovery slightly to 88.0% but lowers benign preservation from 100.0% to 98.6% and increases replay ratio by 25.2%.
-
Selective replay primarily controls cost: without it, replay ratio rises from 12.3% to 75.5% and LLM calls from 5.70 to 24.01, while recovery decreases slightly to 84.0%.
Per-domain results show customer support achieves the highest recovery (94.3%), while shopping is hardest (72.9%). Per-fault-type results show summary-drift has the highest recovery (95.0%), wrong-user faults are hardest (70.6%), and stale faults have high recovery (94.1%) but high recurrence (78.1%). Backbone sensitivity tests with Gemini-3.6-Flash and Qwen3.6-27B confirm the method maintains a stronger balance of recovery, recurrence, preservation, and cost across backbones.
The paper concludes: "Overall, the results do not imply uniformly better trace reconstruction, but show that dependency-guided rollback repair provides a strong recovery–cost trade-off while repairing faulty memory state and preserving benign memory. The authors note that
Remaining controlled-set gaps in recurrence highlight a key direction for future work: improving recurrence robustness and claim-state identification without sacrificing selective repair."
Improvements for AI systems
Improvements to AI Systems:
-
Implement a persistent-memory provenance tracker that logs typed dependency edges (initiate, cite, support, produce, delete, update, consolidate, supersede, derive) between user inputs, execution steps, and memory records in real time, enabling precise fault tracing and selective rollback rather than full-store resets.
-
Add a post-failure recovery module that, upon diagnosing faulty memories (poisoned, stale, wrong-user, summary-drift), automatically constructs an affected-subgraph, applies independent-support checking to preserve benign derived state, and replays only answer-relevant execution steps in trace order—reducing unnecessary computation and LLM calls by 40% while maintaining high recovery.
-
Integrate a rule-guided rollback planner that distinguishes between (a) deleting diagnosed faults, (b) quarantining unsupported derived memories, (c) invalidating unsupported trace outputs, and (d) selecting only the minimal execution closure needed to regenerate the final answer—this planner is the single largest contributor to answer recovery (85.3% vs. 71.3% without it).
-
Enhance memory systems with independent-support validation before preserving any candidate memory during recovery: only keep a memory if it has a sufficient evidentiary predecessor outside the faulty set (e.g., an explicit current-turn instruction, an active non-faulty memory, a successful tool observation, or an unaffected validated claim), preventing over-invalidation while maintaining 100% benign-memory preservation.
-
Add selective replay capability that reuses safe context and replays only the minimal trace subset under the repaired store, cutting replay ratio from 75.5% to 12.3% and LLM calls from 24.0 to 5.7, while still regenerating a correct final answer—enabling cost-efficient recovery in production agents.
What the improved AI system can do:
-
Self-heal persistent memory after failures without losing benign state, by surgically removing only faulty memories and their unsupported descendants, rather than resetting the entire store or replaying the full trace.
-
Recover correct answers from multi-fault scenarios (e.g., 2–4 simultaneous memory faults) with 68% recovery on realistic trajectory-derived cases, outperforming prior methods by 14 percentage points.
-
Operate cost-effectively in long-running agent deployments by minimizing replay steps and LLM invocations during recovery, making it feasible for high-frequency, memory-intensive applications like shopping assistants, travel planners, and customer support bots.
-
Maintain robustness across different backbone models (e.g., GPT-4o, Gemini-3.6-Flash, Qwen3.6-27B), ensuring consistent recovery–cost trade-offs without sacrificing benign memory preservation.
-
Provide explainable recovery actions by exposing the dependency graph, affected subgraph, and rollback plan, allowing operators to audit which memories were deleted, quarantined, or preserved and why.
Sources
- STALE: Can LLM Agents Know When Their Memories Are No Longer Valid?
- HaluMem: Evaluating Hallucinations in Memory Systems of Agents
- Memory Poisoning Attack and Defense on Memory Based LLM-Agents
- MemAudit: Post-hoc Auditing of Poisoned Agent Memory via Causal Attribution and Structural Anomaly Detection
- Voyager: An Open-Ended Embodied Agent with Large Language Models
- AgentTrace: Causal Graph Tracing for Root Cause Analysis in Deployed Multi-Agent Systems
- LongMemEval-V2: Evaluating Long-Term Agent Memory Toward Experienced Colleagues
- An Illusion of Progress? Assessing the Current State of Web Agents
- $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
- ReAct: Synergizing Reasoning and Acting in Language Models
- Useful Memories Become Faulty When Continuously Updated by LLMs
- MEMOREPAIR: Barrier-First Cascade Repair in Agentic Memory
- AMA-Bench: Evaluating Long-Horizon Memory for Agentic Applications
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection