When Errors Become Memories: Causal Pathway Tracing in Multi-Turn Memory-Augmented LLMs
cs.CL
Submitted: 2026-08-31
Updated: 2026-08-31
License: http://creativecommons.org/licenses/by/4.0/
The gist: Long-term memory enables large language models (LLMs) to preserve and reuse information across interactions, but it can also turn localized errors into persistent risks.
Terminology
Abstract
Long-term memory enables large language models (LLMs) to preserve and reuse information across interactions, but it can also turn localized errors into persistent risks. Existing work mainly evaluates whether memory systems store and retrieve information correctly, leaving limited understanding of how errors propagate across responses, memory states, and future interactions. We propose a structural causal model (SCM)-based framework for cross-turn error propagation in memory-augmented LLMs. We model user questions, model responses, and memory states as a dynamic causal process, and identify two entry pathways: internal memory updating and external question feedback. By intervening on these pathways, we construct four counterfactual trajectories and quantify their downstream effects and interaction. Error influence is evaluated at four levels: memory retention, natural responses, targeted diagnostic probing, and probability-level error preference. Experiments show that error influence generally decays with interaction distance, while the memory-update pathway contributes more persistent effects than question feedback; latent errors may remain even after disappearing from natural responses. Propagation patterns also vary across memory categories and memory mechanisms. Pathway-guided restoration further validates this decomposition: Question Repair reduces residual error by 27.5%, Memory Repair by 70.2%, and Joint Repair by 98.3%, nearly eliminating residual propagation.
Sources
- From Untrusted Input to Trusted Memory: A Systematic Study of Memory Poisoning Attacks in LLM Agents
- MemTrace: Tracing and Attributing Errors in Large Language Model Memory Systems
- GLM-5: from Vibe Coding to Agentic Engineering
- Evaluating Memory in LLM Agents via Incremental Multi-Turn Interactions
- LLMs Get Lost In Multi-Turn Conversation
- Olmo 3
- Qwen2.5 Technical Report
- Causal Inference with a Graphical Hierarchy of Interventions
- LongMemEval: Benchmarking Chat Assistants on Long-Term Interactive Memory
- How Memory Management Impacts LLM Agents: An Empirical Study of Experience-Following Behavior
- Useful Memories Become Faulty When Continuously Updated by LLMs
- Towards Best Practices of Activation Patching in Language Models: Metrics and Methods
- How Language Model Hallucinations Can Snowball
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering