Retrieval Observability Bounds on Provenance Detection for Agent Memory Poisoning: Measured Coverage and a Falsified Standalone Detector
cs.CR, cs.LG
Submitted: 2026-06-29
Updated: 2026-09-27
Terminology
Sources
- Enforcing Benign Trajectories: A Behavioral Firewall for Structured-Workflow AI Agents
- From Untrusted Input to Trusted Memory: A Systematic Study of Memory Poisoning Attacks in LLM Agents
- PerD: Perturbation Sensitivity-based Neural Trojan Detection Framework on NLP Applications
- MEMSAD: Gradient-Coupled Anomaly Detection for Memory Poisoning in Retrieval-Augmented Agents
- Injection-Execution Dissociation: A Mechanistic Evaluation of Persistent Memory Attacks and Defenses in Stateful LLM Agents
- VIGIL: Runtime Enforcement of Behavioral Specifications in AI Agent Skills
- TraceAegis: Securing LLM-Based Agents via Hierarchical and Behavioral Anomaly Detection
- MemLineage: Lineage-Guided Enforcement for LLM Agent Memory
- Hidden in Memory: Sleeper Memory Poisoning in LLM Agents
- What If Prompt Injection Never Left? Rethinking Agent Security through Cross-Session Stored Prompt Injection
- Zombie Agents: Persistent Control of Self-Evolving LLM Agents via Self-Reinforcing Injections
- MemMorph: Tool Hijacking in LLM Agents via Memory Poisoning
- Representation Engineering: A Top-Down Approach to AI Transparency
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs