CAPTURE: Disentangling Preference Drift from Memory Poisoning in Personalized LLM Agents
cs.LG
Submitted: 2026-09-02
Updated: 2026-09-02
Terminology
Sources
- Bayesian Online Changepoint Detection
- Constitutional AI: Harmlessness from AI Feedback
- StruQ: Defending Against Prompt Injection with Structured Queries
- AgentPoison: Red-teaming LLM Agents via Poisoning Memory or Knowledge Bases
- AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents
- Defeating Prompt Injections by Design
- Memory Injection Attacks on LLM Agents via Query-Only Interaction
- The Llama 3 Herd of Models
- Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection
- Towards Natural Personalization: Evaluating Long-Horizon Preference Following in Personalized User-LLM Interactions
- Defending Against Indirect Prompt Injection Attacks With Spotlighting
- HorizonBench: Long-Horizon Personalization with Evolving Preferences
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
- MemGPT: Towards LLMs as Operating Systems
- Jatmo: Prompt Injection Defense by Task-Specific Finetuning
- Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
- Safety Alignment Should Be Made More Than Just a Few Tokens Deep
- XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models
- A StrongREJECT for Empty Jailbreaks
- Unveiling Privacy Risks in LLM Agent Memory
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks