Benign Alone, Harmful Together: Exploiting Experience Composition in Self-Evolving LLM Agents
Bingyu Yan, Xiaoming Zhang, Chaozhuo Li, Ziyi Zhou, Yirui Qi, Litian Zhang
cs.CR
Submitted: 2026-08-03
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- A Comprehensive Survey of Self-Evolving AI Agents: A New Paradigm Bridging Foundation Models and Lifelong Agentic Systems
- The Llama 3 Herd of Models
- Controlled Self-Evolution for Algorithmic Code Optimization
- Understanding the planning of LLM agents: A survey
- SE-Agent: Self-Evolution Trajectory Optimization in Multi-Step Reasoning with LLM-Based Agents
- FlipAttack: Jailbreak LLMs via Flipping
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
- ReasoningBank: Scaling Agent Self-Evolving with Reasoning Memory
- MemoryGraft: Persistent Compromise of LLM Agents via Poisoned Experience Retrieval
- OEP: Poisoning Self-Evolving LLM Agents via Locally Correct but Non-Transferable Experiences
- Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory
- Beyond Self-Talk: A Communication-Centric Survey of LLM-Based Multi-Agent Systems
- Zombie Agents: Persistent Control of Self-Evolving LLM Agents via Self-Reinforcing Injections
- Universal and Transferable Adversarial Attacks on Aligned Language Models
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs