When Experience Becomes Instruction: Trajectory Poisoning in Self-Evolving Agent Skill Systems
Jialuo Chen, Lingqi Jiang, Xinhao Deng, Xiaohu Du, Jianan Ma, Yunhao Feng, Yuqi Qing, Zhihao Yuan, Linkang Du, Jingyi Wang
cs.CR, cs.AI
Submitted: 2026-08-07
Updated: 2026-08-10
License: http://creativecommons.org/licenses/by/4.0/
The gist: Self-evolving skill (SES) systems distill agent trajectories into persistent skills, allowing untrusted experience to become trusted instruction.
Terminology
Abstract
Self-evolving skill (SES) systems distill agent trajectories into persistent skills, allowing untrusted experience to become trusted instruction. We introduce PoisonedEvolution, a trajectory-poisoning attack on this promotion process. Our skill-visible black-box attacker can inspect a target skill and contribute bounded evidence, but cannot observe private pools or evolution logic or edit the skill bank. Artifact poisoning requires Inclusion, Evolution Attribution, and Realization. Attribution is the distinctive bottleneck: the target behavior must appear causally useful, recurrent, and generalizable before promotion. We evaluate four representative security-effect families using inert canary specifications. At 10% attacker support, across six mainstream LLM evolvers in SkillClaw, PoisonedEvolution embeds target behaviors in 546/600 trials (91.0% SER). On the structurally different Trace2Skill pipeline at the same ratio, it embeds target behaviors in 369/600 trials (61.5% SER), demonstrating transfer across evolution architectures. In a representative controlled study, three consistent attacker records suffice in a 30-record batch, whereas a single record is much weaker. Ablations identify recurring support, causal framing, and domain-aligned encoding as the main determinants of success. These findings expose evidence promotion as a security boundary for self-evolving agents.
Sources
- SecAlign: Defending Against Prompt Injection with Preference Optimization
- Defeating Prompt Injections by Design
- Poise: Position-Aware One-Instruction Skill Injection for Silent Execution on LLM Agents
- Model Context Protocol Threat Modeling and Analyzing Vulnerabilities to Prompt Injection with Tool Poisoning
- MUSE-Autoskill: Self-Evolving Agents via Skill Creation, Memory, Management, and Evaluation
- SkillClaw: Let Skills Evolve Collectively with Agentic Evolver
- Trace2Skill: Distill Trajectory-Local Lessons into Transferable Agent Skills
- Supply-Chain Poisoning Attacks Against LLM Coding Agent Skill Ecosystems
- Skill-Inject: Measuring Agent Vulnerability to Skill File Attacks
- BadSkill: Backdoor Attacks on Agent Skills via Model-in-Skill Poisoning
- The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions
- Skill-SD: Skill-Conditioned Self-Distillation for Multi-turn LLM Agents
- SkillRL: Evolving Agents via Recursive Skill-Augmented Reinforcement Learning
- RouteGuard: Internal-Signal Detection of Skill Poisoning in LLM Agents
- Agent Skills for Large Language Models: Architecture, Acquisition, Security, and the Path Forward
- AutoSkill: Experience-Driven Lifelong Learning via Skill Self-Evolution
- CoEvoSkills: Self-Evolving Agent Skills via Co-Evolutionary Verification
- MemSkill: Learning and Evolving Memory Skills for Self-Evolving Agents
- Poison Once, Exploit Forever: Environment-Injected Memory Poisoning Attacks on Web Agents
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs