SkillJack: Persistent Skill Backdoors in Self-Evolving Agents
Zonghao Ying, Xiangfan Wu, Huiyu Wu, Xing Zheng, Huangsheng Cheng, Xiaorong Shi, Jing Guo
cs.CR
Submitted: 2026-08-07
Updated: 2026-08-10
Code: https://github.com/Tencent/AI-Infra-Guard
License: http://creativecommons.org/licenses/by/4.0/
The gist: Self-evolving agents increasingly convert interaction histories into reusable skills that persist beyond individual tasks.
Terminology
Abstract
Self-evolving agents increasingly convert interaction histories into reusable skills that persist beyond individual tasks. While prior work studies memory and retrieval poisoning, such attacks only affect agents when poisoned records are retrieved as context. We uncover a new and more fundamental risk: poisoned experiences can be transformed by the agent itself into durable behavioral artifacts. We present SkillJack, the first attack that exploits the experience-to-skill pipeline of self-evolving agents. Instead of directly manipulating runtime context, SkillJack hijacks the agent's own learning process to implant malicious behaviors into its reusable skill repertoire. We identify three key properties of this transformation: sanitization whitewashing, where malicious intent is obscured during skill extraction; cross-layer promotion, where transient experiences become persistent capabilities; and persistence isolation, where the attack survives removal of its original source records. We evaluate SkillJack on two representative systems, SkillX and Anything2Skill, using a shared dataset of 150 trajectories across four policy-risk categories. Results show that skill extraction substantially reduces attack detectability: in SkillX, safety detection drops from 98.5% for poisoned trajectories to 11.4% for extracted skills, while Anything2Skill shows a similar effect. Meanwhile, the implanted skills remain effective, achieving attack success rates of 56.2% and 89.2% on the two systems, respectively. Furthermore, 80.0% of skill-mediated attacks persist after deleting the original poisoned records, and some skills unintentionally activate on benign queries. Our findings reveal skill evolution as a new attack surface and motivate provenance-aware skill lifecycle protection. Our code is available at https://github.com/Tencent/AI-Infra-Guard/research/skilljack.
Sources
- From Untrusted Input to Trusted Memory: A Systematic Study of Memory Poisoning Attacks in LLM Agents
- BadNets: Identifying Vulnerabilities in the Machine Learning Model Supply Chain
- Securing LLM-Agent Long-Term Memory Against Poisoning: Non-Malleable, Origin-Bound Authority with Machine-Checked Guarantees
- Skill-Pro: Learning Reusable Skills from Experience via Non-Parametric PPO for LLM Agents
- Trace2Skill: Distill Trajectory-Local Lessons into Transferable Agent Skills
- Anything2Skill: Compiling External Knowledge into Reusable Skills for Agents
- MemoryGraft: Persistent Compromise of LLM Agents via Poisoned Experience Retrieval
- SkillX: Automatically Constructing Skill Knowledge Bases for Agents
- OEP: Poisoning Self-Evolving LLM Agents via Locally Correct but Non-Transferable Experiences
- Agent Workflow Memory
- AutoSkill: Experience-Driven Lifelong Learning via Skill Self-Evolution
- AgentVisor: Defending LLM Agents Against Prompt Injection via Semantic Virtualization
- Poison Once, Exploit Forever: Environment-Injected Memory Poisoning Attacks on Web Agents
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs