EvoSkill Injection: Red-Teaming Autonomous Skill Generation and Evolution in Self-Evolving Agents
cs.AI, cs.CL
Submitted: 2026-08-31
Updated: 2026-08-31
Comments: Accepted to EMNLP 2026
License: http://creativecommons.org/licenses/by/4.0/
The gist: LLM-based agent systems increasingly adopt skill-based architectures to reduce repetitive reasoning costs and improve stable, efficient task execution.
Terminology
Abstract
LLM-based agent systems increasingly adopt skill-based architectures to reduce repetitive reasoning costs and improve stable, efficient task execution. Recent studies propose self-evolving agents that autonomously generate, refine, and reuse skills from past experiences to enable continuous capability evolution. However, autonomous skill evolution introduces a new attack surface in which malicious capabilities are generated, stored, and reused as legitimate skills. In this paper, we define EvoSkill Injection as a threat model targeting the autonomous skill generation and evolution pipeline of self-evolving agents. We further propose SARGE (Red-teaming Autonomous Skill Generation and Evolution in self-evolving agents), a red-teaming framework for evaluating this threat model through iterative generation, escalation, and reinforcement interactions. To support our framework, we construct EvoSkillBench, a benchmark dataset of malicious interaction trajectories for inducing malicious skill formation in self-evolving agents, and introduce EvoSkillSafetyBench, a post-attack benchmark for evaluating whether injected malicious skills are subsequently retrieved and activated as harmful behaviors. Our evaluation shows that SARGE induces malicious skill formation and that injected skills are persistently stored and repeatedly activated, highlighting the risk of persistent capability corruption.
Sources
- SafeDialBench: A Fine-Grained Safety Evaluation Benchmark for Large Language Models in Multi-Turn Dialogues with Diverse Jailbreak Attacks
- Evaluating Large Language Models Trained on Code
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- SkillTrojan: Backdoor Attacks on Skill-Based Agent Systems
- GPT-4o System Card
- SkillJect: Effectively Automating Skill-Based Prompt Injection for Skill-Enabled Agents
- SoK: Agentic Skills -- Beyond Tool Use in LLM Agents
- "Do Not Mention This to the User": Detecting and Understanding Malicious Agent Skills in the Wild
- Red Teaming the Mind of the Machine: A Systematic Evaluation of Prompt Injection and Jailbreak Vulnerabilities in LLMs
- Supply-Chain Poisoning Attacks Against LLM Coding Agent Skill Ecosystems
- BadSkill: Backdoor Attacks on Agent Skills via Model-in-Skill Poisoning
- BadRAG: Identifying Vulnerabilities in Retrieval Augmented Generation of Large Language Models
- AutoSkill: Experience-Driven Lifelong Learning via Skill Self-Evolution
- AgentTrap: Measuring Runtime Trust Failures in Third-Party Agent Skills
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection