Defense-as-Skill: Evolving Runtime Guard Skill for Skill-Augmented Agents
cs.CR, cs.AI
Submitted: 2026-09-01
Updated: 2026-09-01
Code: https://github.com/pinchbench/skill
License: http://creativecommons.org/licenses/by/4.0/
The gist: Skill-augmented agents load reusable skills as persistent runtime context, improving task performance but also giving malicious skills a durable channel for steering future actions.
Terminology
Abstract
Skill-augmented agents load reusable skills as persistent runtime context, improving task performance but also giving malicious skills a durable channel for steering future actions. Such skills may leak secrets, corrupt code, bypass approvals, or stage data for exfiltration only after a concrete user task and workspace state make the unsafe action appear useful. This makes pre-install vetting insufficient and calls for runtime, task-conditioned protection. We propose Defense-as-Skill, a defense paradigm that implements the runtime guard itself as an installable, inspectable, and editable skill. Our guard, SkillSonar, runs alongside untrusted task skills and checks sensitive actions against the user's task boundary, routing each action to an allow, replan, or confirmation decision without modifying the underlying agent runtime. To study this setting, we construct SCOPE-R, a task-conditioned dataset covering 6 risk families and 21 sub-categories, with 206 attack-confirmed malicious instances and 43 benign tasks. We then improve SkillSonar on the SCOPE-R training subset using runtime guard-skill evolution, a Monte-Carlo Tree Search procedure that evolves the on-disk guard skill from feedback on the rollouts. Across Claude Code and OpenClaw, the evolved guard substantially reduces attack success while maintaining a favorable safety-utility trade-off. On repeated GLM-5 runs, SkillSonar reduces ID ASR from 0.482 to 0.104 and OOD ASR from 0.606 to 0.115. Further analyses demonstrate transfer across victim models, held-out risk families, and external benchmarks, as well as retained protection against adaptive attackers. Ablations further show that explicit safety responsibility assignment and the skill-native representation are both important to the observed gains.
Sources
- WildClawBench: A Benchmark for Real-World, Long-Horizon Agent Evaluation
- SkillAttack: Automated Red Teaming of Agent Skills through Attack Path Refinement
- SkillSafetyBench: Evaluating Agent Safety under Skill-Facing Attack Surfaces
- Agent Skill Acquisition for Large Language Models via CycleQD
- SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks
- "Do Not Mention This to the User": Detecting and Understanding Malicious Agent Skills in the Wild
- Agent Skills in the Wild: An Empirical Study of Security Vulnerabilities at Scale
- SkillClaw: Let Skills Evolve Collectively with Agentic Evolver
- AgentSpec: Customizable Runtime Enforcement for Safe and Reliable LLM Agents
- SkillRL: Evolving Agents via Recursive Skill-Augmented Reinforcement Learning
- SkillPyramid: A Hierarchical Skill Consolidation Framework for Self-Evolving Agents
- Agent Skills for Large Language Models: Architecture, Acquisition, Security, and the Path Forward
- GLM-5: from Vibe Coding to Agentic Engineering
- CoEvoSkills: Self-Evolving Agent Skills via Co-Evolutionary Verification
- SkillFlow:Benchmarking Lifelong Skill Discovery and Evolution for Autonomous Agents
- SkillLearnBench: Benchmarking Continual Learning Methods for Agent Skill Generation on Real-World Tasks
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs