SkillSafetyBench: Evaluating Agent Safety under Skill-Facing Attack Surfaces
cs.CR, cs.AI, cs.CL, cs.LG, cs.MA
Submitted: 2026-05-12
Updated: 2026-08-28
Code: https://github.com/AI45Lab/skill-safety-bench
Project page: https://jinchang1223.github.io/skill-safety-bench-website
Terminology
Sources
- SkillAttack: Automated Red Teaming of Agent Skills through Attack Path Refinement
- SkillTrojan: Backdoor Attacks on Skill-Based Agent Systems
- SkillSieve: A Hierarchical Triage Framework for Detecting Malicious AI Agent Skills
- SoK: Agentic Skills -- Beyond Tool Use in LLM Agents
- SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks
- Towards Secure Agent Skills: Architecture, Threat Taxonomy, and Security Analysis
- ConfusedPilot: Confused Deputy Risks in RAG-based LLMs
- Agent Skills: A Data-Driven Analysis of Claude Skills for Extending Large Language Model Functionality
- Skill-Inject: Measuring Agent Vulnerability to Skill File Attacks
- "Do Not Mention This to the User": Detecting and Understanding Malicious Agent Skills in the Wild
- OpenAI GPT-5 System Card
- Prompt Injection attack against LLM-integrated Applications
- Kimi K2.5: Visual Agentic Intelligence
- Agent Skills in the Wild: An Empirical Study of Security Vulnerabilities at Scale
- BadSkill: Backdoor Attacks on Agent Skills via Model-in-Skill Poisoning
- SkillX: Automatically Constructing Skill Knowledge Bases for Agents
- Voyager: An Open-Ended Embodied Agent with Large Language Models
- SkillWeaver: Web Agents can Self-Improve by Discovering and Honing Skills
- Inducing Programmatic Skills for Agentic Tasks
- Agent Skills for Large Language Models: Architecture, Acquisition, Security, and the Path Forward
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs