Can Agents Trust Their Skills? Uncovering Unsafe Chains of Trust in Skill-Based LLM Agents
cs.CR, cs.AI
Submitted: 2026-09-30
Updated: 2026-09-30
Code: https://github.com/agent0ai/agent-zero
Project page: https://chenbihuan.github.io/paper/ccs18-chen-hawkeye.pdf
Terminology
Sources
- Design Patterns for Securing LLM Agents against Prompt Injections
- Defeating Prompt Injections by Design
- GitSkills: A Dataset of Agent Skills on GitHub
- SkillProbe: Security Auditing for Emerging Agent Skill Marketplaces via Multi-Agent Collaboration
- SkillJect: Effectively Automating Skill-Based Prompt Injection for Skill-Enabled Agents
- ClawHub Security Signals: When VirusTotal, Static Analysis, and SkillSpector Disagree
- SkCC: Portable and Secure Skill Compilation for Cross-Framework LLM Agents
- Skill-Inject: Measuring Agent Vulnerability to Skill File Attacks
- Prompt Injection Attack to Tool Selection in LLM Agents
- ReAct: Synergizing Reasoning and Acting in Language Models
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs