SkillScope: Toward Fine-Grained Least-Privilege Enforcement for Agent Skills
cs.CR
Submitted: 2026-05-07
Updated: 2026-09-10
Comments: 21 pages, 7 figures
Code: https://github.com/alice-dot-io/caterpillar
License: http://creativecommons.org/licenses/by/4.0/
The gist: Agent Skills have become a practical way to extend LLM agents by packaging metadata, natural-language instructions, and executable resources into reusable capability bundles.
Terminology
Abstract
Agent Skills have become a practical way to extend LLM agents by packaging metadata, natural-language instructions, and executable resources into reusable capability bundles. However, this growing Skill ecosystem introduces a new compliance risk: a Skill may perform high-impact actions that fall outside the scope permitted by the user's current request, thereby violating least privilege. Existing skill detection approaches are insufficient for this problem because it is inherently task-conditioned: the same action may be legitimate under one user prompt but over-privileged under another. In this paper, we present SkillScope, a framework for fine-grained least-privilege enforcement in Agent Skills. SkillScope adopts a graph-based analysis approach that models instruction-level procedures and code-level operations as fine-grained action nodes. It extracts potential over-privilege candidates, validates them under graph-instantiated user tasks through runtime analysis, and constrains validated over-privileged actions via control-flow privilege constraining. We evaluate SkillScope through effectiveness experiments and large-scale real-world measurement. SkillScope achieves a 94.53% skill-level F1 score for over-privilege detection. In the wild, SkillScope validates 6,590 of 68,312 valid real-world Skills as exhibiting over-privileged behaviors, showing that least-privilege violations are prevalent in current Skill ecosystems. In the privilege-constraining evaluation, SkillScope reduces triggered over-privileged action-in-task instances by 88.56% while preserving legitimate task completion.
Sources
- Encrypted Prompt: Securing LLM Applications Against Unauthorized Actions
- AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents
- MasterKey: Automated Jailbreak Across Multiple Large Language Model Chatbots
- The Philosopher's Stone: Trojaning Plugins of Large Language Models
- SkillAttack: Automated Red Teaming of Agent Skills through Attack Path Refinement
- SkillProbe: Security Auditing for Emerging Agent Skill Marketplaces via Multi-Agent Collaboration
- SkillJect: Effectively Automating Skill-Based Prompt Injection for Skill-Enabled Agents
- Prompt Flow Integrity to Prevent Privilege Escalation in LLM Agents
- DRIFT: Dynamic Rule-Based Defense with Injection Isolation for Securing LLM Agents
- AgentRaft: Automated Detection of Data Over-Exposure in LLM Agents
- "Do Not Mention This to the User": Detecting and Understanding Malicious Agent Skills in the Wild
- Prompt Injection attack against LLM-integrated Applications
- Agent Skills in the Wild: An Empirical Study of Security Vulnerabilities at Scale
- Agent Skills Enable a New Class of Realistic and Trivially Simple Prompt Injections
- Skill-Inject: Measuring Agent Vulnerability to Skill File Attacks
- Progent: Securing AI Agents with Privilege Control
- "Elementary, My Dear Watson." Detecting Malicious Skills via Neuro-Symbolic Reasoning across Heterogeneous Artifacts
- ChainFuzzer: Greybox Fuzzing for Workflow-Level Multi-Tool Vulnerabilities in LLM Agents
- MANTRA: Enhancing Automated Method-Level Refactoring with Contextual RAG and Multi-Agent LLM Collaboration
- Agent Security Bench (ASB): Formalizing and Benchmarking Attacks and Defenses in LLM-based Agents
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs