SkillGate: Cost Efficient Runtime Malicious Skill File Detection in Coding Agents
Rui Yang, Michael Fu, Kla Tantithamthavorn, Chetan Arora, Joey Chua
cs.SE, cs.CR
Submitted: 2026-07-28
Comments: 10 pages, 5 figures
Code: https://github.com/awsm-research/skillgate
License: http://creativecommons.org/licenses/by/4.0/
The gist: Software engineering teams now deploy AI coding agents (Cursor, Claude Code, GitHub Copilot) as first-class productivity tools, installing domain-specific skill files to tailor agent behavior to
Terminology
Abstract
Software engineering teams now deploy AI coding agents (Cursor, Claude Code, GitHub Copilot) as first-class productivity tools, installing domain-specific skill files to tailor agent behavior to project APIs, framework conventions, and organizational workflows. These complex Markdown files are easily downloaded from public registries with a single npx skills add command and no real security screening, representing a novel supply-chain attack surface: a malicious skill file can silently reprogram agent behavior, exfiltrating credentials, injecting backdoors into generated code, or redirecting agent actions to attacker-controlled endpoints. The threat is not hypothetical: recent reports document hundreds of malicious skill packages in public registries, including organized campaigns that distributed credential-stealing infostealers via fake productivity skills. No systematic toolchain defense exists for this attack surface. We present SkillGate, a deployable security gateway that screens AI skill packages before coding agent installation. SkillGate uses a hybrid regex-prefilter + LLM-judge pipeline: safe-signal files bypass the LLM entirely (skip savings); flagged files have only their matched snippet windows sent to the judge, not the full content (snippet savings). We answer four research questions covering detection effectiveness, screening cost, runtime overhead, and false positive behavior on the SkillsBench benchmark against two existing tools. On SkillsBench (n=1,650, 9.1% malicious), SkillGate achieves F1=0.817, FPR=1.13% while reducing LLM input tokens by 77% vs. full-file screening, and outperforming existing tools by 5-6x on threshold-independent AUPRC (0.830 vs. 0.144/0.162).
Sources
- Vibe Coding vs. Agentic Coding: Fundamentals and Practical Implications of Agentic AI
- Usage, Effects and Requirements for AI Coding Assistants in the Enterprise: An Empirical Study
- Evaluating Large Language Models Trained on Code
- Ignore Previous Prompt: Attack Techniques For Language Models
- "Do Not Mention This to the User": Detecting and Understanding Malicious Agent Skills in the Wild
- Skill-Inject: Measuring Agent Vulnerability to Skill File Attacks
- Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
- Toolformer: Language Models Can Teach Themselves to Use Tools
- Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
- ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs
Related papers
- Falsification-Based Verification of LLM-Generated Optimization Models: Sound Test Batteries and Their Detection Limits
- GitSkills: A Dataset of Agent Skills on GitHub
- SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces
- PackMonitor: Enabling Zero Package Hallucinations Through Decoding-Time Monitoring
- IntentCoding: Amplifying User Intent in Code Generation
- Incentives and Outcomes in Bug Bounties