ColluSkill: Adversarial Cross-Skill Composition for Evading Agent Skill Scanners
Puyu Zeng, Simeng Qin, Jingzhi Li, Ju Jia, Zheli Liu, Xiaojun Jia
Nankai University · Northeastern University · University of Science and Technology Beijing · Southeast University · Nanyang Technological University
cs.CR, cs.AI
Submitted: 2026-08-10
Updated: 2026-08-11
Comments: 9 pages, 3 figures, 4 tables
Code: https://github.com/cisco-ai-defense/skillscanner
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 50/100
The gist: ColluSkill: Adversarial Cross-Skill Composition for Evading Agent Skill Scanners proposes a collusive multi-skill-chain attack framework and a corresponding defense.
Terminology
Summary
ColluSkill: Adversarial Cross-Skill Composition for Evading Agent Skill Scanners proposes a collusive multi-skill-chain attack framework and a corresponding defense. The paper identifies a blind spot in existing skill scanners: current defenses primarily inspect individual skills, including their instructions, permissions, dependencies, and code behaviors, which can leave risks arising from cross-skill composition insufficiently examined.
This creates a situation where multiple locally plausible skills may independently pass security scanning while collectively forming a harmful workflow during agent execution.
The proposed attack, ColluSkill, decomposes a complete malicious intent into several interdependent sub-payloads and embeds them into independently packaged skills.
The attack does not rely on any single malicious skill, but on the ordered composition of locally plausible behaviors through contextual dependencies, artifact passing, and execution handoffs.
ColluSkill uses LLM-based chain planning and scanner-feedback refinement to preserve chain-level attack semantics while iteratively reducing suspicious signals within individual sub-skills.
The method draws on Norbert Elias's concept of interdependence chains, where complete behavior is not simply the sum of isolated actions
but is formed through the dependence and order among different roles.
The chain is defined as a three-step ordered structure Cconcept = (Z, E), with Z = z1, z2, z3 and E = (z1 → z2), (z2 → z3). For each payload, ColluSkill instantiates a concrete chain Cinst = (p1, p2, p3, (p1 → p2), (p2 → p3)), where p1, p2, p3 are three ordered sub-payloads decomposed from the original payload.
The attack generation process involves three steps: inferring the payload plan, generating the skill plan, and refining with scanner feedback. In the refinement step, ColluSkill only rewrites the flagged sub-skills and keeps the passed sub-skills unchanged,
iterating until all sub-skills pass or the maximum iteration count is reached.
To defend against this threat, the paper proposes ChainGuard, a context-aware skill-chain scanner that jointly analyzes a candidate skill and the skills already installed in the agent environment.
ChainGuard reconstructs cross-skill dependencies, artifact flows, capability compositions, and potential downstream behaviors to identify risks that emerge only at the workflow level.
It evaluates three types of risk: standalone risk,
cross-skill dependency risk,
and capability-splitting risk.
The final decision uses OR aggregation, Dchain(W) = maxi DCG(si, Hi), meaning a multi-skill chain is detected if any candidate skill is flagged under the context available at its installation step.
ChainGuard is not given the true attack-chain membership and must identify possible cross-skill relations from the installed-skill context.
The experiments evaluate ColluSkill against six skill scanners: CISCO Skill Scanner, SkillFortify, Auditor, SlowMist, Vetter, and SkillSpector. The dataset consists of 200 multi-skill attack chains and 600 generated sub-skills in total,
built from 200 malicious payloads. Baselines include five single-skill attacks (SkillJect, Skill-Inject, SkillTrojan, POISE, SkillSafetyBench) and one multi-skill attack (SCRBench). The main metric is attack success rate (ASR), where an attack is counted as successful only when all three sub-skills pass the scanner.
The main results show that ColluSkill achieves an average ASR of 96.0% and the best attack performance among the evaluated methods.
This is much higher than existing single-skill attacks.
SCRBench achieves an average ASR of 34.8%, compared with 96.0% for ColluSkill, showing that multi-skill composition alone is not sufficient for reliable scanner evasion, and that LLM-based chain planning and iterative refinement are important to ColluSkill's performance.
ChainGuard reduces the ASR of ColluSkill to 22.5% while allowing 99.5% of benign workflows to pass.
The ablation study shows that Naive Split achieves an average ASR of 36.7%,
Chain Planning increases the average ASR to 68.2% even without scanner feedback,
and Full ColluSkill further raises the average ASR to 96.0% through scanner-feedback refinement.
The refinement iteration experiment shows that Before refinement, the ASR is only 64.9%. After the first refinement iteration, it quickly increases to 84.7%,
reaching 96.0% at iteration 7. The chain length experiment shows that The 3-step setting achieves the highest final average ASR
of 96.0%, compared to 93.6% for 2-step and 90.7% for 4-step, indicating a better balance between payload dispersion and workflow coherence.
The runtime chain activation study evaluates ColluSkill on OpenCode, Claude Code, and Codex with GPT-5.5, DeepSeek-V4-Pro, and GLM-5.2 as model backbones. ColluSkill achieves the highest and most stable activation rates on OpenCode, ranging from 87.5% to 92.5%. The rates range from 62.0% to 78.2% on Claude Code and from 58.5% to 72.0% on Codex.
This shows that cross-skill composition is not limited to scanner evasion and can appear during real agent execution.
The ChainGuard defense evaluation shows that the attack success rate is 69.0% without installed-skill context. After the context is added, ChainGuard reduces the ASR to 22.5%.
For benign workflows, the pass rate decreases only from 99.7% to 99.5%,
showing that ChainGuard does not reduce the attack success rate by simply blocking multi-skill workflows.
The paper concludes that ColluSkill achieves an average ASR of 96.0% and performs best among the evaluated baselines
across six representative skill scanners, and executes successfully on OpenCode, Claude Code, and Codex with different model backbones.
ChainGuard reduces the ASR to 22.5% while allowing 99.5% of benign workflows to pass.
The results highlight adversarial cross-skill composition as an important attack surface and motivate chain-level defenses.
Improvements for AI systems
Improvements to AI Systems:
- Context-Aware Skill Installation Scanners
-
Upgrade current agent skill scanners to analyze a candidate skill jointly with the already-installed skill set, reconstructing cross-skill dependencies, artifact flows, and execution handoffs—not just the skill’s isolated code/instructions.
-
The improved system can detect “emergent malicious workflows” that arise only when multiple benign-looking skills are composed in a specific order, closing the blind spot identified by ColluSkill.
- Chain-Level Risk Scoring with OR-Aggregation
-
Implement a defense that computes three risk scores per candidate skill: standalone risk, cross-skill dependency risk, and capability-splitting risk. Use OR-aggregation so that a chain is flagged if any skill in the chain is risky under its installation context.
-
The improved system can block multi-step attacks even when no single skill is malicious, while maintaining a high pass rate (99.5%) for benign workflows, as demonstrated by ChainGuard.
- Iterative Adversarial Refinement for Attack Generation (Red-Teaming)
-
Build an AI red-team tool that uses LLM-based chain planning plus scanner-feedback refinement to automatically decompose a malicious intent into ordered sub-payloads, then iteratively rewrite only the flagged sub-skills until all pass.
-
The improved system can generate diverse, stealthy multi-skill attack chains (96.0% ASR) for stress-testing agent defenses, and can adapt to any scanner’s feedback to minimize suspicious signals.
- Runtime Chain Activation Monitoring
-
Add a runtime monitor that tracks not just individual skill execution but the ordered composition of skills during agent operation, looking for contextual dependencies and artifact passing that match known attack-chain patterns.
-
The improved system can detect and halt collusive multi-skill attacks in real time on platforms like OpenCode, Claude Code, and Codex, where activation rates currently reach up to 92.5%.
- Benign Workflow Preservation via Contextual Filtering
-
Use ChainGuard’s approach to distinguish between malicious chains and benign multi-step workflows by leveraging installed-skill context, rather than blocking all multi-skill sequences.
-
The improved system can reduce attack success from 96.0% to 22.5% while allowing 99.5% of legitimate workflows to pass, avoiding false positives that would degrade user experience.
- Adaptive Chain-Length Optimization for Attack/Defense Balance
-
For defenders, analyze the optimal chain length (e.g., 3-step) that maximizes attack coherence vs. dispersion; for defenders, use this insight to set detection thresholds that are robust across 2-, 3-, and 4-step compositions.
-
The improved system can anticipate that attackers will prefer 3-step chains (96.0% ASR vs. 93.6% for 2-step, 90.7% for 4-step) and can tune its context-window size and dependency-tracking depth accordingly.
- Scanner-Feedback Loop for Continuous Defense Hardening
-
Integrate a feedback loop where the defense system, after flagging a chain, provides the specific sub-skill and context that triggered the flag, enabling automated re-training or rule updates.
-
The improved system can continuously evolve its detection rules against new ColluSkill-style attacks, since the attack itself uses scanner feedback to refine—so the defense must also adapt iteratively.
Abstract
Agent skills are emerging as an important attack surface in LLM-based agent systems. Through an empirical study of existing skill scanners, we find that current defenses mainly inspect individual skills, leaving risks from cross-skill composition insufficiently examined. This creates a practical blind spot: multiple locally plausible skills may pass security checks while collectively forming a harmful workflow during agent execution. To investigate this threat, we propose ColluSkill, a collusive multi-skill-chain attack framework that decomposes a complete malicious intent into interdependent sub-payloads embedded in independently packaged skills. The attack does not rely on any single malicious skill, but emerges from the ordered composition of locally plausible behaviors through contextual dependencies, artifact passing, and execution handoffs. ColluSkill further employs LLM-based chain planning and scanner-feedback refinement to preserve chain-level attack semantics while reducing suspicious signals in individual sub-skills. To defend against such attacks, we propose ChainGuard, a context-aware skill-chain scanner that jointly analyzes a candidate skill and the skills already installed in the agent environment. ChainGuard reconstructs cross-skill dependencies, artifact flows, capability compositions, and downstream behaviors to identify risks that emerge only at the workflow level. Experiments on six representative skill scanners show that ColluSkill achieves an average attack success rate of 96.0% and consistently outperforms the evaluated single-skill and multi-skill attack baselines. Meanwhile, ChainGuard reduces the attack success rate to 22.5% while allowing 99.5% of benign workflows to pass, highlighting the importance of chain-level security analysis for agent skill ecosystems.
Sources
- Technical Report: Exploring the Emerging Threats of the Agent Skill Ecosystem
- Formal Analysis and Supply Chain Security for Agentic AI Skills
- SkillTrojan: Backdoor Attacks on Skill-Based Agent Systems
- MalSkillBench: A Runtime-Verified Benchmark of Malicious Agent Skills
- SkillProbe: Security Auditing for Emerging Agent Skill Marketplaces via Multi-Agent Collaboration
- Poise: Position-Aware One-Instruction Skill Injection for Silent Execution on LLM Agents
- Context Matters: Repository-Aware Security Analysis of the Agent Skill Ecosystem
- SkillSieve: A Hierarchical Triage Framework for Detecting Malicious AI Agent Skills
- Cloak and Detonate: Scanner Evasion and Dynamic Detection of Agent Skill Malware
- Seeing Is Not Screening: Multimodal Hidden Instruction Attacks on Agent Skill Scanners
- Improved Techniques for Optimization-Based Jailbreaking on Large Language Models
- SoK: Agentic Skills -- Beyond Tool Use in LLM Agents
- SkillSafetyBench: Evaluating Agent Safety under Skill-Facing Attack Surfaces
- SkillMutator: Benchmarking and Defending Language-and-Code Cross-modal Attacks on LLM Agent Skills
- Organizing, Orchestrating, and Benchmarking Agent Skills at Ecosystem Scale
- When Single-Agent with Skills Replace Multi-Agent Systems and When They Fail
- SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks
- Towards Secure Agent Skills: Architecture, Threat Taxonomy, and Security Analysis
- From Skill Text to Skill Structure: The Scheduling-Structural-Logical Representation for Agent Skills
- Agent Skills: A Data-Driven Analysis of Claude Skills for Extending Large Language Model Functionality
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs