ColluSkill: Adversarial Cross-Skill Composition for Evading Agent Skill Scanners

arXiv:2608.09732 · cs.CR, cs.AI · Submitted 2026-08-10 · Read on arXiv

Puyu Zeng, Simeng Qin, Jingzhi Li, Ju Jia, Zheli Liu, Xiaojun Jia

Nankai University · Northeastern University · University of Science and Technology Beijing · Southeast University · Nanyang Technological University

cs.CR, cs.AI

Submitted: 2026-08-10

Updated: 2026-08-11

Comments: 9 pages, 3 figures, 4 tables

Code: https://github.com/cisco-ai-defense/skillscanner

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 50/100

The gist: ColluSkill: Adversarial Cross-Skill Composition for Evading Agent Skill Scanners proposes a collusive multi-skill-chain attack framework and a corresponding defense.

Terminology

Summary

ColluSkill: Adversarial Cross-Skill Composition for Evading Agent Skill Scanners proposes a collusive multi-skill-chain attack framework and a corresponding defense. The paper identifies a blind spot in existing skill scanners: current defenses primarily inspect individual skills, including their instructions, permissions, dependencies, and code behaviors, which can leave risks arising from cross-skill composition insufficiently examined. This creates a situation where multiple locally plausible skills may independently pass security scanning while collectively forming a harmful workflow during agent execution.

The proposed attack, ColluSkill, decomposes a complete malicious intent into several interdependent sub-payloads and embeds them into independently packaged skills. The attack does not rely on any single malicious skill, but on the ordered composition of locally plausible behaviors through contextual dependencies, artifact passing, and execution handoffs. ColluSkill uses LLM-based chain planning and scanner-feedback refinement to preserve chain-level attack semantics while iteratively reducing suspicious signals within individual sub-skills. The method draws on Norbert Elias's concept of interdependence chains, where complete behavior is not simply the sum of isolated actions but is formed through the dependence and order among different roles. The chain is defined as a three-step ordered structure Cconcept = (Z, E), with Z = z1, z2, z3 and E = (z1 → z2), (z2 → z3). For each payload, ColluSkill instantiates a concrete chain Cinst = (p1, p2, p3, (p1 → p2), (p2 → p3)), where p1, p2, p3 are three ordered sub-payloads decomposed from the original payload. The attack generation process involves three steps: inferring the payload plan, generating the skill plan, and refining with scanner feedback. In the refinement step, ColluSkill only rewrites the flagged sub-skills and keeps the passed sub-skills unchanged, iterating until all sub-skills pass or the maximum iteration count is reached.

To defend against this threat, the paper proposes ChainGuard, a context-aware skill-chain scanner that jointly analyzes a candidate skill and the skills already installed in the agent environment. ChainGuard reconstructs cross-skill dependencies, artifact flows, capability compositions, and potential downstream behaviors to identify risks that emerge only at the workflow level. It evaluates three types of risk: standalone risk, cross-skill dependency risk, and capability-splitting risk. The final decision uses OR aggregation, Dchain(W) = maxi DCG(si, Hi), meaning a multi-skill chain is detected if any candidate skill is flagged under the context available at its installation step. ChainGuard is not given the true attack-chain membership and must identify possible cross-skill relations from the installed-skill context.

The experiments evaluate ColluSkill against six skill scanners: CISCO Skill Scanner, SkillFortify, Auditor, SlowMist, Vetter, and SkillSpector. The dataset consists of 200 multi-skill attack chains and 600 generated sub-skills in total, built from 200 malicious payloads. Baselines include five single-skill attacks (SkillJect, Skill-Inject, SkillTrojan, POISE, SkillSafetyBench) and one multi-skill attack (SCRBench). The main metric is attack success rate (ASR), where an attack is counted as successful only when all three sub-skills pass the scanner.

The main results show that ColluSkill achieves an average ASR of 96.0% and the best attack performance among the evaluated methods. This is much higher than existing single-skill attacks. SCRBench achieves an average ASR of 34.8%, compared with 96.0% for ColluSkill, showing that multi-skill composition alone is not sufficient for reliable scanner evasion, and that LLM-based chain planning and iterative refinement are important to ColluSkill's performance. ChainGuard reduces the ASR of ColluSkill to 22.5% while allowing 99.5% of benign workflows to pass.

The ablation study shows that Naive Split achieves an average ASR of 36.7%, Chain Planning increases the average ASR to 68.2% even without scanner feedback, and Full ColluSkill further raises the average ASR to 96.0% through scanner-feedback refinement. The refinement iteration experiment shows that Before refinement, the ASR is only 64.9%. After the first refinement iteration, it quickly increases to 84.7%, reaching 96.0% at iteration 7. The chain length experiment shows that The 3-step setting achieves the highest final average ASR of 96.0%, compared to 93.6% for 2-step and 90.7% for 4-step, indicating a better balance between payload dispersion and workflow coherence.

The runtime chain activation study evaluates ColluSkill on OpenCode, Claude Code, and Codex with GPT-5.5, DeepSeek-V4-Pro, and GLM-5.2 as model backbones. ColluSkill achieves the highest and most stable activation rates on OpenCode, ranging from 87.5% to 92.5%. The rates range from 62.0% to 78.2% on Claude Code and from 58.5% to 72.0% on Codex. This shows that cross-skill composition is not limited to scanner evasion and can appear during real agent execution.

The ChainGuard defense evaluation shows that the attack success rate is 69.0% without installed-skill context. After the context is added, ChainGuard reduces the ASR to 22.5%. For benign workflows, the pass rate decreases only from 99.7% to 99.5%, showing that ChainGuard does not reduce the attack success rate by simply blocking multi-skill workflows.

The paper concludes that ColluSkill achieves an average ASR of 96.0% and performs best among the evaluated baselines across six representative skill scanners, and executes successfully on OpenCode, Claude Code, and Codex with different model backbones. ChainGuard reduces the ASR to 22.5% while allowing 99.5% of benign workflows to pass. The results highlight adversarial cross-skill composition as an important attack surface and motivate chain-level defenses.

Improvements for AI systems

Improvements to AI Systems:

  1. Context-Aware Skill Installation Scanners
  • Upgrade current agent skill scanners to analyze a candidate skill jointly with the already-installed skill set, reconstructing cross-skill dependencies, artifact flows, and execution handoffs—not just the skill’s isolated code/instructions.

  • The improved system can detect “emergent malicious workflows” that arise only when multiple benign-looking skills are composed in a specific order, closing the blind spot identified by ColluSkill.

  1. Chain-Level Risk Scoring with OR-Aggregation
  • Implement a defense that computes three risk scores per candidate skill: standalone risk, cross-skill dependency risk, and capability-splitting risk. Use OR-aggregation so that a chain is flagged if any skill in the chain is risky under its installation context.

  • The improved system can block multi-step attacks even when no single skill is malicious, while maintaining a high pass rate (99.5%) for benign workflows, as demonstrated by ChainGuard.

  1. Iterative Adversarial Refinement for Attack Generation (Red-Teaming)
  • Build an AI red-team tool that uses LLM-based chain planning plus scanner-feedback refinement to automatically decompose a malicious intent into ordered sub-payloads, then iteratively rewrite only the flagged sub-skills until all pass.

  • The improved system can generate diverse, stealthy multi-skill attack chains (96.0% ASR) for stress-testing agent defenses, and can adapt to any scanner’s feedback to minimize suspicious signals.

  1. Runtime Chain Activation Monitoring
  • Add a runtime monitor that tracks not just individual skill execution but the ordered composition of skills during agent operation, looking for contextual dependencies and artifact passing that match known attack-chain patterns.

  • The improved system can detect and halt collusive multi-skill attacks in real time on platforms like OpenCode, Claude Code, and Codex, where activation rates currently reach up to 92.5%.

  1. Benign Workflow Preservation via Contextual Filtering
  • Use ChainGuard’s approach to distinguish between malicious chains and benign multi-step workflows by leveraging installed-skill context, rather than blocking all multi-skill sequences.

  • The improved system can reduce attack success from 96.0% to 22.5% while allowing 99.5% of legitimate workflows to pass, avoiding false positives that would degrade user experience.

  1. Adaptive Chain-Length Optimization for Attack/Defense Balance
  • For defenders, analyze the optimal chain length (e.g., 3-step) that maximizes attack coherence vs. dispersion; for defenders, use this insight to set detection thresholds that are robust across 2-, 3-, and 4-step compositions.

  • The improved system can anticipate that attackers will prefer 3-step chains (96.0% ASR vs. 93.6% for 2-step, 90.7% for 4-step) and can tune its context-window size and dependency-tracking depth accordingly.

  1. Scanner-Feedback Loop for Continuous Defense Hardening
  • Integrate a feedback loop where the defense system, after flagging a chain, provides the specific sub-skill and context that triggered the flag, enabling automated re-training or rule updates.

  • The improved system can continuously evolve its detection rules against new ColluSkill-style attacks, since the attack itself uses scanner feedback to refine—so the defense must also adapt iteratively.

Abstract

Agent skills are emerging as an important attack surface in LLM-based agent systems. Through an empirical study of existing skill scanners, we find that current defenses mainly inspect individual skills, leaving risks from cross-skill composition insufficiently examined. This creates a practical blind spot: multiple locally plausible skills may pass security checks while collectively forming a harmful workflow during agent execution. To investigate this threat, we propose ColluSkill, a collusive multi-skill-chain attack framework that decomposes a complete malicious intent into interdependent sub-payloads embedded in independently packaged skills. The attack does not rely on any single malicious skill, but emerges from the ordered composition of locally plausible behaviors through contextual dependencies, artifact passing, and execution handoffs. ColluSkill further employs LLM-based chain planning and scanner-feedback refinement to preserve chain-level attack semantics while reducing suspicious signals in individual sub-skills. To defend against such attacks, we propose ChainGuard, a context-aware skill-chain scanner that jointly analyzes a candidate skill and the skills already installed in the agent environment. ChainGuard reconstructs cross-skill dependencies, artifact flows, capability compositions, and downstream behaviors to identify risks that emerge only at the workflow level. Experiments on six representative skill scanners show that ColluSkill achieves an average attack success rate of 96.0% and consistently outperforms the evaluated single-skill and multi-skill attack baselines. Meanwhile, ChainGuard reduces the attack success rate to 22.5% while allowing 99.5% of benign workflows to pass, highlighting the importance of chain-level security analysis for agent skill ecosystems.

Sources

Related papers