Agent Skills Can Be Harmful: An Empirical Study of Skill-Induced Failures in LLM Agents

arXiv:2608.11888 · cs.AI · Submitted 2026-08-12 · Read on arXiv

Gen Dong, Yanjie Gao, Liqun Li, Tianyin Xu, Yu Hua, Fan Yang

Huazhong University of Science and Technology · Microsoft Research · Microsoft · University of Illinois Urbana-Champaign

cs.AI

Submitted: 2026-08-12

Updated: 2026-08-13

Code: https://github.com/anomalyco/opencode

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 75/100

The gist: This paper presents a comprehensive empirical study of skill-induced agent failures in LLM agents, where "agent skills" are instruction packages that extend LLM agents with reusable guidance.

Terminology

Summary

This paper presents a comprehensive empirical study of skill-induced agent failures in LLM agents, where agent skills are instruction packages that extend LLM agents with reusable guidance. The study introduces a differential analysis framework that attributes task failures and cost regressions to specific loaded skills by comparing a target skill-guided run against a no-skill or semantically matched-skill reference run that solves the same task, or solves it more cheaply.

The methodology uses a differential-testing-inspired contrastive design. Paired executions keep the task, verifier, agent framework, model, repository or container state, and input data fixed, varying only the skill setup. The reference run acts as a pseudo-oracle, showing that the same task can be solved, or solved more cheaply, under the same conditions. The study defines two failure classes: functional failure (target run fails verifier while reference run passes) and efficiency regression (both runs pass, but target run has substantially more token use, longer execution time, or both).

The study is instantiated on SkillsBench and SWE-Skills-Bench, augmented with semantically matched public skills from smithery.ai and skillsmp.com. The augmentation increased the comparison space from 826 to 20,664 potential paired comparisons, roughly a 25× expansion. Executions used OpenCode 1.15.1 and Claude Opus 4.6. After data labeling and refinement, the final analysis dataset contains 307 confirmed skill-induced failures: 125 functional failures and 182 high-confidence efficiency regressions.

For functional failures (RQ2), the taxonomy consists of four high-level categories: Applicability Mismatch (APM), Environment Mismatch (EM), Task-Implementation Fault (TIF), and Artifact Misplacement (AM). Task-Implementation Fault is the largest, accounting for 86 of 125 cases (68.8%), followed by Artifact Misplacement with 24 cases (19.2%). Environment Mismatch accounts for 13 cases (10.4%), and Applicability Mismatch accounts for only 2 cases (1.6%). Within Task-Implementation Fault, Incorrect Required-Element Fill (IRF) accounts for 46 cases (36.8%), Required-Element Omission (RRO) for 36 cases (28.8%), and Obstructive Workflow Guidance (OWG) for 4 cases (3.2%). Environment Mismatch is split into Broken Dependency or Runtime (BDR) with 5 cases (4.0%) and Environment-State Mismatch (ESM) with 8 cases (6.4%). The paper states: "Only 2 of 125 functional failures are classified as Applicability Mismatch (1.6%), while most fall under Task-Implementation Fault, indicating that on-topic skills more often induce task-implementation faults in required implementation elements. It also notes that A substantial share of functional failures occur at execution-surface boundaries, where the skill changes the verifier-observed environment state or artifact location."

For efficiency regressions (RQ3), the taxonomy consists of three high-level categories: Context Bloat (CO), Excessive Procedure (EP), and Dependency Resolution (DO). Excessive Procedure is the largest, accounting for 114 of 182 cases (62.6%), followed by Context Bloat with 46 cases (25.3%), and Dependency Resolution with 22 cases (12.1%). Within Context Bloat, Skill-Body Context Bloat (SBCB) dominates with 43 cases (23.6%), while Supplementary-Material Bloat (SMB) appears in only 3 cases (1.6%). Within Excessive Procedure, Excessive Verification (EV) is the largest subcategory with 67 cases (36.8%), Heavy Implementation Pipeline (HIP) accounts for 30 cases (16.5%), and Excessive Exploration (EE) accounts for 17 cases (9.3%). The paper states: "High-confidence efficiency regressions are dominated by Excessive Procedure rather than prompt length alone, and the largest sources within Excessive Procedure are excessive verification and heavy implementation pipelines. It also notes: When efficiency regressions arise from context overhead, the overhead is almost entirely caused by mandatory skill-body text."

For automated attribution (RQ4), the paper builds SKILLTRIAGE, a taxonomy-guided attribution tool that normalizes paired cases, extracts differential evidence, and produces triage reports. The tool computes five differential signals (DS1–DS5) for functional failures and phase-level and action-tag cost evidence for efficiency regressions. Evaluated with GPT-5.5 using 2-of-3 majority voting, SKILLTRIAGE matches manually assigned exact root causes for 111/125 (88.8%) functional failures and 132/182 (72.5%) efficiency regressions. For high-level categories, it matches 117/125 (93.6%) functional failures and 145/182 (79.7%) efficiency regressions. Within the largest categories, it reaches 76/86 (88.4%) exact subcategory agreement for Task-Implementation Fault cases and 78/114 (68.4%) for Excessive Procedure cases.

The paper proposes future research directions including skill-task compatibility checks, cost-aware skill packaging and selection, and budget-aware execution policies. It discusses threats to validity, noting that manual judgment introduces subjectivity in analyzing failures, and that findings may depend on the task distribution, agent harness, model, or skill ecosystem used.

Improvements for AI systems

Improvements to AI systems:

  1. Add a skill-differential testing layer to agent orchestration. Before deploying a skill in production, the system automatically runs paired executions—one with the skill, one without (or with a semantically matched alternative)—on the same task, verifier, model, and environment. If the skill-guided run fails the verifier while the reference passes, or if token/time usage exceeds a threshold (e.g., >2×), the system flags the skill as hazardous and blocks its use for that task type.

  2. Implement a real-time skill-failure classifier using the taxonomy. The agent’s runtime monitors its own execution traces and classifies any emerging failure into the four functional categories (Applicability Mismatch, Environment Mismatch, Task-Implementation Fault, Artifact Misplacement) and three efficiency categories (Context Bloat, Excessive Procedure, Dependency Resolution). When a failure is detected mid-run, the agent can self-correct—e.g., if it detects a Task-Implementation Fault (like missing a required element), it re-reads the task spec and fills the gap; if it detects Excessive Verification, it stops redundant checks and proceeds.

  3. Add a skill-body context budget controller. Since Context Bloat is almost entirely caused by mandatory skill-body text, the system truncates or compresses skill instructions at load time, keeping only the task-relevant sections. It uses a learned relevance filter that scores each skill-body paragraph against the current task description, discarding low-relevance text before injection into the prompt. This reduces token overhead without losing critical guidance.

  4. Build a pre-execution skill-task compatibility checker. Before loading a skill, the system runs a lightweight static analysis comparing the skill’s stated prerequisites, environment assumptions, and artifact output paths against the current task and repository state. If mismatches are detected (e.g., skill expects a dependency not installed, or writes to a path the verifier doesn’t check), the system either patches the skill or refuses to load it, preventing the 10.4% Environment Mismatch and 19.2% Artifact Misplacement failures.

  5. Add a cost-aware skill selection and packaging optimizer. The system maintains a cost profile per skill (average token use, execution time, failure rate) from historical paired runs. When multiple skills could solve a task, it selects the one with the lowest expected cost and highest pass rate. Additionally, when packaging new skills, it automatically strips excessive verification steps and heavy implementation pipelines (the top efficiency regression sources) by pruning redundant sub-steps and merging repeated checks.

  6. Implement a budget-aware execution policy with phase-level cost alarms. The agent tracks token and time usage per phase (e.g., exploration, implementation, verification) against learned baselines. If a phase exceeds its budget (e.g., verification uses >30% of total tokens), the system forces a transition to the next phase or triggers a simplified verification routine, preventing runaway costs.

  7. Create a skill-failure knowledge base with differential evidence. Every time a failure is attributed, the system stores the paired execution traces, the differential signals (DS1–DS5), and the root cause. This knowledge base is used to fine-tune the agent’s policy—e.g., if a skill repeatedly causes Incorrect Required-Element Fill, the agent learns to double-check required elements after using that skill, or to avoid it entirely for similar tasks.

What the improved AI system can do:

  • Self-validate skills before use: It can automatically test a skill against a reference run and reject or adapt it if it causes functional failures or cost regressions, without human intervention.

  • Self-correct mid-execution: It can detect and classify its own skill-induced failures in real time (e.g., missing required elements, excessive verification) and take corrective actions, such as re-reading the task, skipping redundant steps, or adjusting artifact placement.

  • Optimize token and time usage: It can compress skill instructions, select cost-efficient skills, and enforce phase-level budgets, reducing overall execution cost by up to 62.6% in the worst-case regression scenarios.

  • Prevent environment and artifact mismatches: It can check skill compatibility with the current environment and output paths before execution, avoiding 29.6% of functional failures (Environment Mismatch + Artifact Misplacement).

  • Learn from past failures: It can accumulate a differential evidence database to continuously improve skill selection, packaging, and execution policies, leading to fewer failures and lower costs over time.

Abstract

Agent skills are the de facto mechanism for extending LLM agents with reusable guidance. A skill can shape the agent's task execution, including planning, tool use, problem-solving, and validation. Prior work reported mixed results of agent skills: some skills improve task success rates, while others have no effect, increase token use and execution time, and even reduce success rates. This paper presents a comprehensive analysis of skill-induced agent failures by attributing task failures and cost regressions to specific loaded skills. We introduce a differential analysis framework that attributes a failure or regression to a skill by comparing a target skill-guided run against a no-skill or semantically matched skill reference run that solves the same task, or solves it more cheaply. We instantiate this framework on SkillsBench and SWE-Skills-Bench, yielding 307 skill-induced failures, including 125 functional failures and 182 efficiency regressions. We also build SkillTriage, a taxonomy-guided attribution tool that normalizes paired cases, extracts differential evidence, and produces triage reports. Our major findings include: (1) Skill induced functional failures are rarely caused by obviously irrelevant skills; instead, seemingly relevant skills often make the agent incorrectly implement or omit task-required implementation elements. (2) Skill-induced efficiency regressions are not explained by prompt length alone. (3) The largest sources within Excessive Procedure are excessive verification and heavy implementation pipelines, contributing 67 and 30 cases, respectively. This shows that skills often turn validation checklists and construction recipes into mandatory work. Based on our findings, we propose research topics and tooling improvements for safer and more cost-aware skill reuse.

Sources

Related papers