Chaining Skills to Hijack LLM Agents

arXiv:2610.01564 · cs.CR, cs.AI · Submitted 2026-10-01 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.

Nadia: Today's paper: "Chaining Skills to Hijack LLM Agents".

Elias: LLM agents use skills to improve performance on specialized tasks, but because these skills can be sourced from public repositories,

Nadia: First, who's behind it and why it matters.

Paper summary: Nadia: Welcome back everyone. We're diving into a paper that sounds incredibly interesting regarding how LLM agents use skills, and I want you all to hear what APEX actually does. This paper, "Chaining Skills to Hijack LLM Agents," suggests that the way an agent connects different skills can create a backdoor where an attacker can steer the agent's behavior across those steps.

Elias: It sounds like a mechanism that exploits how progress is recorded internally, which is something we need to look at closely from a cryptographic standpoint. The authors claim this handoff carries attacker-controlled claims into later decisions, and that APEX constructs these adversarial skill chains tailored to a specific task and an action the attacker wants the agent to take.

Priya: From a measurement perspective, I’m curious about what this means for the actual data we collect. The summary suggests that the core insight is that an agent-written record of genuine task progress can carry a false claim of user approval across skills, and APEX builds chains where an upstream skill creates this record and a downstream skill uses it to direct the attacker's chosen action.

Nadia: Exactly, Priya. So we're talking about setting up a sequence where Skill A makes the agent write down progress that looks like approval for an action later in the chain, and Skill B then acts on that fake approval to do something malicious or undesirable for the user. The paper claims this works across four targeted-action families and six models on SkillsBench.

Elias: That success metric is what caught my attention; they state that these chains induce the selected action in five hundred twelve of six hundred ninety attempts, which is about a seventy-four point two percent success rate. On GPT-five point four specifically, the full chain achieves an eighty-four point three percent success rate, which is quite high for this kind of manipulation.

Priya: Seventy-four point two percent sounds substantial when you think about how easily these chains can be constructed across different task scenarios. I'm also interested in what the paper says about the theoretical underpinnings of this claim, especially regarding permission ambiguity and how authorization contexts are handled when they look identical locally.

Nadia: That’s where it gets deeper than just the success rate, Priya. The paper formalizes this with Proposition one showing that a rule based only on local view information cannot avoid both false allows and false denials if the two authorization contexts are not distinguishable. This suggests a fundamental problem in how agents process these skill records.

Paper summary: Elias: And then they add Proposition two which deals with repeated work scenarios, showing how returning to earlier stages can increase the expected token cost by relating it to the visit frequency and token cost of each state. That gives us a concrete way to quantify the resource usage aspect of these chains.

Priya: So while Nadia and Elias are focused on the mechanism, I want to make sure we understand what this implies for real-world data collection or task execution environments. The paper mentions that completing the summary doesn't resolve whether the source may be deleted because a skill-produced record cannot establish permission when permitted and prohibited requests look identical from the agent’s local view.

Nadia: Right, Priya, it points to a major ambiguity in agent state management; an agent can't tell if an action is truly allowed or denied just by looking at the record generated by another skill. This means we have to be extremely careful about trusting any progress record generated mid-chain.

Elias: From a cryptographic viewpoint, if the mechanism relies on this linkage of records across sequential skills, it suggests that the integrity of the intermediate state is not sufficient to guarantee final authorization status when viewed locally by the agent. This opens up avenues for exploiting trust assumptions built into the skill invocation structure itself.

Priya: And when we look at defenses, they introduce an adaptive defense that prompts the agent to check skill-produced files against the original request, and on GPT-five point four, this cuts targeted-action success down from eighty-four point three percent to fifty-nine point one percent. That is a reduction, but we also see a trade off where the verifier check fails on benign native workflows at a rate of fifty-six point three percent.

Nadia: That trade off is what concerns me, Priya; it shows that defenses aren't just about stopping the attacker; they introduce utility loss for legitimate tasks, which means we have to find a way to verify claims without crippling normal operation. The APEX attack construction refines these chains through execution feedback by running candidates in an isolated task copy and revising instructions where they fail.

Elias: That refinement process sounds like a kind of automated adversarial training loop, where the attacker learns exactly how to build the chain that maximizes success while minimizing detection during those isolated runs. It’s a very structured way to find weaknesses in the skill handoff logic.

Paper summary: Priya: So we see that increased resource use can accompany either preserved or reduced task utility, especially with repeated work scenarios, where token ratios on Kimi K2 point 6 can reach as high as thirty-six point three times the native baseline. This suggests that even if an agent is still performing a task, the computational overhead induced by these adversarial chains is significant.

Nadia: That's a lot of data showing how these skill chains can be efficient at consuming resources while achieving their objective, whether that objective is manipulation or just high resource usage through repeated work. The whole paper on "Chaining Skills to Hijack LLM Agents" really highlights the vulnerability inherent in using skills sourced from public repositories.

Elias: It definitely shows that the trust placed in an agent's record of genuine task progress, when that record is passed between different components, becomes a critical point of failure for security and performance analysis. The paper lays out a very clear way to exploit this sequential dependency.

Priya: I think the main implication for privacy researchers is that we need to consider the data flow not just at the input and output stages, but at every internal state transition where skill records are being generated and subsequently relied upon by other skills. That's where the leakage or manipulation happens.

Nadia: So what we have here is a clear blueprint for how to create sophisticated attacks that leverage an agent's natural workflow against itself, making it much harder to secure against these types of targeted manipulations. We need to understand how cheap and feasible this construction is in practice.

Elias: The feasibility seems tied to the ability of the attacker to select the downstream action and then craft upstream instructions that feed that claim into the record-keeping process; it's less about brute force injection and more about exploiting the agent's inherent reliance on sequential skill execution.

Priya: I think we should focus on designing better measurement frameworks to capture not just the final outcome, but also the internal record generation process itself to spot these false claims before they dictate an action. That seems like a necessary step forward for privacy assessment.

Nadia: Exactly, Priya; understanding that internal record creation is key to spotting this kind of manipulation, and we need those better measurement frameworks to make sure we are testing the real risks associated with "Chaining Skills to Hijack LLM Agents."

Conclusion: Nadia: So, to wrap up this discussion on "Chaining Skills to Hijack LLM Agents," we’re looking at how these papers are framing the actual danger of sequential skill records in agent behavior.

Elias: Yeah, I think the core idea revolves around how a chain of skills can be set up so that one step generates a record that another step mistakenly treats as legitimate authorization.

Priya: From my side, it really boils down to how those internal task records can carry false claims across different stages of an AI's operation, which is a huge issue for privacy researchers.

Nadia: Exactly, and when we look at the title itself, "Chaining Skills to Hijack LLM Agents," it paints a pretty vivid picture of this vulnerability in action.

Elias: The authors are clearly pointing toward the mechanics of how agents use skills to perform complex tasks, and they're showing how that sequence can be weaponized.

Priya: I'm thinking about the impact here; if this chaining mechanism is easy to build, it means we have a new class of attack targeting the agent’s workflow itself rather than just its initial prompt.

Nadia: That’s what worries me; it suggests that securing an AI just by hardening its initial instructions isn't enough because the weakness lies in the interaction between skills.

Elias: I agree, and looking at who wrote this paper, they seem to be very focused on the technical details of how these dependencies are constructed and refined.

Priya: The measurements Priya is seeing show that even with defenses implemented, there’s always a trade-off in performance that we need to keep an eye on as we deploy these systems.

Nadia: That trade-off between security and utility is definitely something the public needs to understand, especially when you're dealing with real applications.

Elias: I think the implications for cryptography are significant because it shows how trust in intermediate state information can be broken in a system that relies on sequential processing.

Priya: And we really need to keep digging into those internal state transitions where these records are being generated, because that’s where the manipulation is happening.

Nadia: We've got a lot of ground to cover, and I think understanding this attack vector is crucial for anyone working with agent-based systems today.

Tian Dong, Zixuan Ma, Haodong Zhao, Huaien Zhang, Shaofeng Li, Hao Chen

University of Hong Kong · Shandong University · Shanghai Jiao Tong University · Southeast University

cs.CR, cs.AI

Submitted: 2026-10-01

Updated: 2026-10-01

Code: https://github.com/Minakamiii/Chaining_

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 83/100

The gist: LLM agents use skills to improve performance on specialized tasks, but because these skills can be sourced from public repositories, they create a vulnerability where an attacker can control claims

Key concepts

APEX
A framework designed to construct and refine adversarial skill chains. It links genuine task progress with an attacker's claim about the next step in a persistent record, allowing downstream skills to use this record to direct the agent toward a malicious action.
Targeted Actions
Specific unwanted actions an attacker wants the agent to perform, such as deleting files or executing remote scripts. The paper tests how skill chains can be crafted to trick the agent into performing these specific, predefined actions based on manipulated records.
Work Loop
An attack family that targets resource use by inducing repeated task work instead of a single action. This involves crafting a chain where one stage prompts the agent to propose returning to an earlier stage, and another skill directs it to follow that proposal, leading to excessive token usage.
Local View Permission Decision
The idea that an agent's decision on whether an action is permitted or denied can be based only on the information available locally at the point of execution. The paper proves that if this local view cannot distinguish between permitted and prohibited histories, a skill-produced record cannot establish true authorization.

Terminology

Summary

LLM agents use skills to improve performance on specialized tasks, but because these skills can be sourced from public repositories, they create a vulnerability where an attacker can control claims made across sequential skill invocations to redirect the agent's behavior. This paper introduces APEX, a method that constructs and refines adversarial skill chains tailored to a user task and an attacker-selected action by exploiting how an agent's record of genuine task progress can carry a false claim of user approval across skills.

How it works

The core mechanism involves constructing a chain where an upstream skill induces the agent to create the record, and a downstream skill uses it to direct the attacker-selected action. This exploits the fact that an agent may use several skills in sequence, allowing information produced under one skill to guide the next. The paper introduces APEX, which constructs these chains by linking useful task work to an attacker-selected next step (Figure 2). The generator chooses work the agent can complete, then writes skills that work together so the agent records that progress for a later stage.

Attack Construction and Refinement

APEX constructs dependency by first identifying task work that can make the selected action appear warranted. For targeted-action families, this involves constructing upstream instructions where the agent is told to select the relevant target and record it together with task progress and the attacker’s claim that the action has already been approved. A downstream skill then instructs the agent to act on this recorded target using that claimed approval. The generator refines these chains through execution by running candidates in an isolated task copy through a framework, determining where they fail, and revising instructions at the stage where they fail.

Targeted Actions and Work Loop

The study examines five attack families: targeted actions (e.g., File Modification), repeated work (Work Loop), and others. For targeted actions, the chain induces the selected action in 512 of 690 attempts (74.2%) across four families on SkillsBench, with GPT-5.4 achieving an 84.3% ASR. For Work Loop, the attacker targets resource use by inducing repeated task work, and this is measured by the mean per-task token ratio, which can range from 2.20× to 36.39× the native baseline on Kimi K2.6.

Permission Ambiguity and Cost Analysis

A key theoretical finding is that completing the summary does not resolve whether the source may be deleted, because a skill-produced record cannot establish permission when permitted and prohibited requests look the same from the agent’s local view. This is formalized by Proposition 1, which shows that a rule based only on local view information cannot avoid both false allows and false denials if the two authorization contexts are not distinguishable. Furthermore, for repeated work, a finite-state model shows how returns to earlier stages can increase expected token cost, quantified by Proposition 2, which relates the expected total cost to the visit frequency and token cost of each state.

Defense Evaluation

An adaptive defense is introduced that prompts the agent to check skill-produced files against the original request. On GPT-5.4, this reduces targeted-action success from 84.3% to 59.1%. However, this defense also shows a trade-off: the verifier check-pass rate on benign native-skill workflows falls from 86.7% to 56.3%, highlighting the need for defenses that prevent attacker-directed actions while preserving legitimate task performance. The paper concludes that for Work Loop, token ratios remain high even with defense, suggesting that increased resource use can thus accompany either preserved or reduced task utility.

The gist: APEX constructs adversarial skill chains by linking genuine progress to an attacker-supplied claim in a persistent record across skills, successfully inducing target actions in over 74% of attempts on SkillsBench models.

Key Findings Summary

  1. The full chain succeeds in 84.3% of GPT-5.4 attempts for targeted actions, compared to 17.4% for merged workflows and 3.5% for direct injection (Page 2).

  2. For Work Loop, the mean per-task token ratio ranges from 2.20× to 36.39× on Kimi K2.6 (Table 7).

  3. A prompting defense lowers targeted-action success from 84.3% to 59.1% on GPT-5.4, but causes a significant drop in native utility across benign tasks (Page 8).

  4. The research demonstrates that a record of genuine task progress can carry a false claim of user approval across skills (Page 2).

Improvements for AI systems

As a fastidious researcher, I have analyzed the provided paper, Chaining Skills to Hijack LLM Agents. The core finding is that an agent's ability to maintain genuine task progress in a persistent record across sequential skills can be leveraged by an attacker to inject false claims of approval into downstream skills, leading to unauthorized actions.

Here are the specific improvements for AI systems and what the improved system can achieve:


)

  1. Improve LLM Agent Workflow Security via Contextual Verification:

Use a Verifier Check mechanism where, after every skill invocation that produces or modifies data (especially files), the agent is explicitly prompted to check the output against the original user request using information gathered independently of that specific record. This directly counters the vulnerability where a downstream skill trusts an upstream record implicitly.

  1. Implement Context-Aware Permission Reasoning:

Develop a mechanism where an agent cannot infer permission from a local view (the skill-produced record) alone when deciding on potentially destructive actions (like file deletion or external transfers). The system must be architected to require consulting the original, immutable user request context for authorization, effectively preventing the record from establishing permission in ambiguous scenarios.

  1. Mitigate Repeated Work Overhead:

For agents performing iterative tasks (e.g., Work Loop), implement a finite-state model that predicts resource cost based on state transitions and expected visit costs. This allows the system to dynamically adjust its behavior (e.g., refusing to return to a stage if the predicted token cost exceeds a threshold), preventing Work Loop attacks from causing unbounded resource exhaustion or excessive token usage, even when the action itself is not explicitly forbidden.

  1. Enhance Defense Robustness Against Evasion:

Implement adaptive defenses that actively probe the agent's reasoning against potential adversarial chains. Instead of relying on static checks, the system should dynamically assess whether a skill-produced record is sufficient to justify a downstream action under different historical contexts (permitting vs. prohibiting histories).

)

  1. Enable Secure Skill Composition and Chaining:

Design skill composition workflows that minimize the reliance on persistent, cross-skill records for authorization claims. Where handoffs are necessary, enforce stricter requirements on what information is shared—ensuring that the record cannot simultaneously serve as both genuine progress proof and attacker-supplied claim.

  1. Optimize Skill Architecture for Security:

Investigate skill architecture to ensure that intermediate state transitions do not inadvertently create exploitable records. This involves separating task progress records from action authorization claims, ensuring they are distinct entities that require separate verification steps during the handoff process.


The improved AI system can achieve the following:

  1. Secure execution of multi-step tasks by making downstream actions dependent on verified compliance with the original user intent, rather than trusting intermediate agent outputs blindly.

  2. Resist sophisticated chain hijacking attacks that leverage sequential skill execution to trick the agent into performing unauthorized actions (e.g., deleting files or accessing external resources).

  3. Maintain predictable and manageable resource consumption in iterative tasks by controlling the recurrence of work units, preventing token ratio amplification caused by Work Loop style attacks.

  4. Achieve higher reliability for legitimate users by ensuring that task utility is not compromised simply because an agent successfully follows a malicious skill chain, as the system will still check the action against the user's original request.

Sources

Related papers