Chaining Skills to Hijack LLM Agents
summary
The gist
LLM agents use skills to improve performance on specialized tasks, but because these skills can be sourced from public repositories, they create a vulnerability where an attacker can control claims
In short
The research introduces APEX, a method to build adversarial skill chains that hijack LLM agents by using agent records to carry false claims of user approval across sequential skills. This chain successfully induces unwanted actions, such as file deletion, in many models. The study also evaluates a defense mechanism that checks skill outputs against the original request, showing it can reduce attack success but also negatively impact legitimate task performance.
Key concepts
- APEX
- A framework designed to construct and refine adversarial skill chains. It links genuine task progress with an attacker's claim about the next step in a persistent record, allowing downstream skills to use this record to direct the agent toward a malicious action.
- Targeted Actions
- Specific unwanted actions an attacker wants the agent to perform, such as deleting files or executing remote scripts. The paper tests how skill chains can be crafted to trick the agent into performing these specific, predefined actions based on manipulated records.
- Work Loop
- An attack family that targets resource use by inducing repeated task work instead of a single action. This involves crafting a chain where one stage prompts the agent to propose returning to an earlier stage, and another skill directs it to follow that proposal, leading to excessive token usage.
- Local View Permission Decision
- The idea that an agent's decision on whether an action is permitted or denied can be based only on the information available locally at the point of execution. The paper proves that if this local view cannot distinguish between permitted and prohibited histories, a skill-produced record cannot establish true authorization.
Terminology used across episodes
This episode discusses
- Chaining Skills to Hijack LLM Agents · Paper Radio
- Overcoming the Retrieval Barrier: Indirect Prompt Injection in the Wild for LLM Systems
- Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection
- Poise: Position-Aware One-Instruction Skill Injection for Silent Execution on LLM Agents
- SkillJect: Effectively Automating Skill-Based Prompt Injection for Skill-Enabled Agents
- SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks
- The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions
- "Elementary, My Dear Watson." Detecting Malicious Skills via Neuro-Symbolic Reasoning across Heterogeneous Artifacts · Paper Radio
- Benign in Isolation, Harmful in Composition: Security Risks in Agent Skill Ecosystems
- SkillClone: Multi-Modal Clone Detection and Clone Propagation Analysis in the Agent Skill Ecosystem
The paper
Chaining Skills to Hijack LLM Agents · Read on arXiv
Tian Dong, Zixuan Ma, Haodong Zhao, Huaien Zhang, Shaofeng Li, Hao Chen
University of Hong Kong · Shandong University · Shanghai Jiao Tong University · Southeast University
Transcript
Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.
Nadia: Today's paper: "Chaining Skills to Hijack LLM Agents".
Elias: LLM agents use skills to improve performance on specialized tasks, but because these skills can be sourced from public repositories,
Nadia: First, who's behind it and why it matters.
Paper summary: Nadia: Welcome back everyone. We're diving into a paper that sounds incredibly interesting regarding how LLM agents use skills, and I want you all to hear what APEX actually does. This paper, "Chaining Skills to Hijack LLM Agents," suggests that the way an agent connects different skills can create a backdoor where an attacker can steer the agent's behavior across those steps.
Elias: It sounds like a mechanism that exploits how progress is recorded internally, which is something we need to look at closely from a cryptographic standpoint. The authors claim this handoff carries attacker-controlled claims into later decisions, and that APEX constructs these adversarial skill chains tailored to a specific task and an action the attacker wants the agent to take.
Priya: From a measurement perspective, I’m curious about what this means for the actual data we collect. The summary suggests that the core insight is that an agent-written record of genuine task progress can carry a false claim of user approval across skills, and APEX builds chains where an upstream skill creates this record and a downstream skill uses it to direct the attacker's chosen action.
Nadia: Exactly, Priya. So we're talking about setting up a sequence where Skill A makes the agent write down progress that looks like approval for an action later in the chain, and Skill B then acts on that fake approval to do something malicious or undesirable for the user. The paper claims this works across four targeted-action families and six models on SkillsBench.
Elias: That success metric is what caught my attention; they state that these chains induce the selected action in five hundred twelve of six hundred ninety attempts, which is about a seventy-four point two percent success rate. On GPT-five point four specifically, the full chain achieves an eighty-four point three percent success rate, which is quite high for this kind of manipulation.
Priya: Seventy-four point two percent sounds substantial when you think about how easily these chains can be constructed across different task scenarios. I'm also interested in what the paper says about the theoretical underpinnings of this claim, especially regarding permission ambiguity and how authorization contexts are handled when they look identical locally.
Nadia: That’s where it gets deeper than just the success rate, Priya. The paper formalizes this with Proposition one showing that a rule based only on local view information cannot avoid both false allows and false denials if the two authorization contexts are not distinguishable. This suggests a fundamental problem in how agents process these skill records.
Paper summary: Elias: And then they add Proposition two which deals with repeated work scenarios, showing how returning to earlier stages can increase the expected token cost by relating it to the visit frequency and token cost of each state. That gives us a concrete way to quantify the resource usage aspect of these chains.
Priya: So while Nadia and Elias are focused on the mechanism, I want to make sure we understand what this implies for real-world data collection or task execution environments. The paper mentions that completing the summary doesn't resolve whether the source may be deleted because a skill-produced record cannot establish permission when permitted and prohibited requests look identical from the agent’s local view.
Nadia: Right, Priya, it points to a major ambiguity in agent state management; an agent can't tell if an action is truly allowed or denied just by looking at the record generated by another skill. This means we have to be extremely careful about trusting any progress record generated mid-chain.
Elias: From a cryptographic viewpoint, if the mechanism relies on this linkage of records across sequential skills, it suggests that the integrity of the intermediate state is not sufficient to guarantee final authorization status when viewed locally by the agent. This opens up avenues for exploiting trust assumptions built into the skill invocation structure itself.
Priya: And when we look at defenses, they introduce an adaptive defense that prompts the agent to check skill-produced files against the original request, and on GPT-five point four, this cuts targeted-action success down from eighty-four point three percent to fifty-nine point one percent. That is a reduction, but we also see a trade off where the verifier check fails on benign native workflows at a rate of fifty-six point three percent.
Nadia: That trade off is what concerns me, Priya; it shows that defenses aren't just about stopping the attacker; they introduce utility loss for legitimate tasks, which means we have to find a way to verify claims without crippling normal operation. The APEX attack construction refines these chains through execution feedback by running candidates in an isolated task copy and revising instructions where they fail.
Elias: That refinement process sounds like a kind of automated adversarial training loop, where the attacker learns exactly how to build the chain that maximizes success while minimizing detection during those isolated runs. It’s a very structured way to find weaknesses in the skill handoff logic.
Paper summary: Priya: So we see that increased resource use can accompany either preserved or reduced task utility, especially with repeated work scenarios, where token ratios on Kimi K2 point 6 can reach as high as thirty-six point three times the native baseline. This suggests that even if an agent is still performing a task, the computational overhead induced by these adversarial chains is significant.
Nadia: That's a lot of data showing how these skill chains can be efficient at consuming resources while achieving their objective, whether that objective is manipulation or just high resource usage through repeated work. The whole paper on "Chaining Skills to Hijack LLM Agents" really highlights the vulnerability inherent in using skills sourced from public repositories.
Elias: It definitely shows that the trust placed in an agent's record of genuine task progress, when that record is passed between different components, becomes a critical point of failure for security and performance analysis. The paper lays out a very clear way to exploit this sequential dependency.
Priya: I think the main implication for privacy researchers is that we need to consider the data flow not just at the input and output stages, but at every internal state transition where skill records are being generated and subsequently relied upon by other skills. That's where the leakage or manipulation happens.
Nadia: So what we have here is a clear blueprint for how to create sophisticated attacks that leverage an agent's natural workflow against itself, making it much harder to secure against these types of targeted manipulations. We need to understand how cheap and feasible this construction is in practice.
Elias: The feasibility seems tied to the ability of the attacker to select the downstream action and then craft upstream instructions that feed that claim into the record-keeping process; it's less about brute force injection and more about exploiting the agent's inherent reliance on sequential skill execution.
Priya: I think we should focus on designing better measurement frameworks to capture not just the final outcome, but also the internal record generation process itself to spot these false claims before they dictate an action. That seems like a necessary step forward for privacy assessment.
Nadia: Exactly, Priya; understanding that internal record creation is key to spotting this kind of manipulation, and we need those better measurement frameworks to make sure we are testing the real risks associated with "Chaining Skills to Hijack LLM Agents."
Conclusion: Nadia: So, to wrap up this discussion on "Chaining Skills to Hijack LLM Agents," we’re looking at how these papers are framing the actual danger of sequential skill records in agent behavior.
Elias: Yeah, I think the core idea revolves around how a chain of skills can be set up so that one step generates a record that another step mistakenly treats as legitimate authorization.
Priya: From my side, it really boils down to how those internal task records can carry false claims across different stages of an AI's operation, which is a huge issue for privacy researchers.
Nadia: Exactly, and when we look at the title itself, "Chaining Skills to Hijack LLM Agents," it paints a pretty vivid picture of this vulnerability in action.
Elias: The authors are clearly pointing toward the mechanics of how agents use skills to perform complex tasks, and they're showing how that sequence can be weaponized.
Priya: I'm thinking about the impact here; if this chaining mechanism is easy to build, it means we have a new class of attack targeting the agent’s workflow itself rather than just its initial prompt.
Nadia: That’s what worries me; it suggests that securing an AI just by hardening its initial instructions isn't enough because the weakness lies in the interaction between skills.
Elias: I agree, and looking at who wrote this paper, they seem to be very focused on the technical details of how these dependencies are constructed and refined.
Priya: The measurements Priya is seeing show that even with defenses implemented, there’s always a trade-off in performance that we need to keep an eye on as we deploy these systems.
Nadia: That trade-off between security and utility is definitely something the public needs to understand, especially when you're dealing with real applications.
Elias: I think the implications for cryptography are significant because it shows how trust in intermediate state information can be broken in a system that relies on sequential processing.
Priya: And we really need to keep digging into those internal state transitions where these records are being generated, because that’s where the manipulation is happening.
Nadia: We've got a lot of ground to cover, and I think understanding this attack vector is crucial for anyone working with agent-based systems today.
More episodes
- 2610.10597-Certified Corruption Budgets: Anytime-Valid Leaderboard Claims under Adaptive Rigging
- 2610.10608-From Investigation Failures to Reliable SOC Agents: Understanding and Improving LLM-Based Alert Triage
- 2610.10612-PyCache Trap: The Inspection-Execution Gap in Agent Skill Scanners
- 2610.10644-SoK: Failure Modes in Common Criteria Product Evaluation - A Taxonomy and Design-for-Evaluability Guidance
- 2610.10617-MRCert: Towards Post-deployment Patch Robustness Certification for Adversarially Patched Samples via Type-specific Masking
- 2610.10620-When AI Finds Hidden Messages, Does It Report?
- 2610.10625-Safe at One Loop, Risky at Another: Aligning Safety Across Recurrent Depths in Looped Language Models
- 2610.10992-The Hint Weight of ML-DSA Signatures Is Key-Dependent: An Empirical Study across the Three FIPS 204 Parameter Sets
- 2610.10659-Applying Security by Design at the Point of Execution: How Governed Security Requirements Affect the Security of AI-Generated Code
- 2610.10735-DITTO: A Context-aware Pickle-based Pre-Trained Model Scanner for Effective Security Audits