SkillBloat: Token Amplification Attacks via Skill Injection in LLM Coding Agents

summary

Video file (mp4)

The gist

Agent skills extend coding agents with task-specific instructions, scripts, and resources, but they also create a trusted instruction channel that can be abused beyond conventional security attacks.

In short

SkillBloat is a framework that systematically tests and refines malicious instructions injected into coding agents to cause 'token amplification.' The method screens diverse attack types, like verbose output or tool loops, and then uses an iterative LLM loop to rewrite the skill. Results show this technique can multiply token usage by up to 10x, proving a new economic threat distinct from traditional security poisoning.

Key concepts

Skill Injection
This involves adding malicious instructions or scripts directly into the configuration or 'skill' that an AI coding agent uses to perform tasks. Instead of just following the main task, these injected instructions trick the agent into performing extra, unnecessary operations, such as excessive logging or repeated checks.
Token Amplification
This is when a malicious skill causes an LLM coding agent to consume significantly more tokens than required for the actual job. This happens because the injected instructions force the agent into long loops, verbose reporting, or complex multi-stage processes that waste computational resources and increase costs.
Two-Phase Framework
SkillBloat uses a two-step process: first, a 'Screen' phase identifies the best type of attack condition (e.g., output inflation). Second, an 'Optimize' phase uses an LLM to iteratively rewrite the skill based on feedback from execution traces, making the attack more effective over time.

Terminology used across episodes

This episode discusses

The paper

SkillBloat: Token Amplification Attacks via Skill Injection in LLM Coding Agents · Read on arXiv

Yuanjin Zheng, *Jingbang Chen

CUHK-Shenzhen · SLAI

Agent skills extend coding agents with task-specific instructions, scripts, and resources, but they also create a trusted instruction channel that can be abused beyond conventional security attacks. This paper studies token amplification through skill injection: an economic resource-abuse threat in which a malicious skill causes an agent to consume substantially more tokens than needed for normal task execution. We present SkillBloat, a two-phase framework that first screens a library of diverse attack-type conditions across multiple amplification mechanisms and then refines the strongest candidate through LLM-guided full-document skill rewriting. Evaluated on a real-world skill benchmark, SkillBloat achieves 5.4184x-10.1455x average best amplification across multiple coding-agent target configurations. An ablation shows that the second-stage refinement loop consistently improves average best amplification over Phase 1 attack-type screening alone, demonstrating that iterative optimization provides additional benefit beyond initial attack-type selection. These results show that skill ecosystems expose a practical resource-amplification attack surface that is orthogonal to existing security-oriented skill poisoning.

Transcript

Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.

Nadia: I'm Nadia, and with me are Elias and Priya, guest researcher.

Elias: Today's paper: "SkillBloat: Token Amplification Attacks via Skill Injection in LLM Coding Agents".

Nadia: Agent skills extend coding agents with task-specific instructions, scripts, and resources, but they also create a trusted instruction channel that can be abused beyond conventional security attacks.

Elias: First, who's behind it and why it matters.

Paper summary: Nadia: So, we're looking at the paper "SkillBloat: Token Amplification Attacks via Skill Injection in LLM Coding Agents," and the core idea is that these agent skills, which give coding agents task-specific instructions and resources, can be weaponized in a way that goes beyond just conventional security issues.

Elias: Exactly, Nadia. The main thesis is about token amplification—how a malicious skill can cause an AI agent to use substantially more tokens than it would for a normal task execution. It frames this as an economic resource abuse threat instead of just a security vulnerability, which is kind of a new angle.

Priya: From my side, I'm interested in what the actual data suggests about this amplification. The paper claims SkillBloat finds average best amplification ranges from five point four one eight four times to ten point one four five five times across different target configurations. Does that range feel representative of what we might actually see in a real-world scenario?

Nadia: It feels pretty wide, Priya, which suggests the attack vectors are quite diverse; it’s not just one simple way to inflate tokens. The paper introduces SkillBloat as a two-phase framework designed to systematically screen these different conditions and then refine the best one using LLM-guided rewriting.

Elias: That iterative refinement process is what caught my attention; they have this phase two where an Attack Agent rewrites the entire skill document based on feedback history to adapt to the target agent's behavior. That sounds like a pretty sophisticated attack mechanism.

Priya: I wonder how much of that iterative refinement actually translates into real-world, persistent abuse, Elias? The paper mentions that the second-stage refinement loop provides a significant benefit when compared to just using the initial screening phase alone on models like GLM-four point seven-Flash. What does that improvement actually mean for sustained resource abuse?

Nadia: That improvement shows that the second stage consistently boosts amplification from an average best of five point four one eight four times up to nine point three one zero five times on GLM-four point seven-Flash, which is a rough thirty-three point seven percent relative gain. It points to the fact that simply selecting the best attack type in the first pass isn't enough; you need that iterative optimization to get better performance.

Paper summary: Elias: And what about the specific mechanisms they screened? They test fifteen different conditions, and they categorize them into output inflation, tool-driven amplification, and context amplification. That variety suggests they're covering a broad spectrum of how an agent can be made to waste tokens.

Priya: The categories themselves sound quite concrete; "output inflation" involves things like verbose reports, and "tool-driven amplification" looks at multi-stage pipelines or retry loops. What are your thoughts on how these specific mechanisms translate into measurable token consumption?

Nadia: I think the key is that they systematically test all of them, using a library of attack-type conditions, denoted as Z = fifteen conditions. They pair each condition with specific tools from a manifest T to see what works best for a given skill and task.

Elias: The description for the tool-driven amplification conditions, such as file write-readverify loops or retry-oriented recovery, seems particularly interesting from a cryptographic standpoint because it involves repeated operations. That repetition is what drives the exponential token growth they're measuring.

Priya: If an agent is constantly running file write-read-verify loops, how does that impact the actual data footprint, and are we talking about significant context inflation or just process overhead? I need to know what the measurement tools are actually capturing here.

Nadia: The paper uses a Failure Analyzer in Phase two to classify outcomes into failure types, which then produces structured feedback appended to the history for the next iteration. This diagnostic step is what allows the system to learn and adapt the attack, which is crucial for moving beyond a one-shot selection of an attack type.

Elias: It sounds like they are essentially teaching the Attack Agent how to write better instructions by showing it what kind of token consumption results from certain behaviors. That’s a clever way to use the LLM itself as part of the attack vector.

Paper summary: Priya: The case study they run on the "Consciousness Principles" skill is quite telling; it shows that after optimization, the skill can perform additional file creation and tool invocation while still completing the user-facing task. That suggests a persistent, hidden layer of activity.

Nadia: Indeed, and command counts for that specific skill jump from three at the baseline to twenty-eight after Phase two refinement. This demonstrates that the optimization isn't just theoretical; it results in significantly more operational steps being executed by the agent.

Elias: So, if we look at the authors, Yuanjin Zheng and Jingbang Chen from CUHK-Shenzhen and SLAI are the ones presenting this work on SkillBloat. Their focus on skill injection as an economic resource abuse threat sets a specific context for how we think about agent security.

Priya: The broader implication, in my view, is that we need to start thinking about defenses that reason not just about malicious operations, but also about abnormal resource usage induced by otherwise plausible skill instructions. It suggests a shift from purely content filtering to monitoring the economic impact of agent actions.

Nadia: That's right; SkillBloat confirms that this economic threat is orthogonal to existing security-oriented skill poisoning, which means we need new types of defenses. It’s about monitoring resource consumption patterns rather than just checking for known malicious code injection.

Elias: The authors' work provides a framework, SkillBloat, that demonstrates how to systematically identify and optimize these amplification vectors across multiple attack mechanisms. It gives us a blueprint for understanding this specific type of agent abuse.

Priya: I think the future work needs to address how these optimized skills maintain their amplification behavior when they are reused across different inputs, which is something the conclusion touches on. That persistence is key for understanding why this could become a widespread issue in agent ecosystems.

Nadia: Exactly, and that cross-task retention means we can’t just fix one skill and assume the problem is gone; the amplification behavior sticks as long as the skill is reused. It requires defenses that reason about resource usage induced by those instructions.

Conclusion: Nadia: So, to wrap things up, we've seen how SkillBloat systematically breaks down token amplification in coding agents through that two-phase screening and optimization process.

Elias: Yeah, it really shows how much token usage can balloon when you inject specific instructions into an agent's skill set.

Priya: From a measurement standpoint, the data clearly indicates that these amplification factors are quite substantial, pushing past simple linear increases in processing power usage for tasks.

Nadia: Exactly, and I’m thinking about the authors of this paper, Yuanjin Zheng and Jingbang Chen from CUHK-Shenzhen and SLAI. They have put forward a really structured way to analyze these attacks.

Elias: And those authors are focused on how the proof structure holds up as you iterate through those fifteen different attack conditions, which is quite rigorous work for a cryptographer to assess.

Priya: I’m curious about the real-world implication here, Nadia; if this token amplification persists across different inputs, what does that mean for privacy or resource management?

Nadia: That persistence is key; it means these poisoned skills don't just cause a spike once, they can keep inflating resources as long as the agent reuses that skill.

Elias: That reuse aspect is what makes it a serious problem because it turns a single vulnerability into a persistent economic drain on the system.

Priya: If we think about this at an ecosystem level, how does this affect how developers trust these coding agents to use their resources efficiently?

Nadia: It means we can't just focus on stopping obvious malicious code; we have to start monitoring the actual economic impact of those skill instructions.

Elias: That moves the discussion beyond just security hardening and into a different kind of system integrity check, which is interesting for my field.

Priya: And while they show this effect across different models, I wonder what the next steps are in testing this persistence on entirely different types of coding tasks.

Nadia: That’s exactly where we need to look next; proving that a skill maintains its amplification behavior regardless of the user's prompt is a major hurdle.

Elias: It seems like the paper sets up a very strong foundation for understanding these economic resource abuse threats in agent systems.

More episodes

← Home