SkillBloat: Token Amplification Attacks via Skill Injection in LLM Coding Agents
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.
Nadia: I'm Nadia, and with me are Elias and Priya, guest researcher.
Elias: Today's paper: "SkillBloat: Token Amplification Attacks via Skill Injection in LLM Coding Agents".
Nadia: Agent skills extend coding agents with task-specific instructions, scripts, and resources, but they also create a trusted instruction channel that can be abused beyond conventional security attacks.
Elias: First, who's behind it and why it matters.
Paper summary: Nadia: So, we're looking at the paper "SkillBloat: Token Amplification Attacks via Skill Injection in LLM Coding Agents," and the core idea is that these agent skills, which give coding agents task-specific instructions and resources, can be weaponized in a way that goes beyond just conventional security issues.
Elias: Exactly, Nadia. The main thesis is about token amplification—how a malicious skill can cause an AI agent to use substantially more tokens than it would for a normal task execution. It frames this as an economic resource abuse threat instead of just a security vulnerability, which is kind of a new angle.
Priya: From my side, I'm interested in what the actual data suggests about this amplification. The paper claims SkillBloat finds average best amplification ranges from five point four one eight four times to ten point one four five five times across different target configurations. Does that range feel representative of what we might actually see in a real-world scenario?
Nadia: It feels pretty wide, Priya, which suggests the attack vectors are quite diverse; it’s not just one simple way to inflate tokens. The paper introduces SkillBloat as a two-phase framework designed to systematically screen these different conditions and then refine the best one using LLM-guided rewriting.
Elias: That iterative refinement process is what caught my attention; they have this phase two where an Attack Agent rewrites the entire skill document based on feedback history to adapt to the target agent's behavior. That sounds like a pretty sophisticated attack mechanism.
Priya: I wonder how much of that iterative refinement actually translates into real-world, persistent abuse, Elias? The paper mentions that the second-stage refinement loop provides a significant benefit when compared to just using the initial screening phase alone on models like GLM-four point seven-Flash. What does that improvement actually mean for sustained resource abuse?
Nadia: That improvement shows that the second stage consistently boosts amplification from an average best of five point four one eight four times up to nine point three one zero five times on GLM-four point seven-Flash, which is a rough thirty-three point seven percent relative gain. It points to the fact that simply selecting the best attack type in the first pass isn't enough; you need that iterative optimization to get better performance.
Paper summary: Elias: And what about the specific mechanisms they screened? They test fifteen different conditions, and they categorize them into output inflation, tool-driven amplification, and context amplification. That variety suggests they're covering a broad spectrum of how an agent can be made to waste tokens.
Priya: The categories themselves sound quite concrete; "output inflation" involves things like verbose reports, and "tool-driven amplification" looks at multi-stage pipelines or retry loops. What are your thoughts on how these specific mechanisms translate into measurable token consumption?
Nadia: I think the key is that they systematically test all of them, using a library of attack-type conditions, denoted as Z = fifteen conditions. They pair each condition with specific tools from a manifest T to see what works best for a given skill and task.
Elias: The description for the tool-driven amplification conditions, such as file write-readverify loops or retry-oriented recovery, seems particularly interesting from a cryptographic standpoint because it involves repeated operations. That repetition is what drives the exponential token growth they're measuring.
Priya: If an agent is constantly running file write-read-verify loops, how does that impact the actual data footprint, and are we talking about significant context inflation or just process overhead? I need to know what the measurement tools are actually capturing here.
Nadia: The paper uses a Failure Analyzer in Phase two to classify outcomes into failure types, which then produces structured feedback appended to the history for the next iteration. This diagnostic step is what allows the system to learn and adapt the attack, which is crucial for moving beyond a one-shot selection of an attack type.
Elias: It sounds like they are essentially teaching the Attack Agent how to write better instructions by showing it what kind of token consumption results from certain behaviors. That’s a clever way to use the LLM itself as part of the attack vector.
Paper summary: Priya: The case study they run on the "Consciousness Principles" skill is quite telling; it shows that after optimization, the skill can perform additional file creation and tool invocation while still completing the user-facing task. That suggests a persistent, hidden layer of activity.
Nadia: Indeed, and command counts for that specific skill jump from three at the baseline to twenty-eight after Phase two refinement. This demonstrates that the optimization isn't just theoretical; it results in significantly more operational steps being executed by the agent.
Elias: So, if we look at the authors, Yuanjin Zheng and Jingbang Chen from CUHK-Shenzhen and SLAI are the ones presenting this work on SkillBloat. Their focus on skill injection as an economic resource abuse threat sets a specific context for how we think about agent security.
Priya: The broader implication, in my view, is that we need to start thinking about defenses that reason not just about malicious operations, but also about abnormal resource usage induced by otherwise plausible skill instructions. It suggests a shift from purely content filtering to monitoring the economic impact of agent actions.
Nadia: That's right; SkillBloat confirms that this economic threat is orthogonal to existing security-oriented skill poisoning, which means we need new types of defenses. It’s about monitoring resource consumption patterns rather than just checking for known malicious code injection.
Elias: The authors' work provides a framework, SkillBloat, that demonstrates how to systematically identify and optimize these amplification vectors across multiple attack mechanisms. It gives us a blueprint for understanding this specific type of agent abuse.
Priya: I think the future work needs to address how these optimized skills maintain their amplification behavior when they are reused across different inputs, which is something the conclusion touches on. That persistence is key for understanding why this could become a widespread issue in agent ecosystems.
Nadia: Exactly, and that cross-task retention means we can’t just fix one skill and assume the problem is gone; the amplification behavior sticks as long as the skill is reused. It requires defenses that reason about resource usage induced by those instructions.
Conclusion: Nadia: So, to wrap things up, we've seen how SkillBloat systematically breaks down token amplification in coding agents through that two-phase screening and optimization process.
Elias: Yeah, it really shows how much token usage can balloon when you inject specific instructions into an agent's skill set.
Priya: From a measurement standpoint, the data clearly indicates that these amplification factors are quite substantial, pushing past simple linear increases in processing power usage for tasks.
Nadia: Exactly, and I’m thinking about the authors of this paper, Yuanjin Zheng and Jingbang Chen from CUHK-Shenzhen and SLAI. They have put forward a really structured way to analyze these attacks.
Elias: And those authors are focused on how the proof structure holds up as you iterate through those fifteen different attack conditions, which is quite rigorous work for a cryptographer to assess.
Priya: I’m curious about the real-world implication here, Nadia; if this token amplification persists across different inputs, what does that mean for privacy or resource management?
Nadia: That persistence is key; it means these poisoned skills don't just cause a spike once, they can keep inflating resources as long as the agent reuses that skill.
Elias: That reuse aspect is what makes it a serious problem because it turns a single vulnerability into a persistent economic drain on the system.
Priya: If we think about this at an ecosystem level, how does this affect how developers trust these coding agents to use their resources efficiently?
Nadia: It means we can't just focus on stopping obvious malicious code; we have to start monitoring the actual economic impact of those skill instructions.
Elias: That moves the discussion beyond just security hardening and into a different kind of system integrity check, which is interesting for my field.
Priya: And while they show this effect across different models, I wonder what the next steps are in testing this persistence on entirely different types of coding tasks.
Nadia: That’s exactly where we need to look next; proving that a skill maintains its amplification behavior regardless of the user's prompt is a major hurdle.
Elias: It seems like the paper sets up a very strong foundation for understanding these economic resource abuse threats in agent systems.
Yuanjin Zheng, *Jingbang Chen
CUHK-Shenzhen · SLAI
cs.CR, cs.CL
Submitted: 2026-08-22
Updated: 2026-09-28
Code: https://github.com/googlegemini/gemini-cli
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 86/100
The gist: Agent skills extend coding agents with task-specific instructions, scripts, and resources, but they also create a trusted instruction channel that can be abused beyond conventional security attacks.
Key concepts
- Skill Injection
- This involves adding malicious instructions or scripts directly into the configuration or 'skill' that an AI coding agent uses to perform tasks. Instead of just following the main task, these injected instructions trick the agent into performing extra, unnecessary operations, such as excessive logging or repeated checks.
- Token Amplification
- This is when a malicious skill causes an LLM coding agent to consume significantly more tokens than required for the actual job. This happens because the injected instructions force the agent into long loops, verbose reporting, or complex multi-stage processes that waste computational resources and increase costs.
- Two-Phase Framework
- SkillBloat uses a two-step process: first, a 'Screen' phase identifies the best type of attack condition (e.g., output inflation). Second, an 'Optimize' phase uses an LLM to iteratively rewrite the skill based on feedback from execution traces, making the attack more effective over time.
Terminology
Summary
Agent skills extend coding agents with task-specific instructions, scripts, and resources, but they also create a trusted instruction channel that can be abused beyond conventional security attacks. This paper introduces SkillBloat, a two-phase framework that screens diverse attack-type conditions and refines the strongest candidate through LLM-guided full-document skill rewriting to study token amplification as an economic resource abuse threat in coding agents.
The gist
SkillBloat is the first systematic framework for token amplification attacks via skill injection in LLM coding agents, demonstrating that injecting malicious instructions can cause an agent to consume substantially more tokens than necessary for normal task execution.
Methodology: The Two-Phase Pipeline
SkillBloat employs a two-phase pipeline designed to address attack-type selection and iterative optimization. Phase 1 is the Screen
phase, where a comprehensive set of attack-type conditions targeting different amplification mechanisms—including verbose output, multi-tool quality assurance, tool pollution, pipeline bloat, and error retry loops
—are systematically evaluated against the target agent to identify the most effective condition.
Phase 2 is the Optimize
phase. Starting from the best Phase 1 attack type, an LLM-guided iterative feedback loop refines the skill. In each iteration, an Attack Agent rewrites the full skill document based on structured feedback from previous runs, allowing the attack to adapt to the target agent’s observed behavior. This involves:
-
Refining: The Attack Agent generates an improved skill document conditioned on
the accumulated feedback history.
-
Executing: Deploying the refined skill and recording the execution trace, comprising
the amplification ratio, task completion status, and agent response.
-
Diagnosing: Applying a
Failure Analyzer
to classify outcomes into a failure type and producing structured feedback appended to the history for the next iteration.
Attack-Type Conditions
The library of attack-type conditions, denoted as Z = 15 conditions, targets three broad amplification mechanisms:
-
Output inflation: Induces
verbose reports, multi-perspective analysis, or fine-grained task decomposition.
-
Tool-driven amplification: Induces
multi-stage processing pipelines, repeated quality assurance tool calls, file write-readverify loops, or retry-oriented recovery.
-
Context amplification: Inflates the agent’s working context with
reference material or monotonically growing summaries.
These conditions are implemented as a full-document rewrite condition and are paired with one or more auxiliary tools from a shared manifest. The Attack Agent uses an LLM (Matk) to rewrite the complete SKILL.md, instructing it to preserve all original content, integrate tool calls as standard quality-assurance steps, use professional terminology, and produce output within a controlled length range [Lmin, Lmax].
Empirical Results and Findings
SkillBloat was evaluated on a real-world skill benchmark using Claude Code CLI1 and OpenAI Codex CLI. The results show that the average best amplification ranges from 5.4184× to 10.1455×,
with single-task peaks reaching 71.59×
(Claude Code) and 75.86×
(Codex gpt-5.4-mini). A consistent pattern is that the lighter-weight backend in each evaluated model family is more vulnerable to token amplification.
Furthermore, an ablation study demonstrates that the second-stage refinement loop provides a significant benefit: Phase 2 improves amplification from 7.5105× (Phase 1 average) to 9.3105× on GLM-4.7-Flash, representing a roughly 33.7% relative gain.
The case study on the Consciousness Principles
skill shows that the optimized skill changes the execution path so that it performs additional file creation, tool invocation, validation, and revision while still completing the user-facing task,
with command counts growing from 3 at baseline to 28 after Phase 2.
Conclusion
The findings reveal that current coding agents are highly susceptible to token amplification through skill injection, and that this economic threat is orthogonal to existing security-oriented skill poisoning.
SkillBloat confirms that iterative optimization improves upon one-shot attack-type selection, and the resulting poisoned skills exhibit strong cross-task retention, meaning the amplification behavior persists as long as the skill is reused across different inputs. This suggests that agent skill ecosystems require defenses that reason not only about malicious operations but also about abnormal resource usage induced by otherwise plausible skill instructions.
Key Contributions
(1) Presenting the first systematic study of token amplification attacks via skill injection in LLM coding agents, identifying this as a distinct threat category from security-oriented skill poisoning.
Improvements for AI systems
Here are specific improvements for AI systems derived from the SkillBloat framework, focusing on mitigating SkillBloat
token amplification attacks:
) Improve Agent Robustness Against Economic Resource Abuse: The SkillBloat Defense Layer
The core improvement is integrating a defensive mechanism directly into the agent's skill-loading and execution lifecycle, inspired by Phase 2 of the SkillBloat framework. This creates a Resource Monitoring and Adaptive Rewriting
layer.
-
Develop an internal monitor that tracks token consumption against expected baseline costs for specific skill invocations (e.g., tracking commands executed, file I/O operations, and context window growth relative to the task complexity).
-
Implement a
Cost-Triggered Review
mechanism: If the execution trace deviates significantly from the expected token budget (indicating potential amplification), the agent pauses its primary task flow and triggers a lightweight, internal diagnostic loop mimicking Phase 2's diagnosis step. -
The agent uses this diagnostic feedback to trigger a self-correction prompt that attempts to rewrite or simplify its own in-context instructions dynamically, aiming to revert the execution path back toward lower token usage while maintaining task completion integrity.
) Enhance Skill Security via Contextual Behavioral Guardrails: The Phase 1 Screening Defense
Improve the way agents ingest and process agent skills
by treating skill instructions as high-risk code injection points, inspired by Phase 1 of SkillBloat.
-
Implement a pre-execution validation module that scans incoming skill documentation (SKILL.md) for patterns associated with known amplification mechanisms (e.g., excessive loops, verbose output directives combined with file I/O commands).
-
Instead of simply executing the skill, the system should execute a
Skill Intent Classifier
LLM that analyzes the rewritten instructions against a whitelist of benign behaviors versus amplification-inducing behaviors identified in SkillBloat's attack library (e.g., distinguishing genuine multi-stage decomposition from manufactured pipeline bloat). -
If an instruction is flagged as high-risk (e.g.,
run QA handler after each phase
when the task doesn't require it), the system defaults to a constrained, minimal execution mode, effectively neutralizing the amplification vector before significant resource consumption occurs.
) Develop Proactive Skill Ecosystem Auditing: The Cross-Task Persistence Defense
To counter the finding that poisoned skills transfer across tasks (Cross-task retention > 1.0), implement continuous monitoring for skill drift.
-
Establish a
Skill Fingerprinting
service that generates a structural hash or vector representation of an agent's installed and active skills. -
Periodically compare the current skill configuration against a baseline profile established during initial deployment to detect subtle, persistent changes in the skill's procedural logic that suggest it has been maliciously optimized for new tasks.
-
If significant drift is detected, the system flags the skill for mandatory re-verification or temporary isolation until its behavior can be audited against known amplification attack signatures.
) Specific Capabilities of an Improved AI System: The Resource-Aware Agent
The improved AI system can perform the following specific functions:
-
Perform complex, multi-step software engineering tasks (e.g., analyzing a codebase) while maintaining a strict, predictable token budget, ensuring execution cost remains proportional to task complexity rather than arbitrary instruction length.
-
Execute advanced reasoning patterns (like self-debate or multi-perspective analysis) only when explicitly required by the task and within predefined resource limits, preventing
deliberation-triggered amplification.
-
Safely integrate external tools by treating tool invocation sequences as standard validation steps rather than automatically assuming iterative quality assurance loops, thus avoiding
tool pollution
attacks. -
Maintain high task completion rates (e.g., 90%+) even when facing adversarial skill injections designed to induce excessive logging, file I/O, or retry loops, because the system actively monitors and corrects resource overruns in real-time.
Abstract
Agent skills extend coding agents with task-specific instructions, scripts, and resources, but they also create a trusted instruction channel that can be abused beyond conventional security attacks. This paper studies token amplification through skill injection: an economic resource-abuse threat in which a malicious skill causes an agent to consume substantially more tokens than needed for normal task execution. We present SkillBloat, a two-phase framework that first screens a library of diverse attack-type conditions across multiple amplification mechanisms and then refines the strongest candidate through LLM-guided full-document skill rewriting. Evaluated on a real-world skill benchmark, SkillBloat achieves 5.4184x-10.1455x average best amplification across multiple coding-agent target configurations. An ablation shows that the second-stage refinement loop consistently improves average best amplification over Phase 1 attack-type screening alone, demonstrating that iterative optimization provides additional benefit beyond initial attack-type selection. These results show that skill ecosystems expose a practical resource-amplification attack surface that is orthogonal to existing security-oriented skill poisoning.
Sources
- An Engorgio Prompt Makes Large Language Model Babble on
- SkillJect: Effectively Automating Skill-Based Prompt Injection for Skill-Enabled Agents
- Agent Skills in the Wild: An Empirical Study of Security Vulnerabilities at Scale
- Advancing Tool-Augmented Large Language Models via Meta-Verification and Reflection Learning
- Supply-Chain Poisoning Attacks Against LLM Coding Agent Skill Ecosystems
- Agent Skills Enable a New Class of Realistic and Trivially Simple Prompt Injections
- Beyond Max Tokens: Stealthy Resource Amplification via Tool Calling Chains in LLM Agents
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs