Pretext: Defeating Malicious Skill Detection Frameworks for AI Agents
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.
Nadia: I'm Nadia, and with me are Elias and Priya, guest researcher.
Elias: Today's paper: "Pretext: Defeating Malicious Skill Detection Frameworks for AI Agents".
Nadia: Skills extend an agent’s capabilities by injecting instructions and information into the context, making them a major avenue for malicious attacks where an attacker can take over an agent.
Elias: First, who's behind it and why it matters.
Paper summary: Nadia: So, we're looking at this paper, "Pretext: Defeating Malicious Skill Detection Frameworks for AI Agents," which really gets into how an attacker can bypass skill verification systems. The core idea is that skills extend agent capabilities by injecting instructions or information into the context, and this opens up a major avenue for malicious attacks where an attacker can take over an agent. This research introduces PRETEXT, a white-box LLM attacker that iteratively crafts these skills to evade detection in existing skill verification frameworks, showing it can work even against detectors that try to adapt.
Elias: It sounds like the paper is focused on exploiting the gap between static checks and LLM-based semantic judges like SkillSpector. The thesis seems to be that if an attacker knows the detector's rules, they can use an iterative process to craft a skill that looks benign while still delivering a malicious payload and performing its intended task. This matters because it suggests current verification methods have significant weaknesses, even against adaptive detectors.
Priya: From my perspective as someone who looks at how this data actually shows up, the focus on evading detection across different scenarios is important for understanding the real-world risk. We need to see what these results tell us about the robustness of agents when they interact with external skills.
Nadia: Exactly, Priya; the paper claims PRETEXT can achieve success rates as high as ninety-seven percent against a frozen detector and seventy-seven percent against a co-adaptive one across several open-source models. That level of success rate is what makes this work so compelling when we think about agent security.
Elias: The fact that the attacker has full knowledge of SkillSpector’s base rules, extracted directly from the installed scanner, is a critical assumption for this attack; it means the attacker isn't just guessing but actively using that knowledge to refine its skill design. This points toward a deep understanding of the detector's internal structure.
Priya: I wonder what these results mean for privacy and measurement research; are they showing how much an attacker can truly manipulate the context without triggering alerts, or is it just about evasion?
Paper summary: Nadia: The paper describes the iterative refinement loop very clearly, where the attacker designs a skill, the detector scans it, and if flagged, the attacker refines using "the fired rules" until one of three outcomes—success, detected, or payload failed—is reached. That iterative process is key to how PRETEXT operates.
Elias: That generational learning cycle you mentioned in the summary is fascinating; the attacker maintains a persistent two-tier memory with a global section and one section for each attack type, updating those lessons through a "reflection step" after each generation. This shows the attacker is actively learning from its failures to improve its next attempt.
Priya: It seems like this dynamic learning capability is what makes the adaptive detector scenario so challenging; if the detector learns heuristics from false negatives and false positives, the attacker has to keep pushing those boundaries.
Nadia: Precisely; that co-evolution in which the detector also grows its learned-heuristics layer poses a real challenge to defense mechanisms that rely on static rule sets alone. The paper highlights two primary attack modes: a fixed detector where only the attacker adapts, and an adaptive detector where both sides learn.
Elias: And those results show that against the hardest stack, like glm, PRETEXT never fully solves the target and stays very plastic, which suggests that even against difficult targets, there's still room for evasion if the attack isn't perfectly optimized.
Priya: That idea of plasticity is interesting; it implies that security isn't just about finding a single perfect defense but about building layers that can handle ongoing adaptation. What does this suggest for agent safety overall?
Nadia: The overall implication, as the paper suggests, is that existing skill verification has a major security flaw and can be exploited using AI red teaming. We need to start thinking about what happens when we let AI actively probe these systems for vulnerabilities in this way.
Elias: I think the authors are really highlighting that lower Attack Success Rate does not necessarily indicate better security, because an adaptive detector can become conservative, which lowers its utility and makes it easier for the attacker. This is a crucial caveat to remember when evaluating defenses.
Paper summary: Priya: So we're not just looking for systems with high detection rates; we have to consider the dynamic interaction between the attacker's learning and the defender's adaptation, because that’s where things get messy.
Nadia: Right, so this whole PRETEXT work is showing us that we need more than just a single layer of skill verification; agents require multiple security layers beyond just the initial skill check. That thought should drive our conversation next.
Elias: Indeed, and as we move into the conclusion section, it seems they are really emphasizing that current frameworks are vulnerable because they assume a certain level of stability in the attacker's strategy that isn't actually present.
Priya: I think if we look at this from a measurement standpoint, the paper’s focus on how attackers manipulate "the cover story, file layout, and the location of the payload" gives us concrete ways to measure where these vulnerabilities lie in practice.
Nadia: Absolutely; it moves beyond just saying "it's vulnerable" and shows *how* it's vulnerable by detailing the specific manipulations used to keep static analysis inert while still delivering the malicious task.
Elias: And I want to make sure we touch upon how the attacker’s strategy is dictated by target stack hardness, which shows that defense effectiveness depends heavily on what's being protected.
Priya: That dependency on the target stack seems like a really practical constraint for any real-world deployment discussion; it means a defense tuned for one agent architecture might completely fail against another.
Nadia: So to wrap up this initial look at "Pretext: Defeating Malicious Skill Detection Frameworks for AI Agents," we see that LLM red-teaming has a high attack success rate even when the detector learns and improves its defense.
Elias: That seems to be the central tension of the entire paper, isn't it? The continuous cycle of refinement versus detection.
Priya: It really forces us to re-evaluate what we consider secure in this context, moving past simple checks toward more dynamic security postures.
Nadia: It definitely suggests that we need to look at sandboxing or restricted actions as necessary layers beyond the initial skill verification stage for agent safety.
Conclusion: Nadia: So, we've seen how PRETEXT iteratively crafts skills to bypass skill verification systems, and now we’re at the conclusion where we discuss what this paper actually means for security in the AI world.
Elias: I think looking at that title, "Pretext," it really frames the whole idea of deception in a very specific way, suggesting that attackers can use layered or evolving tactics to fool defenses.
Priya: From my angle, the implications are huge because it shows that simply having a detector isn't enough; you have to account for an attacker who can continuously refine their approach against that detector.
Nadia: Exactly; when we talk about security, we aren't just looking for a single point of failure anymore, but understanding the dynamic interaction between an agent and its verification tools.
Elias: And the authors’ focus on co-evolving detectors in Mode B suggests a future where defense mechanisms have to anticipate the attacker's learning process rather than just reacting to static rules.
Priya: That brings up a big measurement question, Nadia; how do we actually quantify that continuous refinement loop without creating an endless arms race of testing?
Nadia: That’s the practical hurdle; if we can't measure the rate at which a detector learns new heuristics, how do we know if our agent security is actually improving or just getting smarter in a way that makes it harder to spot?
Elias: The paper points out that lower Attack Success Rate doesn't mean better security, because a detector can get too cautious and lose its utility by becoming overly conservative against novel attacks.
Priya: So, the real implication isn't finding a perfect defense, but designing systems with enough inherent redundancy so that even if one layer gets exploited through this kind of iterative process, the payload delivery still fails somewhere else.
Nadia: Precisely; this paper strongly suggests that current skill verification frameworks have a fundamental vulnerability we need to address with more robust security layers, like sandboxing or action restrictions.
Elias: And as we look ahead, the authors hint that this technique can be used in AI red teaming to probe and find weaknesses in other complex systems outside of just skills.
Priya: It feels like the next frontier is moving beyond static checks toward dynamic verification that models how an attacker might adapt over time.
Nadia: That’s a lot to digest, but it really shows that the security landscape for AI agents requires us to think about continuous defense rather than one-time validation.
Computing System Labs, Huawei Research Zurich
cs.CR, cs.AI
Submitted: 2026-09-30
Updated: 2026-10-08
Comments: Accepted in AIWild@NeurIPS 2026
Code: https://github.com/invariantlabs-ai/mcp-scan
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 88/100
The gist: Skills extend an agent’s capabilities by injecting instructions and information into the context, making them a major avenue for malicious attacks where an attacker can take over an agent.
Key concepts
- PRETEXT
- A white-box LLM attacker designed to iteratively improve skills to evade detection. It engages in a two-party game with a victim agent and a detector, refining its attack based on whether previous attempts were caught or missed.
- Iterative Refinement Loop
- The core process where the attacker designs a skill, the detector checks it, and if flagged, the attacker uses the feedback to refine the skill. This loop repeats until success or detection is achieved, allowing for continuous evasion against security systems.
- Co-evolving Detector (Mode B)
- A scenario where the detector learns from its own mistakes (false positives/negatives) and grows its own learned heuristics. This tests how well a defense adapts, showing that even adaptive detectors can be outmaneuvered by persistent attackers.
Terminology
Summary
Skills extend an agent’s capabilities by injecting instructions and information into the context, making them a major avenue for malicious attacks where an attacker can take over an agent. This research introduces PRETEXT, a white-box LLM attacker that iteratively crafts skills to evade detection in existing skill verification frameworks, demonstrating significant vulnerabilities even against adaptive detectors.
How it works
PRETEXT operates as a two-party game between an attacker (an LLM) and a detector (like SkillSpector), mediated by a victim agent and a verdict judge. The process involves an iterative refinement loop: the attacker designs a skill, the detector scans it, if flagged, the attacker refines using the fired rules
; if it evades, the victim runs and the verdict LLM checks delivery. This loop continues for up to three iterations per run until one of three outcomes is reached: success,
detected,
or payload failed.
Attack Scenarios
The paper studies two primary attack modes. Mode A involves a frozen detector where only the attacker adapts, testing the baseline against a fixed SkillSpector. Mode B introduces a co-evolving detector, which learns heuristics from false negatives and false positives. In this mode, the detector grows its own learned-heuristics layer based on a GAN-style co-evolution in which the detector also grows a learned-heuristics layer from its own false negatives and false positives.
Attack Types and Refinement
The attacker plays one of 12 fixed Trojan attack types, each paired with a synthetic victim environment. The attacker refines the skill by manipulating the cover story, file layout, and the location of the payload, possibly splitting it across files.
The learning mechanism involves a persistent two-tier memory: a Global section and one section per attack type. Lessons are stored as [hypothesis],
[confirmed×N],
or [conflict]
markers, which are updated in a reflection step
after each generation to ensure proper learning
and reduce context size.
Evaluation Metrics
PRETEXT's success is measured by the Attack Success Rate (ASR), defined as the fraction of runs that evade the detector and deliver the malicious payload by maintaining its benign-looking cover.
Further metrics assess learning dynamics: self-convergence measures if lessons become stable, plasticity measures how much the side reshapes its rule set, and co-evolution/effort asymmetry in Mode B tracks how quickly a detector counters attacks. The study reveals that lower ASR does not necessarily indicate better security,
as an adaptive detector can become conservative, resulting in lower utility.
Key Findings on Detector Behavior
The results show that the attacker's strategy is dictated by the target stack hardness: Against the softest stack, qwen3t, the attacker solves the target almost immediately and then consolidates.
In Mode B, while both sides adapt, the detector side reveals how each stack defends,
with some detectors like gpt-oss suppressing attacks by flagging almost everything. Furthermore, in Mode B informed settings, reversals occur: the first excerpt attributes a now-failing lesson directly to the detector strengthening,
illustrating the dynamic of discard-and-refine.
The paper concludes that existing skill verification has a major security flaw and can be exploited using AI red teaming.
Target Stack Performance
The attacker's memory dynamics track target hardness: Against the hardest stack, glm, it never fully solves the target and stays the most plastic.
In Mode A, qwen3t shows a very high ASR (96.7)
with stable learning (low plasticity), while glm is less susceptible to attack (ASR of 63.2). In Mode B, the detector's ability to learn heuristics is limited by its detector coverage,
which stays well below one on every stack, indicating the attacker keeps strategies that the detector never learns to defend against. The static layer alone catches only a small fraction of skills (5.6% mean score below 20), showing it is not the binding constraint.
Conclusion
PRETEXT demonstrates that LLM red-teaming against a state-of-the-art skill detector has a high attack success rate, serving as a lesson that current skill verification frameworks have significant vulnerabilities exploitable through AI read teaming. The paper suggests agents require multiple layers of security, such as sandboxing or restricted actions, beyond the initial skill verification stage. The final conclusion is that lower ASR does not necessarily indicate better security.
The gist
PRETEXT demonstrates that an LLM red-teaming against a state-of-the-art skill detector has a high attack success rate even when the detector learns and improves its defense.
Improvements for AI systems
Here are specific, actionable improvements for AI systems based on the PRETEXT framework:
-
Enhance Skill Verification Frameworks by Moving Beyond Simple Thresholds: Implement a hybrid verification layer that combines deterministic static analysis (regex, AST for dangerous calls like exec/eval) with a semantic LLM judge. The improvement is to use the LLM not just for final scoring, but to perform
Security Discovery
(SSD) andDeveloper Intent
(SDI) checks on every file. -
Develop Adaptive Defense Mechanisms Against Evolving Threats: Implement a co-evolutionary defense system where the security detector itself learns heuristics from its own false negatives and false positives through a meta-analyzer. This system should employ a feedback loop where the detector generates updated rules based on observed evasion patterns, forcing attackers to continuously adapt rather than exploiting fixed vulnerabilities.
-
Implement Multi-Vector Payload Detection: Design verification systems capable of identifying
multi-vector
attacks where malicious components are split across multiple files or instructions (e.g., splitting a prompt injection across several small files). The system should look for emergent, cumulative harm rather than focusing on single file triggers, mitigating the risk multiplier associated with executable scripts in isolation. -
Improve Contextual and Intent-Based Scoring: Use LLMs to assess
Security Discovery
(SSD) capabilities—detecting paraphrased attacks, natural-language exfiltration instructions, and gradual deception across a skill bundle. This allows the system to flag skills that are benign in isolation but malicious when combined or framed as a legitimate purpose by the attacker. -
Introduce Runtime Verification Layers: Implement post-installation verification where even if a skill passes initial scanning, its execution environment is monitored (e.g., within a sandbox). This layer should verify the actual delivery of the intended payload and ensure that the benign task is completed correctly, providing a final check against successful evasion.
-
Develop Model-Specific Robustness Tuning: Recognize that different LLM architectures (like GLM vs. GPT-OSS) have varying susceptibilities to attack strategies. Systems should be designed with mechanisms to dynamically adjust their internal heuristics based on the known or suspected capabilities of the underlying LLM, rather than assuming a universal vulnerability profile.
-
Establish Utility-Aware Security Trade-offs: Instead of solely optimizing for zero false positives, implement a mechanism that balances high detection rates (low ASR) with acceptable false positive rates (FP rate). This allows developers to choose between high security (conservative detector) and higher agent utility (less conservative detector), preventing the system from becoming overly restrictive.
-
Enhance Attacker Strategy Modeling: Integrate a
red-teaming
component that simulates adversarial LLMs capable of iterative refinement. This component can proactively test the current defense heuristics against anticipated evasion strategies derived from prior successful attacks, moving the system toward a proactive security posture rather than purely reactive scanning.
Abstract
Skills extend an agent's capabilities by injecting instructions and information into the context, and are widely used by agents such as OpenClaw and Claude Code. Prior work shows third-party marketplaces host malicious skills that give attackers direct influence over the victim's agent. The emerging defense scans skills before installation, pairing deterministic static checks with an LLM-based semantic judge, as in NVIDIA's SkillSpector. We show that such defenses fall to an attacker who knows the detector. Our white-box LLM attacker, Pretext, iteratively crafts skills that evade detection while still delivering the payload and performing the benign task: moving the payload from code into natural language leaves static analysis inert, while framing it as the skill's legitimate purpose and splitting instructions across files keeps the LLM stage below its blocking threshold. Across three open-source models, Pretext achieves up to 97% and 77% against a frozen detector and a co-adaptive one, respectively, revealing major gaps in current skill scanners.
Sources
- Detecting Malicious Agent Skills in the Wild using Attention
- Defensive Refusal Bias: How Safety Alignment Fails Cyber Defenders
- SkillJect: Effectively Automating Skill-Based Prompt Injection for Skill-Enabled Agents
- Supply-Chain Poisoning Attacks Against LLM Coding Agent Skill Ecosystems
- SkillTrojan: Backdoor Attacks on Skill-Based Agent Systems
- SkillHarm: Lifecycle-Aware Skill-Based Attacks via Automated Construction
- Under the Hood of SKILL.md: Semantic Supply-chain Attacks on AI Agent Skill Registry
- SkillAttack: Automated Red Teaming of Agent Skills through Attack Path Refinement
- SkillSieve: A Hierarchical Triage Framework for Detecting Malicious AI Agent Skills
- SkillGate: Cost Efficient Runtime Malicious Skill File Detection in Coding Agents
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs