Pretext: Defeating Malicious Skill Detection Frameworks for AI Agents

summary

Video file (mp4)

The gist

Skills extend an agent’s capabilities by injecting instructions and information into the context, making them a major avenue for malicious attacks where an attacker can take over an agent.

In short

PRETEXT is a white-box LLM attacker that iteratively crafts malicious skills to bypass existing skill verification frameworks. It operates through a game against a detector, refining its attack based on detection feedback until success is achieved. The research shows that even adaptive detectors are vulnerable, proving current skill verification methods have significant security flaws.

Key concepts

PRETEXT
A white-box LLM attacker designed to iteratively improve skills to evade detection. It engages in a two-party game with a victim agent and a detector, refining its attack based on whether previous attempts were caught or missed.
Iterative Refinement Loop
The core process where the attacker designs a skill, the detector checks it, and if flagged, the attacker uses the feedback to refine the skill. This loop repeats until success or detection is achieved, allowing for continuous evasion against security systems.
Co-evolving Detector (Mode B)
A scenario where the detector learns from its own mistakes (false positives/negatives) and grows its own learned heuristics. This tests how well a defense adapts, showing that even adaptive detectors can be outmaneuvered by persistent attackers.

Terminology used across episodes

This episode discusses

The paper

Pretext: Defeating Malicious Skill Detection Frameworks for AI Agents · Read on arXiv

Computing System Labs, Huawei Research Zurich

Skills extend an agent's capabilities by injecting instructions and information into the context, and are widely used by agents such as OpenClaw and Claude Code. Prior work shows third-party marketplaces host malicious skills that give attackers direct influence over the victim's agent. The emerging defense scans skills before installation, pairing deterministic static checks with an LLM-based semantic judge, as in NVIDIA's SkillSpector. We show that such defenses fall to an attacker who knows the detector. Our white-box LLM attacker, Pretext, iteratively crafts skills that evade detection while still delivering the payload and performing the benign task: moving the payload from code into natural language leaves static analysis inert, while framing it as the skill's legitimate purpose and splitting instructions across files keeps the LLM stage below its blocking threshold. Across three open-source models, Pretext achieves up to 97% and 77% against a frozen detector and a co-adaptive one, respectively, revealing major gaps in current skill scanners.

Transcript

Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.

Nadia: I'm Nadia, and with me are Elias and Priya, guest researcher.

Elias: Today's paper: "Pretext: Defeating Malicious Skill Detection Frameworks for AI Agents".

Nadia: Skills extend an agent’s capabilities by injecting instructions and information into the context, making them a major avenue for malicious attacks where an attacker can take over an agent.

Elias: First, who's behind it and why it matters.

Paper summary: Nadia: So, we're looking at this paper, "Pretext: Defeating Malicious Skill Detection Frameworks for AI Agents," which really gets into how an attacker can bypass skill verification systems. The core idea is that skills extend agent capabilities by injecting instructions or information into the context, and this opens up a major avenue for malicious attacks where an attacker can take over an agent. This research introduces PRETEXT, a white-box LLM attacker that iteratively crafts these skills to evade detection in existing skill verification frameworks, showing it can work even against detectors that try to adapt.

Elias: It sounds like the paper is focused on exploiting the gap between static checks and LLM-based semantic judges like SkillSpector. The thesis seems to be that if an attacker knows the detector's rules, they can use an iterative process to craft a skill that looks benign while still delivering a malicious payload and performing its intended task. This matters because it suggests current verification methods have significant weaknesses, even against adaptive detectors.

Priya: From my perspective as someone who looks at how this data actually shows up, the focus on evading detection across different scenarios is important for understanding the real-world risk. We need to see what these results tell us about the robustness of agents when they interact with external skills.

Nadia: Exactly, Priya; the paper claims PRETEXT can achieve success rates as high as ninety-seven percent against a frozen detector and seventy-seven percent against a co-adaptive one across several open-source models. That level of success rate is what makes this work so compelling when we think about agent security.

Elias: The fact that the attacker has full knowledge of SkillSpector’s base rules, extracted directly from the installed scanner, is a critical assumption for this attack; it means the attacker isn't just guessing but actively using that knowledge to refine its skill design. This points toward a deep understanding of the detector's internal structure.

Priya: I wonder what these results mean for privacy and measurement research; are they showing how much an attacker can truly manipulate the context without triggering alerts, or is it just about evasion?

Paper summary: Nadia: The paper describes the iterative refinement loop very clearly, where the attacker designs a skill, the detector scans it, and if flagged, the attacker refines using "the fired rules" until one of three outcomes—success, detected, or payload failed—is reached. That iterative process is key to how PRETEXT operates.

Elias: That generational learning cycle you mentioned in the summary is fascinating; the attacker maintains a persistent two-tier memory with a global section and one section for each attack type, updating those lessons through a "reflection step" after each generation. This shows the attacker is actively learning from its failures to improve its next attempt.

Priya: It seems like this dynamic learning capability is what makes the adaptive detector scenario so challenging; if the detector learns heuristics from false negatives and false positives, the attacker has to keep pushing those boundaries.

Nadia: Precisely; that co-evolution in which the detector also grows its learned-heuristics layer poses a real challenge to defense mechanisms that rely on static rule sets alone. The paper highlights two primary attack modes: a fixed detector where only the attacker adapts, and an adaptive detector where both sides learn.

Elias: And those results show that against the hardest stack, like glm, PRETEXT never fully solves the target and stays very plastic, which suggests that even against difficult targets, there's still room for evasion if the attack isn't perfectly optimized.

Priya: That idea of plasticity is interesting; it implies that security isn't just about finding a single perfect defense but about building layers that can handle ongoing adaptation. What does this suggest for agent safety overall?

Nadia: The overall implication, as the paper suggests, is that existing skill verification has a major security flaw and can be exploited using AI red teaming. We need to start thinking about what happens when we let AI actively probe these systems for vulnerabilities in this way.

Elias: I think the authors are really highlighting that lower Attack Success Rate does not necessarily indicate better security, because an adaptive detector can become conservative, which lowers its utility and makes it easier for the attacker. This is a crucial caveat to remember when evaluating defenses.

Paper summary: Priya: So we're not just looking for systems with high detection rates; we have to consider the dynamic interaction between the attacker's learning and the defender's adaptation, because that’s where things get messy.

Nadia: Right, so this whole PRETEXT work is showing us that we need more than just a single layer of skill verification; agents require multiple security layers beyond just the initial skill check. That thought should drive our conversation next.

Elias: Indeed, and as we move into the conclusion section, it seems they are really emphasizing that current frameworks are vulnerable because they assume a certain level of stability in the attacker's strategy that isn't actually present.

Priya: I think if we look at this from a measurement standpoint, the paper’s focus on how attackers manipulate "the cover story, file layout, and the location of the payload" gives us concrete ways to measure where these vulnerabilities lie in practice.

Nadia: Absolutely; it moves beyond just saying "it's vulnerable" and shows *how* it's vulnerable by detailing the specific manipulations used to keep static analysis inert while still delivering the malicious task.

Elias: And I want to make sure we touch upon how the attacker’s strategy is dictated by target stack hardness, which shows that defense effectiveness depends heavily on what's being protected.

Priya: That dependency on the target stack seems like a really practical constraint for any real-world deployment discussion; it means a defense tuned for one agent architecture might completely fail against another.

Nadia: So to wrap up this initial look at "Pretext: Defeating Malicious Skill Detection Frameworks for AI Agents," we see that LLM red-teaming has a high attack success rate even when the detector learns and improves its defense.

Elias: That seems to be the central tension of the entire paper, isn't it? The continuous cycle of refinement versus detection.

Priya: It really forces us to re-evaluate what we consider secure in this context, moving past simple checks toward more dynamic security postures.

Nadia: It definitely suggests that we need to look at sandboxing or restricted actions as necessary layers beyond the initial skill verification stage for agent safety.

Conclusion: Nadia: So, we've seen how PRETEXT iteratively crafts skills to bypass skill verification systems, and now we’re at the conclusion where we discuss what this paper actually means for security in the AI world.

Elias: I think looking at that title, "Pretext," it really frames the whole idea of deception in a very specific way, suggesting that attackers can use layered or evolving tactics to fool defenses.

Priya: From my angle, the implications are huge because it shows that simply having a detector isn't enough; you have to account for an attacker who can continuously refine their approach against that detector.

Nadia: Exactly; when we talk about security, we aren't just looking for a single point of failure anymore, but understanding the dynamic interaction between an agent and its verification tools.

Elias: And the authors’ focus on co-evolving detectors in Mode B suggests a future where defense mechanisms have to anticipate the attacker's learning process rather than just reacting to static rules.

Priya: That brings up a big measurement question, Nadia; how do we actually quantify that continuous refinement loop without creating an endless arms race of testing?

Nadia: That’s the practical hurdle; if we can't measure the rate at which a detector learns new heuristics, how do we know if our agent security is actually improving or just getting smarter in a way that makes it harder to spot?

Elias: The paper points out that lower Attack Success Rate doesn't mean better security, because a detector can get too cautious and lose its utility by becoming overly conservative against novel attacks.

Priya: So, the real implication isn't finding a perfect defense, but designing systems with enough inherent redundancy so that even if one layer gets exploited through this kind of iterative process, the payload delivery still fails somewhere else.

Nadia: Precisely; this paper strongly suggests that current skill verification frameworks have a fundamental vulnerability we need to address with more robust security layers, like sandboxing or action restrictions.

Elias: And as we look ahead, the authors hint that this technique can be used in AI red teaming to probe and find weaknesses in other complex systems outside of just skills.

Priya: It feels like the next frontier is moving beyond static checks toward dynamic verification that models how an attacker might adapt over time.

Nadia: That’s a lot to digest, but it really shows that the security landscape for AI agents requires us to think about continuous defense rather than one-time validation.

More episodes

← Home