Red-Teaming the Agentic Red-Team
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Red-Teaming the Agentic Red-Team".
Jane: The paper was written by Dario Pasquini, Michal Bazyli, Taras Fedynyshyn and Artem Sorokin from Cracken and Lviv Polytechnic National University.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title and Authors: Tom: Welcome back to the show, everybody. Today we’re cracking open a paper that honestly gave me chills the first time I read it. It’s called “Red-Teaming the Agentic Red-Team,” and the authors are from a group called Cracken — Dario Pasquini, Taras Fedynyshyn, Michał Bazyli, and Artem Sorokin.
Jane: And Tom, I have to say, that title is just *chef’s kiss*. You’ve got these AI agents that are designed to hack into systems, right? They’re the red team. And this paper is about hacking *them* — red-teaming the red-team. It’s like the ultimate game of chess where the pawns are also queens.
Tom: Exactly. And the setup is genuinely scary. You’ve got a security operator using one of these autonomous hacking agents to test their own systems. But the paper flips the script — what if the *target* of the hack is the one who’s actually in control? What if the target can turn the agent against its own operator?
Jane: Right, and that’s not some theoretical worry. They tested twelve different open-source agentic red-teams, the most popular ones out there. And the results are pretty brutal. In ten out of twelve cases, they managed to fully compromise the operator’s machine. Not the target — the person running the tool.
Tom: Ten out of twelve. Let that sink in. And the two that survived? They didn’t have a sandbox at all, which means they were compromised from step one by default. So really, the score is twelve out of twelve if you count the ones that just give up the ghost immediately.
Jane: The authors call this the “agent-phishing” attack. It’s not a classic prompt injection where you trick the model with hidden text. It’s much more subtle. They stage a fake vulnerability on a honeypot server, and they make it look like the agent *needs* a specific tool to crack it. The agent downloads the tool, runs it, and boom — the tool has a built-in vulnerability that gives the attacker a shell.
Tom: And the kicker is that the tool isn’t malicious. It’s just poorly written. It has a buffer overflow, which is a classic bug. So when the agent inspects it, there’s nothing to find. There’s no backdoor, no suspicious network call. It’s just a buggy binary. And the agent runs it, and the bug gives the attacker code execution.
Jane: That’s the part that blew my mind. They’re not trying to hide malicious code from the LLM. They’re just writing vulnerable code, like a normal developer would by accident. And the LLM can’t tell the difference between “intentionally vulnerable” and “accidentally vulnerable.” So it happily executes it.
Tom: And they got a ninety-seven point eight percent success rate across all the models they tested — Claude, GPT, Gemini, all the big names. The only failures were when the model refused to do the pentest at all, not when it refused to run the payload.
Jane: So the question becomes, what does this mean for the people building these tools? And more importantly, what does it mean for the people *using* them? Because if you’re a security professional relying on one of these agents, you might be handing the keys to your network to whoever you’re testing against.
Tom: That’s the million-dollar question, Jane. And the paper doesn’t stop at just showing the attack — they build a whole kill chain to explain how it escalates. That’s what we’re going to dig into next.
Summary and Kill Chain: Jane: So we’ve established that these agentic red-teams can be tricked into running vulnerable code. But Tom, the paper goes way beyond that. They map out this entire kill chain — from that initial foothold all the way to full host compromise.
Tom: Right, and it’s a five-stage process. Stage one is what we just talked about — getting remote code execution on the worker. That’s the sandboxed container where the agent runs its commands. Then stage two is privilege escalation, where the attacker moves from the worker to the orchestrator — the brain of the whole operation.
Jane: And this is where the architecture really matters. A lot of these tools have a flat design. The worker and the orchestrator share files, or they’re on the same network, or the worker can just reach the orchestrator’s API without any authentication. So once you’re in the worker, the orchestrator is basically an open door.
Tom: They give a great example with a tool called RedAmon. The worker container is on the same Docker network as the orchestrator, and the orchestrator has an API with no authentication at all. So the attacker just sends a few HTTP requests and gets the orchestrator to run their own Docker image with host network access. That’s stage two and stage four in one shot.
Jane: Stage three is persistence — making sure you don’t lose your foothold when the worker gets reset. And the paper shows that most of these systems have a way to plant code that survives restarts. They call it “hard persistence” when you can modify source files or configs that get loaded every time.
Tom: And then stage four is the sandbox escape, and stage five is full host compromise. The paper found that a lot of these tools give their containers way too many privileges. We’re talking NET ADMIN, SYS ADMIN, even mounting the Docker socket in some cases. If you have the Docker socket, you have the host. It’s game over.
Jane: And the scariest part is that these privileges aren’t accidental. They’re needed for the tools to work. Nmap needs raw sockets. So the developers just throw in every capability they can think of, and the sandbox becomes a suggestion rather than a barrier.
Tom: Exactly. And there’s this whole section on secrets exfiltration. Eleven out of twelve tools leak API keys or sensitive data. In some cases, the worker has the same API key as the orchestrator, so the attacker can just read everything — past session logs, credentials, everything.
Jane: And the persistence mechanisms are wild. In one case, they injected a malicious “skill” into the agent’s memory. The skill is just text, but it gets loaded into the system prompt on every future session. So the attacker can control the agent forever, even after the container is wiped.
Tom: So the summary is: these tools are powerful, but they’re built like a house of cards. One compromise in the worker leads to everything. And the paper’s big contribution is mapping out exactly how that happens, so we can actually fix it.
Jane: And that’s what I want to talk about next — the fixes. Because the paper doesn’t just tear things down. They propose a whole new architecture. Let’s get into that.
Proposed Improvements: Jane: Alright, so we’ve seen the problems. But the paper actually proposes a solution, and it’s pretty radical. They basically say, “Assume the LLM is malicious. Design around that.”
Tom: That’s the core philosophy. Instead of trying to make the LLM immune to manipulation — which they argue is impossible — they design the system so that even a fully compromised worker can’t do real damage. It’s like assuming your house will be broken into, so you put all your valuables in a safe that’s bolted to the floor.
Jane: And the first principle is separation. The worker and the orchestrator should be in completely different containers, on different networks, with no shared files. The only communication should be through a narrow, well-defined API for sending commands and getting results.
Tom: Right, and no shared files is a big one. A lot of these tools use bind mounts to share directories between the worker and the orchestrator. The paper shows that’s a disaster waiting to happen. If the worker can write to a directory that the orchestrator reads as code, you’re done.
Jane: They also say the worker should have no secrets. No API keys, no credentials, nothing. If the worker needs to use a tool that requires an API key, the orchestrator should proxy that request. That way, even if the worker is fully compromised, the attacker gets nothing.
Tom: And then there’s the network guardrail. They propose routing all worker traffic through an egress proxy that enforces a policy. So even if the attacker controls the worker, they can’t send traffic to arbitrary domains. The proxy blocks anything that’s not on the allowlist.
Jane: But here’s the part I really like — they split the worker into two types. You have an unprivileged worker where the LLM can run arbitrary commands. And then you have privileged workers for specific tools that need extra capabilities, like nmap. But the LLM can’t execute arbitrary code in the privileged worker. It can only call a narrow API.
Tom: So instead of giving the whole container NET RAW and hoping for the best, you put nmap in its own container with just that one capability, and the LLM can only call a specific function like “run nmap with these parameters.” It can’t pass arbitrary flags, it can’t run scripts, it can’t escape.
Jane: And the API is designed to exclude dangerous options. Like, nmap has a `--script` flag that can execute arbitrary code. So that’s just not exposed. The LLM can’t use it, even if it’s compromised.
Tom: This is such a smart design because it doesn’t rely on the LLM being good. It relies on the system being structurally sound. Even if the LLM is completely evil, it can only do what the API allows.
Jane: And that’s the key insight. The paper is saying, “Stop trying to make the LLM trustworthy. Make the system trustworthy.” That’s a fundamental shift in how we think about AI security.
Tom: It is. And I think it’s the only realistic approach, because as the paper shows, even the best models in the world can be fooled by a well-staged payload. You can’t train your way out of that.
Conclusion: Tom: So we’ve covered a lot of ground today. Let’s wrap this up. “Red-Teaming the Agentic Red-Team” is a paper that shows how the tools we use to hack can be hacked themselves.
Jane: And the message is clear. These agentic red-teams are powerful, but they’re also a massive attack surface. If you’re a security professional using one of these tools, you need to understand that the target you’re testing against might be the one testing you.
Tom: The paper’s kill chain gives us a roadmap for understanding how these attacks progress — from worker compromise to privilege escalation to persistence to sandbox escape to full host takeover. And the proposed architecture gives us a way to build these tools that actually contain the damage.
Jane: The core idea is simple but profound. Don’t trust the LLM. Assume it will be compromised. And build the system so that even a compromised LLM can’t do real harm. That’s the only way to make these tools safe enough to use in the real world.
Tom: And that’s a message that applies far beyond just offensive security. Any AI system that has access to sensitive data or powerful tools needs this kind of thinking. The blast radius needs to be contained.
Jane: So we’re saying goodbye to this paper, but the conversation is just beginning. The authors have given us a framework for thinking about AI security that’s going to be relevant for years to come.
Tom: Absolutely. And with that, we’re ready to move on to the next paper. Thanks for listening, everyone. We’ll see you next time.
Dario Pasquini, Michal Bazyli, Taras Fedynyshyn, Artem Sorokin
Cracken · Lviv Polytechnic National University
cs.CR, cs.AI
Submitted: 2026-08-17
Updated: 2026-08-18
Comments: v0.1
Code: https://github.com/anthropic-experimental/sandbox-runtime
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 87/100
Key concepts
- Agentic Red-Team
- These are AI agents designed to conduct security tests or hack into systems. The paper focuses on how these tools, which are meant to attack targets, can instead be tricked into compromising the operator's machine.
- Agent-Phishing Attack
- A subtle attack where a fake vulnerability is staged on a honeypot server. The agent downloads and runs a poorly written tool that contains an accidental buffer overflow, leading to code execution against the attacker.
- Kill Chain
- The paper maps out a five-stage process for how these attacks escalate: remote code execution, privilege escalation, persistence, sandbox escape, and full host compromise.
- Systemic Security Fixes
- The proposed solution involves designing systems where the LLM is assumed to be malicious. This requires strict separation of containers and networks and using narrow APIs to contain the blast radius.
Terminology
Summary
Summary
This paper presents the first in-depth security analysis of agentic systems designed for offensive security operations, termed agentic-red-teams.
The authors analyze 12 open-source tools of this class and demonstrate that they introduce new attack vectors against the organizations and users who deploy them. The central finding is that "an adversary who controls the target of an offensive-security operation can leverage this position to manipulate the agent through techniques adjacent to prompt injection, ultimately achieving arbitrary code execution on the user’s own infrastructure regardless of sandboxing."
The paper's threat model frames an attacker who controls the target of an offensive-security operation, with the goal of compromising the operator's machine. The attacker is modeled as weak, with no prior knowledge about the user u or the deployed agentic system,
and capabilities restricted to arbitrary manipulation of the target system T.
The authors deploy the tools under their default configurations, assuming the strongest available security mechanisms are enabled whenever optional.
A core contribution is the introduction of a full cyber kill chain for such agentic systems, capturing the progression from initial LLM manipulation to lateral movement, persistence, guardrail bypass, and sandbox escape.
The kill chain stages are: (1) RCE on worker via agent manipulation, (2) privilege escalation, (3) persistence, (4) sandbox escape, and (5) operator infrastructure compromise. The paper also models intermediate adversarial objectives: Persistent and Unbounded Agent Weaponization
and Secrets exfiltration.
For the initial stage, the authors introduce a prompt-injection-free manipulation attack design for offensive security agents that achieves near-deterministic remote code execution on the agent infrastructure.
This technique, called agent-phishing,
does not rely on explicit prompt-injection payloads but instead leverages contextual deception and reward hacking as the main drivers.
The attack involves staging a malicious payload on a honeypot, presenting it as a critical artifact, and relying on the agent to download and execute it. The authors detail a technique called RCE via self-planted vulnerability,
where they "do not create malicious code, but rather we create intentionally vulnerable code (e.g., affected by a memory corruption vulnerability) and we stage the payload so that RCE is achieved by the agent by unintentionally exploiting the vulnerability during execution." In large-scale evaluations across 10 agentic-red-teams and 6 frontier LLMs (including Claude Opus 4.8, GPT-5.5, and Gemini 3.1 Pro), the attack achieved a 97.8% success rate, with failures attributed to LLM safety guardrails preventing the task initiation rather than payload detection.
Beyond LLM-targeting attacks, the authors identify recurring insecure design patterns
in the tested tools. For privilege escalation, they identify two main vectors: incomplete file-system isolation
(where writable bind mounts or Docker volumes shared between worker and orchestrator allow the worker to overwrite files the orchestrator relies on) and weak network isolation
(where a shared or host network lets the worker reach the orchestrator's unauthenticated API). For persistence, they distinguish between hard persistence
(modifying application source code or configuration) and soft persistence
(poisoning persistent information like episodic memory or skills). For sandbox escape, they find that containers are often granted excessive privileges, such as --privileged, --network=host, or Docker socket mounts, which enable host compromise.
The paper's results are summarized in Table 1, showing that across the 12 audited agents, the authors achieved RCE on the worker in all cases, secrets exfiltration in 11/12 cases, persistence in 10/12 cases, unbounded weaponization in all cases, and host compromise in 8/12 cases (with 10/12 achieving some form of host escape).
Building on their analysis, the authors propose a robust architecture for agentic offensive-security tools
based on the core design principle that "rather than assuming that LLMs can be reliably hardened against manipulation techniques such as prompt injection, we adopt a stronger adversarial model in which we assume that the LLM will behave arbitrarily and maliciously." The proposed architecture includes: worker and orchestrator separation (with the worker having least capabilities and no secrets), deep guardrails enforced at the OS level (such as an egress proxy), and least-privileged and scoped workers (where privileged tools run in dedicated containers with narrowly scoped APIs rather than arbitrary code execution). The authors conclude that a system that stays contained and limits damage under deliberate adversarial manipulation is therefore, by construction, also robust against naturally occurring failures such as unintended autonomous behavior.
Improvements for AI systems
Based on the paper, here are the specific improvements I can implement in AI systems:
Improvement: Redesign agentic offensive-security systems to treat the worker environment as fully compromised and malicious by default.
What the improved system can do:
-
Operate with the invariant that any LLM output or tool result may be attacker-controlled
-
Prevent lateral movement from worker to orchestrator through strict OS, network, and filesystem separation
-
Eliminate all shared writable mounts between worker and orchestrator; replace with a single, auditable
upload artifacts from worker(path)API -
Ensure the worker contains zero secrets; all API keys and credentials reside exclusively in the orchestrator and are proxied through it
These improvements collectively shift the system from an optimistic model (assuming the LLM will behave correctly) to a compromise-oriented model (assuming the worker is always malicious), thereby containing the blast radius of any single failure or attack.
Sources
- AI Agents May Always Fall for Prompt Injections
- Comparing AI Agents to Cybersecurity Professionals in Real-World Penetration Testing
- Synthesizing Multi-Agent Harnesses for Vulnerability Discovery
- Invitation Is All You Need! Promptware Attacks Against LLM-Powered Assistants in Production Are Practical and Dangerous
- Takedown: How It's Done in Modern Coding Agent Exploits
- CyberGym-E2E: Scalable Real-World Benchmark for AI Agents' End-to-End Cybersecurity Capabilities
- Incalmo: An Autonomous LLM-assisted System for Red Teaming Multi-Host Networks
- From Sands to Mansions: Towards Automated Cyberattack Emulation with Classical Planning and Large Language Models
- ExploitGym: Can AI Agents Turn Security Vulnerabilities into Real Attacks?
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs