No One Architecture Fits All: A Cross-Environment Evaluation of Hierarchical Red Team Agents

summary

Video file (mp4)

The gist

Autonomous red team agents increasingly stress-test AI-enabled cyber defenses by planning strategy and executing multistage attacks, and this study addresses whether observed architectural advantages

In short

The study compared two red team architectures—RL+RL and LLM+LLM—across different cyber environments (CAGE-4 and Cyberwheel at 100/1010 hosts). Results showed that RL+RL performed better in compact networks like CAGE-4, while LLM+LLM excelled in larger networks. This indicates that no single architecture is universally superior; effectiveness depends on the specific network structure and environment.

Key concepts

RL+RL Architecture
This setup uses Reinforcement Learning (RL) for both planning and execution. The planner learns a strategy based on past rewards, and the executor translates that strategy into actions. It is effective when the reward signal is dense and escalation paths are clearly defined within the network.
LLM+LLM Architecture
This setup uses a large language model (LLM) for both planning and execution. The LLM acts as a sophisticated planner, leveraging its vast pre-trained knowledge to formulate complex attack plans. It is favored in larger environments where prior tactical knowledge is needed to overcome complex challenges.
Environment-Dependent Inversion
This describes the finding that the best architecture changes depending on the network. RL+RL wins in CAGE-4, but LLM+LLM wins in Cyberwheel's 1010-host setting. This inversion proves that architectural success is not fixed but is dictated by environmental properties like network size and connectivity.
Privilege Escalation Bottleneck
This refers to the stage where the attack stalls, regardless of the initial success. In large networks, RL agents fail here because they cannot convert access into operational impact. In compact networks, LLM agents succeed in gaining access but struggle to weaponize it further.

Terminology used across episodes

This episode discusses

The paper

No One Architecture Fits All: A Cross-Environment Evaluation of Hierarchical Red Team Agents · Read on arXiv

Ayan Javeed Shaikh, Arunesh Sinha, Nathaniel D. Bastian, *Ankit Shah

Indiana University · Rutgers University · Johns Hopkins University

Autonomous red team agents increasingly stress-test AI-enabled cyber defenses by planning strategy and executing multistage attacks. Reinforcement learning (RL) and large language models (LLMs) offer complementary mechanisms for the planning and execution such agents require, and prior work has combined them in hybrid hierarchies. Yet a given architecture is typically developed and evaluated within a single environment, leaving open whether an observed advantage reflects a generally stronger decision mechanism or merely alignment with a particular setting. We address this gap with a controlled cross-environment comparison of two homogeneous hierarchical red team architectures: an RL planner with an RL executor (RL+RL) and an LLM planner with an LLM executor (LLM+LLM). We evaluate both against expert autonomous defenders in CybORG CAGE-4 and in Cyberwheel at two network scales, across 18 configurations under one unified disruption metric. We find a pronounced environment-dependent inversion. RL+RL wins the compact, densely rewarded CAGE-4 (78.5% disruption success versus 18.0% for the strongest LLM configuration) and the 100-host Cyberwheel network (81.0% versus 50.5%), while a pretrained cybersecurity LLM agent wins the larger, escalation-gated 1010-host Cyberwheel network (55.0% versus 0.0% for RL). A kill-chain analysis explains the inversion through architecture-specific bottlenecks that aggregate success rates conceal.In the 1010-host Cyberwheel network, RL discovers and compromises hosts but stalls at privilege escalation, whereas in CAGE-4, LLM agents obtain privileged access but rarely convert it into operational impact. These results indicate that conclusions drawn in a single environment may not generalize, and that hybrid planner-executor designs should be motivated by specific failure modes rather than the assumption that one architecture is universally preferable.

Transcript

Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.

Nadia: I'm Nadia, and with me are Elias and Priya, guest researcher.

Elias: Today's paper: "No One Architecture Fits All".

Nadia: Autonomous red team agents increasingly stress-test AI-enabled cyber defenses by planning strategy and executing multistage attacks, and this study addresses whether observed architectural advantages generalize across different cyber environments.

Elias: First, who's behind it and why it matters.

Title and authors: Nadia: So, we're looking at this paper titled "No One Architecture Fits All: A Cross-Environment Evaluation of Hierarchical Red Team Agents," which seems to be digging into whether the success of an AI red team approach depends on the specific network it's running in.

Elias: I agree, Nadia, and the authors are testing two very different ways of using reinforcement learning and large language models to plan and execute attacks across varying environments. This whole idea suggests that we can't just pick one architecture and assume it works everywhere without checking its limits against different network structures.

Priya: From a measurement standpoint, I'm curious what this means for the data we collect; are they just looking at raw success rates, or is there some deeper behavioral data showing *why* one architecture performs better than the other in different settings?

Nadia: Exactly, Priya. The core idea of this paper is to compare two specific setups—an RL planner with an RL executor called RL+RL, and an LLM planner with an LLM executor called LLM+LLM—across two distinct environments: CybORG CAGE-four and Cyberwheel, testing at both a one hundred-host scale and a one thousand ten-host scale.

Elias: That setup is what makes it interesting; they are comparing the learned policy approach of RL against the knowledge-based reasoning of LLMs in these different contexts to see which mechanism is more robust.

Priya: And I'm looking at how those environments differ structurally, especially since CAGE-four is described as a "compact, densely rewarded network," while Cyberwheel has a one thousand ten-host topology that "has no cross-subnet interfaces," contrasting with the one hundred-host setup which "permits lateral movement."

Nadia: Right, that structural difference is key because it seems to dictate where each architectural approach finds its footing. The authors are looking at metrics like disruption success, attack chain progression, behavioral efficiency, and planning cost across all these configurations.

Elias: It really gets down to whether the advantage observed in one environment is just a fluke based on how well that specific architecture aligns with the rewards or knowledge available in that particular setting.

Priya: I'm also interested in the results they found regarding performance inversion; it sounds like there's a significant difference between environments when you look at who wins.

Nadia: Well, according to this paper, there's a pronounced environment-dependent inversion in architectural effectiveness when we look at the outcomes. In CybORG CAGE-four the RL+RL setup achieves a disruption success rate of seventy-eight point five percent, which is significantly higher than the eighteen point zero percent seen for their strongest LLM configuration on that same environment.

Title and authors: Elias: That disparity in CAGE-four shows that when the reward signal is dense and immediate, like in CAGE-four the learned policy from RL seems much more effective at achieving high disruption success right out of the gate.

Priya: But then we look at Cyberwheel, and things get complicated because it's tested at two scales. The paper shows that RL+RL still leads in the one hundred-host environment with an eighty-one point zero percent success rate compared to fifty point five percent for LLM+LLM, but the relationship flips dramatically in the larger one thousand ten-host environment where LLM+LLM achieves fifty-five point zero percent and RL+RL achieves a complete failure at zero percent.

Nadia: That reversal in Cyberwheel is quite striking; it shows that what works well for one architecture can completely fail when the network topology or scale changes, which is exactly what the title of "No One Architecture Fits All" is pointing to.

Elias: The authors explain this inversion by looking at architecture-specific bottlenecks that are hidden when you only look at the overall success rates. They pinpoint privilege escalation as a key area where these different architectures struggle in different ways.

Priya: That's interesting because they characterize the failure modes differently too; for instance, in the one thousand ten-host Cyberwheel network, RL agents "discover and compromise hosts but stall at privilege escalation," while in CAGE-four LLM agents "obtain privileged access but rarely convert it into operational impact."

Nadia: It really highlights that the failure isn't just about getting in; it’s about what happens next based on the specific decision-making mechanism of the agent. So, what do we learn from these specific failure modes regarding how we should design these systems?

Elias: The paper suggests that hybrid planner-executor designs shouldn't be motivated by assuming one architecture is universally better, but rather by identifying which failure mode you want to mitigate. For example, if the reward is dense and escalation is reachable from the reward signal, then an RL learned policy seems preferable.

Priya: And conversely, if the network is large and escalation depends on prior tactical knowledge that immediate rewards don't supply, then a capable language model seems like the better choice for that stage.

Nadia: So, in practical terms for defense research and development, the paper suggests we need to prefer an RL learned policy when the reward is dense and escalation is reachable from the reward signal.

Title and authors: Elias: That directly informs how we think about system design; you should look at where your specific success criteria align with either immediate feedback loops or long-horizon tactical knowledge retrieval.

Priya: And this leads to another important point: the shared failure stage, privilege escalation, is identified as being directly actionable for defense, suggesting that hardening efforts should concentrate there rather than just on initial access.

Nadia: That's a clear directive for defensive engineering; instead of focusing solely on getting past the perimeter, we should be focusing our resources where the agents consistently stall across different environments.

Elias: Exactly; the environment properties that cause this inversion also predict the type of threat you're dealing with—compact networks favor learned agents, while large enterprise networks favor language models supplying prior knowledge.

Priya: It’s a very pragmatic conclusion for researchers looking at autonomous offensive capabilities, showing that generalization isn't automatic and requires context-aware design.

Nadia: So, to wrap up this discussion on "No One Architecture Fits All: A Cross-Environment Evaluation of Hierarchical Red Team Agents," the main implication is that the choice between an RL+RL architecture and an LLM+LLM architecture isn't universal.

Elias: We've established that the effectiveness of these architectures shifts depending on whether we are in a compact, densely rewarded setting or a large, sparsely connected one like Cyberwheel.

Priya: I think what this paper really adds is the framework for building more resilient agents that can adapt their planning mechanism based on the environment they encounter.

Nadia: It moves us away from assuming one learning paradigm is inherently superior and pushes us toward designing hybrid systems motivated by specific failure modes, like privilege escalation.

Elias: Indeed, the paper shows that understanding these architecture-specific bottlenecks is crucial for creating effective AI agents in complex cyber landscapes.

Priya: It’s important to remember the limitation they mentioned: rewards are architecture-specific and excluded from cross-architecture performance rankings, meaning we can't just compare raw scores without accounting for those hidden reward structures.

Nadia: That’s a fair limitation to keep in mind when interpreting these results; we have to be careful about how we weigh the performance metrics they report.

Elias: And it sets up the next big question for us: how do we build the mechanisms that allow an AI agent to dynamically decide which architectural approach—RL or LLM based—is best suited for its current operational context?

The paper's summary: Nadia: So, to recap, this paper is really showing that there isn't just one way to build these kinds of autonomous red team agents; their effectiveness totally depends on whether they are in a compact network or a much larger one, and that same architecture might fail spectacularly depending on the environment.

Elias: Exactly what Nadia means; it’s about how the underlying decision-making structure, whether it's reinforcement learning or a language model, performs differently when faced with different network properties like density or scale.

Priya: And from my side, I'm focused on what the actual data reveals about this dependency; it seems like the results show a real inversion in success rates depending on whether we look at CybORG CAGE-four versus Cyberwheel, which is pretty telling for our measurement work.

Nadia: Right, and that inversion isn't just a statistical fluke; it’s tied to fundamental differences in how those two environments reward or constrain the agents' actions.

Elias: I agree; the paper suggests we need to stop treating RL+RL or LLM+LLM as universal solutions and start looking at which one fits the specific constraints of the network topology.

Priya: That makes me think about the real-world impact; if this is true, it means defensive AI systems can't just be built once and deployed everywhere; they have to be tailored to anticipate these environmental shifts.

Nadia: Precisely, Priya; it implies that building a single "best" autonomous attacker is a mistake, and instead, we need flexible frameworks that can dynamically switch between learned policy and language-model reasoning based on the context.

Elias: That flexibility means we have to design systems where the planner itself can evaluate which architectural approach is best suited for the current stage of the attack.

Priya: And this points toward a huge implication for privacy researchers, because if these agents are being used by attackers, understanding how they navigate different network structures could help us build better defenses against those varied threats.

Nadia: It gives us a concrete target; instead of trying to make one agent perfect for everything, we can focus on designing mechanisms that handle the known failure modes—like privilege escalation—differently in each setting.

Elias: And focusing on those specific bottlenecks, rather than just optimizing overall success rates, is where the real cryptographic and security insight lies for us.

Priya: It’s exciting because it shows that context is a critical variable in AI performance, which is something we need to factor into every privacy and measurement study moving forward.

Nadia: So, the big picture here is that future work needs to focus less on comparing architectures in isolation and more on developing dynamic, hybrid systems capable of adapting their planning strategy in real time.

The paper's improvements: Nadia: So, to sum up what we've heard about this paper, the authors are pushing for a more flexible approach where AI red team systems don't rely on just one fixed architecture but instead adapt their planning strategy based on the specific network environment they find themselves in.

Elias: That flexibility means moving away from rigid structures and toward hybrid designs that can switch between different decision-making mechanisms, like using reinforcement learning when immediate rewards are dense and relying on language models when long-term knowledge is more important.

Priya: I think the most important part of their suggestions is the idea to use learned critics to rank subgoals based on observed performance shortfalls in a specific environment, which sounds like a way to make the system self-correct its architectural choice.

Nadia: Exactly, Priya; it’s about giving the AI system a mechanism to judge its own strategic path and choose whether it needs more immediate tactical feedback from an RL planner or broader prior knowledge from an LLM planner.

Elias: That brings up a fascinating point for me regarding the underlying assumptions of these systems; if we can build in that value-guided decoding, we have to ensure the reward function itself is robust enough to guide that choice effectively.

Priya: And I see a massive implication for measurement research; if you can quantify these "plan load-bearing" versus "non-load-bearing" scenarios, it provides a new way to measure the efficacy of different AI reasoning styles in real attacks.

Nadia: Right, and on the security side, this suggests that we should be designing agents with built-in logic to detect when they've hit a known failure mode—like being stuck at privilege escalation—and immediately pivot to a different planning strategy.

Elias: That’s a practical application; it means we’re looking for "escape hatches" in the agent's architecture that allow it to change its fundamental operating model when things go sideways.

Priya: It really shifts the focus from just building a powerful agent to building an intelligent system that understands *when* and *why* a certain type of reasoning is failing in a specific context.

Nadia: That’s the core message, and it tells us that the future isn't about picking one ultimate AI architecture; it's about engineering AI systems with the self-awareness to choose their tool based on the situation.

Conclusion: Tom: So, to wrap things up on "No One Architecture Fits All: A Cross-Environment Evaluation of Hierarchical Red Team Agents," we've seen how these hierarchical red team architectures perform very differently depending on whether they are operating in a compact or a larger network environment.

Nadia: It really boils down to the idea that you can't just deploy one AI planning system and expect it to work everywhere, which is what this paper demonstrates through those stark performance differences across CybORG CAGE-four and Cyberwheel.

Elias: I agree; the core finding is that the architectural strength of an RL planner versus an LLM planner isn't fixed, but depends entirely on the specific constraints of the network topology and how rewards are structured there.

Priya: And from a measurement standpoint, it’s fascinating because it proves that our data needs to be highly contextual; you can't just look at average performance without knowing the environment details first.

Nadia: Exactly, Priya; this research tells us that for defense researchers, the focus needs to shift from finding a universally superior AI architecture to designing systems that are aware of their operational context.

Elias: The implications for cryptographers are interesting because it highlights how architectural assumptions—like the planner's decision-making mechanism—directly impact the success rate of an attack chain, which is what we look for when testing security protocols.

Priya: I think this means that future privacy research has to incorporate environmental modeling; if we want to understand how these agents behave in a real setting, we have to model the environment's structure as much as the agent itself.

Nadia: It’s exciting because it suggests we can start designing more resilient AI defenses that are context-aware and can choose their strategy on the fly based on what they encounter.

Elias: That dynamic capability is key; if an AI system can adapt its planning method, it becomes much harder for us to rely on static security assumptions when testing complex protocols.

Priya: I'm just curious how this feeds into the other papers we've been looking at, like the ones discussing prompt injection or data poisoning; does this environmental sensitivity apply across all those AI safety and resilience studies?

Nadia: It seems to be a common thread; whether it’s an attacker adapting its plan based on network size, or a defense system needing to adapt its model based on the threat environment, the principle of context-dependent effectiveness is central.

Elias: Indeed, that contextual dependency is what separates theoretical models from real-world vulnerabilities in complex systems like these hierarchical red teams.

Priya: So we're looking at a future where AI agents are less about executing a single pre-programmed strategy and more about intelligently selecting the right reasoning tool for the immediate task.

Nadia: That’s a big shift, and I think it’s something that will seriously influence how we approach building any autonomous system in cybersecurity.

Elias: And that's where we need to keep our eyes on; understanding these architectural trade-offs is crucial for figuring out what parameters might cause those very inversions we saw in the experiment.

More episodes

← Home