Towards Hierarchical Cyber Defense with Large Language Models: From Planning to Execution
summary
The gist
An autonomous cyber defender trained with reinforcement learning (RL) is typically tied to its training network, limiting its generalization, and this research investigates whether frozen, zero-shot
In short
The research tested if frozen, zero-shot Large Language Models (LLMs) could control a cyber defense system across different network sizes without retraining. The study introduced a hierarchical framework separating strategic planning from tactical execution. Findings showed that while planning alone is limited, extending LLM control to tactical execution significantly improved performance across large networks.
Key concepts
- Hierarchical Framework
- This design separates the defense into two distinct levels: a planner that decides which subnet to defend (strategic targeting) and an executor that performs the actual defensive actions within that chosen area (tactical execution). This separation allows researchers to test how well different control methods work at each level.
- Planner-Executor Hierarchy
- A structure where the planner makes high-level decisions every 'k' steps, selecting a goal for a subnet. The executor then takes over, making step-by-step defensive choices conditioned on that specific goal. This abstraction helps analyze how strategic goals translate into concrete actions.
- Cross-Scale Performance
- This refers to measuring the defense system's effectiveness as the network size changes—from small to large. The study found that a frozen LLM control method maintained strong defensive performance across these different scales, demonstrating its robustness without needing new training for each size.
Terminology used across episodes
This episode discusses
- Towards Hierarchical Cyber Defense with Large Language Models: From Planning to Execution · Paper Radio
- Llama-3.1-FoundationAI-SecurityLLM-Base-8B Technical Report
- On Autonomous Agents in a Cyber Defence Environment
The paper
Towards Hierarchical Cyber Defense with Large Language Models: From Planning to Execution · Read on arXiv
Harshith Doppalapudi, Nathaniel D. Bastian, Ankit Shah
Indiana University · Johns Hopkins University
An autonomous cyber defender trained with reinforcement learning (RL) is typically tied to the network on which it was trained, limiting its ability to generalize as network scale changes. Hierarchical RL reduces decision complexity by separating strategic targeting from tactical execution, but it does not eliminate this retraining dependence. We investigate whether frozen, zero-shot large language models (LLMs) can provide retraining-free control in hierarchical cyber defense and how performance changes as LLM control is extended from planning to execution. We formulate a controller-agnostic planner-executor hierarchy in which the planner selects a subnet to defend over a fixed horizon and the executor selects defensive actions within that subnet. Using the high fidelity Cyberwheel environment, with its built-in automated red team agent mapped to the MITRE ATT&CK framework, we compare RL+RL, LLM+RL, and LLM+LLM configurations using six models ranging from 3B to 70B parameters, including two cybersecurity-specialized models, across small, medium, and large networks. Replacing only the planner with an LLM yields limited gains as network size increases. In contrast, extending LLM control to execution produces notable improvements for sufficiently capable models. For instance, a frozen general purpose 70B model holds successful lateral movement to approximately 1% of steps and attacker impact near zero across all three network scales using the same model weights, while the RL baseline is retrained for each scale. Our results show that sufficiently capable frozen LLMs can maintain strong defensive performance across the evaluated network scales without task-specific retraining, while also indicating that strong tactical execution is important to realizing the benefits of LLM-based control.
Transcript
Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.
Nadia: Today's paper: "Towards Hierarchical Cyber Defense with Large Language Models".
Elias: An autonomous cyber defender trained with reinforcement learning (RL) is typically tied to its training network, limiting its generalization, and this research investigates whether frozen,
Nadia: First, who's behind it and why it matters.
Title and authors: Nadia: Moving on to the specific architectural suggestions they offer for improvement in "Towards Hierarchical Cyber Defense with Large Language Models: From Planning to Execution," the authors propose a clear structural division between the planner and executor roles.
Elias: They suggest formulating a controller-agnostic planner-executor hierarchy, which basically means you have one part of the system choosing where to focus its defense—the strategic targeting—and another part picking exactly what actions to take within that chosen area.
Priya: From a measurement standpoint, this separation is valuable because it allows researchers to isolate whether improvements come from better high-level decision-making or better low-level action selection during the evaluation process.
Nadia: Precisely, Priya; this design lets them test three distinct architectures: RL plus RL for a fully trained hierarchy, LLM plus RL where the LLM plans and an RL component executes, and then the all-LLM version where both are based on frozen LLMs.
Elias: The paper points out that this temporal abstraction is what allows them to control how often the planner acts—every 'k' environment steps—while the executor responds every single environment step conditioned on that selected goal.
Priya: I find that controlling the horizon, which they set at k=five for their evaluation, is a very practical constraint because it prevents the system from getting overwhelmed by too much immediate complexity while still allowing for strategic foresight.
Nadia: That fixed horizon seems to be a necessary simplification to manage the decision space before you even consider how an LLM handles those different scales.
Elias: And when we look at their comparison protocol, they run two complementary tests: first checking performance across different network sizes using frozen LLMs, and second comparing the LLM+RL setup against the LLM+LLM setup to see how extending control helps.
Priya: That two-pronged approach is smart because it addresses both the generalization aspect across scales and the mechanism of control extension simultaneously, which gives us a really complete picture of the system's capabilities.
Nadia: So, in short, they advocate for this structured hierarchy as the way to effectively leverage frozen LLMs for cyber defense without getting locked into task-specific retraining requirements.
Elias: It really emphasizes that the planner is a natural fit in a hierarchical network defender because it handles the coordination of large policy spaces better than trying to manage everything at once.
Priya: And the limitation they flag, which I think is important, is that while hierarchy helps reduce complexity, it doesn't entirely eliminate that dependence on retraining whenever the underlying network structure fundamentally changes.
Nadia: So they’re saying the structure helps manage the complexity of a single environment but doesn't solve the bigger problem of adapting to completely new environments without updates.
Elias: That distinction between managing complexity within a known environment versus achieving true cross-scale adaptability is what makes this paper so relevant for cryptography and security research.
The paper's summary: Nadia: So we've covered the core of "Towards Hierarchical Cyber Defense with Large Language Models: From Planning to Execution," which shows how extending LLM control from planning to tactical execution yields substantially stronger cross-scale performance.
Elias: To wrap up, the main implication for us is that a frozen, zero-shot LLM can provide retraining-free control in hierarchical cyber defense across various network scales when it's given the right architectural framework.
Priya: For me, the real impact is seeing how this translates into deployable tools; if we can achieve those high performance metrics without needing continuous retraining for every new network size, that drastically lowers the barrier to deploying sophisticated defenses.
Nadia: It really highlights that strong tactical execution is important for realizing the benefits of LLM-based control, as they noted when moving from planning alone to both planning and execution.
Elias: And we have to keep in mind their caveat: cybersecurity specialization doesn't automatically guarantee robustness as the network scales; smaller or some specialized models failed to maintain that cross-scale generalization.
Priya: That limitation is important because it tells us that relying solely on domain specialization isn't a guaranteed fix for scaling challenges in autonomous systems.
Nadia: Indeed, so the paper concludes that adding pretrained reasoning only at the top of a hierarchical defender might be insufficient without that strong tactical execution component.
Elias: We should definitely keep an eye on future work mentioned, like testing generalization across multiple independent seeds and unseen topologies and attacker strategies to see if this holds up further.
Priya: It’s exciting because it gives us a concrete path forward for how we can build more resilient AI defenders that are less brittle when they encounter unexpected network conditions.
Nadia: That’s the essence of what this paper on "Towards Hierarchical Cyber Defense with Large Language Models: From Planning to Execution" shows us, a structured approach coupled with capable frozen models can offer significant potential for scalable autonomous cyber defense.
The paper's improvements: Nadia: So, we've established that this paper proposes separating strategic planning from tactical execution in cyber defense to handle complexity better.
Elias: Right, and we saw how they use that planner-executor hierarchy to test different control architectures, like the LLM+RL versus the all-LLM setup.
Nadia: Exactly, and now what's really interesting are these specific improvements they suggest for building these systems.
Elias: They advocate for a controller-agnostic design where you can swap out the planner or the executor with either a trained reinforcement learning policy or a frozen, zero-shot LLM.
Nadia: That flexibility is huge because it means we don't have to commit to one specific type of controller; we can use whatever works best for the task.
Elias: And they suggest a system where the LLM handles both high-level planning and low-level tactical action selection based on real-time observations.
Nadia: That sounds like it could mean the defense doesn't just decide *where* to look, but also *exactly what* to do in that spot at every step.
Elias: It moves the decision-making from a broad strategic goal toward fine-grained execution, which is exactly what they were testing when they compared LLM+RL with LLM+LLM.
Nadia: The implication here is that for the AI to be truly useful in this domain, it needs that bridge between the big picture and the small, precise actions.
Elias: I agree; if the LLM only plans, it’s like having a brilliant general who can’t actually lift anything heavy or pick up a specific tool on the ground.
Nadia: It seems like this paper is pushing us toward building defensive AI that isn't just good at thinking, but also good at doing, which is a significant step for autonomous systems.
Elias: And we have to remember their limitation mentioned in the text; they flag that this structure helps manage complexity in a known environment but doesn't solve the problem of adapting to entirely new network layouts without retraining.
Nadia: That means while we get these improvements for existing setups, we still face the challenge of making them robust when things change completely, which is a fair point.
Elias: So the next thing we should consider is how to test this system's performance when those underlying network assumptions are completely different from what it was trained on.
Conclusion: Nadia: So, to wrap up this session on "Towards Hierarchical Cyber Defense with Large Language Models: From Planning to Execution," we've seen how separating strategy from tactics really helps make these systems more effective across different network sizes.
Elias: Indeed, and the core finding is that extending the LLM control into execution significantly boosts performance when compared to just having it plan everything.
Priya: From my side, what truly stands out is how this framework allows us to measure tangible results like compromise rates and impact across those different scales without needing massive retraining efforts for every single variation.
Nadia: It really shows that we can get substantial defensive gains by making the AI better at the actual doing part of its job rather than just the thinking part.
Elias: I agree; that shift in focus is what makes a difference when you’re dealing with complex, real-world adversarial scenarios where you need precise action selection.
Priya: And it's fascinating how much the data actually shows in terms of those normalized defense rates, which gives us concrete metrics on how well this architecture holds up under stress.
Nadia: It's exciting because this means we can build autonomous defenders that are not just theoretically sound but actually perform well when facing large-scale threats.
Elias: And we should keep in mind the authors did flag a limitation, which is that while it works within its tested parameters, it still faces challenges when the network topology itself shifts drastically outside of those initial conditions.
Priya: That's important because it tells us that while the control mechanism is improved, true resilience against completely unknown environments still requires more research.
Nadia: So, we can be optimistic about using this hierarchical approach for building next-generation cyber defense tools right now.
Elias: We definitely can; it gives us a solid architectural blueprint for how to integrate large language models into layered defense strategies effectively.
Priya: It’s a great step forward in making autonomous cyber defenses more practical and measurable for real-world application.
Nadia: That’s the big picture here, showing how structure matters when building sophisticated AI defenses.
Elias: We'll take this structural separation into account as we look at other papers that explore control mechanisms in agent frameworks next week.
More episodes
- 2610.10644-SoK: Failure Modes in Common Criteria Product Evaluation - A Taxonomy and Design-for-Evaluability Guidance
- 2610.10617-MRCert: Towards Post-deployment Patch Robustness Certification for Adversarially Patched Samples via Type-specific Masking
- 2610.10620-When AI Finds Hidden Messages, Does It Report?
- 2610.10625-Safe at One Loop, Risky at Another: Aligning Safety Across Recurrent Depths in Looped Language Models
- 2610.10992-The Hint Weight of ML-DSA Signatures Is Key-Dependent: An Empirical Study across the Three FIPS 204 Parameter Sets
- 2610.10659-Applying Security by Design at the Point of Execution: How Governed Security Requirements Affect the Security of AI-Generated Code
- 2610.10735-DITTO: A Context-aware Pickle-based Pre-Trained Model Scanner for Effective Security Audits
- 2610.10742-BRANCH: Bypassing Multi-Scanner AI Guardrails
- 2610.10752-Detection-Guided Adaptive Purification with Diffusion Models for Robust Audio Deepfake Detection
- 2610.10766-CPU-Auth: Device Fingerprinting for Authentication via DVFS Side-Channel