Crossing the Cyber Divide: Sim-to-Sim and Sim-to-Real Transfer for RL Agents

arXiv:2610.00759 · cs.CR, cs.LG · Submitted 2026-09-30 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.

Nadia: I'm Nadia, and with me are Elias and Priya, guest researcher.

Elias: Today's paper: "Crossing the Cyber Divide".

Nadia: Cyber attack agents are typically trained and evaluated within a single simulator, making it unclear whether learned policies transfer beyond the environments in which they were developed.

Elias: First, who's behind it and why it matters.

Title and authors: Nadia: We’ve seen that they propose a framework with an intent interface and a DAPN encoder to handle state and action translation, and now we need to look at how this actually performs across different scenarios. The paper summarizes their evaluation across four environments: CyberBattleSim, NetSecGame, CyberWheel, and NASim.

Elias: That summary section is where they lay out the concrete evidence of their methodology in action; it shows them testing this framework against those specific platforms to see what happens when you try to move a policy from one place to another.

Priya: I'm interested in the specific findings regarding the "narrow domain gap" versus a "large domain gap," because that distinction seems critical for determining whether simple feature engineering is enough or if the full DAPN alignment is necessary.

Nadia: Right, Priya, because their results show different outcomes depending on how similar or different those source and target environments are; it’s not a one-size-fits-all approach.

Elias: They found that on a narrow domain gap, like moving from CyberWheel to NetSecGame, simple feature engineering alone can lift the win rate to forty-seven point four percent, which is a decent improvement over the baseline.

Priya: That suggests that when the schemas are relatively similar, aligning the features might be enough to get good results without needing the full complexity of a domain adversarial network.

Nadia: But then they test a larger gap, moving from CyberWheel to CyberBattleSim, and that’s where things get interesting because feature engineering alone fails completely in that case.

Elias: On the large domain gap, the paper shows that zero-shot transfer with just feature engineering results in the mean nodes owned collapsing to zero point zero zero and a mean return of nineteen which is pretty bad performance.

Priya: That outcome really hammers home their point: when the state distributions are substantially different, you absolutely need that DAPN encoder to bridge the gap for successful transfer.

Nadia: Precisely, because only when they applied the DAPN did they achieve a one hundred percent win rate on those large gaps, confirming that aligning the state distribution is what bridges those domains.

Elias: It’s interesting how their methodology directly addresses that distributional gap by forcing alignment in the latent space, which is a very direct solution to the problem they set up earlier.

Priya: And on top of all that, they also looked at simulation-to-real transfer, and what did they find regarding the behavior when moving into an emulated virtual machine environment?

Nadia: They found that for sim-to-real transfer, the policies exhibit strong behavioral similarity, specifically showing a Jensen–Shannon divergence of zero point zero eight five from native policies in those emulated VM environments.

Elias: That divergence figure is quite low, suggesting that the policy's behavior when deployed in a simulated real environment is very close to its original performance, which supports the plausibility of this transfer mechanism.

Priya: So, to summarize this section: they show that success depends heavily on the gap size and that alignment tools are necessary when the data distributions diverge significantly.

Nadia: Exactly, and this whole exercise with "Crossing the Cyber Divide" shows how structural misalignment in cyber simulations is a real bottleneck for practical AI deployment.

The paper's summary: Nadia: Moving past what they found, we need to talk about the actual proposed improvements of their framework, because it’s not just about reporting results, but showing *how* they achieved this transfer capability.

Elias: The authors suggest breaking the problem into two distinct sub-problems—state transfer and action transfer—and solving them separately using a shared abstraction layer rather than trying to do everything at once.

Priya: So the primary improvement is this separation: learning a mapping for states and then learning a separate mapping for actions, which sounds like they are tackling the complexity piece by piece.

Nadia: That’s right; they introduce the kill-chain intent interface as this shared abstraction layer, which is central to Phase one because it constructs both mappings without having to modify the underlying simulators.

Elias: The action abstraction part of that interface is particularly clever: it exposes a discrete action space focused on selecting a host and advancing along the kill-chain, plus a noop, and then the mapping psi resolves those selections into simulator-native actions.

Priya: That decoupling of intent from mechanical details sounds like it makes the policy itself much more portable because it’s focused on high-level goals rather than low-level simulator specifics.

Nadia: And for the state abstraction, they project raw observations into a fixed-size vector organized by kill-chain stage per tracked host, capturing things like phase and reachability. This standardized projection is computed deterministically via environment-specific wrappers.

Elias: That deterministic projection method is important because it ensures that even though the raw inputs differ, the agent always sees a consistent structure based on its current kill-chain stage.

Priya: So, if we look at the DAPN encoder in Phase two what’s the specific mechanism they use to ensure that numerical values still align after the interface has done its job?

Nadia: The DAPN encoder uses a lightweight MLP trained with adversarial domain confusion and reconstruction losses to align those latent representations of phi(S B) and phi(S A).

Elias: And they use an input partitioning strategy by splitting the observation into two streams: features that vary across simulators pass through the encoder for alignment, while semantically identical features bypass the encoder entirely.

Priya: That input partitioning sounds like a smart way to ensure that they are only aligning what matters—the parts of the data that cause numerical differences—while preserving the core semantic meaning of the observation.

Nadia: So, in short, their improvement is a two-stage approach: first, abstracting intent and state through an interface, and second, using a specialized adversarial network to align the resulting distributions when they diverge.

Elias: That sequence addresses the fundamental mismatch by handling structure first and then fine-tuning the distribution alignment for robustness across different cyber environments.

The paper's improvements: Nadia: So we’ve covered a lot, and I think it’s time to wrap up this discussion on "Crossing the Cyber Divide: Sim-to-Sim and Sim-to-Real Transfer for RL Agents." The paper successfully outlines a method for policy transfer by separating state alignment from action translation using an interface and then closing the distributional gap with a DAPN encoder.

Elias: It seems the main conclusion is that when source and target environments share a similar feature schema, just the translator component is sufficient to bridge the gap, but if schemas diverge significantly, that alignment encoder becomes critical for success.

Priya: From a data perspective, I see this as proving that policy transfer isn't just about brute force retraining; it’s about intelligently aligning the underlying semantic representations of the environment itself.

Nadia: That’s a really important way to frame it, Priya; we aren't just teaching an agent a new game, we are teaching it how to interpret the structure of different games.

Elias: And regarding real-world deployment, they showed that this framework enables zero-shot execution across simulators and emulated systems while maintaining strong behavioral similarity in the sim-to-real transfer tests.

Priya: I think the implication here is that we can move toward deploying more versatile offensive tools because we don't have to build entirely new agents for every specific simulation platform available.

Nadia: Exactly, so the "Crossing the Cyber Divide" paper gives us a concrete framework for making our RL agents much more adaptable to different cyber scenarios, which is a huge step forward.

Elias: It’s solid work that tackles the inherent brittleness of current cyber reinforcement learning agents by providing a principled way to manage the differences between simulators.

Priya: I think this work on policy transfer provides a solid foundation for future research in making AI agents more robust and deployable across the entire spectrum of simulation tools.

Conclusion: Nadia: So we’ve gone through the whole process of looking at "Crossing the Cyber Divide: Sim-to-Sim and Sim-to-Real Transfer for RL Agents," and honestly, this framework seems to offer a real path forward for making our offensive AI tools more versatile.

Elias: I agree, Nadia; separating state alignment from action translation is a clever structural move that gives us a handle on the problem, even if I still have questions about the underlying assumptions in those mappings.

Priya: From my side as someone focused on measurement, what really stood out to me was how they quantified that distribution gap and how their DAPN encoder directly addressed that numerical difference across different simulators.

Nadia: That's exactly it, Priya; the results clearly show that feature engineering alone isn't always enough when the environments diverge, which is a crucial piece of information for us researchers.

Elias: And from a cryptographer’s viewpoint, I’m interested in how robust those mappings phi and psi are against adversarial manipulation; if the interface itself can be gamed, the whole transfer mechanism falls apart.

Priya: The data suggests that when the gap is large, like between CyberWheel and CyberBattleSim, zero-shot transfer without that alignment component yields almost no useful results for the agent's performance.

Nadia: That’s a pretty stark finding; it really shows us exactly where our current methods are hitting a wall when we try to jump between different simulation setups.

Elias: The implication is that we need more sophisticated ways to ensure that the latent space alignment isn't just superficial, but truly captures the necessary operational semantics for a policy to work in a new domain.

Priya: And for the privacy aspect, I see this as a way to reduce our reliance on collecting massive amounts of real-world data just to validate agent performance across different simulated threats.

Nadia: It sounds like we’re talking about making our agents much more scalable and deployable by proving they can handle more diverse environments without needing a total overhaul every time.

Elias: Exactly, Nadia; the paper suggests that when the source and target environments share a similar feature schema, you can rely on that translator alone, which simplifies things for deployment.

Priya: So overall, "Crossing the Cyber Divide" gives us a practical toolkit to manage environment heterogeneity in our AI agents by handling both structural and distributional mismatches systematically.

Nadia: It's a solid piece of research that provides concrete steps for how we can make these advanced security agents more flexible and reliable across different platforms.

Elias: Indeed, the separation of concerns is a useful architectural pattern that we should certainly keep in mind as we explore more complex agent architectures.

Sabrina Saika, Yinuo Du, Aritran Piplai

The University of Texas at El Paso

cs.CR, cs.LG

Submitted: 2026-09-30

Updated: 2026-09-30

Comments: 13 pages, 3 figures, 1st Workshop on Real-world AI Security and Engineering for Cybersecurity Systems (RAISE) 2026

Code: https://github.com/stratosphereips/NetSecGame

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 83/100

The gist: Cyber attack agents are typically trained and evaluated within a single simulator, making it unclear whether learned policies transfer beyond the environments in which they were developed.

Key concepts

State Transfer
This involves creating a mapping between the state representations of two different environments. The goal is to turn an observation from one simulator into a meaningful input for the policy trained in another simulator. This ensures the agent understands what is happening in a new setting.
Action Translation
This focuses on converting the abstract intent of an action into a specific, executable command within a target environment. The framework uses an abstraction layer to decouple the policy's high-level goal from the low-level mechanics of any specific simulator, allowing it to execute actions correctly.
Kill-Chain Intent Interface
This is a shared abstraction layer designed to bridge the gap between different simulators. It creates a standardized action space and projects raw observations into a fixed vector organized by kill-chain stages. This allows the policy's intent to be independent of the simulator's specific details.
DAPN Encoder
This modified Domain-Adversarial Network aligns the latent representations of states from different simulators. It uses adversarial training and input partitioning to ensure that features varying across simulators are aligned, while semantically identical features bypass the encoder, effectively bridging distributional gaps.

Terminology

Summary

Cyber attack agents are typically trained and evaluated within a single simulator, making it unclear whether learned policies transfer beyond the environments in which they were developed. The gist: A framework separating state alignment from action translation enables a policy trained in one environment to operate in another without retraining. This work studies policy transfer across cyber environments and argues that simulator-to-simulator and simulator-to-real transfer are instances of the same underlying alignment problem.

Problem Statement

The core difficulty lies in the fact that SA ̸= SB and AA ̸= AB in general, simulators differ in how they represent host state, network topology, and available actions. A policy trained on one environment cannot be directly executed on another without a bridge between the two representations. The problem is decomposed into two sub-problems: (1) State transfer: Learn or construct a mapping ϕ: SB → SA such that πA(ϕ(sB)) produces meaningful actions when the agent observes sB ∈ SB, and (2) Action transfer: Learn or construct a mapping ψ: AA → AB such that the action selected by πA can be executed in MB.

Kill-Chain Intent Interface

To address the mismatch between state and action spaces, the authors introduce a kill-chain intent interface, which acts as a shared abstraction layer. This interface is central to Phase 1 of their framework, where it constructs both mappings without modifying the underlying simulators. Specifically:

  1. The Action abstraction (ψ) exposes a shared discrete action space in which each action selects a host to advance along the kill-chain, plus a noop. The mapping ψ resolves these selections into simulator-native actions, decoupling the policy’s intent from the mechanical details of any given simulator.

  2. The Observation abstraction (ϕ) projects raw observations into a fixed-size vector organized by kill-chain stage per tracked host, capturing phase, reachability, attacker presence, and target designation. This projection is computed deterministically via environment-specific wrappers.

DAPN Encoder for Distributional Alignment

Even after the interface constructs the mappings ϕ and ψ, a distributional gap remains: even after applying ϕ, numerical values in the mapped state differ across simulators due to different network sizes, discovery rates, and reward scales. This gap is closed by a modified Domain-Adversarial Network (DAPN) encoder (Phase 2), which aligns the latent representations of ϕ(SB) and ϕ(SA). The encoder employs a lightweight MLP trained with adversarial domain confusion and reconstruction. Crucially, it uses Input partitioning, splitting the observation into two streams: features that vary across simulators pass through the encoder for alignment, while semantically identical features bypass the encoder entirely. The training objective includes three losses: L = λadv Ladv + λrec Lrec + λalign Lalign.

Transfer Regimes and Findings

The paper evaluates transfer across four environments: CyberBattleSim, NetSecGame, CyberWheel (source), and NASim (emulation). The results show that transfer is conditional on the gap size:

  1. On a narrow domain gap (e.g., CW to NSG), Feature engineering lifts the win rate to 47.4% (+30.9 return), while DAPN achieves 45.2%, demonstrating that feature alignment alone can be competitive when schemas align but fails when they diverge, as seen in CyberBattleSim where matching observation shape yields only no-ops (0%).

  2. On a large domain gap (e.g., CW to CBS), the state distribution is the critical transferable aspect; zero-shot transfer with feature engineering alone results in mean nodes owned collapsing to 0.00 and mean return to 19. Only when DAPN is applied does it achieve a 100% win rate, confirming that aligning the state distribution bridges the two domains for large gaps.

  3. Regarding simulation-to-real transfer (RQ3), transferred policies exhibit strong behavioral similarity, with a Jensen–Shannon divergence of 0.085 from native policies in emulated virtual machine environments, supporting sim-to-real plausibility.

Conclusion and Future Work

The framework successfully bridges both structural and distributional gaps by combining the kill-chain interface with the DAPN encoder, enabling zero-shot execution across simulators and emulated systems. The study concludes that when the source and target share a similar feature schema, the translator alone is sufficient; the encoder becomes critical when schemas diverge. Future work focuses on closing the sim-to-emulation gap by addressing practical obstacles like exploit version mismatches, indirect or delayed observation of action outcomes, and non-deterministic service behavior.


The gist

A framework separating state alignment from action translation enables a policy trained in one environment to operate in another without retraining.

Improvements for AI systems

Here are specific improvements for AI systems based on the findings in Crossing the Cyber Divide: Sim-to-Sim and Sim-to-Real Transfer for RL Agents:

  1. Improve autonomous red-teaming agents by enabling them to operate reliably across diverse cyber simulation environments (e.g., moving from one simulator to another or into an emulated real environment) without requiring complete retraining.

  2. Develop a robust policy transfer mechanism that separates state alignment from action translation. This allows an agent trained on a specific simulator's state representation to function effectively in another environment by learning mappings (like the DAPN encoder) that align the underlying semantic kill-chain features, rather than relying on exact input/output matching.

  3. Enhance agent generalization by implementing a domain-adversarial encoder that learns a shared latent feature space for cyber environments. This ensures that the policy's decision-making remains consistent even when the visual or numerical representation of network states changes substantially between simulators (addressing the large domain gap problem).

  4. Improve sim-to-real deployment feasibility by utilizing controllable simulation settings (like NASim emulation) as a bridge. The system can transfer policies to emulated VM environments, showing behavioral similarity with high fidelity (JS divergence of 0.085), significantly reducing the reliance on expensive, unsafe real-world data collection for initial policy validation.

  5. Create modular agent architectures where an Intent Interface acts as a shared abstraction layer. This interface decouples the high-level policy intent (e.g., advance to exploitation phase 2) from the low-level simulator specifics (e.g., specific host slots or service enumeration counts), allowing the core RL policy to remain stable while only the mapping components change during transfer.

  6. Implement a diagnostic tool that measures the transfer regime in real-time during training or deployment, automatically switching between simple feature engineering and full DAPN alignment based on detected domain divergence, thereby optimizing computational cost while ensuring policy robustness across both narrow and large environment gaps.

These improvements enable AI systems to be significantly more scalable, deployable, and versatile in offensive cybersecurity applications by overcoming the fundamental brittleness caused by simulator heterogeneity.

Sources

Related papers