Sim-to-Real RL for ASVs using SysID

arXiv:2610.12202 · cs.RO · Submitted 2026-10-08 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: I'm Rosa, and with me are Dev and Taro, guest researcher.

Dev: Today's paper: "Sim-to-Real RL for ASVs using SysID".

Rosa: The gist: This work presents an ASV simulator and accompanying pipeline that enables training policies starting from unknown vehicle dynamics.

Dev: First, who's behind it and why it matters.

Title and authors: Rosa: So we're looking at this paper now called "Sim-to-Real RL for ASVs using SysID." It tackles a big problem: training policies for autonomous surface vehicles in dynamic ocean environments when you don't have perfect knowledge of the vehicle's exact dynamics.

Dev: Exactly, and what they’re proposing is a way to train those reinforcement learning policies even when we start with totally unknown vehicle dynamics, which usually means we need tons of data from towing tanks or CFD simulations just to set up the initial physics parameters.

Taro: It sounds like they're trying to cut out that expensive, slow process of getting perfect hydrodynamic models before you can even start training the AI controller.

Rosa: Right. Essentially, this framework uses just a CAD model and a small set of open-water field trajectories to figure out both the hydrodynamic and thruster parameters needed for the simulation.

Dev: They use system identification, specifically something called CMAES, within their simulator to refine those parameters based on what they observe in real-world conditions.

Taro: That sounds like a smart way to bridge that gap between simulation and reality by grounding the model in actual behavior rather than just theoretical assumptions.

Rosa: And the results they show are pretty compelling, because they demonstrate successful zero-shot transfer for both path following and station-keeping tasks right on a real BlueBoat ASV in a river environment.

Dev: That's what I find interesting from an engineering standpoint, because being able to deploy a policy without that prior hydrodynamic information is a huge deal for practical application.

Taro: It means we don't have to spend months building the perfect CFD model before we can even test the autonomy part of the system.

Rosa: They also detail how they model things like propeller thrust, including spool-up lag and asymmetries between different motors, which are often tricky in real life.

Dev: That level of detail on the propulsion modeling is important because those small performance issues can really throw off a control loop if you're not accounting for them.

Taro: And they show how they handle the environment too, using things like a specific wave spectrum derived from empirical data to make the simulation realistic.

Rosa: So, let's talk about how they suggest improving this approach, because as we know, sim-to-real transfer isn't always a straight line.

Dev: They focus on grounding the simulated dynamics using CMAES over trajectories to make sure the parameters are optimized against observed behavior, which they call Grounded Simulation Learning.

Taro: That’s important because just having a set of parameters isn't enough; you need them tuned correctly to minimize that gap between simulation and reality.

Rosa: They also suggest using different observation spaces for different tasks, like a specific one for path following that includes terms like cross-track error and heading error.

Dev: Having task-specific observations makes sense because the control signals needed to follow a line are fundamentally different from the signals needed just to stay still.

Taro: And I also noticed they mention using an asymmetric actor-critic architecture, which maintains separate weights for each network, which lets them specialize in those different observation spaces.

Rosa: That specialization helps the system learn more effectively when it has to handle both navigation and station-keeping simultaneously.

Dev: They also shape the reward function by adding penalties for the first and second derivatives of thrust commands, which should encourage smoother control inputs during operation.

Taro: Those smoothing penalties are important because they translate directly into less wear on the physical vehicle when it's actually running in the water.

Rosa: So, to wrap up this paper, "Sim-to-Real RL for ASVs using SysID," it presents a simulator with hydrodynamics, waves, and propeller modeling that allows for zero-shot deployment of reinforcement learning policies.

Dev: The main implication here is that we can train these policies in a simulation and deploy them on the actual vehicle without needing prior hydrodynamic or propeller information.

Taro: It means autonomy researchers don't have to be stuck waiting for perfect physical modeling before they can start testing control strategies.

Rosa: They also show how crucial it is to ground the simulation using real-world data, and that the CMAES approach with those trajectories is a key part of making that happen.

Dev: The paper points out a limitation, though, which is that while this pipeline works well in this river environment, it’s still focused on open-water conditions as defined by their trajectory set.

Taro: And they flag that if you introduce dynamics not covered by those trajectories—like certain types of severe environmental disturbances—the performance might degrade because the grounding wasn't comprehensive enough for those specific scenarios.

Rosa: It’s a good point, because it shows that the success is highly dependent on how well you ground the simulation to represent what’s actually happening in the real world.

Dev: So, looking at all this, "Sim-to-Real RL for ASVs using SysID" gives us a concrete toolset to move RL policies from theory into actual field testing with less upfront modeling work.

Taro: It moves the focus from just building models to building robust pipelines that use data to inform and refine those models as they go.

Rosa: That’s the big picture for autonomous systems today, moving toward these types of data-driven grounding methods.

Dev: We'll keep an eye on how this method performs when we push it into more complex, non-linear dynamics that aren't just simple open-water conditions.

The paper's summary: Rosa: So basically, this paper shows how you can train an AI controller for a surface vehicle even when you don't know its exact physics beforehand, using just a CAD drawing and some field test data to figure out the parameters yourself.

Dev: Right. It’s about building this whole pipeline so that you can get your reinforcement learning policy onto a real vehicle without needing months of expensive towing tank tests or CFD simulations first.

Taro: The main idea here is using system identification, specifically something called CMAES, to ground the simulation dynamics in actual observed behavior from open-water trajectories.

Rosa: That’s the grounding part. They use those trajectories to refine both the hydrodynamic forces and the thruster parameters within their simulator until it matches what a real vehicle actually does.

Dev: And they show that this allows for zero-shot transfer, meaning you can deploy that trained policy on a physical BlueBoat ASV for path following and station-keeping right away.

Taro: I think what’s interesting is how they handle the complexity of the marine environment, like modeling wave effects using a spectrum derived from empirical data instead of just picking one generic model.

Rosa: Exactly, they’re not just throwing random noise at it; they’re incorporating realistic environmental disturbances into the simulation during training to make sure it's robust.

Dev: Now, as an engineer, I look at the numbers they give you about tracking error reduction—they showed a sixty-three point nine percent drop in position error for station-keeping when they used their proposed configuration compared to a real-world setup.

Taro: That’s huge because it proves that this specific grounding method actually helps the policy perform better in the messy real world, not just in a perfect lab setting.

Rosa: It moves the conversation away from needing perfect initial models and toward building data-driven methods that can adapt as they learn.

Dev: But there are caveats, though. They mention that if you use domain randomization—basically adding noise and varying dynamics during training—it can actually hurt the performance compared to their best configuration in real-world current conditions.

Taro: So it’s a trade-off, right? You get better grounding by using real data, but you have to be careful not to overdo the randomization if you want good sim-to-real transfer.

Rosa: And they also point out that incomplete dynamic grounding can make maneuvers brittle when the vehicle encounters dynamics it hasn't seen before in the training data.

Dev: That’s a failure mode we need to watch for, because if your controller relies on a dynamic interaction that wasn't properly grounded, it might fail catastrophically in an unexpected situation.

Taro: It shows that even with smart RL techniques, if the simulation doesn't accurately reflect the physical limits of the vehicle—like those propeller asymmetries they model—the policy won't generalize well.

Rosa: So this paper is really about building a smarter way to teach robots to move in water by using real data to calibrate their training ground.

Dev: It gives us a concrete framework for how we can get autonomous systems moving from the lab bench into the actual ocean, without needing that massive upfront cost of perfect physics modeling.

The paper's improvements: Rosa: So we’ve covered how they use system identification to ground their simulation, but now let's look at what they think could make this whole process even better.

Dev: They suggest a few specific improvements to the framework, focusing on making the observation spaces and the reward function more specialized for different tasks.

Taro: I saw something about task-specific observation spaces, which means instead of one general set of inputs for everything, you get tailored inputs for path following versus station-keeping.

Rosa: Right. They’re talking about having distinct observation vectors, like one that specifically tracks cross-track error and another that focuses on heading error and waypoints.

Dev: And they pair that up with an asymmetric actor-critic architecture, which means the policy network gets its own separate weights for the actor and critic networks for each task.

Taro: That specialization should help the AI learn more efficiently when it has to juggle both navigation and station-keeping goals at the same time.

Rosa: Then there’s a tweak to how they calculate rewards, adding penalties on the first and second derivatives of thrust commands.

Dev: That smooth control shaping is important because it directly translates into less physical wear and tear on the actual vehicle when it’s operating in real conditions, not just simulation.

Taro: It's about making sure the AI learns to be a gentle operator rather than jerky one.

Rosa: So, they are focusing on refining the learning structure itself—making the observation inputs smarter and the reward signals more nuanced.

Dev: They want that refined observation and shaping to work together so that when you transfer this policy to reality, it’s not just a rough guess but a finely tuned controller.

Taro: What they don't really focus on in this paper is how to handle things completely outside of their specific open-water trajectory set, which is where the real limitations lie for broader deployment.

Rosa: True. They are very successful in this open-water context, but the authors themselves flag that if you introduce severe environmental disturbances not covered by those trajectories, performance might degrade because the grounding wasn't comprehensive enough for those specific scenarios.

Dev: So they’re setting expectations: this pipeline works well here and now, but it’s not a universal solution for every kind of rough sea or unexpected current.

Taro: It means the next big step isn't just better RL algorithms; it’s better data collection and more robust ways to model those extreme, unobserved dynamics.

Rosa: Exactly. So this paper gives us a solid starting point, but for wide-scale use, we need to figure out how to make that grounding method even more adaptable when the world throws something totally new at it.

Conclusion: Rosa: So we're wrapping up this look at "Sim-to-Real RL for ASVs using SysID" by Cody Sheltraw and his team, which is basically about building a simulator and pipeline that lets you train policies for surface vehicles even without knowing the vehicle's exact physics beforehand.

Dev: It’s a solid piece of work because it tackles the sim-to-real gap head-on by using real-world data to calibrate those simulation parameters, which is crucial for getting reliable control loops running in practice.

Taro: I think what this means for autonomy is that we can train these complex policies in a simulated environment and then deploy them on a physical vehicle without having to spend months building a perfect CFD model first.

Rosa: It’s about moving the barrier away from needing perfect physics models right at the start, toward using observed behavior to guide the learning process.

Dev: And I have to stress that this relies heavily on those specific open-water field trajectories they use for system identification, so it’s really tied to what kind of environment you’re simulating.

Taro: That’s a valid point; if you want it to work in a different ocean or even near the coast, you have to ground it with different types of data, otherwise the generalization won't hold up when the world misbehaves.

Rosa: It shows that grounding simulations in observed data is a powerful way to build better control policies for things like path following and station-keeping.

Dev: The numbers they shared on tracking error reduction were pretty good, but as an engineer, I’m always wondering how stable those learned parameters are when the vehicle experiences sudden changes in dynamics that weren't in their training set.

Taro: That’s where the next big challenge is—making sure the system can handle those unexpected events without completely breaking down, which requires even more robust data augmentation.

Rosa: So, "Sim-to-Real RL for ASVs using SysID" gives us a tool to start training policies faster and with less upfront modeling work.

Dev: It’s a practical tool for getting a head start on real deployment, provided the environment you're simulating is similar enough to the data used during the identification phase.

Taro: We need to keep pushing on those environmental boundaries, because this method is very successful in open water, but it doesn't solve every hydrodynamic problem out there.

Rosa: Exactly. It’s a great step forward for deploying autonomous agents in marine settings without waiting around for perfect physical models.

Cody Sheltraw, Tsimafei Lazouski, Maani Ghaffari, Alan Papalia

cs.RO

Submitted: 2026-10-08

Updated: 2026-10-08

The gist: The gist: This work presents an ASV simulator and accompanying pipeline that enables training policies starting from unknown vehicle dynamics.

Key concepts

System Identification (SysID)
This process refines the unknown physical dynamics of the vehicle by fitting a mathematical model to observed data. The authors use the Covariance Matrix Adaptation Evolution Strategy (CMAES) within the simulator to extract accurate hydrodynamic damping and thruster coefficients from real-world trajectories, bridging the gap between CAD models and actual performance.
Domain Randomization (DR)
To improve sim-to-real transfer, this technique introduces variability into the simulation environment during training. It involves randomly sampling parameters like starting twists, wind disturbances (using an Ornstein-Uhlenbeck process), and wave conditions to ensure the trained policy is robust against variations encountered in reality.
Markov Decision Process (MDP) Framework
This defines how the RL agent interacts with the environment. The vehicle's actions are continuous control vectors, which are mapped to a normalized thrust command. The reward function is designed to balance achieving specific tasks like path following or station-keeping with penalties on rapid changes in control inputs.
Zero-Shot Transfer
This refers to the ability of a policy trained entirely within the simulation environment (using the refined dynamics) to perform successfully in a real-world setting without needing further retraining or extensive adaptation. The framework aims to achieve this by grounding simulation dynamics using real data before deployment.

Terminology

Summary

The gist: This work presents an ASV simulator and accompanying pipeline that enables training policies starting from unknown vehicle dynamics.

How it works

The framework uses only a CAD model and brief set of open-water field trajectories for approximating and refining both hydrodynamic and thruster parameters Initial hydrodynamic parameters are derived from a hull CAD model via a boundary element method (BEM) solver alongside a set of open-water trajectories from the physical autonomous surface vehicle (ASV). System identification (SysID) is performed in the simulator using the covariance matrix adaptation evolution strategy (CMAES) to obtain refined hydrodynamic and thruster parameters. Reinforcement learning (RL) policies are trained in the grounded simulator, enabling zero-shot transfer for path following and station-keeping tasks.

Simulation and Environment Design

The project uses NVIDIA Isaac Lab 2.3.2 built on Isaac Sim 5.1.0 using Omniverse PhysX for physics. Hydrodynamics are modeled as a 6-degree-of-freedom (6-DOF) system in PhysX, while SysID and maneuvering are modeled in 3-DOF based on Fossen’s relative motion model. The wave implementation uses NVIDIA Warp following an internal Isaac Sim example ocean based on empirical directional wave spectra defined in Simulated waves are derived using the Texel-Marsen-Arsloe (TMA) spectrum which generalizes the Joint North Sea Wave Project (JONSWAP) spectrum to finite water depths. Buoyant forces are governed by hydrostatic principles, with the buoyant force proportional to the submerged hull volume Vsub.

Propeller Thrust Modeling

Open-water propeller thrust T is determined by the standard marine propeller thrust equation. For simulation purposes, motor spool-up lag τlag (nominally 0.15 s) and a starboard motor imbalance factor cstbd (nominally 1.0) are introduced to capture propeller performance asymmetries. The net active thrust is modeled as T ≈ T0(1 − cuua). Motor spool-up lag is modeled as a first-order low-pass filter applied to the commanded thrust Tcmd with time constant τlag. Physical thrust asymmetries between port/starboard thrusters and forward/reverse operation are defined via dimensionless scaling ratios cstbd and crev.

Real-World System Identification

The parameters ultimately fine-tuned in the simulation include hydrodynamic damping (Xu, Xuu, Yv, Yvv, Nr), yaw inertia (Iz), and thruster coefficients cu and crev. For the real-world analytical baseline, the hydrodynamic damping and thruster coefficients are estimated using trajectories obtained from a physical BlueBoat in open-water conditions with minimal environmental disturbances. The remaining baseline parameters are established physically, while yaw inertia (Iz) and added mass terms (Xu˙, Yv˙, Nr˙) are obtained from a CAD model using a BEM solver.

Markov Decision Process Framework

The action space at time step t is defined as a continuous control vector at ∈ R N where N represents the number of propellers on the ASV. Raw actions at,i are mapped to a normalized control space xt,i = tanh(at,i/5), which is then filtered through a universal deadband threshold fdead to yield the target thrust command Tcmd,i[t]. The observation space for path following (PF) includes the core vehicle state vector obase,πt and path-specific terms like cross-track error and heading error. The reward function combines task-specific objectives with shaping penalties applied to the thruster commands’ first and second derivatives.

Domain Randomization

To bridge the sim-to-real gap, domain randomization is applied during training The domain randomization parameters used are summarized in Table III. Dynamics include uniformly sampled starting twist, surge velocity, sway velocity, and heading rate for both tasks. Wind disturbance is applied via an Ornstein-Uhlenbeck (OU) process to evolve ambient wind speed and direction over time. Wave disturbance is introduced via a curriculum schedule after 64,000 steps. Observation noise is applied to all components of the actor’s observation vectors, with the exception of the processed action history.

Experiments and Results

To evaluate the impact of our simulation grounding and isolate the CMA-ES curriculum’s contribution, we train and evaluate RL policies across four distinct configurations. The Proposed configuration achieved the lowest tracking error across all tasks, with a 0.417 m cross-track RMSE for path following. Notably, during station-keeping (SK), the Proposed policy reduced position RMSE by 63.9% compared to the Real-World configuration (1.074 m to 0.388 m RMSE). The Zigzag ablation proved unstable in real-world testing after missing a turn, demonstrating that incomplete dynamic grounding causes brittle performance during maneuvers reliant on those missing dynamics. Introducing CMA-ES DR degraded tracking performance compared to the Proposed configuration, resulting in increased lateral drift in real-world currents for both tasks.

Conclusion

This paper presents an ASV simulator with hydrodynamics, waves, and thruster modeling, along with a real-to-sim-to-real pipeline for zero-shot deployment of RL policies. Using a set of four open-water trajectories, we ground simulator dynamics and evaluate our pipeline on a real-world BlueBoat in path following and station-keeping tasks. Future work will focus on the importance of individual hydrodynamic terms and their domain randomization for sim-to-real transfer.

--- Page 1 ---

Sim-to-Real RL for ASVs using SysID

Cody Sheltraw1, Tsimafei Lazouski1, Maani Ghaffari1, and Alan Papalia1

Abstract—Autonomous Surface Vehicles (ASVs) operating in dynamic marine environments require robust control policies for tasks such as path following and station keeping, making reinforcement learning (RL) a promising alternative to classical controllers. However, existing ASV simulators rarely support parallel environments for RL training. Such existing simulators require accurate hydrodynamic modeling from computational fluid dynamics solvers or towing tank tests for setting hydrodynamic parameters to address the sim-to-real gap. To address these challenges, we present an ASV simulator and accompanying pipeline that enables training policies starting from unknown vehicle dynamics. Our framework uses only a CAD model and brief set of open-water field trajectories for approximating and refining both hydrodynamic and thruster parameters. Real-world deployments on a BlueBoat ASV demonstrate successful zero-shot sim-to-real transfer in path following and station-keeping tasks without prior hydrodynamic and propeller information.

Index Terms—Marine Robotics, Simulation and Animation, Reinforcement Learning

--- Page 2 ---

I. INTRODUCTION

Autonomous surface vehicles (ASVs) operating in marine environments require robust control despite non-linear hydrodynamic interactions, uncertain propulsion characteristics, and environmental disturbances, such as wind, waves, and currents [1], [2]. Since these hydrodynamic forces depend directly on the vehicle’s operating state and surrounding sea conditions, accurate modeling and controller tuning are challenging for tasks such as path following and station keeping. Reinforcement learning (RL) offers an attractive alternative for such tasks by learning control policies directly from interactions, without explicit modeling of each relevant dynamic interaction, necessitating the availability of RL-capable simulators. In other robotic domains, RL policies have demonstrated the ability to control complex non-linear systems in challenging operating conditions [3], [4]. However, transferring policies trained in simulation to real-world deployment remains challenging for ASVs due to the sim-to-real gap between simulated and real marine dynamics, [6].

To reduce this sim-to-real gap, an adequate representation of the vessel dynamics is needed. Such high-fidelity models traditionally rely on towing tank experiments or computational fluid dynamics (CFD) solvers for accurate hydrodynamic system identification (SysID), creating a high barrier to entry. Furthermore, using real-world SysID parameters in a simulator does not guarantee improved sim-to-real performance. The mismatch between real-world and simplified simulator dynamics means parameters identified from the real vessel may not align with the simulator parameters that best reproduce its observed behavior To compensate for this, other works in the marine domain rely on heavy domain randomization (DR), which can lead to conservative policies and decreased realworld performance. To address these limitations, this work presents a simulator and an accompanying real-to-sim-to-real pipeline that facilitates the deployment of simulation-trained policies on a physical ASV using little real-world data. Our approach reduces the sim-to-real gap by grounding the simulated dynamics in observed real-world ASV behavior using the covariance matrix adaptation evolution strategy (CMAES) over a set of trajectories. Starting with unknown vessel dynamics, our method reduced station-keeping position error by 63.9% (1.

Improvements for AI systems

  1. Bold Header: ASV Simulator Grounding Capabilities

The improved system can perform zero-shot sim-to-real transfer in path following and station-keeping tasks without prior hydrodynamic and propeller information by using only a CAD model and brief set of open-water field trajectories for approximating and refining both hydrodynamic and thruster parameters.

  1. Bold Header: Robust Parameter Grounding via CMA-ES

The system can achieve better parameter grounding by employing Grounded Simulation Learning (GSL) through the covariance matrix adaptation evolution strategy (CMA-ES) over a set of trajectories, which is specifically designed to minimize simulation-to-real gaps by optimizing parameters against observed behavior.

  1. Bold Header: Task-Specific Observation Spaces

The system can handle complex control tasks by using distinct observation spaces, such as the Path Following observation space o PF,π t, which includes p wp t, ecross t, sin(eψ,path t), cos(eψ,path t) to track waypoints and cross-track error.

  1. Bold Header: Asymmetric Actor-Critic Architecture

The system can effectively learn complex control policies by utilizing an asymmetric actorcritic architecture, maintaining separate weights for each network for the policy (actor) and critic, enabling specialized learning for different observation spaces.

  1. Bold Header: Enhanced Reward Function Shaping

The reward function can be shaped to promote smooth control and minimize physical wear by including shaping penalties like r act t = −1.5∥a˙ t∥2 − 0.5∥a¨t∥2 to penalize the first and second derivatives of thrust commands.

Related papers