OGR-MARL: Option-Guided Residual Multi-Agent Reinforcement Learning for Heterogeneous USV Cooperative Pursuit in Constrained Port Waterways
Jiayang Mao, Lanfeng Wang, Zhao-Han Peng
Sichuan Agricultural University · Tsinghua University
cs.AI, cs.MA
Submitted: 2026-08-13
Updated: 2026-08-17
Comments: 6 pages,5 figures, accepted by ICUS 2026
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 95/100
The gist: OGR-MARL: Option-Guided Residual Multi-Agent Reinforcement Learning for Heterogeneous USV Cooperative Pursuit in Constrained Port Waterways Abstract Summary: The paper addresses heterogeneous USV
Terminology
Summary
OGR-MARL: Option-Guided Residual Multi-Agent Reinforcement Learning for Heterogeneous USV Cooperative Pursuit in Constrained Port Waterways
Abstract Summary:
The paper addresses heterogeneous USV cooperative pursuit in constrained port waterways, which requires evader interception under navigation, traffic, and role constraints. The authors propose OGR-MARL, an option-guided residual multi-agent reinforcement learning framework that is decoupled from a specific MARL algorithm. OGR-MARL integrates shared evader belief, role-conditioned option targets, adaptive rule penalties, and residual policy learning, allowing different MARL algorithms to learn corrective actions on top of rule-guided behaviors rather than exploring constrained port environments from scratch. The framework is instantiated with representative continuous-control MARL backbones including MADDPG, MATD3, MAPPO, and MASAC, yielding OGR-MADDPG, OGR-MATD3, OGR-MAPPO, and OGR-MASAC. Experiments in an abstract Xiazhimen port-waterway scenario show that the OGR-MASAC instantiation achieves a 75.0% capture rate, promising mission-effective rule compliance, and the best heterogeneous coordination among the tested methods. Without retraining, zero-shot transfer to a QGIS/AIS-informed Xiazhimen map achieves promising results, demonstrating the generalization potential of OGR-MARL in more complex port scenarios.
Introduction Summary:
The paper notes that unmanned surface vehicles (USVs) are increasingly vital for maritime surveillance, environmental monitoring, and harbor security. Multi-USV coordination has received sustained attention, with cooperative pursuit being a significant class of issues. However, pursuit in constrained port-waterways is more demanding than in open environments because pursuers must navigate narrow lanes, avoid shores, anchorages, and moving traffic while capturing an evader. Existing pursuit studies often focus on simplified or open scenarios. Classical model-based methods such as artificial potential fields (APF), model predictive control, and path planning provide interpretable guidance and COLREGs-aware collision avoidance but require scenario-specific tuning and struggle with partial observability, role heterogeneity, and online role switching. Conversely, MARL offers a data-driven approach, but directly applying MARL to constrained port scenarios from scratch remains challenging because a pure policy-gradient learner may suffer from sparse rewards, frequent rule violations, and exploration inefficiency. A pure rule-based controller is interpretable but lacks adaptability against tactical evader replanning. To bridge this gap, the authors propose OGR-MARL, an algorithm-agnostic framework for heterogeneous multi-USV cooperative pursuit in constrained port-waterways. The pursuing team is heterogeneous: interceptors have capture capability with higher speed, while scouts have a larger sensing radius for target visibility. A* search and APF are utilized to model evader behavior and generate high-level option targets. The approach integrates an option guidance module, which provides high-level modes and geometric targets, with a general MARL algorithm that learns residual continuous actions and coordination patterns for the heterogeneous pursuing team.
Main Contributions:
-
The authors establish a benchmark scenario for heterogeneous multi-USV cooperative pursuit in constrained port waterways, incorporating shores, anchorage areas, bidirectional traffic lanes, buffer regions, and moving cargo vessels.
-
They propose OGR-MARL, an algorithm-agnostic option-guided residual MARL framework that integrates shared evader belief, role-conditioned option targets, and adaptive rule penalties, and can be instantiated with different MARL backbones.
-
They evaluate the proposed framework through backbone comparison, reward ablation, and zero-shot transfer to a QGIS/AIS-informed Xiazhimen map.
Problem Formulation Summary:
The simulation scenario is an abstracted Xiazhimen port waterway (in Zhejiang Province of China) on a 1000 × 400 domain. The channel is bounded by the north and south islands, with a USV base located at x ∈ [400, 600]. Two anchorage areas occupy [100, 300]×[330, 370] and [700, 900] × [230, 270]. The outbound and inbound traffic lanes occupy y ∈ [310, 350] and [250, 290], respectively, and a buffer area occupies y ∈ [290, 310]. The line x = 980 serves as the eastern escape boundary. The pursuer team is denoted by N = P1, P2, P3, P4. Pursuers P1 and P2 are fast interceptors and the only USVs with capture capability, while P3 and P4 are slower scouts with larger sensing radius. Four cargo vessels move along the center lines of the traffic lanes. Each pursuer is modeled by a three-degree-of-freedom (3-DOF) motion model with a fourth-order Runge–Kutta integrator for inner-loop dynamics and a low-level speed-heading controller. The high-level action of pursuer i is ai = [avi, aψi]T. The evader is modeled as a mass point with constrained acceleration and yaw rate, using an A*-based global planner and an APF-based local collision-avoidance module. The task is partially observable because only scouts have long-range sensing capability: interceptors have a sensing radius of 8 simulation units, whereas scouts have a sensing radius of 180. The team maintains a shared evader belief from the latest evader observation, updated by dead reckoning when no detection occurs. Each pursuer receives a 61-dimensional observation vector.
OGR-MARL Algorithm Summary:
The cooperative pursuit task is formulated as a partially observable multi-agent Markov decision process, adopting the centralized training and decentralized execution (CTDE) paradigm. During training, a centralized critic utilizes global information, while each actor executes using only local observations during decentralized execution. Each actor outputs a two-dimensional continuous residual control action. Due to heterogeneous dynamics, sensing ranges, and functional roles, independent actor–critic networks without parameter sharing are adopted. The option set is O = search, track, intercept, block, recover. The option layer assigns each pursuer an execution mode according to its role, rule-compliance state, evader visibility, and belief confidence. Scouts select the search or track option to maintain channel coverage and update the shared evader belief. Interceptors select the intercept or block option to form a pincer maneuver, close in on the evader, and complete the capture. The recover option guides a pursuer away from shores, anchorage areas, and buffer regions, or prevents violations of traffic lane direction constraints. The final executed action is obtained via a bounded blending mechanism: ai = clip((1 − βi)arl i + βi aopt i, −1, 1). The training design combines structured reward shaping with curriculum learning. The individual reward is decomposed into four components: task-progress reward, rule-compliance reward, heterogeneous-role reward, and terminal reward. The training reward blends individual feedback with team-level feedback. A five-stage curriculum learning strategy is adopted, where across stages the evader becomes more difficult to capture, the capture condition becomes stricter, and the rule-compliance penalty is gradually strengthened.
Experiments Summary:
The algorithms are trained using eight parallel environments. MARL baselines execute only learned actions without option-guided action blending; option-related observation entries are kept only to preserve the same input schema. Evaluation takes 100 test episodes with 100 seeds and a decision frame skip of 4. Metrics include successful capture rate (SR), mean task time cost (TC), mission-effective rule compliance (MRC), and heterogeneous coordination score (HCS). MRC is computed as MRC = SR · RCR · ET, where RCR is the rule compliance rate and ET is the time-efficiency score. HCS is computed as HCS = 0.30SR + 0.25ET + 0.25Sv + 0.20Ic, where Sv is the fraction of simulation steps in which at least one scout can sense the evader, and Ic is the ratio of positive distance-closing contribution provided by the interceptors.
Baseline Comparison Results:
The results show that OGR-MARL provides a broadly applicable option-guided residual learning interface for different MARL algorithms. The pure expert-rule controller reaches a success rate of 35.0%, showing that rule-based option guidance alone is useful but insufficient against a fast and replanning evader. Pure MARL baselines perform poorly because direct continuous-control learning from scratch does not reliably discover the required cooperative behaviors under port-waterway constraints. Compared with their corresponding pure MARL baselines, the OGR variants generally improve SR, MRC, and HCS. In particular, OGR-MATD3 improves the SR from 1.0% to 59.0%, and OGR-MASAC improves the SR from 14.0% to 75.0%. Among the tested instantiations, OGR-MASAC achieves the best overall performance, with the highest SR (75.0%), the highest HCS (0.6802), and promising TC (121.94 s) and MRC (0.4283). The remaining failures are mainly caused by inefficient terminal close-in interception rather than complete loss of the evader. Successful episodes finish with a mean closest interceptor–evader distance of 4.17, while failed episodes have a mean closest distance of 8.19. Scouts observe the evader for a larger fraction of time in failed episodes than in successful episodes, 0.738 versus 0.512, suggesting that the primary bottleneck is not target visibility but the final interception efficiency of the interceptors.
Reward Ablation Results:
The reward ablation study removes the rule reward, role reward, or both. Removing the rule reward or role reward reduces the performance, and removing both performs worst. OGR-MASAC with the full reward achieves SR of 75.0%, while w/o rule achieves 31.0%, w/o role achieves 51.0%, and w/o rule+role achieves 28.0%. The full reward maintains the most stable improvement by 700k training steps.
Zero-Shot Real-Map Transfer Results:
The zero-shot real-map transfer experiment uses a map of Xiazhimen in Zhejiang Province of China, where the OGR-MARL algorithms are evaluated without retraining. The shores and anchorage areas are drawn in QGIS, while the traffic data is derived from AIS data. Four cargo vessels from AIS-C1 to AIS-C4 are considered as dynamic obstacles. OGR-MASAC achieves 66.67% of SR with [41.7, 84.8]% of CI. Successful captures finish in 31.65±11.68 s, while escaped episodes end in 126.30±2.71 s. The SR in realistic maps is lower compared to the simulation experiments on the abstract map, attributed to factors such as curved coastlines, irregular port layouts, irregularly moving cargo vessels, and initial position offsets.
Conclusion Summary:
The paper proposes OGR-MARL, an algorithm-agnostic option-guided residual MARL framework for heterogeneous USV cooperative pursuit in constrained port waterways. Experiments demonstrate that it can be instantiated with different continuous-control MARL algorithms. In the abstract Xiazhimen scenario, the MASAC-based instantiation, namely OGR-MASAC, achieves a 75.0% capture rate and obtains promising mission time cost, mission-effective rule compliance, and the highest heterogeneous coordination score among the tested methods. The QGIS/AIS-informed zero-shot real-map transfer experiment further demonstrates the potential generalization capability of the proposed framework without retraining. Future directions include improving sample efficiency using world models and conducting further validation in more realistic or real-world maritime environments.
Improvements for AI systems
Improvements to AI Systems:
- Algorithm-Agnostic Residual Policy Learning with Rule-Based Priors:
-
Integrate a modular
option-guided residual
layer into any MARL backbone (e.g., MADDPG, MATD3, MAPPO, MASAC) that blends learned actions with rule-based high-level targets (search, track, intercept, block, recover). -
The improved system can bootstrap from domain knowledge (A*, APF, COLREGs) instead of exploring from scratch, drastically reducing sample complexity and improving convergence in constrained, safety-critical environments.
- Role-Conditioned Heterogeneous Coordination:
-
Implement shared evader belief propagation and role-specific option targets (e.g., interceptors for capture, scouts for sensing) with independent actor–critic networks per role.
-
The improved system can coordinate heterogeneous agents with asymmetric capabilities (speed, sensing range) in real time, achieving higher capture rates and better team-level efficiency than homogeneous or rule-only baselines.
- Adaptive Rule-Compliance Reward Shaping with Curriculum Learning:
-
Use a decomposed reward (task progress, rule compliance, role contribution, terminal) and a five-stage curriculum that progressively increases evader difficulty, capture strictness, and penalty strength.
-
The improved system can learn policies that respect navigation, traffic, and role constraints (e.g., avoid shores, anchorage, wrong-way lanes) while maintaining mission effectiveness, reducing violations and improving safety in dynamic port waterways.
- Zero-Shot Generalization to Real-World Maps via Belief and Option Abstraction:
-
Train on abstracted maps but transfer to realistic GIS/AIS-informed environments without retraining, using shared evader belief and option targets that are map-agnostic.
-
The improved system can generalize to unseen, complex port layouts (curved coastlines, irregular obstacles, moving vessels) and still achieve high capture rates, making it deployable in real-world scenarios with minimal adaptation.
- Bounded Action Blending for Safe Exploration:
-
Use a blending coefficient (β) to combine learned residual actions with rule-based option actions, clipping the final action to safe bounds.
-
The improved system can maintain safety during training and execution, preventing erratic or dangerous maneuvers, and can be tuned to trust learned policies more as they mature.
- Improved Final-Phase Interception via Residual Learning:
-
Address the bottleneck of terminal close-in interception (mean closest distance 8.19 in failures vs. 4.17 in successes) by learning residual corrections on top of option-guided intercept maneuvers.
-
The improved system can execute precise, high-speed capture maneuvers even when the evader replans, increasing success rates from 35% (rule-only) to 75% (OGR-MASAC) in constrained environments.
What the Improved AI System Can Do:
-
Deploy a team of heterogeneous USVs (or other multi-agent systems) to cooperatively intercept a fast, evasive target in constrained, rule-heavy environments (e.g., ports, urban canyons, airspace) with high success rates and safety compliance.
-
Operate with partial observability, leveraging shared beliefs and role-based coordination to maintain situational awareness and adapt to dynamic obstacles and traffic.
-
Transfer learned policies to new, realistic maps without retraining, enabling rapid deployment in real-world maritime, aerial, or ground scenarios.
-
Provide a plug-and-play framework that can upgrade existing MARL algorithms with rule-based priors, improving their performance in safety-critical tasks without redesigning the underlying learning method.
Abstract
Heterogeneous USV cooperative pursuit in constrained port waterways requires evader interception under navigation, traffic, and role constraints. This paper proposes OGR-MARL, an option-guided residual multi-agent reinforcement learning framework that is decoupled from a specific MARL algorithm. OGR-MARL integrates shared evader belief, role-conditioned option targets, adaptive rule penalties, and residual policy learning, allowing different MARL algorithms to learn corrective actions on top of rule-guided behaviors rather than exploring constrained port environments from scratch. We instantiate OGR-MARL with representative continuous-control MARL backbones, including MADDPG, MATD3, MAPPO, and MASAC, yielding OGR-MADDPG, OGR-MATD3, OGR-MAPPO, and OGR-MASAC. Experiments in an abstract Xiazhimen port-waterway scenario show that the OGR-MASAC instantiation achieves a 75.0% capture rate, promising mission-effective rule compliance, and the best heterogeneous coordination among the tested methods. Without retraining, zero-shot transfer to a QGIS/AIS-informed Xiazhimen map achieves promising results, demonstrating the generalization potential of OGR-MARL in more complex port scenarios.
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection