IoT-Enabled Autonomous Maritime Navigation in Smart Ports: A Curriculum-Guided Shared Policy Learning Framework

arXiv:2608.11597 · cs.RO, cs.LG · Submitted 2026-08-12 · Read on arXiv

Yuqing Lin, Rangya Zhang, Kum Fai Yuen

Nanyang Technological University

cs.RO, cs.LG

Submitted: 2026-08-12

Updated: 2026-08-13

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 85/100

The gist: This paper investigates onboard autonomous navigation for IoT-enabled autonomous surface vehicles operating in congested smart port environments under partial observability and dense traffic

Terminology

Summary

This paper investigates onboard autonomous navigation for IoT-enabled autonomous surface vehicles operating in congested smart port environments under partial observability and dense traffic conditions. A curriculum-guided reinforcement learning framework with a shared recurrent policy is developed to enhance temporal reasoning, deployment scalability, and robustness of edge-level decision-making. Centralized training is adopted as an offline design-time strategy, while all navigation actions are executed fully onboard, consistent with IoT edge intelligence paradigms.

The navigation task is formulated as a POMDP, and a shared recurrent PPO architecture with LSTM units embedded in both actor and critic networks is adopted to capture temporal dependencies and infer latent environmental states from observation histories. The observation space includes radar-based detection of static and dynamic obstacles, relative coordinates to the navigation target, normalized distances to environmental boundaries, relative positions of nearby vessels, and ego-motion variables. The continuous action space consists of commanded yaw rate and forward surge velocity.

A composite reward function is designed with four dense components: heading alignment, distance progress, regulation-aware interaction reward, and proximity-based safety penalty, supplemented by sparse terminal rewards. The regulation-aware interaction reward categorizes encounters (head-on, crossing, overtaking, parallel) based on relative bearing and ship domain location, embedding COLREGs as hard regulatory constraints within the environment and reward formulation.

A hybrid curriculum learning strategy combines designer-specified task staging with self-paced progression driven by policy performance. Task difficulty is gradually increased across three dimensions: deployment scale, traffic interaction complexity, and environmental complexity. Three representative port environments of increasing operational complexity are used: Los Angeles (open water, low traffic), Singapore Pasir Panjang (semi-structured, dense traffic), and Rotterdam (hybrid layout with complex transitions). Curriculum progression is triggered by either a fixed number of training episodes or performance thresholds (e.g., success rate ≥ 80% over a moving evaluation window).

Extensive simulations in multiple realistic port environments demonstrate that the proposed approach improves navigation reliability, collision avoidance, and training stability compared with standard baseline methods (DDPG, SAC, and standard PPO), and generalizes effectively to previously unseen high-density scenarios. In Stage 1 (Los Angeles), the proposed PPO achieves 96.2% success rate and 2.1% collision rate. In Stage 2 (Singapore), it maintains 90.5% success and 4.0% collisions. In Stage 3 (Rotterdam), under dense interactions and disturbances, it achieves 85.6% success and 6.3% collisions. Generalization scores follow a similar trend, with the proposed method consistently achieving strong performance across stages.

Ablation studies show that the combination of shared recurrent policy learning, staged curriculum training, and composite reward design improves robustness, safety, and reliability of onboard navigation under partial observability. Removing curriculum training or using individual policies increases collisions by approximately 40% and significantly increases risk. Training directly in Stage 3 without curriculum results in substantially higher collision rates and over 20% reduction in success rate.

Behavioral analysis reveals that the learned policy develops consistent navigation strategies, including early conflict anticipation, smooth overtaking, structured head-on avoidance, and stable separation in multi-vessel interactions. These behaviors persist in progressively complex port scenarios and under environmental disturbances, indicating that the policy captures transferable navigation patterns rather than scenario-specific behaviors. The results indicate that curriculum-guided shared learning provides a practical solution for scalable deployment of IoT-enabled autonomous maritime devices in smart port operations.

Improvements for AI systems

Improvements to AI Systems:

  1. Recurrent Policy Architecture for Partial Observability – Integrate LSTM-based actor-critic networks into any sequential decision-making system (e.g., robotics, autonomous driving) to infer hidden states from observation histories, enabling robust operation when sensors are noisy or occluded.

  2. Curriculum-Guided Reinforcement Learning – Implement a hybrid curriculum that combines fixed task staging with self-paced progression based on performance thresholds (e.g., success rate ≥80%). This allows AI agents to learn complex tasks incrementally, reducing catastrophic forgetting and improving final policy stability.

  3. Regulation-Aware Reward Shaping – Embed domain-specific rules (e.g., COLREGs for maritime, traffic laws for driving) directly into the reward function via encounter categorization (head-on, crossing, overtaking). This ensures the AI system learns to comply with hard constraints while optimizing for efficiency, not just raw reward maximization.

  4. Shared Policy for Multi-Agent Scalability – Use a single recurrent policy shared across multiple agents (vessels, vehicles, drones) trained centrally but executed edge-locally. This reduces deployment memory, enables seamless scaling to new environments without retraining, and improves generalization to unseen dense traffic.

  5. Composite Dense Reward Design – Combine multiple dense reward components (heading alignment, distance progress, interaction-aware penalties, safety proximity) with sparse terminal rewards. This accelerates early learning and prevents reward hacking, leading to smoother, safer trajectories.

  6. Performance-Triggered Curriculum Progression – Automatically advance task difficulty based on moving-window success rates rather than fixed episode counts. This adapts training to the agent’s current capability, improving sample efficiency and final robustness.


What the Improved AI System Can Do:

  • Navigate autonomously in congested, partially observable environments (e.g., smart ports, urban intersections, warehouse floors) with high reliability (85–96% success) and low collision rates (2–6%), even under dense traffic and environmental disturbances.

  • Generalize to unseen high-density scenarios without retraining, thanks to shared recurrent policies and curriculum learning that capture transferable navigation patterns (early conflict anticipation, smooth overtaking, structured avoidance).

  • Comply with regulatory and safety constraints (e.g., maritime COLREGs, traffic rules) by learning encounter-specific interaction behaviors, reducing legal and operational risks.

  • Deploy on edge devices with limited compute (IoT-enabled hardware) because all inference is onboard, while training remains centralized offline—enabling real-time decision-making with low latency.

  • Scale to multiple agents or vessels with a single policy, reducing memory footprint and enabling coordinated behavior in multi-agent settings without explicit communication.

  • Train more stably and faster than standard RL baselines (DDPG, SAC, vanilla PPO), with 40% fewer collisions and >20% higher success rates when curriculum and shared recurrent learning are used, making it suitable for safety-critical applications.

Sources

Related papers