Multi-AUV Ad-hoc network-based Target Tracking: A Value Gradient Guidance Multi-Agent Diffusion Reinforcement Learning Approach
Jiaao Ma, Chuan Lin, Guangjie Han, Shengchao Zhu, Qian Zhu, Ying Liu, Zhenyu Wang
Northeastern University · Hohai University
cs.LG, cs.MA, cs.NI
Submitted: 2026-08-12
Updated: 2026-08-14
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 75/100
The gist: This paper investigates cooperative target tracking in multi-AUV ad-hoc networks and proposes the MDCA hierarchical control architecture together with the VGG-MADiffRL algorithm.
Terminology
Summary
This paper investigates cooperative target tracking in multi-AUV ad-hoc networks and proposes the MDCA hierarchical control architecture together with the VGG-MADiffRL algorithm. The paper states: This paper investigated cooperative target tracking in multi-AUV ad-hoc networks and proposed the MDCA hierarchical control architecture together with the VGG-MADiffRL algorithm.
The problem addressed is that "Multi-AUV ad-hoc network-based target tracking requires networked autonomous underwater vehicles (AUVs) to cooperatively track maneuvering targets under constrained acoustic communication, dynamic topology, and uncertain ocean disturbances. The paper notes that
although multi-agent reinforcement learning (MARL) enables decentralized coordination through centralized training, existing methods suffer from high-dimensional joint state-action modeling, noise-sensitive policy generation, leading to unstable training and degraded tracking."
To address these issues, the authors propose VGG-MADiffRL, a value-gradient-guided multi-agent diffusion RL algorithm, and MDCA, a diffusion-based hierarchical control architecture.
The paper explains that Leveraging underwater mission characteristics, we model sonar detection mechanisms and ocean current disturbances, formulating cooperative tracking for multi-AUV ad-hoc networks as an MDP.
The proposed MDCA "constitutes a three-tier closed-loop control framework: a global intelligent control layer, a local online training layer, and a physical action execution layer. This structure enables synergistic optimization across task allocation, local decision processes, and execution feedback. The paper describes that
Within MDCA, the local online training layer is the policy learning framework; VGG-MADiffRL builds on diffusion policies and incorporates value gradients to guide action generation in the reverse denoising process, steering the generated actions towards higher expected returns. It employs twin value networks with joint optimization and soft target updates to mitigate overestimation and training oscillations, promoting more stable convergence."
The paper identifies three main challenges in existing MARL methods for Multi-AUV ad-hoc network cooperative tracking: "(1) Insufficient modeling of complex continuous action distributions: traditional deterministic policies struggle to capture the strong interdependencies and couplings of continuous joint actions needed for collaborative decision-making under dynamic topologies, often leading to limited expressiveness and suboptimal decisions; (2) Training instability caused by relying on a single optimization objective: when policy updates depend solely on a single reward signal or local supervisory signal, they become highly sensitive to intermittent communication failures, which can cause oscillations in policy learning, slow convergence, and performance fluctuations; (3) Lack of value guidance during sampling: in the reverse denoising process of diffusion models, the absence of explicit value constraints allows sampled actions to drift away from regions of high expected return and introduces ineffective noise, reducing decision consistency, a problem that is especially harmful in low-bandwidth underwater environments."
The core components of VGG-MADiffRL are described as: "(1) A diffusion policy architecture that improves the modeling of complex, interdependent cooperative actions. (2) A dual-objective framework that combines value signals with policy gradients to stabilize training when the topology changes. (3) A sampling mechanism that uses value gradients to refine denoising trajectories, prioritizing actions with high expected returns."
The paper details the three key innovations: "1) Framework for policy learning via diffusion models: VGG-MADiffRL replaces the conventional deterministic actor with a diffusion model to construct a generative policy architecture closely integrated with value estimation. This enables stable and robust policy optimization in multi-AUV ad-hoc networks, thereby enhancing the policy capability to model complex, continuous, and highly interdependent action distributions inherent in dynamic cooperative scenarios. 2) Combined optimization mechanism for actor networks: VGG-MADiffRL incorporates a dual objective optimization scheme that jointly minimizes a loss function informed by value estimates alongside a policy gradient loss. The component guided by value signals leverages global value estimates to constrain the diffusion policy update direction, effectively mitigating policy oscillations caused by single objective optimization under volatile network topologies. 3) Diffusion sampling mechanism guided by value gradients: VGG-MADiffRL integrates value signals into the reverse denoising process to iteratively refine the denoising trajectory, steering action sampling toward regions of high expected return from the outset. This mechanism reduces the influence of suboptimal or irrelevant actions and enhances training stability in underwater networks with limited bandwidth and autonomously organized topologies."
The paper models the underwater environment, stating: "Given the uncertain and dynamic underwater environment, sonar is employed for precise relative positioning between AUVs and targets. The AUV emits acoustic waves to scan its surroundings, and a sector coverage receiver array captures echoes from multiple directions, enabling target localization via intensity analysis." The target detection process is modeled using the active sonar equation. Ocean currents are modeled using the Navier–Stokes equation, and a Kalman filter is used for state estimation.
The cooperative tracking problem is formulated as a Markov Decision Process (MDP) defined by the tuple M = (S, A, P, R, γ). The state space contains the states of all AUVs, the action space represents continuous actions of the multi-AUV system, and the reward function accounts for several factors: it explicitly incorporates tracking accuracy, collision avoidance, and environmental constraints to improve cooperative multi-AUV tracking performance.
The proposed reward function is composite, given by Ri = α · rpos + β · rcol + δ · rland, comprising three reward components for tracking fidelity, inter-agent safety, and environmental constraints.
The target position reward is defined based on distance to the target, the collision penalty is based on relative inter-agent distances, and the obstacle avoidance penalty uses a Sigmoid function for smooth modeling of landmark region constraints.
The paper presents the complete algorithm in Algorithm 1, which integrates guided diffusion sampling, twin-critic learning, and soft target updates into a unified multi-AUV loop.
The algorithm initializes the environment, replay buffer, and global timestep counter, iterates over episodes, and at each timestep generates actions by either value-gradient-guided diffusion sampling or unguided diffusion policy sampling. The joint action is executed, transitions are stored in the replay buffer, and when sufficient samples are available, the twin critics and diffusion actors are updated using the respective loss functions.
Experimental results are presented across four scenarios: 10 AUVs tracking 3 targets, 8 AUVs tracking 3 targets, 6 AUVs tracking 2 targets, and 4 AUVs tracking 2 targets.
The paper compares VGG-MADiffRL against seven baseline methods: MASAC, MAPPO, MAAC, MATD3, and MADDPG
from the general-purpose MARL group, and DSBM and MA-A3C
from the underwater-specific group.
The results show that VGG-MADiffRL converges faster than all compared methods across all scenarios.
The paper explains: "During early training, the critic value-guided mechanism drives rapid policy improvement: the diffusion policy generates actions via differentiable sampling, and a joint loss combining policy gradient objectives with Q-guidance terms from global dual-Q network outputs steers updates toward high-value regions. Batch updates that start after the replay buffer reaches a minimum size suppress small-sample bias and improve sample reuse. In later training, the algorithm remains smooth and stable. The dual-Q target networks take the minimum of two independent estimates to reduce overestimation, while soft updates avoid abrupt parameter shifts."
For tracking accuracy, the paper states: "VGG-MADiffRL achieves the highest tracking accuracy under this scenario, significantly outperforming existing baseline methods. These results fully verify the robustness and reliability of the proposed algorithm in achieving high-precision, sustained, and stable tracking within complex, dynamic underwater ad-hoc networks."
For mean tracking error (MTE), "Table III shows that VGG-MADiffRL achieves the lowest MTE value among all compared methods, demonstrating a more pronounced error advantage. These results indicate the proposed method enables higher-precision and more robust continuous cooperative tracking in complex dynamic underwater ad-hoc networks."
For error standard deviation, "Table IV shows that VGG-MADiffRL achieves the lowest Error Std among all compared methods, demonstrating that the proposed algorithm not only maintains a low mean tracking error but also exhibits superior error suppression and more stable dynamic tracking performance in complex underwater ad-hoc networks."
The paper also evaluates the number of diffusion time steps, stating: "Experimental results demonstrate that the adopted diffusion step configuration achieves a better balance between tracking effectiveness and computational cost, while maintaining satisfactory cooperative tracking performance in multi-AUV ad-hoc networks."
Ablation studies are conducted with two variants: "(1) removal of the value-gradient-guided reverse diffusion mechanism (excluding value function guidance during action sampling); and (2) replacement of the diffusion policy module with a conventional deterministic policy network to examine the individual impact of diffusion modeling on performance. The paper reports:
Fig. 6 shows the complete method consistently outperforms both ablated variants in convergence stability and cumulative return. These results demonstrate that both the value-gradient guidance mechanism and the diffusion-based policy module play critical roles in enhancing cooperative tracking performance, thereby validating the effectiveness of the proposed architectural design."
An availability evaluation is also conducted using a high-fidelity underwater simulation environment using the 3D modeling and physics engine of Unity,
which "reproduces complex underwater dynamics (acoustic communication constraints, ocean current disturbances, and time-varying network topologies) and provides a reliable platform for validating the stability and effectiveness of the algorithm."
The paper concludes: "Extensive experiments across four multi-AUV tracking scenarios demonstrate that VGG-MADiffRL consistently outperforms seven state-of-the-art MARL algorithms in convergence speed, tracking accuracy, mean tracking error, and error stability. Ablation studies confirm that both the value-gradient guidance and the diffusion policy module contribute substantially to overall performance."
Future research directions identified include: "(1) optimizing underwater obstacle avoidance to reduce potential AUV damage; (2) balancing energy consumption among AUVs to extend system endurance; and (3) designing robust control frameworks that explicitly account for unstable underwater acoustic communication."
Improvements for AI systems
Improvements to AI Systems:
- Value-Gradient-Guided Diffusion Sampling for Continuous Action Generation
-
Integrate value gradients into the reverse denoising process of diffusion policies, steering sampled actions toward high-expected-return regions. This reduces ineffective noise and improves decision consistency in low-bandwidth, dynamic environments.
-
Capability: AI systems can generate more precise, interdependent continuous actions (e.g., multi-robot coordination, autonomous navigation) with fewer wasted samples and higher task success rates.
- Dual-Objective Actor Optimization with Twin Value Networks
-
Combine policy gradient loss with value-estimate-guided loss, using twin critics and soft target updates to mitigate overestimation and training oscillations.
-
Capability: AI systems achieve stable convergence under intermittent communication or topology changes, making them robust for decentralized multi-agent tasks in real-world networks (e.g., underwater, space, or disaster-response).
- Hierarchical Three-Tier Control Architecture (MDCA)
-
Separate global task allocation, local online learning, and physical action execution into closed-loop layers. This enables synergistic optimization across planning, decision-making, and feedback.
-
Capability: AI systems can scale to large teams of agents (e.g., 10+ AUVs) while maintaining real-time adaptability, reducing computational burden on individual agents, and improving overall mission reliability.
- Composite Reward Modeling for Safety and Environmental Constraints
-
Use a multi-component reward function (tracking accuracy, collision avoidance, obstacle penalties via smooth Sigmoid functions) to guide learning under physical constraints.
-
Capability: AI systems can learn safe, collision-free behaviors in cluttered or hazardous environments, reducing damage risk and improving long-term operational endurance.
- Physics-Informed Environment Modeling for Realistic Training
-
Incorporate active sonar equations, Navier–Stokes ocean current models, and Kalman-filter-based state estimation into the MDP formulation.
-
Capability: AI systems trained in simulation can transfer more effectively to real-world conditions with uncertain dynamics, improving generalization and reducing sim-to-real gaps.
- Diffusion Policy as a Generative Alternative to Deterministic Actors
-
Replace deterministic policies with diffusion models to capture complex, high-dimensional joint action distributions, especially when actions are strongly interdependent.
-
Capability: AI systems can handle tasks requiring coordinated multi-agent maneuvers (e.g., formation control, cooperative transport) with higher expressiveness and fewer suboptimal decisions.
- Ablation-Informed Modular Design for Performance Tuning
-
Use ablation studies (e.g., removing value guidance or diffusion module) to isolate and validate the contribution of each component, enabling systematic optimization of AI architectures.
-
Capability: AI developers can efficiently identify critical modules, reduce unnecessary complexity, and tailor systems to specific operational constraints (e.g., limited compute or bandwidth).
- Unified Algorithm Loop for Sample-Efficient Learning
-
Implement a replay buffer with minimum-size thresholding and batch updates to suppress small-sample bias, improving sample reuse and early-training stability.
-
Capability: AI systems learn faster with fewer interactions, crucial for applications where data collection is expensive or dangerous (e.g., underwater missions, robotic surgery).
What the Improved AI System Can Do:
-
Coordinate large teams of autonomous agents (e.g., AUVs, drones, ground robots) to track or intercept moving targets with high precision, even under constrained communication and dynamic topology.
-
Maintain stable learning and performance despite noisy, intermittent, or delayed feedback, making it suitable for real-world ad-hoc networks.
-
Generate smooth, safe, and energy-efficient actions in complex environments, reducing collisions and extending mission duration.
-
Adapt quickly to new scenarios (e.g., varying team sizes, target counts, or environmental disturbances) with minimal retraining, thanks to hierarchical control and generative policies.
-
Provide a reliable, physics-grounded training framework that reduces the gap between simulation and deployment, enabling faster field adoption.
Sources
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks