AgenticDiffusion: Multi-View Reasoning with View-Conditioned Diffusion Planning for Vision-Based UAV Navigation
cs.RO, cs.AI, cs.SY, eess.SY
Submitted: 2026-06-02
Updated: 2026-09-22
Code: https://github.com/openclaw/openclaw
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
The gist: Vision-based UAV navigation becomes challenging when navigation targets are distributed across complementary camera views and cannot be reliably observed from a single viewpoint.
Terminology
Abstract
Vision-based UAV navigation becomes challenging when navigation targets are distributed across complementary camera views and cannot be reliably observed from a single viewpoint. We propose AgenticDiffusion, an agentic multi-view UAV navigation framework that semantically coordinates first-person-view (FPV) and top-view observations for mission-level navigation. Given a natural-language instruction, AgenticDiffusion identifies the requested targets, selects the most appropriate camera view for each navigation task, determines the corresponding navigation goal, and invokes the appropriate view-conditioned diffusion planner for trajectory generation. The resulting trajectories are executed using Nonlinear Model Predictive Control (NMPC). AgenticDiffusion was evaluated in four real-world indoor scenarios, achieving an overall mission success rate of 80% across 40 physical-flight trials. In mixed-visibility scenarios, where the requested targets were distributed across FPV and top-view observations, coordinated multi-view navigation reduced average mission time by 50.8% relative to FPV-only navigation and by 26.8% relative to Top-only navigation. The semantic view-selection mechanism was also robust to lexical variation in target descriptions, achieving 100% accuracy across 66 test cases, compared with 63.64% for a confidence-based view-selection baseline. In a substantially larger Gazebo environment, AgenticDiffusion achieved a 90% mission success rate and completed the multi-stage mission, whereas the FPV-only and Top-only variants were unable to complete all requested navigation tasks.
Sources
- NavDP: Learning Sim-to-Real Navigation Diffusion Policy with Privileged Information Guidance
- 3D RL-DWA: A Hybrid Reinforcement Learning and Dynamic Window Approach for Goal-Directed Local Navigation in Multi-DoF Robots
- Deep Reinforcement Learning Based Navigation with Macro Actions and Topological Maps
- MM-Nav: Multi-View VLA Model for Robust Visual Navigation via Multi-Expert Learning
- MapNav: A Novel Memory Representation via Annotated Semantic Maps for Vision-and-Language Navigation
- Fast-SmartWay: Panoramic-Free End-to-End Zero-Shot Vision-and-Language Navigation
- AgentVLN: Towards Agentic Vision-and-Language Navigation
- A Hierarchical Agentic Framework for Autonomous Drone-Based Visual Inspection
- Action Agent: Agentic Video Generation Meets Flow-Constrained Diffusion
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving