Air-Ground Collaborative Vision-and-Language Navigation via Shared Bird's-Eye Maps
cs.RO, cs.AI
Submitted: 2026-09-03
Updated: 2026-09-08
Comments: 8 pages, 5 figures
Code: https://github.com/ZSN2024/AGC-VLN
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Uni-LaViRA: Language-Vision-Robot Actions Translation for Unified Embodied Navigation
- Fly0: Persistent Metric Anchoring for Zero-Shot Aerial Vision-Language Navigation
- CARLA-Air: Fly Drones Inside a CARLA World -- A Unified Infrastructure for Air-Ground Embodied Intelligence
- AirGroundBench: Probing Spatial Intelligence in Multimodal Large Models under Heterogeneous Multi-View Embodied Collaboration
- Can Aerial VLA Models Cooperate? Evaluating Closed-Loop Air-Ground Coordination with CARLA-Air
- ABot-N1: Toward a General Visual Language Navigation Foundation Model
- LongNav-R1: Horizon-Adaptive Multi-Turn RL for Long-Horizon VLA Navigation
- OmniVLN: Omnidirectional 3D Perception and Token-Efficient LLM Reasoning for Visual-Language Navigation across Air and Ground Platforms
- CLOSER-VLN: Closed-Loop Self-Verified Retrieval-Augmented Reasoning for Aerial Vision-Language Navigation
- FSD-VLN: Fast-Slow Dual-System Modeling for Aerial Long-Horizon Vision-Language Navigation
- SpatialUAV: Benchmarking Spatial Intelligence for Low-Altitude UAV Perception, Collaboration, and Motion
- Towards Reliable Aerial Ground Vehicle Collaboration: An Integrated Planning and Autonomy Framework for Field Deployment
- Vision-and-Language Navigation for UAVs: Progress, Challenges, and a Research Roadmap
- Gemini: A Family of Highly Capable Multimodal Models
- See-and-Reach: Precise Vision-Language Navigation for UAVs within the Field of View
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving