HA-VLN 2.0: An Open Benchmark and Leaderboard for Human-Aware Navigation in Discrete and Continuous Environments with Dynamic Multi-Human Interactions
Yifei Dong, Fengyi Wu, Qi He, Lingdong Kong, Heng Li, Minghan Li, Zebang Cheng, Yuxuan Zhou, Jingdong Sun, Qi Dai, Alexander G Hauptmann, Zhi-Qi Cheng
cs.AI, cs.CV, cs.RO
Submitted: 2026-08-18
Updated: 2026-08-19
Comments: Accepted to IROS 2026. 35 pages, 20 figures, website: https://uwmilab.github.io/HA-VLN-webpage/
Code: https://github.com/F1y1113/HA-VLN
Project page: https://uwmilab.github.io/HA-VLN-webpage
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- SHIELD: LLM-Driven Schema Induction for Predictive Analytics in EV Battery Supply Chain Disruptions
- BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
- Data Augmentation using Large Language Models: Data Perspectives, Learning Paradigms and Challenges
- Towards Unified World Models for Visual Navigation via Memory-Augmented Planning and Foresight
- Language-Conditioned World Modeling for Visual Navigation
- Stay on the Path: Instruction Fidelity in Vision-and-Language Navigation
- AI2-THOR: An Interactive 3D Environment for Visual AI
- Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal Grounding
- Human-Aware Vision-and-Language Navigation: Bridging Simulation to Reality with Dynamic Human Interactions
- VLNVerse: A Benchmark for Vision-Language Navigation with Versatile, Embodied, Realistic Simulation and Evaluation
- Self-Monitoring Navigation Agent via Auxiliary Progress Estimation
- Habitat 3.0: A Co-Habitat for Humans, Avatars and Robots
- Learning to Navigate Unseen Environments: Back Translation with Environmental Dropout
- Human Motion Diffusion Model
- Towards Versatile Embodied Navigation
- GoViG: Goal-Conditioned Visual Navigation Instruction Generation via Multimodal Reasoning
- Building Generalizable Agents with a Realistic and Rich 3D Environment
- Natural Language Can Help Bridge the Sim2Real Gap
- NaVid: Video-based VLM Plans the Next Step for Vision-and-Language Navigation
- Vision-and-Language Navigation Today and Tomorrow: A Survey in the Era of Foundation Models
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection