EAGOR: Embodied Reasoning in Omni-direction
cs.RO
Submitted: 2026-07-07
Updated: 2026-09-25
Terminology
Sources
- Embodied Question Answering
- Beyond the Nav-Graph: Vision-and-Language Navigation in Continuous Environments
- Mind Meets Space: Rethinking Agentic Spatial Intelligence from a Neuroscience-inspired Perspective
- Vision-and-Language Navigation: Interpreting visually-grounded navigation instructions in real environments
- PanoContext-Former: Panoramic Total Scene Understanding with a Transformer
- Learning Spherical Convolution for Fast Features from 360{\deg} Imagery
- Visual Question Answering on 360{\deg} Images
- Object Goal Navigation using Goal-Oriented Semantic Exploration
- Pano-AVQA: Grounded Audio-Visual Question Answering on 360$^\circ$ Videos
- Tangent Images for Mitigating Spherical Distortion
- SphereUFormer: A U-Shaped Transformer for Spherical 360 Perception
- Bending Reality: Distortion-aware Transformers for Adapting to Panoramic Semantic Segmentation
- DA$^{2}$: Depth Anything in Any Direction
- Are Multimodal Large Language Models Ready for Omnidirectional Spatial Reasoning?
- Spatial Reasoning in Multimodal Large Language Models: A Survey of Tasks, Benchmarks and Methods
- Multimodal Spatial Reasoning in the Large Model Era: A Survey and Benchmarks
- ImagineNav: Prompting Vision-Language Models as Embodied Navigator through Scene Imagination
- PanoNav: Mapless Zero-Shot Object Navigation with Panoramic Scene Parsing and Dynamic Memory
- VLFM: Vision-Language Frontier Maps for Zero-Shot Semantic Navigation
- GridMM: Grid Memory Map for Vision-and-Language Navigation
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving