Towards the Vision-Sound-Language-Action Paradigm: The HEAR Framework for Sound-Centric Manipulation
cs.RO, cs.AI, cs.CV, cs.SD
Submitted: 2026-03-17
Updated: 2026-09-18
Comments: Accepted by The International Journal of Robotics Research (IJRR 2026). Project page: https://hear.irmv.top
DOI: 10.1177/02783649261477778
Project page: https://hear.irmv.top
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation
- Moshi: a speech-text foundation model for real-time dialogue
- Compact 3D Gaussian Splatting For Dense Visual SLAM
- What Is The Best 3D Scene Representation for Robotics? From Geometric to Foundation Models
- $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization
- Mobile ALOHA: Learning Bimanual Mobile Manipulation with Low-Cost Whole-Body Teleoperation
- Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning
- OmniVLA: Physically-Grounded Multimodal VLA with Unified Multi-Sensor Perception for Robotic Manipulation
- Memory-based control with recurrent neural networks
- Tactile-VLA: Unlocking Vision-Language-Action Model's Physical Knowledge for Tactile Generalization
- MLA: A Multisensory Language-Action Model for Multimodal Understanding and Forecasting in Robotic Manipulation
- VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning
- The Sound of Simulation: Learning Multimodal Sim-to-Real Robot Policies with Generative Audio
- RoboOmni: Proactive Robot Manipulation in Omni-modal Context
- Audio-VLA: Adding Contact Audio Perception to Vision-Language-Action Model for Robotic Manipulation
- ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents
- Qwen3-Omni Technical Report
- Qwen3 Technical Report
- VibeCheck: Using Active Acoustic Tactile Sensing for Contact-Rich Manipulation
- VLAS: Vision-Language-Action Model With Speech Instructions For Customized Robot Manipulation
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving