Supervising Sound Localization by In-the-wild Egomotion
cs.CV, cs.AI, cs.MM, cs.SD
Submitted: 2026-10-01
Updated: 2026-10-01
Terminology
Sources
- A multi-room reverberant dataset for sound event localization and detection
- Audio-Visual Synchronisation in the wild
- Structure from Silence: Learning Scene Structure from Ambient Sound
- BatVision with GCC-PHAT Features for Better Sound to Vision Predictions
- FMA: A Dataset For Music Analysis
- Self-Supervised Video Forensics by Audio-Visual Anomaly Detection
- Synchformer: Efficient Synchronization from Sparse Cues
- Self-Supervised Generation of Spatial Audio for 360 Video
- A Dataset of Reverberant Spatial Sound Scenes with Moving Sources for Sound Event Localization and Detection
- A Dataset of Dynamic Reverberant Sound Scenes with Directional Interferers for Sound Event Localization and Detection
- STARSS22: A dataset of spatial recordings of real scenes with spatiotemporal annotations of sound events
- Habitat-Matterport 3D Dataset (HM3D): 1000 Large-scale 3D Environments for Embodied AI
- BAT: Learning to Reason about Spatial Sounds with Large Language Models
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models