CTAN: Cycle-Temporal Attention Network for Embodied Audio-Visual Navigation
cs.MM, cs.AI, cs.SD, eess.SP
Submitted: 2026-07-25
Updated: 2026-07-25
Comments: Main paper (6 pages). Accepted for publication by IEEE International Conference on Systems, Man, and Cybernetics 2026 (IEEE SMC 2026)
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
The gist: Audio-visual embodied navigation equips robots with the capability to infer the locations of sound sources by integrating visual inputs and acoustic information (e.g., depth observations and binaural
Terminology
Abstract
Audio-visual embodied navigation equips robots with the capability to infer the locations of sound sources by integrating visual inputs and acoustic information (e.g., depth observations and binaural audio cues). The core challenge lies in establishing effective semantic interactions across heterogeneous modalities (which exhibit distinct feature distributions). Existing feature fusion strategies, however, often rely on simple multimodal aggregation and therefore fail to capture the underlying geometric and semantic relationships, leading to information degradation in complex environments. To overcome these limitations, this work presents the Cycle-Temporal Attention Network (CTAN), a framework designed for active semantic-enhanced fusion (rather than straightforward multimodal combination). Specifically, the proposed Audio-Visual Reconstruction Cross-Attention (AVRCA) module employs a bidirectional cycle-consistency constraint (between visual and acoustic representations) to reinforce the spatial semantic attributes of both modalities, thereby facilitating more robust cross-modal interaction. Additionally, we design a Temporal Cross-Modal Memory (TCMM) mechanism to dynamically integrate real-time enhanced multimodal features with historical context, reducing performance drops caused by auditory dead zones. Experimental results obtained on the Replica and Matterport3D benchmarks indicate that the proposed approach achieves superior performance over previous audio-visual navigation methods in terms of success rate (SR), success weighted by path length (SPL), and scene navigation accuracy (SNA).
Related papers
- ControlFoley: Unified and Controllable Video-to-Audio Generation with Cross-Modal Conflict Handling
- Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models
- Mitigating GenAI-powered Evidence Pollution for Out-of-Context Multimodal Misinformation Detection
- A Rate-Distortion-Classification Approach for Lossy Image Compression