If You Hear It, Help Find It: Orthogonal Knowledge Distillation for Open-Vocabulary Audio-Visual Event Localization
cs.MM, cs.CV, cs.LG, cs.SD
Submitted: 2026-09-20
Updated: 2026-09-20
Comments: Accepted to ACM Multimedia 2026 (poster). 9 pages, 5 figures
Code: https://github.com/ScottBlizzard/OV-OrthKD
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- BEATs: Audio Pre-Training with Acoustic Tokenizers
- Decoupled Weight Decay Regularization
- OrthKD: Extracting Generalized Clinical Knowledge from Heterogeneous Teachers for Lightweight Deployment
Related papers
- ControlFoley: Unified and Controllable Video-to-Audio Generation with Cross-Modal Conflict Handling
- Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models
- Mitigating GenAI-powered Evidence Pollution for Out-of-Context Multimodal Misinformation Detection
- A Rate-Distortion-Classification Approach for Lossy Image Compression