OP-CAD: On-Policy Clean-Audio Distillation for Robust Audio-Visual Reasoning
cs.CV, cs.SD
Submitted: 2026-09-30
Updated: 2026-09-30
Terminology
Sources
- VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
- VITA: Towards Open-Source Interactive Omni Multimodal LLM
- Distilling the Knowledge in a Neural Network
- CORD: Bridging the Audio-Text Reasoning Gap via Weighted On-policy Cross-modal Distillation
- VideoMem: Enhancing Ultra-Long Video Understanding via Adaptive Memory Management
- Baichuan-Omni Technical Report
- VISD: Enhancing Video Reasoning via Structured Self-Distillation
- EchoDistill: Robust Large Audio Language Models via Noisy-to-Clean Self-Distillation
- VidForensics-M1: Meta-Detection Reinforcement Learning with Verifiable Temporal Grounding for AI-Generated Video Forensics
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- MUSAN: A Music, Speech, and Noise Corpus
- AdaThinkV: Adaptive Thinking for Token-Efficient Video Reasoning
- Qwen2.5-Omni Technical Report
- Qwen3-Omni Technical Report
- Focus Then Listen: An Empirical Study of Plug-and-Play Audio Enhancer for Noise-Robust Large Audio Language Models
- Generative Universal Verifier as Multimodal Meta-Reasoner
- Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models
- Daily-Omni: Towards Audio-Visual Reasoning with Temporal Alignment across Modalities
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models