Do Audio LLMs Listen Before They Act? Diagnosing Acoustic-Context Gating in Voice Agents
cs.SD, cs.AI, cs.CL, cs.MM, eess.AS
Submitted: 2026-09-26
Updated: 2026-09-26
Terminology
Sources
- Nemotron 3 Nano Omni: Efficient and Open Multimodal Intelligence
- Kimi-Audio Technical Report
- Reinforcement Learning Outperforms Supervised Fine-Tuning: A Case Study on Audio Question Answering
- Device-directed Utterance Detection
- SpeakerLLM: A Speaker-Specialized Audio-LLM for Speaker Understanding and Verification Reasoning
- Audio2Tool: Speak, Call, Act -- A Dataset for Benchmarking Speech Tool Use
- Do Audio LLMs Listen or Read? Analyzing and Mitigating Paralinguistic Failures with VoxParadox
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- MSU-Bench: Towards Speaker-Centric Understanding in Conversational Multi-Speaker Scenarios
- Step-Audio-R1 Technical Report
- SARI: Structured Audio Reasoning via Curriculum-Guided Reinforcement Learning
- AudioToolAgent: An Agentic Framework for Audio-Language Models
- Escape the Language Prior: Mitigating Late-Stage Modality Collapse in Audio Reasoning via Modality-Aware Policy Optimization
- Thinking with Sound: Audio Chain-of-Thought Enables Multimodal Reasoning in Large Audio-Language Models
- Qwen3-Omni Technical Report
- From Reactive to Proactive: Assessing the Proactivity of Voice Agents via ProVoice-Bench
- AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning
- WildSpeech-Bench: Benchmarking End-to-End SpeechLLMs in the Wild
Related papers
- Few-Shot Open-Set Audio Classification via Transductive Prototype Refinement and Class Logit Enhancement
- Spectral Masking and Interpolation Attack (SMIA): A Black-box Adversarial Attack against Voice Authentication and Anti-Spoofing Systems
- AVMeme Exam: A Multimodal Multilingual Multicultural Benchmark for LLMs' Contextual and Cultural Knowledge and Thinking
- SoundWeaver: Compositional Warm-Starting for Text-to-Audio Diffusion Serving
- WASIL: In-the-Wild Arabic Spoken Interactions with LLMs
- Efficient Audiovisual Speech Processing via MUTUD: Multimodal Training and Unimodal Deployment