How Contrastive Decoding Enhances Large Audio Language Models
cs.SD, cs.CL, eess.AS
Submitted: 2026-03-10
Updated: 2026-09-13
Comments: Accepted at IEEE SLT 2026. Code is available at: https://github.com/nervjack2/LALM-Contrastive-Decoding-Error-Profiles
Code: https://github.com/nervjack2/LALM-Contrastive-Decoding-Error-Profiles
License: http://creativecommons.org/licenses/by/4.0/
The gist: While Contrastive Decoding (CD) has been proposed to enhance Large Audio Language Models (LALMs), it has not been evaluated at scale, and the underlying mechanisms driving its success remain unclear.
Terminology
Abstract
While Contrastive Decoding (CD) has been proposed to enhance Large Audio Language Models (LALMs), it has not been evaluated at scale, and the underlying mechanisms driving its success remain unclear. This study systematically evaluates four distinct CD strategies across diverse LALM architectures. We identify Audio-Aware Decoding and Audio Contrastive Decoding as the most effective methods. However, their impact varies significantly across models. To explain this variability, we profile the baseline error composition of each model and measure how readily contrastive decoding corrects each error type. Our analysis demonstrates that CD reliably rectifies errors in which models falsely claim an absence of audio or resort to uncertainty-driven guessing, but is relatively poor at correcting flawed reasoning or confident misassertions. Crucially, CD's benefit closely tracks the composition of a model's baseline error profile: when errors caused by audio ignorance or uncertainty-driven guessing constitute only a small fraction of a model's errors, gains are marginal or even negative. A token-level analysis reveals the underlying mechanism: when the amateur's output is dominated by hesitation markers, CD's suppression naturally targets uncertainty-driven errors while having limited effect on confident misassertions.
Sources
- Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models
- Qwen2-Audio Technical Report
- Qwen2.5-Omni Technical Report
- DeSTA2.5-Audio: Toward General-Purpose Large Audio Language Model with Self-Generated Cross-Modal Alignment
- WavLLM: Towards Robust and Adaptive Speech Large Language Model
- Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs
- Distilling an End-to-End Voice Assistant Without Instruction Training Data
- Less is More: Improving LLM Reasoning with Minimal Test-Time Intervention
Related papers
- Few-Shot Open-Set Audio Classification via Transductive Prototype Refinement and Class Logit Enhancement
- Spectral Masking and Interpolation Attack (SMIA): A Black-box Adversarial Attack against Voice Authentication and Anti-Spoofing Systems
- AVMeme Exam: A Multimodal Multilingual Multicultural Benchmark for LLMs' Contextual and Cultural Knowledge and Thinking
- SoundWeaver: Compositional Warm-Starting for Text-to-Audio Diffusion Serving
- WASIL: In-the-Wild Arabic Spoken Interactions with LLMs
- Efficient Audiovisual Speech Processing via MUTUD: Multimodal Training and Unimodal Deployment