MRMAD: A Multi-Round Multi-Audio Benchmark for Evaluating Acoustic Degradation Perception in Large Audio-Language Models
cs.SD, cs.LG, eess.AS
Submitted: 2026-08-23
Updated: 2026-08-31
Comments: EMNLP 2026, Code and benchmark: https://github.com/Bose/MRMAD
Code: https://github.com/Bose/MRMAD
License: http://creativecommons.org/licenses/by/4.0/
The gist: Large audio-language models (LALMs) have shown promising progress in understanding speech, music, and general sound events, yet their ability to reason about how audio signals are degraded remains
Terminology
Abstract
Large audio-language models (LALMs) have shown promising progress in understanding speech, music, and general sound events, yet their ability to reason about how audio signals are degraded remains underexplored. Existing benchmarks primarily evaluate semantic understanding, event recognition, or high-level audio reasoning, leaving a basic question unanswered: Do LALMs understand the differences in audio quality? We introduce MRMAD, a Multi-Round Multi-Audio Degradation benchmark for evaluating audio degradation perception and understanding in LALMs. MRMAD spans speech, music, and sound, and frames evaluation as multi-turn dialogues across multiple audio inputs, requiring models to identify types of degradation, compare severity, and perceive corruption changes across turns. Unlike current single-turn audio-language benchmarks, MRMAD evaluates whether LALMs can maintain consistent degradation hypotheses with new evidence and comprehend low-level acoustic phenomena over multi-turn dialogues. Through a systematic evaluation of 18 representative LALMs from non-thinking to reasoning and Omni models, we find that current models often recognize coarse content while failing to diagnose, compare, or reason about degradations reliably. Human evaluations further reveal a significant perception gap between LALMs and human listeners. MRMAD thus exposes a critical yet overlooked aspect of audio-language understanding and provides a diagnostic foundation for building future LALMs that are robust to real-world acoustic conditions.
Sources
- Hi-Fi Multi-Speaker English TTS Dataset
- Qwen2-Audio Technical Report
- Audio Flamingo Next: Next-Generation Open Audio-Language Models for Speech, Sound, and Music
- LoC-Path: Learning to Compress for Pathology Multimodal Large Language Models
- GPT-4o System Card
- Token Reduction Should Go Beyond Efficiency in Generative Models -- From Vision, Language to Multimodality
- BigVGAN: A Universal Neural Vocoder with Large-Scale Training
- Voxtral
- SEGAN: Speech Enhancement Generative Adversarial Network
- SpeechBrain: A General-Purpose Speech Toolkit
- MiMo-Audio: Audio Language Models are Few-Shot Learners
- Gemini: A Family of Highly Capable Multimodal Models
- Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
- Step-Audio-R1 Technical Report
- Step-Audio 2 Technical Report
- Why Can't They Remember? Uncovering Representation and Retrieval Bottlenecks in Multi-Turn Acoustic Memory
- Adaptive Information Control for Search-Augmented LLM Reasoning
- Qwen2.5-Omni Technical Report
- Qwen3-Omni Technical Report
- A Survey of Advancing Audio Super-Resolution and Bandwidth Extension from Discriminative to Generative Models
Related papers
- Few-Shot Open-Set Audio Classification via Transductive Prototype Refinement and Class Logit Enhancement
- Spectral Masking and Interpolation Attack (SMIA): A Black-box Adversarial Attack against Voice Authentication and Anti-Spoofing Systems
- AVMeme Exam: A Multimodal Multilingual Multicultural Benchmark for LLMs' Contextual and Cultural Knowledge and Thinking
- SoundWeaver: Compositional Warm-Starting for Text-to-Audio Diffusion Serving
- WASIL: In-the-Wild Arabic Spoken Interactions with LLMs
- Efficient Audiovisual Speech Processing via MUTUD: Multimodal Training and Unimodal Deployment