EchoDistill: Robust Large Audio Language Models via Noisy-to-Clean Self-Distillation

arXiv:2605.23954 · cs.CL, cs.AI, cs.SD · Submitted 2026-05-11 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "EchoDistill: Robust Large Audio Language Models via Noisy-to-Clean Self-Distillation".

Jane: Audio Large Language Models (ALLMs) are highly vulnerable to real-world noise, which often induces severe semantic drift and hallucinations.

Tom: First, who's behind it and why it matters.

Paper summary: Tom: Exactly, so EchoDistill proposes using an information gap to do self-distillation, meaning the student takes noisy audio while the teacher model itself gets access to the corresponding clean audio. This lets them align candidate responses with clean semantic evidence even under corrupted conditions.

Jane: It sounds like they are trying to address a problem where existing robustness methods mostly focus on waveform enhancement or answer-level supervision, which they feel doesn't intrinsically improve robustness through the training process itself.

Lu: The paper sets up noisy student rollouts to expose the model's test-time behavior under noisy distributions, and then they convert clean-audio preferences into token-level supervision using a distillation loss called L(i)distill.

Meng: So, instead of just giving the model a reward based on whether its final answer is right or wrong in a noisy setting, they are looking at the consistency between what the student says and what the clean teacher would say at the token level.

Lalam: That token-level consistency seems like it gives us a much finer grain to tune how robust these models are, moving beyond just overall accuracy scores.

Conclusion: Tom: So, looking at the title, "EchoDistill: Robust Large Audio Language Models via Noisy-to-Clean Self-Distillation," it really summarizes how they are using that alignment trick to build audio language models that can handle real-world noise better.

Jane: I think for our listeners, the big picture is that this framework shows a way to train AI systems so they develop a stronger grasp of what the actual meaning is, rather than just memorizing patterns in perfect audio.

Lu: The implication here is that we can achieve better semantic grounding for audio tasks because the distillation process actively suppresses semantic drift during training.

Meng: From an engineering standpoint, it suggests that we don't necessarily need complex external noise removal modules to get better results; instead, this internal alignment mechanism makes the model inherently more resilient to artifacts.

Lalam: The impact on culture is interesting because if these models are more robust, we could deploy them in environments where audio quality is unpredictable, like real-time voice assistants in noisy cafes or field recordings.

Tom: It sounds like EchoDistill isn't just about fixing a bug; it's about fundamentally reshaping the training dynamics of audio AI to prioritize clean semantic consistency over raw acoustic fidelity under duress.

NTU 2SHU 3 ICT CAS 4 HDU 5 BUPT 6 USTC SKL-NST BUPT

cs.CL, cs.AI, cs.SD

Submitted: 2026-05-11

Updated: 2026-10-05

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 92/100

The gist: Audio Large Language Models (ALLMs) are highly vulnerable to real-world noise, which often induces severe semantic drift and hallucinations.

Key concepts

Noisy Student Rollouts
The model generates multiple candidate responses from noisy audio, each getting a reward based on matching a target answer. This forces the student model to learn how to generate plausible answers even when the input is corrupted by noise, exposing errors that are linguistically correct but acoustically unsupported.
Noisy-to-Clean Evidence Alignment
This technique converts clean audio preferences into token-level supervision. It calculates a distillation loss between the next-token distributions of the clean teacher and the noisy student on the same text, ensuring the noisy input still follows the semantic rules established by clean audio.
Audio-Aware Reward Shaping
The alignment signal is used to modify rewards for candidates. This creates a shaped reward that rewards responses that show high similarity between what the noisy student generates and what the clean teacher would prefer, guiding the model toward cleaner semantic matches.

Terminology

Summary

Audio Large Language Models (ALLMs) are highly vulnerable to real-world noise, which often induces severe semantic drift and hallucinations. The gist: EchoDistill, an alignment-based noisy-to-clean self-distillation framework, leverages a frozen clean-audio teacher to provide semantic references for an inference-time noisy-audio student by aligning its candidate responses with clean semantic evidence under corrupted conditions.

The Problem and Motivation

Real-world audio is often corrupted by device artifacts and environmental noise, which not only distorts the waveform but also degrades ALLMs generation quality. Under severe noise, models may misinterpret acoustic evidence and produce unstable responses, leading to semantic drift and hallucinations. Existing research on ALLM denoising has predominantly focused on inference-stage interventions or feature-level modifications, which fundamentally fall short of intrinsically enhancing the model’s robustness through the training process itself. This study seeks to propose a novel training paradigm: EchoDistill, which leverages an information gap to conduct self-distillation—where the student model takes noisy audio as input and the teacher model (the model itself) is granted access to the corresponding clean audio.

Key Components of EchoDistill

EchoDistill combines several mechanisms to achieve robust generation:

  1. Noisy student rollouts: The inference-time noisy student samples a group of candidate responses from the noisy audio input, where each candidate receives a task reward based on target-answer or choice matching. This step exposes optimization to the student’s actual noisy-input distribution, including linguistically plausible but acoustically unsupported errors.

  2. Noisy-to-clean evidence alignment: To address the coarseness of sequence-level rewards, EchoDistill converts clean-audio preferences into token-level supervision. It computes the next-token distributions of the clean teacher and noisy student on the same continuation, and then calculates a distillation loss, denoted as L(i)distill, which is a masked teacher-to-student KL between these distributions over response tokens. This encourages the noisy-input distribution to preserve the clean-audio semantic preference while generating the answer.

  3. Audio-aware reward shaping: To differentiate between correct and grounded responses, the alignment signal is used to guide reward shaping. The method converts noisy-to-clean discrepancy into an audio-aware similarity score and applies it only to task-positive candidates, yielding a shaped reward: "r¯(k)i = clip r(k)i + β 1[r(k)i > 0] si, -1, 2," where larger scores indicate closer agreement between the noisy student and the clean teacher.

Optimization Strategy

The framework optimizes the sampled candidates using a modified policy optimization approach. Instead of relying on a generic old-policy reference, EchoDistill uses the clean teacher as a detached reference scorer for the same sampled response to compute group-relative advantages: ρ(k)i = expl(k)θ − sg[l(k)ϕ]. The final training loss is a combination of policy and distillation losses: LEchoDistill = 1/N Σ λpolicyL(i)policy + λdistillL(i)distilli. This objective favors responses that both receive task reward under noisy input and remain plausible under cleanaudio evidence.

Experimental Results and Contributions

Extensive experiments on Qwen-Omni, MiniCPM-o, and StepAudio demonstrate significant improvements. EchoDistill achieves an average improvement of 4.18% in GSR compared to the strongest baseline. Ablation studies confirm that noisy-to-clean distillation provides the core semantic anchor for robust generation, and policy optimization brings complementary gains, showing that distillation aligns the noisy student with clean semantic evidence, while policy optimization further encourages task-correct and reward-aligned reasoning trajectories. The method shows particularly remarkable gains in the Sound and Speech domains. Furthermore, analysis of training dynamics reveals that noisy-to-clean alignment progressively suppresses semantic drift, as both StepAudio and MiniCPM show increasing noisy-to-clean consistency throughout training, demonstrating that EchoDistill reshapes the student’s generation behavior toward clean-audio semantic consistency rather than merely removing waveform noise. The framework is noted for being independent of specific acoustic enhancement modules, making it compatible with external denoising methods.

Limitations

Despite its effectiveness, EchoDistill still depends on the reliability of the clean-audio teacher and requires extra training-time computation for noisy rollouts, teacher scoring, and distribution-level alignment. Moreover, the current framework mainly focuses on single-audio understanding, suggesting future work may extend noisy-to-clean alignment to other modalities.

Improvements for AI systems

Based on the analysis of the EchoDistill: Alignment Noisy-to-Clean Self-Distillation for Robust Audio LLMs paper, here are specific improvements for AI systems and what those improved systems can achieve:


The core improvement lies in shifting from noise suppression (which often only fixes waveform fidelity) to robust semantic alignment during the reasoning phase. The improved system can be described as an Alignment-Aware Reasoning Engine.

Here are the specific improvements and capabilities:

  1. A. Alignment-Based Robustness Training via Noisy-to-Clean Self-Distillation:

  2. B. Audio-Aware Reward Shaping for Grounded Reasoning Trajectories:

  3. C. Synergistic Optimization of Policy and Distillation Losses (EchoDistill Framework):

Specific Improvements and Capabilities:

  1. The system will incorporate a frozen, high-fidelity Clean-Audio Teacher model that provides privileged semantic references during training. The student model operates under severe acoustic noise at inference time.

  2. This student is trained to sample candidate responses under noisy conditions and optimize its trajectories using Group-Relative Policy Optimization (GRPO), where the reward is augmented by a token-level consistency bonus with the teacher's clean semantic evidence.

  3. The system will specifically leverage Noisy-to-Clean KL Divergence loss, masking only non-prompt tokens, to force the noisy student’s output distribution to preserve the semantic preferences learned from clean audio, rather than just minimizing waveform error.

  4. The reward function will be dynamically shaped using an Audio-Aware Similarity Score derived from the distillation loss. This score acts as a bonus for candidates that are both task-correct and highly consistent with what the clean teacher would produce, effectively penalizing responses that are linguistically plausible but acoustically unsupported by the ground truth semantics.

Specific Capabilities of the Improved AI System:

The improved system (EchoDistill) can perform the following tasks robustly under severe acoustic corruption:

  1. Answering complex questions based on audio inputs (e.g., What is producing the sound in this audio?) with significantly higher reliability than current models, even when SNR is below-10 dB.

  2. Maintaining high generation stability across diverse domains (Music, Sound, Speech), showing particularly strong performance in the challenging Speech domain where fine-grained phonetic cues are often masked by noise.

  3. Reducing semantic drift and hallucinations by ensuring that the model's reasoning trajectories remain firmly anchored in reliable acoustic evidence rather than drifting toward spurious language priors when the input is corrupted.

  4. Achieving superior Generation Success Rate (GSR) compared to traditional denoising methods (like STFT or DFL), demonstrating a 4%+ improvement in robustness, meaning it can maintain successful generation even when the input audio is severely degraded.

Abstract

Large Audio Language Models (LALMs) remain vulnerable to acoustic noise, which can obscure task-relevant evidence and produce unreliable responses. We propose EchoDistill, a noisy-to-clean self-distillation framework that uses clean audio as privileged information during post-training. A noisy-input student samples candidate responses reflecting its inference-time behavior, while a frozen copy of the same backbone processes the corresponding clean audio. EchoDistill combines masked response-token distillation, task-gated consistency shaping, and teacher-referenced group-relative optimization to align noisy-input generation with clean-conditioned semantics. Only the student is retained at inference time, introducing no additional inference cost. Across three LALM backbones and three audio domains at-10dB, EchoDistill improves average noisy-input accuracy by 1.63 percentage points over the strongest baseline. On Qwen2.5-Omni, it raises noisy-input accuracy from 59.33% to 62.94%, while clean-audio accuracy increases from 76.56% to 77.56%. Replacing matched audio with random, shuffled, or silent inputs reduces accuracy by 3.08-6.42 points, confirming that matched acoustic evidence contributes to its predictions. Additional evaluations show improvements on held-out additive noises and external benchmarks, while revealing that these gains do not reliably extend to non-additive distortions. These results demonstrate robust post-training improvements under severe additive noise without sacrificing clean-audio capability across diverse tasks.

Sources

Related papers