Closing the Speech-Text Gap with Limited Audio for Effective Domain Adaptation in LLM-Based ASR
cs.CL
Submitted: 2026-04-07
Updated: 2026-10-01
Comments: Accepted at Interspeech
License: http://creativecommons.org/licenses/by/4.0/
The gist: Conventional end-to-end automatic speech recognition (ASR) systems rely on paired speech-text data for domain adaptation.
Terminology
Abstract
Conventional end-to-end automatic speech recognition (ASR) systems rely on paired speech-text data for domain adaptation. Recent LLM-based ASR architectures connect a speech encoder to a large language model via a projection module, enabling adaptation with text-only data. However, this introduces a modality gap, as the LLM is not exposed to the noisy representations produced by the speech projector. We investigate whether small amounts of speech can mitigate this mismatch. We compare three strategies: text-only adaptation, paired speech-text adaptation, and mixed batching (MB), which combines both. Experiments in in-domain and out-of-domain settings show that even limited speech consistently improves performance. Notably, MB using only 10% of the target-domain (less than 4 hours) speech achieves word error rates comparable to, or better than, conventional ASR fine-tuning with the full dataset, indicating that small amounts of speech provide a strong modality-alignment signal.
Sources
- An Embarrassingly Simple Approach for LLM with Strong ASR Capacity
- Bridging the Modality Gap: Softly Discretizing Audio Representation for LLM-based Automatic Speech Recognition
- Cross-Language Transfer Learning, Continuous Learning, and Domain Adaptation for End-to-End Automatic Speech Recognition
- Low-Resource Domain Adaptation for Speech LLMs via Text-Only Fine-Tuning
- Text-only adaptation in LLM-based ASR through text denoising
- The Llama 3 Herd of Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering