Encoder Awakening via Adapters: Effective Domain-Adaptive Fine-tuning of Speech-LLMs
eess.AS, cs.CL
Submitted: 2026-09-16
Updated: 2026-09-16
Comments: Accepted to IEEE SLT 2026
Code: https://github.com/morganshi/EAVA
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- The Llama 3 Herd of Models
- Phi-4 Technical Report
- Qwen3 Technical Report
- An Embarrassingly Simple Approach for LLM with Strong ASR Capacity
- Gumbel-BEARD: Automatic Layer Selection for Self-Supervised Adaptation of Whisper in Low-Resource Domains
- Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs
- Searching for Activation Functions
- GC-LoRA: Gated Convolutional LoRA for Parameter-Efficient Acoustic Adaptation
Related papers
- X-VC: Zero-shot Streaming Voice Conversion in Codec Space
- Autoregressive Guidance of Deep Spatially Selective Filters using Bayesian Tracking for Efficient Extraction of Moving Speakers
- Anonymization, Not Elimination: Utility-Preserved Speech Anonymization
- Towards Audio Token Compression in Large Audio Language Models
- WaveScat: Wavelet Scattering Front-Ends with Self-Supervised Features for Speech Deepfake Detection
- ProPS: Prompted Profile Synthesis for Natural Language-Conditioned Speaker Embedding Distributions