Reducing the Output-Mode Gap in Speech Language Models via Joint-Output On-Policy Distillation
eess.AS, cs.CL, cs.SD
Submitted: 2026-09-14
Updated: 2026-09-14
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models
- SALMONN: Towards Generic Hearing Abilities for Large Language Models
- SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities
- Moshi: a speech-text foundation model for real-time dialogue
- GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
- Qwen3-Omni Technical Report
- Step-Audio 2 Technical Report
- Baichuan-Audio: A Unified Framework for End-to-End Speech Interaction
- Understanding the Modality Gap: An Empirical Study on the Speech-Text Alignment Mechanism of Large Speech Language Models
- Closing the Gap Between Text and Speech Understanding in LLMs
- X-OPD: Cross-Modal On-Policy Distillation for Capability Alignment in Speech LLMs
- CORD: Bridging the Audio-Text Reasoning Gap via Weighted On-policy Cross-modal Distillation
- Closing the Modality Reasoning Gap for Speech Large Language Models
- Minimizing Modality Gap from the Input Side: Your Speech LLM Can Be a Prosody-Aware Text LLM
- Spirit LM: Interleaved Spoken and Written Language Model
- MiMo-Audio: Audio Language Models are Few-Shot Learners
- LongCat-Next: Lexicalizing Modalities as Discrete Tokens
- Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation
- X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment
- Tulu 3: Pushing Frontiers in Open Language Model Post-Training
Related papers
- X-VC: Zero-shot Streaming Voice Conversion in Codec Space
- Autoregressive Guidance of Deep Spatially Selective Filters using Bayesian Tracking for Efficient Extraction of Moving Speakers
- Anonymization, Not Elimination: Utility-Preserved Speech Anonymization
- Towards Audio Token Compression in Large Audio Language Models
- WaveScat: Wavelet Scattering Front-Ends with Self-Supervised Features for Speech Deepfake Detection
- ProPS: Prompted Profile Synthesis for Natural Language-Conditioned Speaker Embedding Distributions