Rethinking Speech-LLM Integration for ASR: Effective Joint Speech-Text Training by Interleaving
cs.CL, eess.AS
Submitted: 2026-07-02
Updated: 2026-09-20
Comments: Accepted to SLT 2026, camera-ready version
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- GPT-4 Technical Report
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- DeepSeek-V3 Technical Report
- Gemma 3 Technical Report
- Qwen3 Technical Report
- Moshi: a speech-text foundation model for real-time dialogue
- Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs
- Kimi-Audio Technical Report
- Step-Audio 2 Technical Report
- Qwen3-Omni Technical Report
- SLM-S2ST: A multimodal language model for direct speech-to-speech translation
- Fun-Audio-Chat Technical Report
- FireRedASR: Open-Source Industrial-Grade Mandarin Speech Recognition Models from Encoder-Decoder to LLM Integration
- Seed-ASR: Understanding Diverse Speech and Contexts with LLM-based Speech Recognition
- Fun-ASR Technical Report
- Index-ASR Technical Report
- Efficient Scaling for LLM-based ASR
- Qwen3-ASR Technical Report
- Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities
- Train Short, Infer Long: Speech-LLM Enables Zero-Shot Streamable Joint ASR and Diarization on Long Audio
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering