Interleaved Speech Language Models Latently Work In Text
cs.CL, cs.LG, cs.SD, eess.AS
Submitted: 2026-06-21
Updated: 2026-09-06
Comments: Preprint. 23 pages, 20 figures, 5 tables
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Beyond Transcription: Mechanistic Interpretability in ASR
- On The Landscape of Spoken Language Models: A Comprehensive Survey
- The Cascade Equivalence Hypothesis: When Do Speech LLMs Behave Like ASR$\rightarrow$LLM Pipelines?
- Do Audio LLMs Really LISTEN, or Just Transcribe? Measuring Lexical vs. Acoustic Emotion Cues Reliance
- VoiceBench: Benchmarking LLM-Based Voice Assistants
- Anatomy of the Modality Gap: Dissecting the Internal States of End-to-End Speech LLMs
- HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units
- Moshi: a speech-text foundation model for real-time dialogue
- The Llama 3 Herd of Models
- The People's Speech: A Large-Scale Diverse English Speech Recognition Dataset for Commercial Usage
- Looking Beyond The Top-1: Transformers Determine Top Tokens In Order
- Improving Multilingual Language Models by Aligning Representations through Steering
- Scaling Open Discrete Audio Foundation Models with Interleaved Semantic, Acoustic, and Text Tokens
- Qwen2.5 Technical Report
- Scaling Language Models: Methods, Analysis & Insights from Training Gopher
- ALAS: An Automatic Latent Alignment Score for Audio Language Models
- VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation
- The Semantic Hub Hypothesis: Language Models Share Semantic Representations Across Languages and Modalities
- Understanding the Modality Gap: An Empirical Study on the Speech-Text Alignment Mechanism of Large Speech Language Models
- Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering