Quantifying the Generation Modality Gap in Speech-Text Language Models
cs.CL
Submitted: 2026-09-13
Updated: 2026-10-06
Comments: Accepted to SLT 2026, extended version with appendix
Code: https://github.com/jjery2243542/flow-slm
License: http://creativecommons.org/licenses/by/4.0/
The gist: Pure speech language models often lag behind text and speech-text language models in generating coherent content, but this gap is difficult to quantify because speech and text systems are typically
Terminology
Abstract
Pure speech language models often lag behind text and speech-text language models in generating coherent content, but this gap is difficult to quantify because speech and text systems are typically evaluated with different metrics and trained on different data. We study the speech-text modality gap in a family of spoken language models, based on flow matching for continuous acoustic feature generation. We construct a unified generation-based evaluation suite that compares speech-only, text-only, and speech-text language models trained on matched data distributions and evaluated in matched generation settings. We evaluate generated continuations along multiple dimensions: semantic coherence, measured by transcribing generated speech and scoring it with a reference language model; local phonetic structure, measured by phone n-gram distributional statistics; speaker consistency and acoustic quality; and emotion-based distributional metrics. Across datasets, we find that joint speech-text modeling substantially improves semantic coherence. However, the improvement is not uniform across metrics: phone-level metrics change only modestly, speaker similarity and predicted quality are lower for speech-text continuations, while emotion-based distributional metrics improve. Compared with larger-scale speech-only models, our speech-text model closes much of the scaling gap in transcript-based semantic coherence, suggesting that text provides an efficient semantic training signal for spoken language modeling.
Sources
- Llama-Mimi: Exploring the Limits of Flattened Speech Language Modeling
- Moshi: a speech-text foundation model for real-time dialogue
- Scaling Properties of Continuous Diffusion Spoken Language Models
- Anatomy of the Modality Gap: Dissecting the Internal States of End-to-End Speech LLMs
- Olmo 3
- Scaling Laws for Neural Language Models
- The False Resonance: A Critical Examination of Emotion Embedding Similarity for Speech Generation Evaluation
- Ministral 3
- Gemma: Open Models Based on Gemini Research and Technology
- Qwen2.5 Technical Report
- The Llama 3 Herd of Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering