SPAR-K: Scheduled Periodic Alternating Early Exit for Spoken Language Models
cs.CL, eess.AS
Submitted: 2026-03-10
Updated: 2026-08-27
Comments: 8 pages, 1 figure, 2 tables
License: http://creativecommons.org/licenses/by-sa/4.0/
The gist: Interleaved spoken language models (SLMs) alternately generate text and speech tokens, but decoding at full transformer depth for every step becomes costly, especially due to long speech sequences.
Terminology
Abstract
Interleaved spoken language models (SLMs) alternately generate text and speech tokens, but decoding at full transformer depth for every step becomes costly, especially due to long speech sequences. We propose SPAR-K, a modality-aware early exit framework designed to accelerate interleaved SLM inference while preserving perceptual quality. SPAR-K introduces a speech alternating-depth schedule: most speech positions exit at a fixed intermediate layer, while periodic full-depth "refresh" steps mitigate distribution shift due to early exit. We evaluate our framework using Step-Audio-2-mini and GLM-4-Voice across four datasets spanning reasoning, factual QA, and dialogue tasks, measuring performance in terms of ASR transcription accuracy and perceptual quality. Experimental results demonstrate that SPAR-K largely preserves question-answering accuracy with a maximum accuracy drop of 0.82% while reducing average speech decoding depth by up to 11% on Step-Audio-2-mini and 5% on GLM-4-Voice, both with negligible changes in MOS and WER and no auxiliary computation overhead. We further demonstrate that confidence-based early exit strategies, widely used in text LLMs, are suboptimal for SLMs, highlighting that the unique statistical nature of speech tokens necessitates a specialized early exit design.
Sources
- CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens
- On The Landscape of Spoken Language Models: A Comprehensive Survey
- Scaling Speech-Text Pre-training with Synthetic Interleaved Data
- Moshi: a speech-text foundation model for real-time dialogue
- Kimi-Audio Technical Report
- Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis
- GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
- Accelerating Large Language Model Inference with Self-Supervised Early Exits
- SimLens for Early Exit in Large Language Models: Eliciting Accurate Latent Predictions with One More Token
- An Efficient Inference Framework for Early-exit Large Language Models
- A Survey of Early Exit Deep Neural Networks in NLP
- EE-LLM: Large-Scale Training and Inference of Early-Exit Large Language Models with 3D Parallelism
- Depth-Adaptive Transformer
- Accelerating Large Language Model Decoding with Speculative Sampling
- Step-Audio 2 Technical Report
- Eliciting Latent Predictions from Transformers with the Tuned Lens
- Baichuan-Audio: A Unified Framework for End-to-End Speech Interaction
- Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming
- STITCH: Simultaneous Thinking and Talking with Chunked Reasoning for Spoken Language Models
- CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering