Stride-k Subsampling: Train-Free Audio Token Reduction for Whisper
cs.SD, cs.AI
Submitted: 2026-08-31
Updated: 2026-08-31
Comments: Accepted EMNLP 2026 Main
License: http://creativecommons.org/licenses/by/4.0/
The gist: Whisper exposes speech through a fixed 1500-token encoder interface, now a default representation for ASR decoders and Whisper-based speech language models (SpeechLMs), yet its redundancy remains
Terminology
Abstract
Whisper exposes speech through a fixed 1500-token encoder interface, now a default representation for ASR decoders and Whisper-based speech language models (SpeechLMs), yet its redundancy remains largely unexamined. We propose stride-k subsampling, a deterministic indexing operation that retains every k-th token after the convolutional stem or encoder transformer. Across five Whisper scales, k=2 preserves baseline WER at both positions, with CKA attributing this stability to acoustic overlap at the stem and attention-induced redistribution at the encoder output. Applying stride-2 at both positions cuts audio tokens by 75% and total GFLOPs by 52-58%, with small WER costs on most ASR benchmarks and larger costs on harder ones. The same configuration extends to three Whisper-based SpeechLMs, yielding modest accuracy drops on stronger baselines and larger drops on weaker ones, while reducing end-to-end latency by 19.6-27.4%. Requiring no training or auxiliary computation, stride-k subsampling exploits Whisper's preprocessing redundancy, indicating that its audio-token interface carries more capacity than downstream tasks require.
Sources
- Towards Audio Token Compression in Large Audio Language Models
- Qwen2-Audio Technical Report
- Distil-Whisper: Robust Knowledge Distillation via Large-Scale Pseudo Labelling
- Pruning as Regularization: Sensitivity-Aware One-Shot Pruning in ASR
- FastAV: Efficient Token Pruning for Audio-Visual Large Language Model Inference
- Whisfusion: Parallel ASR Decoding with Masked Diffusion
- Token Pruning in Audio Transformers: Optimizing Performance and Decoding Patch Importance
- Locality Matters for Training-Free Audio Token Compression in Audio-Language Models
- WhisperKit: On-device Real-time ASR with Billion-Scale Transformers
- Do We Need Distinct Representations for Every Speech Token? Unveiling and Exploiting Redundancy in Large Speech Language Models
- Early Attentive Sparsification Accelerates Neural Speech Transcription
- LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech
Related papers
- Few-Shot Open-Set Audio Classification via Transductive Prototype Refinement and Class Logit Enhancement
- Spectral Masking and Interpolation Attack (SMIA): A Black-box Adversarial Attack against Voice Authentication and Anti-Spoofing Systems
- AVMeme Exam: A Multimodal Multilingual Multicultural Benchmark for LLMs' Contextual and Cultural Knowledge and Thinking
- SoundWeaver: Compositional Warm-Starting for Text-to-Audio Diffusion Serving
- WASIL: In-the-Wild Arabic Spoken Interactions with LLMs
- Efficient Audiovisual Speech Processing via MUTUD: Multimodal Training and Unimodal Deployment