Scaling Audio Models Efficiently: Joint Optimization of Scale, Resolution, Adaptation, Precision, and Sparsity
cs.SD, cs.AI
Submitted: 2026-06-22
Updated: 2026-09-19
Code: https://github.com/vyomya/SAME
License: http://creativecommons.org/licenses/by/4.0/
The gist: Large automatic speech recognition (ASR) models such as Whisper must be deployed across hardware with widely varying memory and inference-speed constraints.
Terminology
Abstract
Large automatic speech recognition (ASR) models such as Whisper must be deployed across hardware with widely varying memory and inference-speed constraints. We present a compression framework that jointly parametrizes Whisper deployment along six dimensions: model size x N, temporal resolution x T, encoder token stride x V, low-rank adaptation capacity x R, weight precision x Q and sparsity pattern x P. All axes are jointly optimized against three deployment objectives (word error rate, inference FLOPs, and memory footprint) using a non-dominated sorting genetic evolutionary search (NSGA). Across 50 of the 1,680 candidate configurations evaluated, we measure the marginal effect of each axis on the three objectives and identify compression combinations that dominate naive single-axis scaling, and report a consistent negative result: 1:4 structured sparsity fails to recover acceptable accuracy under any tested recovery budget. We report real, measured memory and accuracy figures for genuinely quantized deployment artifacts, and provide a lookup table mapping deployment scenarios (cloud, server, edge, ultra-constrained) to specific axis configurations with their measured accuracy/memory/compute trade-offs
Related papers
- Few-Shot Open-Set Audio Classification via Transductive Prototype Refinement and Class Logit Enhancement
- Spectral Masking and Interpolation Attack (SMIA): A Black-box Adversarial Attack against Voice Authentication and Anti-Spoofing Systems
- AVMeme Exam: A Multimodal Multilingual Multicultural Benchmark for LLMs' Contextual and Cultural Knowledge and Thinking
- SoundWeaver: Compositional Warm-Starting for Text-to-Audio Diffusion Serving
- WASIL: In-the-Wild Arabic Spoken Interactions with LLMs
- Efficient Audiovisual Speech Processing via MUTUD: Multimodal Training and Unimodal Deployment