Scaling Audio Models Efficiently: Joint Optimization of Scale, Resolution, Adaptation, Precision, and Sparsity

arXiv:2606.22790 · cs.SD, cs.AI · Submitted 2026-06-22 · Read on arXiv

cs.SD, cs.AI

Submitted: 2026-06-22

Updated: 2026-09-19

Code: https://github.com/vyomya/SAME

License: http://creativecommons.org/licenses/by/4.0/

The gist: Large automatic speech recognition (ASR) models such as Whisper must be deployed across hardware with widely varying memory and inference-speed constraints.

Terminology

Abstract

Large automatic speech recognition (ASR) models such as Whisper must be deployed across hardware with widely varying memory and inference-speed constraints. We present a compression framework that jointly parametrizes Whisper deployment along six dimensions: model size x N, temporal resolution x T, encoder token stride x V, low-rank adaptation capacity x R, weight precision x Q and sparsity pattern x P. All axes are jointly optimized against three deployment objectives (word error rate, inference FLOPs, and memory footprint) using a non-dominated sorting genetic evolutionary search (NSGA). Across 50 of the 1,680 candidate configurations evaluated, we measure the marginal effect of each axis on the three objectives and identify compression combinations that dominate naive single-axis scaling, and report a consistent negative result: 1:4 structured sparsity fails to recover acceptable accuracy under any tested recovery budget. We report real, measured memory and accuracy figures for genuinely quantized deployment artifacts, and provide a lookup table mapping deployment scenarios (cloud, server, edge, ultra-constrained) to specific axis configurations with their measured accuracy/memory/compute trade-offs

Related papers