Performance and Complexity Trade-off Optimization of Speech Models During Training
cs.SD, cs.AI, cs.LG, eess.AS
Submitted: 2026-01-20
Updated: 2026-09-15
Comments: This work has been submitted to the IEEE for possible publication
Code: https://github.com/stefanwebb/open-voice-activity-detection
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: In speech machine learning, neural network models are typically designed by choosing an architecture with fixed layer sizes and structure.
Terminology
Abstract
In speech machine learning, neural network models are typically designed by choosing an architecture with fixed layer sizes and structure. These models are then trained to maximize performance on metrics aligned with the task's objective. While the overall architecture is usually guided by prior knowledge of the task, the sizes of individual layers are often chosen heuristically. However, this approach does not guarantee an optimal trade-off between performance and computational complexity; consequently, post hoc methods such as weight quantization or model pruning are typically employed to reduce computational cost. This occurs because stochastic gradient descent (SGD) methods can only optimize differentiable functions, while factors influencing computational complexity, such as layer sizes and floating-point operations per second (FLOP/s), are non-differentiable and require modifying the model structure during training. We propose a reparameterization technique based on feature noise injection that enables joint optimization of performance and computational complexity during training using SGD-based methods. Unlike traditional pruning methods, our approach allows the model size to be dynamically optimized for a target performance-complexity trade-off, without relying on heuristic criteria to select which weights or structures to remove. We demonstrate the effectiveness of our method through three case studies, including a synthetic example and two practical real-world applications: voice activity detection and audio anti-spoofing. The code related to our work is publicly available to encourage further research.
Sources
- Quantizing deep convolutional networks for efficient inference: A whitepaper
- Compressing Deep Neural Networks via Layer Fusion
- Distilling the Knowledge in a Neural Network
- Neural Architecture Search: Insights from 1000 Papers
- The State of Sparsity in Deep Neural Networks
- Distil-Whisper: Robust Knowledge Distillation via Large-Scale Pseudo Labelling
- A Survey on Speech Deepfake Detection
Related papers
- Few-Shot Open-Set Audio Classification via Transductive Prototype Refinement and Class Logit Enhancement
- Spectral Masking and Interpolation Attack (SMIA): A Black-box Adversarial Attack against Voice Authentication and Anti-Spoofing Systems
- AVMeme Exam: A Multimodal Multilingual Multicultural Benchmark for LLMs' Contextual and Cultural Knowledge and Thinking
- SoundWeaver: Compositional Warm-Starting for Text-to-Audio Diffusion Serving
- WASIL: In-the-Wild Arabic Spoken Interactions with LLMs
- Efficient Audiovisual Speech Processing via MUTUD: Multimodal Training and Unimodal Deployment