On the Interpretability of Whisper Encodings Using Sparse Autoencoders
cs.CL
Submitted: 2026-05-12
Updated: 2026-10-01
Comments: Accepted to the IEEE Real-Time Communications Conference (RTC) 2026
Code: https://github.com/openai/sparse
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
The gist: While deep transformer-based models have advanced rapidly, their internal mechanisms remain largely a mystery.
Terminology
Abstract
While deep transformer-based models have advanced rapidly, their internal mechanisms remain largely a mystery. Recent work has prioritized understanding text-based transformer models, leaving ASR systems largely unexplored. In order to address this gap, we examine the internal representations of Whisper's encoder using a sparse autoencoder. We find diverse monosemantic features across linguistic and non-linguistic boundaries, spanning a hierarchy from phonetic to semantic representations, and conduct a causal feature-steering campaign across this hierarchy, including cross-lingual steering. We further find that steering is more reliable for higher-level features than lower-level ones, an asymmetry that may reflect redundant encoding of lower-level information. Altogether, this work demonstrates that Whisper's encoder represents a surprisingly rich hierarchy of linguistic information that extends well beyond what is strictly necessary for transcription.
Sources
- The Holistic Storage of Verb+Up Phrases in Text-based and Audio-based Language Models
- MUSAN: A Music, Speech, and Noise Corpus
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering