[b] = [d] - [t] + [p]: Self-supervised Speech Models Discover Phonological Vector Arithmetic
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "
b: =
d: -
t: +
p: ".
Tom: Self-supervised speech models (S3Ms) are shown to encode speech using phonologically interpretable and compositional vectors, demonstrating phonological vector arithmetic.
Jane: First, who's behind it and why it matters.
Paper summary: Tom: So we just covered how these models learn to use phonological vectors in a compositional way, and now we need to make sure everyone has a solid grasp of the core thesis presented in "b = d - t + p: Self-supervised Speech Models Discover Phonological Vector Arithmetic."
Jane: Right, Tom; so basically, the main point is that self-supervised speech models are not just learning random acoustic patterns but are actually encoding speech using vectors that have a structure based on phonology.
Lu: The central claim is that there exist linear directions within the model’s representation space that correspond directly to phonological features, which allows for the discovery of compositional vectors and demonstrates this phonological vector arithmetic.
Meng: This means we can map abstract linguistic concepts onto concrete, manipulable directions in the model's internal structure, which moves beyond just correlation to actual structural understanding.
Lalam: It’s significant because it suggests that the AI is developing an internal representation that mirrors human linguistic organization, which is a really deep step for how we conceptualize language processing in artificial intelligence.
Tom: That’s right, Lalam; it means we are seeing the structure of speech as something mathematically structured and predictable rather than just a complex jumble of numbers.
Jane: And what matters is that this isn't just about finding one feature; they tested nineteen different PanPhon features and found that analogies based on all of them consistently hold in the S3M representations across multiple datasets <ref:2602.18899#pg1,that analogies based on all>.
Lu: That consistency across so many features confirms that the discovered linear relationships are not coincidental or specific to a single sound but represent a general rule for how S3Ms structure speech representations.
Meng: From an engineering viewpoint, if this structural understanding is general, it gives us confidence in using these models as robust tools for language tasks because we know their underlying structure aligns with established linguistic knowledge.
Lalam: For the future of AI culture, this means we can start to build systems where the internal logic is inherently more linguistically aware, allowing for richer and more nuanced interactions that feel genuinely communicative rather than just robotic.
Tom: Exactly; it sets a foundation for building AI that can truly grasp the underlying grammar and structure of language through these compositional vector relationships. So, what's our next step in understanding this concept?
Jane: Our next step is looking at how these vectors function in practice, specifically examining the scale aspect of these vectors to see how they relate to continuous acoustic reality.
Conclusion: Tom: Alright team, we’ve looked at the mechanics of these compositional vectors and their structure, and now it’s time to summarize the broader meaning of "b = d - t + p: Self-supervised Speech Models Discover Phonological Vector Arithmetic."
Jane: We’re talking about how authors Kwanghee Choi, Eunjung Yeo, Cheol Jun Cho, David Harwath, and David R. Mortensen found that these self-supervised speech models discover phonological vector arithmetic.
Lu: The implication is that these models can be seen as sophisticated systems with an internal logic that is directly tied to linguistic principles; they aren't just statistical black boxes but are structured according to speech science.
Meng: This suggests a future where we could design AI architectures where the representation space itself is guided by known acoustic and phonological rules, rather than relying solely on massive data-driven statistical learning for every single detail.
Lalam: For culture, this points toward a future where language understanding in AI isn't just about pattern matching but about building systems that have an internalized, structured knowledge of how language is organized at a fundamental level.
Tom: It’s about moving from mere statistical modeling to building representations that are inherently meaningful to the structure of human language itself. We’re seeing a new way to understand what speech really is in computational terms.
Jane: So, in simple terms, this paper shows that these models discover mathematical relationships between phonological features within their own internal structure, which gives us a blueprint for making AI representations that are more linguistically sound and controllable.
Kwanghee Choi, Eunjung Yeo, Cheol Jun Cho, David Harwath, David R. Mortensen
UT Austin · UC Berkeley · CMU
eess.AS, cs.CL, cs.LG, cs.SD
Submitted: 2026-02-21
Updated: 2026-10-06
Code: https://github.com/juice500ml/phonetic-arithmetic
Importance score: 73/100
The gist: Self-supervised speech models (S3Ms) are shown to encode speech using phonologically interpretable and compositional vectors, demonstrating phonological vector arithmetic.
Key concepts
- Phonological Vector Arithmetic
- This refers to the ability of S3M representations to follow mathematical rules based on phonology. Researchers tested whether relationships between different speech sounds (like [b], [p], [d], and [t]) could be expressed through simple vector addition or subtraction, proving the model understands how phonemes relate to each other in a structured way.
- Compositional Vectors
- These are vectors within the S3M representation space that correspond directly to specific phonological features, like voicing or place of articulation. Instead of just representing a sound as a single blob, these vectors allow researchers to isolate and manipulate specific acoustic characteristics by operating on them.
- Scaling ($\lambda$)
- The study introduced a scalar value called $\lambda$ to control the continuous scale of the phonological vectors. This scaling allows for more nuanced control over acoustic properties like the degree of voicing or burst characteristics, demonstrating that S3M representations encode features as a continuum rather than just binary distinctions.
Terminology
Summary
Self-supervised speech models (S3Ms) are shown to encode speech using phonologically interpretable and compositional vectors, demonstrating phonological vector arithmetic. This research reveals that S3Ms learn linear directions within their representation space that correspond to phonological features, allowing for the discovery of compositional vectors and the ability to control speech synthesis along these dimensions.
How it works
The study investigates whether S3Ms represent phonology in an analogous compositional manner by testing for linear phonological analogies. The authors first establish that there exist two symmetric phonological analogies based on a phone quadruplet, such as [b], [p], [d], and [t]: [b]: [p] = [d]: [t] (voicing)
and [b]: [d] = [p]: [t] (POA).
These analogies yield compositional phonological vectors, specifically a voicing vector, defined as vvoi. = r[d − r[t]] in eq. (4),
and a change of POA vector, vPOA = r[p − r[t]] in eq. (5).
The hypothesis is tested by finding that analogies based on all 19 PanPhon phonological features consistently hold in S3M representations across TIMIT and VoxAngeles datasets.
Direction of Phonological Vectors
The core finding regarding the direction of these vectors is that S3M representations exhibit phonological vector arithmetic, i.e., existence of compositional vectors that align with phonological features.
This is demonstrated by analyzing the consistency of analogies based on all 19 PanPhon features. The success rate, defined as the proportion of quadruplets whose similarity scores satisfy the ordering in eq. (13) such that phonological analogies hold,
is calculated to evaluate this directionality. Empirical results show that S3Ms consistently outperform spectral representations (MFCC and MelSpec) in preserving these linear relationships, with WavLM achieving success rates up to 93% on TIMIT. Furthermore, the analysis of layerwise behavior suggests that deeper layers are more likely to leverage broader contextual information when forming abstract phonological vectors.
Scale of Phonological Vectors
The second major contribution is demonstrating that the scale of these phonological vectors corresponds to acoustic measurements in a continuous manner. The authors introduce a scalar λ into the analogy equation: r[b] ≃ r[p] + λ · (r[d] − r[t]).
They hypothesize that this scale controls acoustic characteristics continuously, such as voicing degree. To validate this, they train a vocoder to approximate the inverse of the S3M function and modify the representation using scaled vectors: "R˜ t = (Rt + λv (t′s ≤ t < t′e) Rt (otherwise), followed by resynthesis. Acoustic measurements from these modified representations show that
acoustic measurements strongly correlate with the scale λ, for both interpolation (λ ≤ 1) and extrapolation (λ > 1) of phonological vectors. This confirms that S3M representations encode features
not as purely binary distinctions but as a continuum through specific vector directions and scales."
Acoustic Realization and Generalization
The study further explores the acoustic consequences of scaling these vectors. For instance, applying the voicing vector to [b] shows that increasing λ for positive values decreases the voice onset time (VOT),
while for strident features, increasing λ removes the burst characteristics of plosives.
Qualitative analysis confirms these effects: increasing stridency introduces frication above 4kHz and removes burst characteristics. The results indicate that S3M representations encode not only static spectral envelopes, but also internal temporal structure,
as the vectors modulate both spectral envelopes and temporal cues. Moreover, the findings demonstrate generalization to unseen phones; VoxAngeles analogies contain a significant proportion of phones not present in the English TIMIT set, and WavLM achieves high success rates on these cross-linguistic samples.
Vector Comparison and Limitations
The analysis of pairwise cosine similarities between different phonological vectors confirms meaningful relationships: vowel-related and consonant-related phonological vectors exhibit near-orthogonal similarities,
while within groups, such as consonants, nasal, sonorant, and voice vectors exhibit positive similarity.
However, the study acknowledges limitations. It notes that the results are dependent on the specific feature system assumed by PanPhon and that different S3Ms behave differently. Additionally, it points out that synthesis results are influenced by both the S3Ms and the vocoder used for resynthesis, suggesting some observed behaviors may be vocoder-specific characteristics rather than properties of the S3Ms alone.
The paper concludes that these findings provide intuitive interpretations of S3M representations and fine-grained control of speech synthesis along phonological dimensions.
The gist: Self-supervised speech models learn linearly composable and scalable phonological vectors.
Improvements for AI systems
Here are specific, actionable improvements for AI systems derived from this research:
)Phonologically-Grounded Speech Synthesis and Control Systems:
-
Acoustic Parameter Manipulation via Vector Arithmetic: Implement a
Phonological Steering Layer
in Text-to-Speech (TTS) or End-to-End Speech Models (like WavLM). This layer would allow users to manipulate speech synthesis by adding scaled phonological vectors (e.g., voicing vector, rounding vector) to the intermediate representation matrix before the vocoder decoder. -
Fine-Grained Acoustic Realization Control: Instead of binary control over acoustic features, the system can achieve continuous control over acoustic realizations. For example, instead of simply toggling
voiced
orunvoiced,
a user could input a scalar value for the voicing vector to smoothly transition from one voicing state to another, allowing for nuanced prosodic and phonetic variation (e.g., subtle changes in Voice Onset Time (VOT) or F1/F2 frequencies). -
Cross-Lingual Phonological Transfer: Develop a
Phonological Vector Translator
module. Because the research shows that S3M representations learn universal, linearly composable phonological vectors, this system could allow an AI model trained on English speech to apply phonologically grounded modifications to speech from a low-resource language (like those in VoxAngeles) by leveraging the shared vector space structure, even if those specific phones weren't explicitly seen during training. -
Robustness and Generalization: Integrate the learned phonological vectors as an explicit regularization term during training. This ensures that the model's representations are not only phonetically accurate but also lie within a
phonologically interpretable
subspace, leading to better generalization to unseen languages and novel acoustic conditions (extrapolation beyond training data). -
Interpretability for Auditory Design: Create an AI interface where designers can visualize how specific phonological features (like nasalization or stridency) manifest in the latent space of the S3M. This moves beyond simple spectral analysis to show exactly which vector directions correspond to specific articulatory/acoustic properties, enabling more intentional and linguistically informed control over synthesized output.
Sources
- Opening the Black Box of wav2vec Feature Encoder
- Toy Models of Superposition
- Efficient Estimation of Word Representations in Vector Space
- The Origins of Representation Manifolds in Large Language Models
- Steering Language Models With Activation Engineering
Related papers
- X-VC: Zero-shot Streaming Voice Conversion in Codec Space
- Autoregressive Guidance of Deep Spatially Selective Filters using Bayesian Tracking for Efficient Extraction of Moving Speakers
- Anonymization, Not Elimination: Utility-Preserved Speech Anonymization
- Towards Audio Token Compression in Large Audio Language Models
- WaveScat: Wavelet Scattering Front-Ends with Self-Supervised Features for Speech Deepfake Detection
- ProPS: Prompted Profile Synthesis for Natural Language-Conditioned Speaker Embedding Distributions