A mathematical model of the vowel space

arXiv:2111.00868 · cs.SD, eess.AS, physics.class-ph, q-bio.PE · Submitted 2021-10-19 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Genomics Radio. Generated commentary on the latest computational biology and genomics papers.

Ines: Today's paper: "A mathematical model of the vowel space".

Marcus: The gist The aim of this paper is to revisit the roots in order to construct a link with articulatory modelling and to understand which causality and structure these models…

Ines: First, who's behind it and why it matters.

Title and authors: Ines: So we're looking at this paper called "A mathematical model of the vowel space," and it tackles this huge problem in speech science: how we can actually link the movements you make in your mouth to the sounds you produce.

Marcus: Exactly. The core issue they point out is that there's this messy, many-to-one relationship between how much space your vocal tract takes up and the specific acoustic measurements like f1 and f2. It’s really hard to get a reliable way to go backward from sound to actual physical movements, no matter what kind of vocal tract you’re looking at <ref:2111.00868#pg4>.

Yuki: From a population genetics view, this difficulty is huge because it prevents us from tracing how speech evolved across different species or even different human populations. If we can't map the physical shape to the sound, we can't really connect movement to evolutionary pressures <ref:2111.00868#pg2>.

Ines: The paper proposes a way around that mess by setting up a bijection between the vowel space—which is just those two formant frequencies, f1 and f2—and the parametric space of different vocal tract models. They build a generic model using mixtures of cosines to generate eight vowels with just two variables and the length being static <ref:2111.00868#pg5>.

Marcus: That's an interesting simplification, but they move beyond that basic idea by constructing a generative function P theta based on something called the Gibbs triangle used in chemistry to mix components <ref:2111.00868#pg4>. This function lets them generate closed-open tube shapes for vowels like u, a, and i by mixing vectors i, j, and k at different angles theta <ref:2111.00868#pg3>.

Yuki: So they're essentially using this mixing function to define how different vocal tract shapes correspond to the acoustic space of vowels in a controlled way. I wonder how robust that is when you move from a simple model to real human speech <ref:2111.00868#pg4>.

Ines: They take that initial function and transform it into something more practical, they call it the coordination function, which expresses the coordination between the articulatory parameters along that theta cycle <ref:2111.00868#pg5>. This new formula, P theta = Omega + Psi one cos(Psi two - theta), is designed to handle those parameters directly <ref:2111.00868#pg6>.

Marcus: The authors show that this coordination function works similarly to other models they’ve looked at before, like Fant’s model and the four-tube DRM derived from their generic model <ref:2111.00868#pg6>. When they test it against real vocal tract data from reference seventeen, it generates a mapping for three extreme vowels including that length component <ref:2111.00868#pg6>.

Title and authors: Yuki: That's cool because it suggests this mathematical structure can actually handle the complexity of the real vocal tract shapes, which is where many of these models usually fall apart <ref:2111.00868#pg4>.

Ines: They compare this coordination function against two specific four-tube models, the generic model and the four-tube DRM, to show how they both manage to reduce that many-to-one problem <ref:2111.00868#pg6>. They do this by calculating P3 and P4 from P1 and P2 so that it explicitly cuts down the number of free parameters to just two <ref:2111.00868#pg6>.

Marcus: And they show that this coordination function only generates the internal antisymmetry, which is something both models already do on their own <ref:2111.00868#pg6>. For both the generic model and the reduced DRM, they find that the intrinsic dimension of these vowel spaces drops to two <ref:2111.00868#pg6>.

Yuki: So, despite using different parameter sets in condition C1 and C2, the resulting vowel spaces end up looking very similar <ref:2111.00868#pg6>. That consistency across different models is a big hint about the underlying structure of speech production <ref:2111.00868#pg4>.

Ines: The crucial observation they make is that for condition C1, you still see that many-to-one relationship, but for condition C2, that relationship disappears because of this coordination function <ref:2111.00868#pg6>. This suggests the coordination function is a very effective way to take the potential of a vocal tract and create a vowel space with a one-to-one correspondence <ref:2111.00868#pg4>.

Marcus: I think what they really highlight is that for non-optimal vocal tract forms, this coordination function generates the regular relation we already see in the generic model and the reduced DRM <ref:2111.00868#pg6>. It’s a mechanism that brings them together structurally <ref:2111.00868#pg4>.

Yuki: What this means for us is that we might be able to use this coordination function as a test—a criterion to judge whether certain anatomical ranges are actually ready for speech production, since they show a regular relation <ref:2111.00868#pg6>.

Ines: Exactly. They also point out that the anatomical ranges studied by reference ten are quite small, and this gives us a new objective criterion to judge if those shapes are truly speech ready <ref:2111.00868#pg9>.

Marcus: So, to wrap up on the paper "A mathematical model of the vowel space," the main point is that this coordination function is an optimal way to use a vocal tract's potential to create a vowel space and establish a one-to-one correspondence <ref:2111.00868#pg4>. It shows how it generates the regular relation seen in both the generic model and the reduced DRM <ref:2111.00868#pg6>.

Yuki: I think what this implies is that we can start thinking about causal links between anatomy and sound more clearly than before, even with these mathematical models <ref:2111.00868#pg4>. It opens up new avenues for connecting physical constraints to the patterns we see in animal vocalizations <ref:2111.00868#pg2>.

Title and authors: Ines: It definitely shifts the focus from just trying to map every single acoustic measurement to finding a structure that organizes those measurements into a manageable, one-to-one system <ref:2111.00868#pg4>. This is a step forward for understanding the deep causality here.

Marcus: For anyone working with speech synthesis or modeling, this suggests using these coordination functions might be a better starting point than trying to invert every single vocal tract measurement directly <ref:2111.00868#pg4>. It provides a way to exploit the articulatory potential systematically.

Yuki: And it makes the historical debate about speech models less about two separate approaches and more about finding which mathematical structures share the same underlying logic <ref:2111.00868#pg2>.

Ines: So, in summary, we’ve looked at how this paper uses a generative function to build a coordination function that successfully handles the many-to-one relationship by reducing the complexity down to two intrinsic dimensions <ref:2111.00868#pg6>. It’s an interesting structural result.

Marcus: And it shows that for certain conditions, like condition C2, this reduction in dimensionality is consistent across different models we've tested <ref:2111.00868#pg6>. That consistency is statistically meaningful for our analysis of the cohort data.

Yuki: It’s a nice piece of mathematical machinery that helps ground the biological intuition we have about vocal tract constraints <ref:2111.00868#pg4>. It connects the abstract geometry back to something tangible in biology.

Ines: So, moving forward, this paper suggests we should look for other generative functions like this one when trying to build models that bridge anatomy and sound <ref:2111.00868#pg3>. The focus is on finding that link between the physical movement and the resulting acoustic pattern.

Marcus: And it sets up a new way to test those links by looking for this coordination property in different vocal tract setups <ref:2111.00868#pg6>. It gives us a concrete function to use when we’re dealing with complex parameter sets.

Yuki: I hope these mathematical tools help us finally get a clearer picture of how speech systems are constrained by the physical reality of the vocal tract <ref:2111.00868#pg4>. That's a big goal for understanding language evolution.

Ines: We’ll keep watching how this concept evolves in other modeling efforts, because establishing that causal link between movement and sound is still the major hurdle in speech science <ref:2111.00868#pg4>.

Marcus: Right, let's see what other papers are coming out that try to tackle these kinds of inherent non-linearities in vocal tract modeling.

Yuki: We'll keep following the work on these mathematical models because they might provide the framework we need for a deeper understanding of how vocal systems function at a fundamental level <ref:2111.00868#pg2>.

The paper's summary: Ines: So, what we're looking at now is actually the paper's summary of this vowel space model and what they think it means for how we study speech.

Marcus: Basically, they’re saying the main goal was to take that really messy relationship between mouth shape and sound—that many-to-one problem—and build a coordination function that fixes it.

Ines: Right, so instead of trying to map every single acoustic measurement directly to a physical movement, they use this coordination function as an optimal way to exploit the vocal tract's potential to create a vowel space with a one-to-one correspondence.

Marcus: That’s the core idea, and it’s what makes it interesting statistically because they show that for condition C2, that many-to-one relationship just disappears when you use this function.

Ines: So, to summarize the analysis, they are showing how this generative function P theta is designed to lock all those different parameters together in a continuous domain so the shape and sound really match up.

Marcus: And what I find crucial for the cohort data is that both models—the generic one and the reduced DRM—end up having an intrinsic dimension of two for those vowel spaces, regardless of the initial parameter sets they start with.

Ines: That means even if you change how you define your vocal tract parameters in condition C1 versus C2, the resulting vowel space structure stays similar because of this underlying mathematical constraint.

Marcus: It suggests that the coordination function is a way to systematically reduce the complexity down to just two important dimensions for describing vowel production.

Ines: And they use this as a criterion to judge whether an anatomical range is truly ready for speech, because those ranges that work well have a regular relation, whereas others don't.

Marcus: That connects back to their finding that the coordination function generates the standard relation you already see in their simpler models for non-optimal forms.

Ines: It’s a structural result that tells us how the physical constraints of a vocal tract organize itself into an acoustic space we can actually map reliably.

Marcus: But they also flag something important—the method doesn't work perfectly for every single vocal tract form, so we have to be careful about where this mapping is most accurate.

Ines: So, it’s a powerful tool for generating a one-to-one correspondence and finding underlying structure in the way speech sounds are produced physically.

Marcus: Next up, we're going to look at how this structural thinking relates to the broader population genetics work they did on evolutionary constraints in species.

The paper's improvements: Ines: We just talked about how they built this coordination function to manage that messy many-to-one relationship in vowel modeling.

Marcus: Yeah, and now we're moving on to what the authors suggest as improvements for their original framework.

Ines: So, the suggestion is basically that this coordination function isn't just a one-off fix; it’s an optimal way to exploit the vocal tract’s potential to create a vowel space while simultaneously generating a one-to-one correspondence.

Marcus: That means they want us to use this as a baseline for how we should structure our models, rather than just treating it as an add-on piece of math.

Ines: Right, so the improvement is using this framework to give us that structural understanding of how the acoustic space is organized by physical constraints.

Marcus: And they point out that for those non-optimal vocal tract forms—the ones that aren't perfect—this coordination function actually generates the regular relation we already see in their simpler generic model and the reduced DRM.

Ines: So, it’s a way to bridge the gap between an imperfect physical system and a clean mathematical description of sound.

Marcus: What this implies for us is that when we're looking at real-world data, especially from cohorts with variations in vocal tract shape, this function will be really useful for identifying which forms are structurally viable for speech.

Ines: Exactly. It gives us a new objective criterion to judge if an anatomical range is truly speech ready because it shows a predictable relationship between the form and the sound it makes.

Marcus: But they also acknowledge that the method itself has limitations; it only generates that regular relation when dealing with non-optimal forms, so we need to keep that in mind.

Ines: So, while it’s not a perfect universal solution for every single vocal tract shape, it serves as a powerful diagnostic tool for understanding the underlying organization of those shapes.

Marcus: It shifts the focus from just trying to get an exact acoustic-to-articulatory map to finding these structural relationships that hold true across different models.

Ines: And that leads us right into how this structural thinking connects with the broader population genetics work they did on evolutionary constraints in species.

Conclusion: Ines: So we’re wrapping up this look at "A mathematical model of the vowel space." The main thing to remember is that they found a way to use a coordination function to fix that messy many-to-one relationship between mouth shape and sound.

Marcus: Yeah, and they proved that this approach can reduce the complexity down to two dimensions for those vowel spaces, which is what we need when looking at large cohorts.

Ines: It means we can use this structural framework to judge if certain vocal tract shapes are actually speech ready because they show a predictable relationship between the form and the sound it makes.

Marcus: That’s significant for how we analyze our data, because it gives us a way to cut through the noise and see what's structurally sound versus what's just noisy variation.

Ines: So, this paper suggests that for modeling speech production, we should be looking at these coordination functions as a baseline before trying to invert every single acoustic measurement directly.

Marcus: And they flag that the method doesn't work perfectly for every single vocal tract form, so we have to keep that in mind when applying it to real data.

Ines: Right, the caveat is important—it’s a diagnostic tool, not a perfect inversion technique across the board.

Marcus: It’s still incredibly useful though because it gives us that structural organization for how vocal tract parameters interact in a continuous domain.

Ines: So, to wrap up on "A mathematical model of the vowel space," it’s about finding an optimal way to exploit physical constraints to create a reliable one-to-one correspondence in speech modeling.

Marcus: It gives us a concrete function that generates the regular relation we see in simpler models when those forms are not perfectly optimized.

Ines: And that structural understanding is what opens up new avenues for connecting anatomy and sound patterns more clearly than before.

Marcus: Next, we’ll be looking at how this kind of structural thinking relates to the population genetics work they did on evolutionary constraints in species.

Yuki: From my side, I see this mathematical structure as a way to test hypotheses about how speech systems are constrained by physical reality across different evolutionary histories.

Ines: It really does bridge that gap between the abstract math and the biological reality of what a vocal tract can actually do.

Marcus: And for us on the data side, it means we have a better way to handle those complex parameter sets without getting lost in intractable non-linearities.

Yuki: I think connecting this structural organization to species-level constraints could tell us more about the fundamental biological mechanisms driving speech differences between populations.

Ines: Exactly. It’s a step toward building models that aren't just descriptive but truly explanatory regarding how speech evolved within those species.

Univ. Grenoble Alpes, CNRS, Grenoble INP

cs.SD, eess.AS, physics.class-ph, q-bio.PE

Submitted: 2021-10-19

Updated: 2026-10-05

Comments: 11 pages, 7 figures

DOI: 10.5281/zenodo.23241627

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 73/100

The gist: The gist The aim of this paper is to revisit the roots in order to construct a link with articulatory modelling and to understand which causality and structure these models share<ref:2111.00868#pg6>

Key concepts

Generative Function P(θ)
This is a mathematical formula used to generate different vocal tract shapes. It uses a principle similar to the Gibbs triangle from chemistry to mix components, where the angle θ controls how these components are mixed. This function is used to create closed-open tube shapes corresponding to specific vowels.
Coordination Function P(θ)
This transformed function represents the coordination between human-like articulators along a cycle defined by θ. It takes the initial generative function and expresses how the parameters interact, resulting in a formula like P(θ) = Ω + Ψ1 cos(Ψ2 − θ). This function is key to creating a one-to-one mapping.
Many-to-One Relationship
This refers to the problem where many different physical vocal tract shapes can produce the same vowel sound. The paper shows that certain models, like the generic model, exhibit this relationship. The coordination function is proposed as a solution because it reduces this relationship to a one-to-one correspondence.
Intrinsic Dimension Reduction
This concept describes how the complexity of the vocal tract space is simplified. Both models discussed in the paper reduce the intrinsic dimension to 2. This means that despite having many possible physical shapes, they can be effectively represented within a two-dimensional structure for vowel production.

Terminology

Summary

The gist The aim of this paper is to revisit the roots in order to construct a link with articulatory modelling and to understand which causality and structure these models share<ref:2111.00868#pg6>

The Problem

Establishing a causal relationship between articulatory movements and speech sounds is a major goal of the speech sciences<ref:2111.00868#pg4> This problem remains poorly understood even when limited to vowels and their first two formants (f1, f2)<ref:2111.00868#pg4> This difficulty stems from the many-to-one relation between the very large space of VT forms and the two-dimensional space of formants as well as non-linearities<ref:2111.00868#pg4> De facto, there is no reliable inversion technique applicable to any vocal tract, even with modern Bayesian or neural network approaches<ref:2111.00868#pg4>

The Generic Model and Generative Function

The paper proposes a simplification by setting a bijection between the vowel space (f1, f2) and the parametric space of different vocal tract models<ref:2111.00868#pg4> The generic model allows for the synthesis of 8 vowels with two concise formulas having 2 variables and length as the only static parameter<ref:2111.00868#pg5> A generative function P (θ) is constructed with the principle of the Gibbs triangle used in chemistry for the representation of ternary mixtures of components<ref:2111.00868#pg4> This function is given by P (θ) = 1/3(i + j + k) + 2/3(i cos(θ − π/3) + j cos(θ − π) + k cos(θ − 5π/3))<ref:2111.00868#pg5> The angle θ varies from 0 to 2π in order to mix these components<ref:2111.00868#pg5> This function is used as a generator of closed-open tube shapes, considering that the vectors defining the shapes of the tubes producing [u,a,i] placed at given angles θ ∈ π/3, π, 5π/3

The Coordination Function

The first generative function is transformed into a more comprehensible coordination function which is the coordination of human-like articulators<ref:2111.00868#pg5> This transformation leads to the form P (θ) = omega + Ψ1 cos(Ψ2 − θ), which expresses the coordination between the parameters along the θ cycle<ref:2111.00868#pg6> The values of the vectors Ψ1, Ψ2 are derived from those of i, j, k through equations involving atan and cos−1<ref:2111.00868#pg6> This formula is easy to use and if applied to real VT data published by [17], it generates a mapping from 3 vectors corresponding to extreme vowels including the length component<ref:2111.00868#pg6> The coordination function acts similarly with the Fant’s model and with the 4-Tube DRM derived from the generic model<ref:2111.00868#pg6>

Comparison of Models and Findings

The paper compares two 4-tube models, the generic model and the 4-tube DRM, to show how they reduce the many-to-one relationship<ref:2111.00868#pg6> The reduction is achieved by systematically calculating P3, P4 from P1, P2 so that this property explicitly reduces the number of free parameters to 2<ref:2111.00868#pg6> The coordination function only generates the internal antisymmetry as does the explicit reduction<ref:2111.00868#pg6> For both models, the intrinsic dimension is reduced at 2 for both models The two vowel spaces acquire a similar structure despite different parameter sets in condition C1 and C2 The key point is that we observe for C1 the many to one relationship whereas it disappears for C2 This suggests that the coordination function is an optimal way to exploit the articulatory potential of a vocal tract to create a vowel space as well as to generate a one-to-one correspondence The coordination function has the main property of generating for the non-optimal VT forms the regular relation existing with the generic model and the reduced DRM

Conclusion

The coordination function is an optimal way to exploit the articulatory potential of a vocal tract to create a vowel space as well as to generate a one-to-one correspondence The anatomical ranges observed by [10] are quite small and this provides a new objective criterion for judging whether they are truly speech ready The coordination function has the main property of generating for the non-optimal VT forms the regular relation existing with the generic model and the reduced DRM The coordination function is an optimal way to exploit the articulatory potential of a vocal tract to create a vowel space as well as to generate a one-to-one correspondence

References

[1] B. S. Atal, J. J. Chang, M. V. Mathews, and J. W. Tukey Inversion of articulatory-to-acoustic transformation in the vocal tract by a computersorting technique The Journal of the Acoustical Society of America 63(5):1535–1555, 1978<ref:2111.00868#pg9>

[2] P. Badin and G. Fant Notes on vocal tract computations Stl- qpsr 2-3/1984, Royal Institute of Technology, Stockholm, Sweden, 1984

[3] P. Badin, P. Perrier, L.-J. Boë, and C. Abry Vocalic nomograms: Acoustic and articulatory considerations upon formant convergences The Journal of the Acoustical Society of America 87(3):1290–1300, 1990

[4] Louis-Jean Boë and Pascal Perrier Comments on “distinctive regions and modes: A new theory of speech production” by m. mrayati, r. carré and b. guérin Speech Communication, 9(3):217–230, 1990

[5] Louis-Jean Boë, Pascal Perrier, and Gérard Bailly The geometric vocal tract variables controlled for vowel production: proposals for constraining acoustic-to-articulatory inversion Journal of Phonetics, 20(1):27–38, 1992

[6] Louis-Jean Boë, Thomas R. Sawallis, Joël Fagot, Pierre Badin, Guillaume Barbier, Guillaume Captier, Lucie Ménard Jean-Louis Heim and Jean-Luc Schwartz Which way to the dawn of speech?: Reanalyzing half a century of debates and data in light of speech science Science Advances 5(12):eaaw3916, 2019

[7] René Carré Dynamic properties of an acoustic tube: Prediction of vowel systems Speech Communication, 51(1):26–41, 2009

[8] G. Fant Acoustic Theory of Speech Production Mouton, 1960

[9] G. Fant and S. Pauli Spatial characteristics of vocal tract resonance modes In Proceeding of the Speech Communication Seminar, 1974

[10] W. Tecumseh Fitch, Bart de Boer, Neil Mathur, and Asif A. Ghazanfar Monkey vocal tracts are speech-ready Science Advances 2(12):e1600723, 2016

[11] J. M. Heinz Perturbation functions for the determination of vocal-tract area functions from vocal-tract eigenvalues Stl- qpsr 1/1967, Royal Institute of Technology, Stockholm, Sweden, 1967

[12] P. Mermelstein Determination of the vocal-tract shape from measured formant frequencies The Journal of the Acoustical Society of America 41(5):1283–1294, 1967

[13] M. Mrayati, R. Carré, and B. Guérin Distinctive regions and modes: A new theory of speech production Speech Commun., 7(3):257–286, October 1988

[14] M. Mrayati, R. Carré, and B.

Improvements for AI systems

  1. The AI system can perform a bijection between high-dimensional acoustic space (f1, f2) and a constrained domain in vocal tract shape space by using the coordination function, which is an optimal way to exploit the articulatory potential of a vocal tract to create a vowel space as well as to generate a one-to-one correspondence.

  2. The system can effectively model and predict vowel production by utilizing the coordination function (Eq. 7) derived from the three-phase mixing function, which explains how the phases and amplitudes of the different parameters of a model are locked together in a continuous domain.

  3. The AI can distinguish between generative models based on parameter control: for condition C1, it can observe the many to one relationship, whereas under condition C2 (using the coordination function), this relationship is reduced, showing that the intrinsic dimension is reduced at 2 for both models.

  4. The system can establish a baseline for analyzing implicit nonlinearities by using the generic model and its derived forms, noting that the coordination function has the main property of generating for the non-optimal VT forms the regular relation existing with the generic model and the reduced DRM.

Related papers