A Situational Speech Synthesizer for Yoruba: System Design, Phonological Rule Architecture, and Orthographic Extensions for Contour
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "A Situational Speech Synthesizer for Yoruba".
Jane: TTSYoruba presents a rule-based concatenative diphone speech synthesizer for Yorùbá, deployed online to generate audio for personal names within the YorubaName.com open dictionary.
Tom: First, who's behind it and why it matters.
Title and authors: Tom: So, we're looking at this paper titled "A Situational Speech Synthesizer for Yoruba: System Design, Phonological Rule Architecture, and Orthographic Extensions for Contour Tones," which really shows a complete approach to solving pronunciation issues in this language.
Jane: That title tells us they didn't just throw some model at the problem; they focused on the design and the specific rules needed to handle the phonology correctly, which is super helpful for understanding how these systems work under the hood.
Lu: The authors, Túbọ̀sún, Adédayọ̀ Olúòkun, Hafiz Adéwuyì, and Dadépọ̀ Adérẹ̀mí YorubaName.com, clearly have deep knowledge of both linguistics and speech technology to create such a detailed architecture.
Meng: It’s interesting that they focused on the orthographic extensions for contour tones using the caron and circumflex marks; that shows they weren't just treating it as a purely phonetic problem but also considering how people actually write the language.
Lalam: I see a lot of potential in this because they are establishing a public record and a standard for future Yorùbá speech technology development, which is huge for the community.
The paper's summary: Tom: To summarize what they did with TTSYoruba, it's a rule-based system that takes tone-marked text as input and uses their hand-crafted phonological rules to generate audio from their recorded inventory of six hundred fifty-one diphone units.
Jane: They detailed the entire architecture, including how they select which file to use based on the graphemic tone marking on the current syllable and the preceding syllable's tone, which is a very precise way to manage tonal transitions.
Lu: The summary also highlights their specific treatment of nasal disambiguation—how they distinguish between an oral onset, a nasality marker, and a syllabic nasal for consonants like /n/ and /m/.
Meng: They also explained how they derived contextual rising and falling tones from the level-tone input, which is the mechanism that allows the system to handle those subtle contour tones without needing massive amounts of training data.
Lalam: And their contribution regarding orthography is significant; they adopted the caron and circumflex marks as standard single-vowel contour tone markers integrated into their normalization pipeline.
The paper's improvements: Tom: The main improvement they focus on is moving away from just general text-to-speech toward a specialized system that explicitly incorporates phonological rule architecture to handle the specific tonal and nasal complexities of Yoruba.
Jane: They also suggest a structured framework where the file selection logic is determined by the combination of graphemic tone marking on the current syllable and the tone of the immediately preceding syllable, which provides a clear decision path for synthesis.
Lu: They also introduced an orthographic extension system that uses caron and circumflex marks as standard markers for contour tones, which resolves issues where conventional spellings might omit the vowel gemination that would normally represent a contour tone in single syllables.
Meng: The paper also addresses the three-way nasal disambiguation problem by clearly defining three roles for /n/, which should make their system much more robust when dealing with complex name structures that involve nasals.
Lalam: Their work on the orthographic extensions is really important because it gives users a way to represent contour tones in a Unicode-compatible, keyboard-accessible manner without having to change established spellings.
Conclusion: Tom: So, wrapping up this paper on "A Situational Speech Synthesizer for Yoruba: System Design, Phonological Rule Architecture, and Orthographic Extensions for Contour," the core implication is that a complete rule-based system can successfully synthesize correct pronunciation in a low-resource language like Yoruba.
Jane: They've shown how meticulous design around phonological rules and specific orthographic adaptations allows for high intelligibility across various categories, even if the naturalness scores are moderate due to the concatenative synthesis method.
Lu: The possibility of using this architecture as a template for other tonal languages, given the documented phonological stability in things like high tones being resistant to deletion, opens up a lot of avenues for future research.
Meng: From an engineering viewpoint, it confirms that you can build a high-quality domain-specific system for personal names without needing massive neural datasets if you have a solid rule set and a small, well-defined inventory.
Lalam: I think the most profound cultural implication is giving communities access to automated pronunciation tools for their personal names, which validates the importance of their existing written orthography and helps preserve linguistic accuracy.
cs.SD, cs.CL, eess.AS
Submitted: 2026-07-17
Updated: 2026-10-01
Comments: Currently under review at Speech Communication
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 82/100
The gist: TTSYoruba presents a rule-based concatenative diphone speech synthesizer for Yorùbá, deployed online to generate audio for personal names within the YorubaName.com open dictionary.
Key concepts
- Diphone Speech Synthesizer
- This is a method of speech synthesis that combines two adjacent phonemes (a consonant and a vowel) into a single recorded unit or 'diphone.' TTSYoruba uses these specific, pre-recorded units from its corpus to build the final spoken word.
- Phonological Rule Architecture
- This is the core logic of the system, consisting of four sequential steps: normalizing text, breaking it into syllables and handling nasal sounds, selecting the correct tonal file based on neighboring syllables, and finally stitching these units together.
- Nasal Disambiguation System
- This addresses the difficulty in determining how to pronounce /n/ or /m/. The system distinguishes between three roles: when 'n' starts a syllable (oral onset), when it marks a vowel as nasalized, or when it forms its own complete syllable with its own tone.
- Orthographic Extensions for Contour Tones
- This feature allows users to use the caron (ˇ) and circumflex (ˆ) marks in writing to indicate contour tones. The system expands these marks into their 'geminated equivalents' before processing, mapping them specifically to rising or falling diphone files.
Terminology
Summary
TTSYoruba presents a rule-based concatenative diphone speech synthesizer for Yorùbá, deployed online to generate audio for personal names within the YorubaName.com open dictionary. This system is significant because it provides a complete, documented framework for synthesizing correct pronunciation from tone-marked text in a low-resource language where neural methods are currently infeasible due to lack of annotated data.
The gist
We present TTSYoruba, a rule-based concatenative diphone speech synthesizer for Yorùbá, deployed online as part of the YorubaName.com open dictionary of Yorùbá personal names.
Corpus and Inventory Construction
The system is built upon a 651-unit diphone corpus covering all phonologically predicted syllable-tone combinations in Yorùbá, recorded under consistent acoustic conditions.
The inventory is organized using a specific file-naming convention: l.wav Low Default; not after H or R tone,
m.wav Mid Anywhere,
h.wav High Default; not after L or F tone,
and f.wav Falling (H→L) After H or R tone; or circumflex mark (â),
and r.wav Rising (L→H) After L or F tone; or caron mark (ǎ).
The total inventory includes 126 CV diphones
derived from 18 consonants × 7 vowels,
plus 7 standalone vowels
with three variants each, totaling 651 recorded units.
Phonological Rule Architecture
The core of the system is a four-module pipeline: (1) Unicode normalization, (2) syllabification and nasal disambiguation, (3) tonal file selection, and (4) diphone concatenation. The tonal file selection
rule is determined by the combination of the graphemic tone marking on the current syllable and the tone of the immediately preceding syllable.
This logic incorporates contextual rules:
-
Contextual rising tone: A high-tone grapheme is assigned a rising (r) diphone file when it
immediately follows a syllable assigned a low or falling tone.
-
Contextual falling tone: A low-tone grapheme is assigned a falling (f) diphone file when it
immediately follows a syllable assigned a high or rising tone.
Nasal Disambiguation System
The system explicitly addresses the three-way nasal disambiguation problem
for the oral consonant /n/ (and by extension /m/). The three roles are:
-
Oral onset: /n/ begins a CV syllable (e.g., File: noo h.wav).
-
Nasalization marker: /n/ following a vowel nasalizes that vowel and is not itself a separate syllable (e.g., File: preceding vowel-consonant pair is replaced by the corresponding nasal diphone).
-
Syllabic nasal: /n/ or /m/ constitutes a complete syllable bearing its own tone (e.g., File: n r.wav if it follows a low-tone syllable).
Orthographic Extensions for Contour Tones
A key contribution is the adoption of the caron (ˇ) and circumflex (ˆ) as standard single-vowel contour tone markers.
This resolves cases where conventional spellings omit the vowel gemination that standard Yorùbá orthography would use to represent a contour tone within a single syllable.
The system implements this by:
-
Normalization: A preprocessing step expands caron-marked and circumflex-marked vowels to their
geminated equivalents
before G2P conversion, ensuringinput-form equivalence.
-
Mapping: The caron maps to the rising file (r) and the circumflex maps to the falling file (f). For instance,
ǎ ≡ àá (rising on /a/)
andâ ≡ áà (falling on /a/).
This convention allows for a Unicode-compatible, keyboard-accessible way to represent contour tones without altering established spellings.
Evaluation Results
A listener study with N=50 participants assessed naturalness and intelligibility across four categories: Standard tonal (Category A), Syllabic nasal (Category B), Geminated contour (Category C), and Caron/circumflex contour (Category D). The results showed that Intelligibility was uniformly high across all four categories
but Naturalness scores were more moderate,
consistent with concatenative synthesis limitations. Crucially, a paired-samples t-test comparing Category C (geminated) and Category D (caron/circumflex) found "neither dimension showed a significant difference: naturalness t(99)=−1.03, p=.307, d=0.10; intelligibility t(99)=−1.65, p=.101, d=0.
Improvements for AI systems
Here are specific improvements for AI systems based on the TTSYoruba paper, categorized by application:
)AI System Improvements & Capabilities
-
Acoustic Modeling Layer Integration (Hybrid Architecture):
-
Prosodic Synthesis Module: The rule-based phonological layer (TTSYoruba's core) should serve as a high-precision text normalization front-end for a subsequent neural TTS model.
-
Contour Tone Encoding Pipeline: Implement the Unicode caron/circumflex mapping directly into the Grapheme-to-Phoneme (G2P) normalization step. This allows the system to process
orthographically ambiguous
inputs (e.g., Níkẹ̀ẹ́ vs. Níkẹ̌) and map them consistently to the correct phonological units before acoustic modeling, ensuring that contour tones are synthesized accurately without needing separate audio files for every possible orthographic variant. -
Nasal Disambiguation Engine: Incorporate a machine learning classifier or an expanded rule-based system (extending Rules N1–N4) trained on larger datasets to handle the
morpheme-boundary nasal errors
(e.g., in Ìtànìfẹ́). This would allow the system to resolve ambiguous nasal forms with higher accuracy than purely heuristic rules, reducing synthesis failures for complex name structures. -
Low-Resource Language Adaptation Framework: Utilize the documented rule architecture and the small diphone inventory as a template for
few-shot
ortransfer learning
systems in other low-resource tonal languages. The explicit documentation of phonological stability (e.g., High tones are most stable) provides a principled framework for building robust synthesis rules where data is scarce. -
Domain-Specific TTS Generation: The system can be specialized for generating high-quality audio specifically within the domain of personal names, as demonstrated by its consistent performance across four distinct name categories (Standard, Syllabic Nasal, Geminated Contour, and Caron/Circumflex). This ensures superior output quality for a specific use case compared to general language TTS models.
-
Perceptual Equivalence Validation Module: Integrate a module that performs paired-sample statistical tests (like the t-tests in Section 6.4) on synthesized output against target orthographic forms (geminated vs. caron). This allows the AI system to provide confidence scores regarding the perceptual equivalence of different written notations, addressing
unfamiliarity responses
identified in the study. -
Cross-Lingual Rule Transfer: Develop a mechanism to transfer the core phonological rule system (tone file selection logic) from Yorùbá to other tonal languages by mapping their respective tonal inventories and phonological constraints, accelerating TTS development for those target languages.
Abstract
We present TTSYoruba, a rule-based concatenative diphone speech synthesizer for Yoruba, deployed at online as part of the YorubaName.com open dictionary of Yoruba personal names. The system takes tone-marked Yoruba text as input and produces audio output by applying a hand-crafted phonological rule system to a recorded inventory of 651 diphone units spanning five tonal variants of every consonant-vowel combination in the language. We describe the phonological architecture of the system in detail, including our complete tonal file-selection logic, our treatment of the three-way nasal disambiguation problem (oral /n/, nasalized vowel, and syllabic nasal), and the derivation of contextual rising and falling tones from level-tone input. We also present, as an orthographic contribution, the adoption of the caron and circumflex, which are symbols with prior standing in Yoruba phonological transcription, as standard single-vowel contour tone markers, integrated into the TTS normalization pipeline and the WriteYoruba keyboard input tool. The system's performance was evaluated through a listener study (N=50), with detailed results on Mean Opinion Scores (MOS) presented in Section 6. Keywords: Yoruba, text-to-speech, low-resource languages, diphone synthesis, contour tones, African language NLP, rule-based synthesis
Sources
- \`{I}r\`{o}y\`{i}nSpeech: A multi-purpose Yor\`{u}b'{a} Speech Corpus
- Geometry of nowhere vanishing, point separating sub-algebras of $\mathcal{H}ol(\Gamma\cup\text{Int}(\Gamma))$ and zeros of Holomorphic functions
Related papers
- Few-Shot Open-Set Audio Classification via Transductive Prototype Refinement and Class Logit Enhancement
- Spectral Masking and Interpolation Attack (SMIA): A Black-box Adversarial Attack against Voice Authentication and Anti-Spoofing Systems
- AVMeme Exam: A Multimodal Multilingual Multicultural Benchmark for LLMs' Contextual and Cultural Knowledge and Thinking
- SoundWeaver: Compositional Warm-Starting for Text-to-Audio Diffusion Serving
- WASIL: In-the-Wild Arabic Spoken Interactions with LLMs
- Efficient Audiovisual Speech Processing via MUTUD: Multimodal Training and Unimodal Deployment