Pronunciation Editing for Finnish Speech using Phonetic Posteriorgrams

arXiv:2507.02115 · eess.AS, cs.CL · Submitted 2025-07-02 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Pronunciation Editing for Finnish Speech using Phonetic Posteriorgrams".

Jane: Synthesizing second-language (L2) speech is potentially highly valued for L2 language learning experience and feedback,

Tom: First, who's behind it and why it matters.

Title and authors: Tom: So, let’s talk about who wrote this and what that title really means for us in the context of AI research. The paper "Pronunciation Editing for Finnish Speech using Phonetic Posteriorgrams" was put out by Zirui Li, Lauri Juvela, and Mikko Kurimo from Aalto University.

Jane: It’s interesting to see researchers from different backgrounds collaborating on something so specialized; it shows how interdisciplinary work is becoming more common in this area.

Lu: The authors are clearly working at the intersection of deep learning and speech processing, which is where the most exciting things happen right now.

Meng: I'm curious about their approach because they specifically mention using Phonetic Posteriorgrams instead of just phoneme sequences as input for editing.

Lalam: That focus on "soft, explicit alignment information" mentioned in the summary seems like a sophisticated way to handle pronunciation variations that traditional methods might miss.

The paper's summary: Tom: So, what’s the core idea here? Basically, the paper explains how PPG2Speech is a diffusion-based model that can edit a single phoneme in native speech to make it sound like an L2 learner said it, all without needing any text alignment.

Jane: That’s a big deal because traditional speech editing often relies on masking and external aligners, which can actually make the resulting synthesized speech sound less natural.

Lu: The paper introduces PPGs as time-varying categorical distributions over phonemes, which they argue is a richer representation for assessment and detection than just the sequence of phonemes.

Meng: From a practical standpoint, if you can edit just one phoneme without complex alignment steps, that significantly simplifies the workflow for people who actually need this capability.

Lalam: It’s about taking what they call Matcha-TTS’s flow-matching decoder and strengthening it with Classifierfree Guidance and Sway Sampling to get high quality output from these PPG inputs.

The paper's improvements: Tom: They introduce a new way to measure success called the Phonetic Aligned Consistency, or PAC, which is designed specifically for evaluating how well the editing worked on the phoneme level.

Jane: That sounds like a very targeted metric because it moves beyond just listening tests and tries to quantify the actual phonetic correction achieved.

Lu: They define PAC by calculating a DTW between the edited region of those PPGs and the corresponding region in the synthetic speech's PPG, using Jensen-Shannon Distance for discrepancy measurement.

Meng: I see why they use DTW to handle length differences between the edited and synthetic regions; it’s smart because it normalizes the cost by dividing by the length of the edited region.

Lalam: And they found that when they strengthened their model with Classifierfree Guidance, their PAC score reached zero point seven zero nine, which is better than what a text-based editing method achieved with Matcha-TTS and CFG, which scored zero point eight zero four.

Conclusion: Tom: So, to wrap up the "Pronunciation Editing for Finnish Speech using Phonetic Posteriorgrams," this model shows that you can edit speech at the phoneme level using PPGs to approximate L2 pronunciations without needing text alignment.

Jane: It’s a powerful tool because it keeps the naturalness of the synthesized speech high while allowing precise control over pronunciation errors.

Lu: The implication is that we have a more interpretable way to manipulate speech representations, moving toward editing based on phonetic structure rather than just sequence alignment.

Meng: For practical deployment, this means we could potentially create tools that instantly correct common L2 pronunciation mistakes in recorded speech for language learners or even improve the quality of training data.

Lalam: Ultimately, PPG2Speech is a diffusion-based model that uses the Matcha-TTS flow-matching decoder enhanced with Classifierfree Guidance to achieve this phoneme-level editing.

Tom: That’s a great summary of what they achieved with PPG2Speech and how it works for Finnish speech.

Jane: It really puts into perspective how much information we can extract from representations like the Phonetic Posteriorgrams.

Lu: I think this work opens up new avenues for studying how different types of generative models handle specific, localized modifications in complex data like speech.

Meng: From my side, the ability to edit a single phoneme without text alignment is what makes it practically useful for integrating into existing speech processing pipelines.

Lalam: It’s inspiring because it shows that by focusing on richer input representations, we can solve real problems in low-resourced language synthesis.

Zirui Li, Lauri Juvela, Mikko Kurimo

Department of Information and Communication Engineering, Aalto University, Finland

eess.AS, cs.CL

Submitted: 2025-07-02

Updated: 2026-02-07

Comments: Accepted by Proceeding of 13th edition of the Speech Synthesis Workshop; 5 pages, 1 figure

DOI: 10.21437/SSW.2025-36

Code: https://github.com/aalto-speech/PPG2Speech

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 82/100

The gist: Synthesizing second-language (L2) speech is potentially highly valued for L2 language learning experience and feedback, but synthesizing L2 speech for low-resourced languages like Finnish is

Key concepts

Phonetic Posteriorgrams (PPGs)
These are time-varying categorical distributions that represent rich pronunciation information across phonemes. They are ideal for assessing and detecting mispronunciations because they capture the probability of different phonemes at each point in time, providing a detailed view of how speech sounds change.
Diffusion-based Model
This is a type of generative AI model that works by gradually adding noise to data until it becomes pure noise. The model learns to reverse this process, starting from random noise and iteratively refining it into high-quality speech based on the input PPG information.
Classifierfree Guidance (CFG)
CFG is a technique used during the generation process to guide the model toward specific desired outputs. It helps improve the quality and adherence of the synthesized speech, particularly in complex tasks like editing, by conditioning the generation on external speaker and pitch embeddings.

Terminology

Summary

Synthesizing second-language (L2) speech is potentially highly valued for L2 language learning experience and feedback, but synthesizing L2 speech for low-resourced languages like Finnish is difficult due to a lack of suitable datasets. This paper presents PPG2Speech, a diffusion-based multispeaker PhoneticPosteriorgrams-to-Speech model that enables editing native speech to approximate L2 pronunciations without text alignment.

The gist

PPG2Speech is a diffusion-based multispeaker PhoneticPosteriorgrams-to-Speech model that is capable of editing a single phoneme without text alignment.

Background and Motivation

The scarcity of L2 speech synthesis data necessitates exploring methods such as phoneme-level speech editing, where native speech is systematically modified to approximate specific L2 pronunciations. Current speech editing models often adopt a mask-and-infill approach that relies on external aligners, which can hurt the naturalness of synthesized speech. To gain more control over pronunciation information, the authors use Phonetic Posteriorgrams (PPGs) as input instead of the phoneme sequence. PPGs are described as time-varying categorical distributions over phonemes, containing rich pronunciation information ideal for pronunciation assessment and mispronunciation detection and being interpretable representations for neural speech editing.

Model Architecture (PPG2Speech)

The PPG2Speech model is a diffusion-based multispeaker PPG-to-speech model that utilizes Matcha-TTS’s flowmatching decoder, enhanced with Classifierfree Guidance (CFG) and Sway Sampling. The architecture consists of a PPG encoder and a flow-matching decoder. The PPG encoder uses a prenet, followed by a stack of Transformer/Conformer layers, with an upsampling layer to map latent PPG representations to the target mel-spectrogram resolution based on the nearest neighbor method. For the flow-matching decoder, they extend the 1-D U-Net based decoder in Matcha-TTS with CFG and condition it on external speaker embeddings and pitch embedding sequences. During inference, they use Sway Sampling with a sway coefficient of s = −1 to transform speech from noise using smaller steps in initial stages to improve quality. The synthesized mel-spectrogram is then converted to waveform using a pretrained HiFi-GAN vocoder.

Editing Evaluation Metric (PAC)

The authors propose a new task-specific objective evaluation metric called the Phonetic Aligned Consistency (PAC) between the edited PPGs and the PPGs extracted from the synthetic speech for editing effects. The PAC calculation is defined as:

  1. Calculate DTW between the edited region of the edited PPG and the corresponding region in the synthetic speech's PPG:

&textstyle PAC1:m,1:n = &frac1mDTWJSD(PPGmedited, PPGnsyn), (7)

  1. The JSD is used because PPG is a time-varying distribution over phoneme categories to measure discrepancy between frames.

  2. DTW tackles length differences between the edited and synthetic regions, and normalization by the length of the edited region compensates for larger costs from longer edits, ensuring comparability across different utterances.

Experimental Validation

The experiments were conducted on a combination of two Finnish datasets: Perso Synteesi [26] and Finsyn [27]. They extracted PPGs using a Kaldi HMM-TDNN-medium model and used pretrained SimAMResNet341 for speaker embeddings and PENN for pitch/periodicity estimation. The model was trained from scratch on the training set, with a 10% probability of dropping PPG latent and conditional input to train the unconditioned model. Objective evaluations compared PPG2Speech with Matcha-TTS and Matcha-TTS-CFG, showing that while Matcha-TTS outperforms it on seen data, PPG2Speech performs better on unseen speaker similarity. Strengthening the PPG2Speech model with CFG resulted in an unseen SECS of 0.86 and a PAC of 0.709, demonstrating its effectiveness in editing tasks compared to the text-based editing methods which achieved a lower PAC of 0.804 for Matcha-TTS with CFG. Subjective evaluations showed that PPG2Speech-CFG achieved the highest Naturalness MOS (NMOS) of 3.51 ± 0.07, and the highest Editing MOS (EMOS) NMOS of 3.66 ± 0.06, indicating its superior performance in both synthesis and editing tasks despite some timbre leakage affecting speaker similarity scores.

Conclusions

PPG2Speech is a diffusion-based multispeaker PPG-to-Speech model capable of phoneme-level pronunciation editing to approximate L2 language learners by strengthening the Matcha-TTS flow-matching decoder with Classifierfree Guidance (CFG).

Improvements for AI systems

Here are the specific improvements that can be made to AI speech synthesis systems based on this paper, and what those improved systems will be capable of doing:


  1. Improved L2 Speech Synthesis for Low-Resourced Languages (like Finnish):

  2. Enhanced Pronunciation Editing Capabilities: The system can now perform precise, phoneme-level pronunciation corrections on native speech to approximate specific L2 pronunciations without requiring text alignment or complex masking procedures common in traditional speech editing models.

  3. Naturalness and Speaker Similarity Preservation during Editing: The improved system will synthesize edited speech that maintains high naturalness (measured by NMOS) and speaker similarity (SMOS), overcoming the typical degradation seen in existing editing methods.

  4. Robust Generalization to Unseen Speakers: By leveraging Classifier-Free Guidance (CFG) in the diffusion decoder, the system will show improved speaker similarity metrics (SECS) on unseen speakers compared to baseline models, allowing for better handling of diverse voice identities.

  5. Task-Specific Evaluation and Control: The introduction of Phonetic Aligned Consistency (PAC) provides a new objective metric specifically tailored to measure the success of phoneme-level editing effects, offering a more accurate assessment than traditional metrics alone.

  6. Enhanced Synthesis Fidelity via Flow Matching: Utilizing the flow-matching decoder (Matcha-TTS backbone) with Classifier-Free Guidance and Sway Sampling will result in higher quality, faster inference synthesis compared to standard NAR models or simpler diffusion samplers.

These improvements enable the following specific capabilities for the improved AI system:

  1. A researcher or language learner can input native speech and a target L2 pronunciation (defined via PPGs) and receive an output audio file where only the targeted phoneme is corrected, resulting in fluent, natural-sounding L2 speech.

  2. The system can be used in personalized language learning tools to provide immediate feedback on pronunciation errors by synthesizing corrected examples based on user input, leading to a richer and more effective learning experience.

  3. The system can be deployed in applications requiring voice cloning or style transfer where the speaker identity is preserved while the prosody and specific phoneme realizations are modified for L2 target sounds.

  4. The system offers a method to create high-quality, edited training data for other speech models by systematically correcting common L2 pronunciation errors in existing datasets, thereby improving the robustness of downstream TTS systems.

Abstract

Synthesizing second-language (L2) speech is potentially highly valued for L2 language learning experience and feedback. However, due to the lack of L2 speech synthesis datasets, it is difficult to synthesize L2 speech for low-resourced languages. In this paper, we provide a practical solution for editing native speech to approximate L2 speech and present PPG2Speech, a diffusion-based multispeaker Phonetic-Posteriorgrams-to-Speech model that is capable of editing a single phoneme without text alignment. We use Matcha-TTS's flow-matching decoder as the backbone, transforming Phonetic Posteriorgrams (PPGs) to mel-spectrograms conditioned on external speaker embeddings and pitch. PPG2Speech strengthens the Matcha-TTS's flow-matching decoder with Classifier-free Guidance (CFG) and Sway Sampling. We also propose a new task-specific objective evaluation metric, the Phonetic Aligned Consistency (PAC), between the edited PPGs and the PPGs extracted from the synthetic speech for editing effects. We validate the effectiveness of our method on Finnish, a low-resourced, nearly phonetic language, using approximately 60 hours of data. We conduct objective and subjective evaluations of our approach to compare its naturalness, speaker similarity, and editing effectiveness with TTS-based editing. Our source code is published at https://github.com/aalto-speech/PPG2Speech.

Sources

Related papers