EmphTTS: an emphasis-control TTS with reinforcement learning
summary
The gist
Generating controllable and human-like emphasis remains an open challenge in text-to-speech, even when explicit emphasis control signals are provided in the text input, limiting the communicative
In short
EmphTTS is a non-autoregressive text-to-speech system that uses Reinforcement Learning to control word-level emphasis. It optimizes a separate duration predictor using Group Relative Policy Optimization (GRPO) guided by an emphasis localization reward. This method successfully improves objective metrics like WER and F1, and yields strong subjective preference for controllable prosody.
Key concepts
- EmphTTS
- A non-autoregressive TTS system designed to generate speech with precise control over word emphasis. It achieves this by optimizing the duration prediction component using Reinforcement Learning to directly target emphasis localization.
- GRPO
- Group Relative Policy Optimization is a specific Reinforcement Learning technique used here. It optimizes the duration predictor model while keeping the main TTS model fixed, guiding it toward generating more accurate speech lengths based on an emphasis-specific reward function.
- Emphasis Localization Reward
- This is a custom reward function used during training that balances three goals: minimizing Word Error Rate (WER), maximizing speaker similarity (SIM), and maximizing Balanced Accuracy. This reward directly incentivizes the duration predictor to learn how to adjust speech length specifically for emphasis.
- Duration Predictor MDP
- The Markov Decision Process (MDP) component within the system predicts the remaining speech length at each step as a probability distribution. The GRPO optimization trains this MDP to sample target durations that align with the desired emphasis, rather than just predicting a global speaking rate.
Terminology used across episodes
This episode discusses
- EmphTTS: an emphasis-control TTS with reinforcement learning · Paper Radio
- CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training
- Fish Audio S2 Technical Report
- Qwen3-TTS Technical Report
- Knowing What to Stress: A Discourse-Conditioned Text-to-Speech Benchmark
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Seed-TTS: A Family of High-Quality Versatile Speech Generation Models
- Qwen2-Audio Technical Report
The paper
EmphTTS: an emphasis-control TTS with reinforcement learning · Read on arXiv
Department of Information and Communications Engineering, Aalto University
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "EmphTTS: an emphasis-control TTS with reinforcement learning".
Tom: Generating controllable and human-like emphasis remains an open challenge in text-to-speech, even when explicit emphasis control signals are provided in the text input,
Jane: First, who's behind it and why it matters.
Paper summary: Tom: So Jane, let's get started on this paper, "EmphTTS: an emphasis-control TTS with reinforcement learning." The main idea here is that controlling emphasis in text speech is still a tough challenge even when we give the system explicit control signals.
Jane: Right, Tom? It seems the authors are tackling this by using reinforcement learning to optimize something specific called word-level emphasis directly.
Lu: I find the idea of applying Group Relative Policy Optimization, or GRPO, specifically to the duration predictor with an emphasis localization reward pretty fascinating; it suggests a very targeted way to improve that control Meng I'm curious how they manage this optimization without just making the speech longer overall?
Jane: That's exactly what makes it interesting, Lu. The paper claims that EmphTTS is a non-autoregressive TTS system where they apply GRPO to the duration predictor using an emphasis localization reward, which allows for direct optimization at the word level Tom Essentially, they are training the duration prediction part of the system to make sure emphasis sounds right for each word individually.
Lu: I think that's a key distinction; instead of just tweaking a global speed factor, which is what some other systems do, they are learning how to adjust the length based on where those emphasis markers appear in the text Jane Which is powerful because it moves beyond just utterance-level changes.
Meng: From an engineering standpoint, that implies a lot of fine-grained adjustments happening at each step of the duration prediction process, which sounds computationally intensive but potentially very precise Tom I mean, if they can optimize for word-level emphasis without losing intelligibility, that's a big deal for real applications.
Lalam: It’s really about the cultural impact here; if we can generate speech where emphasis is perfectly localized and controllable, it opens up incredible possibilities for personalized communication and accessibility Lu Imagine how this could help people communicate nuance in complex scenarios or even tailor digital assistants to convey exactly the right tone Meng This level of control could fundamentally alter how we interact with synthetic media.
Paper summary: Jane: Exactly, Lalam. The paper states that their evaluation shows EmphTTS achieves the best emphasis controllability and performs best in emphasis objective evaluation Tom So, they've set up a system where they can actually measure how well it controls that emphasis compared to other systems out there.
Lu: And the authors show strong results in subjective preference tests, noting that EmphTTS is significantly preferred over synthetic groundtruth and most baselines Jane That suggests their objective improvements translate into something listeners actually perceive as better.
Tom: It sounds like they've really managed to bridge that gap between technical metrics and actual human perception, which is often where these types of models fall short Meng We need to keep an eye on how robust this emphasis control is when we move it out of the controlled testing environment.
Jane: That’s a fair point, Tom. The paper does mention some limitations, specifically that while they have a global speed factor option, they acknowledge that fixed global scaling is inherently coarse for word-level emphasis control Lu They also note that the performance depends on balancing three different reward components: minimizing Word Error Rate, maximizing speaker similarity, and maximizing Balanced Accuracy Meng That balancing act sounds tricky to get right in practice.
Tom: That's the complexity I was thinking about; managing that trade-off between accuracy, naturalness, and emphasis control simultaneously is a genuine hurdle Jane So, moving from the SFT stage to this GRPO stage with that specific reward function is where they really put their innovative spin on things.
Lu: The way they define the clipped surrogate objective in the GRPO loss seems designed to handle those mismatches between the independently trained TTS model and duration predictor Tom It’s not just optimizing one part in isolation; it’s coordinating them through that RL framework.
Meng: Coordinating two separate components via reinforcement learning is sophisticated stuff, but I wonder about the practical deployment; how easy is it for a developer to set up an emphasis-specific reward function if they don't have deep reinforcement learning expertise?
Jane: The paper shows their ablation studies compare several configurations, including MTTS-SFT without duration prediction and MTTS-DP-SFT, which combines the two separately trained systems Tom This suggests they are testing different ways to combine the components to find the best setup.
Paper summary: Tom: And what they found was that their proposed EmphTTS system—MTTS-SFT with MDP-GRPO—outperformed those other setups, showing improvements in emphasis realization beyond what simple global speed reduction could achieve Lu That’s a solid finding for demonstrating the value of this approach.
Jane: It really shows that the GRPO optimization learns input-dependent duration adjustments specifically for emphasis rather than just increasing the utterance duration globally Tom Which is a much more nuanced way to think about it.
Lu: And I'm excited about what this implies for future work; they suggest that subjectively preferred emphasis can be influenced by overall speech quality, which motivates using both objective and subjective measures together Meng That points toward a more holistic view of TTS evaluation.
Tom: So, to wrap up the essence of this paper, "EmphTTS: an emphasis-control TTS with reinforcement learning," it’s about showing that you can use reinforcement learning to fine-tune the duration predictor using an emphasis reward to get direct control over word-level emphasis Jane It's a system that successfully addresses the difficulty of creating controllable and human-like emphasis in text speech.
Jane: And the authors conclude that this method enables fine-grained prosodic control, suggesting Reinforcement Learning has potential for this kind of control beyond just traditional emphasis marking Tom It really points toward a way to make synthetic speech feel much more natural when it comes to conveying subtle emotional cues.
Lu: The implications stretch into how we build more expressive and context-aware AI systems; if the underlying TTS can handle this level of detail, the applications for voice synthesis become much richer Meng We could see this applied in areas where tone and emphasis are critical for understanding intent.
Tom: Absolutely, Lu. This work moves the needle on how we define and realize prosodic control in neural TTS systems Jane It shows that targeted RL optimization on a specific component, like the duration predictor here, can yield tangible improvements in real-world applicability.
Jane: So as we wrap up this discussion on "EmphTTS: an emphasis-control TTS with reinforcement learning," the main thing is that they successfully showed how GRPO can coordinate components to achieve better word-level emphasis than previous methods Tom It’s a strong step forward in making synthetic speech more controllable.
Conclusion: Tom: So we've been diving deep into how EmphTTS uses reinforcement learning to fine-tune the duration predictor for word-level emphasis control in TTS, and now it's time to look at what this actually means for us as a team and the world.
Jane: I think the title itself, "EmphTTS: an emphasis-control TTS with reinforcement learning," really tells the story of this paper, focusing on how they used RL to get that precise control over emphasis in text speech.
Lu: From my perspective as a researcher, this work is significant because it shows we can move beyond just general prosody and start tackling such fine-grained acoustic control directly through optimization.
Meng: I'm thinking about the practical side here; if this works reliably, it means the AI outputs won't just sound smooth, but they can convey specific intent like emphasis in a way that matters for real-world applications.
Lalam: I see this as a major step forward because if we can generate speech with truly controllable nuance, it opens up incredible possibilities for how people communicate and how we interact with technology.
Tom: Exactly, Lalam, and that’s what I want to emphasize—this isn't just about making the speech sound better; it’s about giving the AI more expressive power. The authors managed to coordinate a duration predictor with the main TTS model using this specific reward function to achieve word-level emphasis optimization.
Jane: It really is a clever mechanism, Tom; they aren't just tweaking one part in isolation, but they are making sure the duration prediction learns how to adjust based on where emphasis signals appear in the text.
Lu: That coordination through GRPO is what makes it so interesting; it’s about learning input-dependent adjustments rather than just applying a blanket speed reduction across the whole utterance.
Meng: I'm curious about how stable this optimization is when we move it from the controlled test sets to more unpredictable, real-world text inputs, because that's where engineers usually hit their first roadblock.
Lalam: And if we can achieve this level of precision, it could profoundly affect culture by allowing for much richer and more nuanced digital interactions and content creation.
Tom: That’s the big picture, Lalam—imagine the future applications when we can generate speech that perfectly conveys subtle emotional cues or specific rhetorical emphasis. This paper shows a viable path toward that level of control.
More episodes
- 2610.10857-Self-Supervised Keyframe Discovery for Horizon-Invariant Behavior Cloning
- 2610.10768-Strategic Investment Decision Making for Value Creation in Energy Transition: A Reinforcement Learning Approach
- 2610.10858-RFChipAgent: Multi-Agentic AI Flow for Analog/RF Chip Design
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization