HPRO: Hierarchical Progressive Reward Optimization via Preference Extraction for Emotional Text-to-Speech
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "HPRO: Hierarchical Progressive Reward Optimization via Preference Extraction for Emotional Text-to-Speech".
Jane: Hierarchical Progressive Reward Optimization (HPRO) is a novel framework designed to enhance emotional expressiveness in Large Language Model-based Text-to-Speech (TTS) systems by overcoming structural mismatches between content and emotion optimization,…
Tom: First, who's behind it and why it matters.
Title and authors: Tom: Let’s look at who wrote this piece; it’s Sihang Nie, Xiaofen Xing, Rui Xing, Haoming Li, Ruitong Xiao, Jingyuan Xing, Baiji Liu, and Xiangmin Xu from South China University of Technology. It shows a really collaborative effort across different institutions.
Jane: Their affiliations suggest a strong foundation in both core AI research and the practical application side; it’s interesting to see that mix of academic rigor and industry involvement represented in the team behind HPRO.
Lu: The diversity of expertise in the authorship is exactly what you need when dealing with something as complex as optimizing emotional expression across multiple levels, which is what HPRO aims to do.
Meng: I wonder if having researchers from different places helps them spot those structural mismatches better; sometimes a fresh set of eyes on the problem can reveal hidden constraints in the architecture.
Lalam: It's impressive to see such a broad team coming together to tackle this specific challenge; it shows how many different perspectives are needed when you’re pushing the boundaries of what language models can do.
The paper's summary: Tom: Now, let’s get into the actual meat of what HPRO proposes; essentially, they introduce a hierarchical progressive reward optimization framework to handle those two issues we talked about earlier.
Jane: That sounds like they are building a system that doesn't just try to optimize emotion all at once, but instead builds it up in stages, addressing the content versus style conflict and the scale gap simultaneously.
Lu: The core mechanism involves using something called the HD-Emo codec as a differentiable reward model; this codec projects speech tokens into separate spaces for content and style preferences, which is a clever way to isolate those conflicting objectives.
Meng: Isolating them sounds promising for stability; if we can keep the acoustic structure tied strictly to the text, we reduce the risk of that semantic degradation they mentioned.
Lalam: If this method works well, it means we could move past just sounding fluent and start capturing genuine human emotional tones in our AI outputs.
The paper's improvements: Tom: The improvements they detail focus on resolving the information conflict by using that HD-Emo codec to separate content and style tokens, and then they tackle the scale gap by progressively aligning rewards across frame, word, and sentence levels.
Jane: That progressive alignment strategy is where I see the real innovation; it’s not just one big optimization goal, but a structured path where they start with smaller steps—frame level—and build up to global sentence-level consistency.
Lu: The way they handle semantic alignment is particularly interesting; they feed latent representations into a content adapter supervised by an ASR objective, specifically the auto-regressive negative log-likelihood of target text tokens, which is formulated as LASR = −∑j=one log P(yj y<j, Tc) (one).
Meng: That stop-gradient mechanism they use to prevent acoustic leakage from affecting the content extractor is a critical detail; it ensures that optimizing the style doesn't accidentally mess up what the model learns about the actual words.
Lalam: It’s really smart how they set up word-level constraints using metrics like Concordance Correlation Coefficient for Valence–Arousal–Dominance dimensions, which gives us very precise control over the emotional aspects of speech.
Conclusion: Tom: So, to wrap up, HPRO essentially takes those two major hurdles—information conflict and scale gap—and solves them by using a hierarchical progressive reward optimization guided by that HD-Emo codec.
Jane: They’ve shown that this approach significantly improves fine-grained emotional expressiveness while still keeping the linguistic intelligibility intact through their multi-scale supervision strategy.
Lu: The paper suggests that by structuring the optimization this way, they can achieve robust emotional control at multiple granularities, from individual speech frames up to the entire sentence context.
Meng: From a practical standpoint, the structured training path is important because it prevents reward hacking by ensuring that every stage of optimization serves a specific purpose in building that acoustic foundation.
Lalam: I think the biggest implication is that we are moving closer to AI systems where emotional expression isn't just an afterthought but a reliably controlled feature, which could really enhance the richness of our digital interactions.
Tom: Absolutely; HPRO gives us a clear blueprint for how to guide these complex generative models toward more human-like emotional output without sacrificing accuracy. We’ve been talking about how this paper tackles the core issues in preference-driven TTS.
Sihang Nie, Xiaofen Xing, Rui Xing, Haoming Li, Ruitong Xiao, Jingyuan Xing, Baiji Liu
South China University of Technology, China · Huya Inc., China · Tongyi Fun Team, Alibaba Group, China · Foshan University, China
eess.AS, cs.CL, cs.SD
Submitted: 2026-06-26
Updated: 2026-09-28
Comments: 7 pages, 3 figures, 3 tables; Accepted to IEEE SLT 2026
Code: https://github.com/microsoft/DNS-Challenge
Project page: https://xxh333.github.io/hpro-demo
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 75/100
The gist: Hierarchical Progressive Reward Optimization (HPRO) is a novel framework designed to enhance emotional expressiveness in Large Language Model-based Text-to-Speech (TTS) systems by overcoming
Key concepts
- HD-Emo Codec
- This is a novel differentiable reward model that projects speech tokens into separate subspaces for content and style preferences. It helps resolve information conflict by structurally isolating stylistic optimization from the semantic content of the speech, ensuring both are optimized effectively.
- Progressive Optimization
- HPRO uses a three-stage strategy to bridge the gap between sparse sentence-level rewards and dense frame-level generation. It starts with frame alignment, moves to word refinement, and finally aligns global sentence emotion, ensuring stable convergence and preventing reward hacking through careful annealing of optimization stages.
- LASR Loss
- This is an auto-regressive negative log-likelihood loss used to enforce strict semantic alignment. It supervises the content extraction by ensuring the latent representations match the target text tokens, while a stop-gradient mechanism prevents acoustic leakage from corrupting this content supervision.
Terminology
Summary
Hierarchical Progressive Reward Optimization (HPRO) is a novel framework designed to enhance emotional expressiveness in Large Language Model-based Text-to-Speech (TTS) systems by overcoming structural mismatches between content and emotion optimization, and between sparse rewards and dense generation. The core finding is that HPRO significantly improves fine-grained emotional expressiveness while effectively preserving linguistic intelligibility through a hierarchical progressive optimization strategy guided by a novel differentiable reward model.
The Gist
HPRO proposes a hierarchical progressive reward optimization framework that introduces the HD-Emo codec as a differentiable reward model to resolve information conflict, and bridges the scale gap by progressively aligning frame-, word-, and sentence-level objectives, leading to enhanced emotional expressiveness while effectively preserving linguistic intelligibility.
How it works: Resolving Information Conflict via HD-Emo Codec
The framework first addresses the Information Conflict,
where semantic content and emotion in a shared latent space produce conflicting gradients. To resolve this, HPRO introduces the HD-Emo codec as a novel differentiable reward model that projects speech tokens into distinct content and style preference subspaces. This mechanism structurally isolates stylistic optimization from semantic content. Specifically:
-
The codec processes discrete speech tokens through
dual preference extractors with FSQ bottlenecks
to obtaincontent and style preference tokens Tc and Ts.
-
To ensure strict semantic alignment, the latent representations prior to quantization are fed into a content adapter supervised by an ASR objective, formulated as the auto-regressive negative log-likelihood of target text tokens: "LASR = −∑j=1 log P(yj y<j, Tc) (1).
Crucially, a
stop-gradient mechanism" is applied to prevent acoustic leakage from influencing content extraction. -
The style tokens are optimized via hierarchical supervision: at the sentence level, a pre-trained emotion2vec model supervises the predicted distribution pˆ from the emotion decoder through a CE loss (Equation 2). For fine-grained control, word-level constraints use metrics like the Concordance Correlation Coefficient (CCC) to enforce consistency across Valence–Arousal–Dominance dimensions.
How it works: Bridging the Scale Gap via Progressive Optimization
To bridge the Scale Gap
between sparse sentence-level rewards and dense frame-level generation, HPRO employs a progressive optimization strategy that constructs a continuous gradient bridge.
This is achieved by organizing supervision at three distinct levels:
-
Frame-level reward: This involves aligning generated pre-quantization representations (Zˆc, Zˆs) directly with ground-truth discrete preference tokens (Tc, Ts) using L1 regression losses:
Lcp = ∥Zˆc − Tc∥1, Lsp = ∥Zˆs − Ts∥1 (6).
-
Word-level reward: This level incorporates boundary-aware constraints to refine local emotional trajectories. It utilizes a
wVAD CCC loss LwV AD
to match predicted wVAD trajectories with target signals, alongside the ASR loss (LASR) for semantic consistency. -
Sentence-level reward: This ensures global affective alignment by aligning the predicted emotion distribution with target soft labels using a CE loss (LSER).
How it works: Progressive Optimization Strategy
The framework employs a three-stage progressive optimization strategy to ensure stable convergence and prevent reward hacking.
The optimization is structured as follows:
-
Stage I (Frame-level warm-up): Initially, only frame-level preference alignment losses (Lcp and Lsp) and the KL regularization term are optimized, with initial weights set to λKL = 0.05, λsp = 2, and λcp = 1. The Gumbel temperature is initialized to τ = 2.
-
Stage II (Word-level refinement): Word-level supervision through LwV AD and LASR is introduced to refine local trajectories while preserving semantics, with weights adjusted to λKL = 0.02, λsp = 2, λcp = 1, λASR = 5, and λwVAD = 1. The Gumbel temperature is annealed to τ = 1.
-
Stage III (Sentence-level alignment): Finally, the sentence-level emotion classification loss (LSER) is incorporated with λSER = 0.5 and the temperature annealed to τ = 0.8, unifying the global affective style under comprehensive hierarchical supervision.
How it works: Experimental Validation
Experiments utilize LibriSpeech for ASR supervision and LSSED/EmoVoice-DB for emotional modeling, comparing HPRO against baselines like CosyVoice2 and HD-PPT. The objective metrics evaluated include MOSN (naturalness), MOS-E (consistency), WER, wVAD-CCC, EMO-SIM, and DNSMOS. The results demonstrate that HPRO achieves the highest MOSN and the second-highest MOS-E in subjective evaluation.
Improvements for AI systems
As a diligent AI researcher, I have analyzed the HPRO framework presented in this paper. The core innovation lies in its hierarchical, progressive reward optimization designed specifically to solve the information conflict
and scale gap
challenges inherent in preference-driven emotional Text-to-Speech (TTS).
Here are the specific improvements you can implement by adopting or extending the HPRO methodology, and what these improved AI systems can achieve:
) Hierarchical Content/Style Decoupling via HD-Emo Codec
The system should be upgraded from monolithic latent spaces to a structured preference space using the HD-Emo codec.
-
A transformer backbone should be augmented with a dual preference extractor (Content Token Extractor and Style Token Extractor).
-
Implement Finite Scalar Quantization (FSQ) to map continuous speech representations into discrete, separable tokens: Content Tokens (for semantic accuracy) and Style Tokens (for prosodic/emotional features).
-
The system can achieve high fidelity in emotional synthesis while strictly guaranteeing linguistic intelligibility.
) Multi-Scale Gradient Bridging via Progressive Optimization
Instead of relying on single-scale rewards, the training objective must be structured across three distinct levels: Frame, Word, and Sentence.
-
Implement a progressive optimization schedule: Start by optimizing frame-level alignment (dense acoustic structure), then refine with word-level constraints (temporal alignment + semantic integrity), and finally unify with sentence-level affective consistency.
-
This allows the model to build a robust acoustic foundation before imposing global emotional constraints, effectively preventing early reward hacking.
-
The system can generate speech where emotional nuances are captured at both the micro (word boundary) and macro (sentence context) levels, leading to highly coherent and expressive output.
) Fine-Grained Emotional Control via Hierarchical Loss Functions
The loss function must be explicitly multi-faceted, ensuring that different aspects of emotion are optimized independently yet coherently.
-
Incorporate Content Loss: Use an ASR objective (LASR with stop-gradient mechanism) to strictly anchor the content tokens to the ground truth text, preventing semantic degradation.
-
Incorporate Style Loss: Use hierarchical objectives—a global emotion distribution loss (LSER) for sentence context, and a fine-grained Valence–Arousal–Dominance (wVAD) constraint loss for word-level prosody alignment via CCC.
-
The system can produce speech that is not only emotionally rich but also precisely controlled in terms of specific acoustic parameters like intensity, energy, and pacing.
) Robustness Against Reward Hacking via Structural Isolation
The HD-Emo codec ensures that the optimization targets are structurally isolated, mitigating the risk of optimizing for superficial emotional markers at the expense of linguistic quality.
-
By decoupling content tokens from style tokens in the latent space, any gradient pushing for higher emotion is constrained to only affect style features, leaving content features largely protected by ASR supervision.
-
The system can maintain high semantic accuracy (low WER) even when maximizing complex emotional scores, a critical capability that monolithic reward systems fail to provide.
This improved AI system can perform:
-
Generate highly expressive Text-to-Speech (TTS) output that captures complex human emotions with high perceptual quality (MOSN and MOS-E maximization).
-
Produce speech with superior linguistic intelligibility (low WER), ensuring that emotional modulation does not corrupt the underlying text meaning.
-
Achieve fine-grained prosodic control, allowing for precise manipulation of acoustic features like word boundary emphasis and sentence-level affective tone, surpassing the capabilities of stereotyped emotion models (like IndexTTS2).
-
Operate reliably in complex optimization landscapes, avoiding reward hacking by using a structured, progressive path that balances dense local acoustic alignment with global affective consistency.
Sources
- Seed-TTS: A Family of High-Quality Versatile Speech Generation Models
- FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot
- Qwen3-TTS Technical Report
- Emo-LiPO: Listwise Preference Optimization for Fine-Grained Emotion Intensity Control in LLM-based Text-to-Speech
- RLAIF-SPA: Structured AI Feedback for Semantic-Prosodic Alignment in Speech Synthesis
- GLM-TTS Technical Report
- CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models
- Qwen2.5 Technical Report
- CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training
Related papers
- X-VC: Zero-shot Streaming Voice Conversion in Codec Space
- Autoregressive Guidance of Deep Spatially Selective Filters using Bayesian Tracking for Efficient Extraction of Moving Speakers
- Anonymization, Not Elimination: Utility-Preserved Speech Anonymization
- Towards Audio Token Compression in Large Audio Language Models
- WaveScat: Wavelet Scattering Front-Ends with Self-Supervised Features for Speech Deepfake Detection
- ProPS: Prompted Profile Synthesis for Natural Language-Conditioned Speaker Embedding Distributions