EmoTra-TTS: Smooth Intra-Utterance Emotion Transitions for Speech Synthesis
eess.AS, cs.AI
Submitted: 2026-08-24
Updated: 2026-08-24
Comments: Accepted to EMNLP 2026 (Main Conference)
Code: https://github.com/Liu-Tianchi/EmoTra-TTS
Project page: https://liu-tianchi.github.io/EmoTra_DemoPage
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Understanding intermediate layers using linear classifier probes
- Seed-TTS: A Family of High-Quality Versatile Speech Generation Models
- MINT-Bench: A Comprehensive Multilingual Benchmark for Instruction-Following Text-to-Speech
- GLM-TTS Technical Report
- Qwen3-TTS Technical Report
- InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems
- Scores Know Bobs Voice: Speaker Impersonation Attack
- CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training
- CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models
- ReStyle-TTS: Relative and Continuous Style Control for Zero-Shot Speech Synthesis
- EMORL-TTS: Reinforcement Learning for Fine-Grained Emotion Control in LLM-based TTS
- TED-TTS: Training-Free Intra-Utterance Emotion and Duration Control for Text-to-Speech Synthesis
- Ctrl-Adapter: An Efficient and Versatile Framework for Adapting Diverse Controls to Any Diffusion Model
- CoCoEmo: Composable and Controllable Human-Like Emotional TTS via Activation Steering
- Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
- EmoS: A Theory-Grounded Framework for Evaluating and Aligning Emotional Intelligence in Spoken Language Models
- Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
- Shallow Flow Matching for Coarse-to-Fine Text-to-Speech Synthesis
- Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis
- Emotional Dimension Control in Language Model-Based Text-to-Speech: Spanning a Broad Spectrum of Human Emotions
Related papers
- X-VC: Zero-shot Streaming Voice Conversion in Codec Space
- Autoregressive Guidance of Deep Spatially Selective Filters using Bayesian Tracking for Efficient Extraction of Moving Speakers
- Anonymization, Not Elimination: Utility-Preserved Speech Anonymization
- Towards Audio Token Compression in Large Audio Language Models
- WaveScat: Wavelet Scattering Front-Ends with Self-Supervised Features for Speech Deepfake Detection
- ProPS: Prompted Profile Synthesis for Natural Language-Conditioned Speaker Embedding Distributions