Sequential Trajectories and Simultaneous Blending: Multi-Emotion Modeling for Instruction-Following TTS
cs.CL, cs.SD
Submitted: 2026-08-31
Updated: 2026-08-31
Code: https://github.com/ictnlp/HybridEmo
Terminology
Sources
- FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs
- Seed-TTS: A Family of High-Quality Versatile Speech Generation Models
- FlexiVoice: Enabling Flexible Style Control in Zero-Shot TTS with Natural Language Instructions
- CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training
- CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models
- Qwen3-TTS Technical Report
- Group Relative Policy Optimization for Text-to-Speech with Large Language Models
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- F5R-TTS: Improving Flow-Matching based Text-to-Speech with Group Relative Policy Optimization
- Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
- Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
- Qwen3-Omni Technical Report
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering