Self-EmoQ: Plutchik-Guided Value-based Planning to Drive Streaming Emotional TTS
cs.HC, cs.AI
Submitted: 2026-04-21
Updated: 2026-09-01
Comments: ACL 2026 Findings
Project page: https://sixingdeguo.github.io/EmoQ-page
License: http://creativecommons.org/licenses/by/4.0/
The gist: Emotional interaction is increasingly crucial for conversational AI, yet current systems lack a self-emotion determination mechanism to drive the streaming text-to-speech (TTS) synthesis.
Terminology
Abstract
Emotional interaction is increasingly crucial for conversational AI, yet current systems lack a self-emotion determination mechanism to drive the streaming text-to-speech (TTS) synthesis. We propose an emotion-planning framework that determines the emotion prior to the textual generation, grounding the downstream emotional TTS in a streaming manner. The framework is implemented by a plug-and-play LLM module, initialized from pretrained LLMs, and trained by reinforcement learning (RL) with emotions as the actions. A hybrid reward is employed which combines imitation signals with theory-driven scoring, in which the theory of Plutchik's wheel of emotions is adopted. By experiments on DailyDialog, EmoryNLP, IMEOCAP, and MELD, our method outperforms prompting and finetuning baselines on both emotion determination and response quality. We finally implement an entire streaming pipeline for real-time deployment, with the speech quality confirming the framework's emotional alignment, contextual coherence, and expressive fluency. Codes, cases, and demos are available in https://sixingdeguo.github.io/EmoQ-page/.
Sources
- CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training
- CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models
- Prompt-Unseen-Emotion: Zero-shot Expressive Speech Synthesis with Prompt-LLM Contextual Knowledge for Mixed Emotions
- Making Flow-Matching-Based Zero-Shot Text-to-Speech Laugh as You Like
- InstructERC: Reforming Emotion Recognition in Conversation with Multi-task Retrieval-Augmented Large Language Models
- A Diversity-Promoting Objective Function for Neural Conversation Models
- DailyDialog: A Manually Labelled Multi-turn Dialogue Dataset
- SpeechBERTScore: Reference-Aware Automatic Evaluation of Speech Generation Leveraging NLP Evaluation Metrics
- Enhancing Emotional Generation Capability of Large Language Models via Emotional Chain-of-Thought
- Emotional Cues Extraction and Fusion for Multi-modal Emotion Prediction and Recognition in Conversation
- Word-Level Emotional Expression Control in Zero-Shot Text-to-Speech Synthesis
- Towards Interpretable Mental Health Analysis with Large Language Models
- RLAIF-SPA: Structured AI Feedback for Semantic-Prosodic Alignment in Speech Synthesis
- EmoFSM: A Finite State Machine for Emotional Support Conversation
Related papers
- EduGage: A Multimodal Dataset and Benchmark for Sensor-Based Momentary Assessment of Engagement in Self-Guided Video Learning
- EvoDesign: Agentic Editable Diagram Creation via Design Expertise Evolution
- HAGI++: Head-Assisted Gaze Imputation and Generation
- Linking Behaviour and Perception to Evaluate Meaningful Human Control over Partially Automated Driving
- Review of Explainable Decision Support and Adaptive Human-Machine Interfaces for Automation Transparency in Maritime Autonomous Surface Ships
- Towards Cognitive Process-Aware Proactive Writing Support