COT-TTS: Audio Context-Aware Text-to-Speech with Chain-of-Thought Reasoning
cs.CL, cs.SD, eess.AS
Submitted: 2026-09-19
Updated: 2026-09-25
Comments: 13 pages, 6 figures, 6 tables
Project page: https://luckybian.github.io/COT-TTS
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Fish Audio S2 Technical Report
- VoxCPM: Tokenizer-Free TTS for Context-Aware Speech Generation and True-to-Life Voice Cloning
- Borderless Long Speech Synthesis
- SoulX-Podcast: Towards Realistic Long-form Podcasts with Dialectal and Paralinguistic Diversity
- ActorMind: Emulating Human Actor Reasoning for Speech Role-Playing
- Audio-Reasoner: Improving Reasoning Capability in Large Audio Language Models
- VoxCPM2 Technical Report
- VoiceSculptor: Your Voice, Designed By You
- DuplexCascade: Full-Duplex Speech-to-Speech Dialogue with VAD-Free Cascaded ASR-LLM-TTS Pipeline and Micro-Turn Optimization
- FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs
- CapTalk: Unified Voice Design for Single-Utterance and Dialogue Speech Generation
- Harness TTS: Towards Context-Aware Expressive Speech Synthesis with Harness Layer
- Qwen-Music Technical Report
- Beyond Semantic Dominance: Cognitive Affective Reasoning and Empathetic Response Alignment in Audio Language Models
- OSUM-EChat: Enhancing End-to-End Empathetic Spoken Chatbot via Understanding-Driven Spoken Dialogue
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- OV-MER: Towards Open-Vocabulary Multimodal Emotion Recognition
- Kimi-Audio Technical Report
- BUT System Description to VoxCeleb Speaker Recognition Challenge 2019
- Qwen3-ASR Technical Report
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering