UltraVoice: Scaling Fine-Grained Style-Controlled Speech Conversations for Spoken Dialogue Models
eess.AS, cs.CL
Submitted: 2025-10-26
Updated: 2026-08-29
Code: https://github.com/bigai-nlco/UltraVoice
Terminology
Sources
- FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs
- Seed-TTS: A Family of High-Quality Versatile Speech Generation Models
- SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training
- Audio-Aware Large Language Models as Judges for Speaking Styles
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- Recent Advances in Speech Language Models: A Survey
- Moshi: a speech-text foundation model for real-time dialogue
- Enhancing Chat Language Models by Scaling High-quality Instructional Conversations
- CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens
- CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models
- CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training
- LLaMA-Omni: Seamless Speech Interaction with Large Language Models
- LLaMA-Omni2: LLM-based Real-time Spoken Chatbot with Autoregressive Streaming Speech Synthesis
- OSUM-EChat: Enhancing End-to-End Empathetic Spoken Chatbot via Understanding-Driven Spoken Dialogue
- Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction
- InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems
- GPT-4o System Card
- WavChat: A Survey of Spoken Dialogue Models
- emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation
- UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022
Related papers
- X-VC: Zero-shot Streaming Voice Conversion in Codec Space
- Autoregressive Guidance of Deep Spatially Selective Filters using Bayesian Tracking for Efficient Extraction of Moving Speakers
- Anonymization, Not Elimination: Utility-Preserved Speech Anonymization
- Towards Audio Token Compression in Large Audio Language Models
- WaveScat: Wavelet Scattering Front-Ends with Self-Supervised Features for Speech Deepfake Detection
- ProPS: Prompted Profile Synthesis for Natural Language-Conditioned Speaker Embedding Distributions