Preference Optimization for Non-Verbal Vocalization Synthesis
eess.AS, cs.AI, cs.LG
Submitted: 2026-08-25
Updated: 2026-08-25
Code: https://github.com/resemble-ai/Resemblyzer
Project page: https://nvvspeech-challenge.github.io
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Qwen3-TTS Technical Report
- NonverbalTTS: A Public English Corpus of Text-Aligned Nonverbal Vocalizations with Emotion Annotations for Text-to-Speech
- A Scalable Pipeline for Enabling Non-Verbal Speech Generation and Understanding
- NV-Bench: Benchmark of Nonverbal Vocalization Synthesis for Expressive Text-to-Speech Generation
- NVV-SuperBench: Beyond Words, Beyond Quality-Benchmarking Nonverbal Vocalizations in Speech Generation
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Improving Code-Switching ASR with Code-Mixing Guided Synthetic Speech
- Fish Audio S2 Technical Report
- CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models
- CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training
- FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs
Related papers
- X-VC: Zero-shot Streaming Voice Conversion in Codec Space
- Autoregressive Guidance of Deep Spatially Selective Filters using Bayesian Tracking for Efficient Extraction of Moving Speakers
- Anonymization, Not Elimination: Utility-Preserved Speech Anonymization
- Towards Audio Token Compression in Large Audio Language Models
- WaveScat: Wavelet Scattering Front-Ends with Self-Supervised Features for Speech Deepfake Detection
- ProPS: Prompted Profile Synthesis for Natural Language-Conditioned Speaker Embedding Distributions