Distributional Metrics for Evaluating Spoken Conversational Systems
eess.AS, cs.CL, cs.SD
Submitted: 2026-09-21
Updated: 2026-09-21
Code: https://github.com/snakers4/silero-vad
Terminology
Sources
- A Survey on Neural Speech Synthesis
- M3-DuplexBench: A Multi-Turn, Multilingual, Multidomain Benchmark for Full-Duplex Spoken Dialogue Models
- SPEARBench: A Benchmark for Naturalness Evaluation in Streaming Speech-to-Speech Language Models
- Moshi: a speech-text foundation model for real-time dialogue
- Reference-Based Prosody and Rhythm Evaluation for Spoken Dialogue Systems
Related papers
- X-VC: Zero-shot Streaming Voice Conversion in Codec Space
- Autoregressive Guidance of Deep Spatially Selective Filters using Bayesian Tracking for Efficient Extraction of Moving Speakers
- Anonymization, Not Elimination: Utility-Preserved Speech Anonymization
- Towards Audio Token Compression in Large Audio Language Models
- WaveScat: Wavelet Scattering Front-Ends with Self-Supervised Features for Speech Deepfake Detection
- ProPS: Prompted Profile Synthesis for Natural Language-Conditioned Speaker Embedding Distributions