SpokenUS: A Spoken User Simulator for Task-Oriented Dialogue
cs.CL
Submitted: 2026-03-17
Updated: 2026-09-01
Comments: Accepted at EMNLP 2026 (Main)
Code: https://github.com/holi-lab/SpokenUS
License: http://creativecommons.org/licenses/by/4.0/
The gist: Robust voice agents require exposure to the full diversity of how people interact through speech.
Terminology
Abstract
Robust voice agents require exposure to the full diversity of how people interact through speech. However, obtaining enough spoken interactions is prohibitively expensive. Building spoken user simulators that address this requires large-scale spoken task-oriented dialogue (TOD) data encompassing spoken user behaviors, yet existing datasets are limited in scale and domain coverage, with no systematic pipeline for augmenting them. To address this, we introduce SpokenTOD, a spoken TOD dataset of 52,390 dialogues and 1,034 hours of speech augmented with four spoken user behaviors---cross-turn slots, barge-in, disfluency, and emotional prosody---across diverse speakers and domains. Building on SpokenTOD, we present SpokenUS, a spoken user simulator grounded in TOD that decides when to speak through a dedicated turn-taking head. SpokenUS achieves comparable goal coverage to much larger models while substantially outperforming all baselines in human MOS, disclosing slot values gradually across the dialogue as humans do rather than front-loading them. Further analysis confirms that SpokenUS's spoken behaviors pose meaningful challenges to voice agents, making it a practical tool for evaluating more robust spoken dialogue systems. Our code is available at https://github.com/holi-lab/SpokenUS.
Sources
- CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training
- Qwen3-TTS Technical Report
- Moshi: a speech-text foundation model for real-time dialogue
- VocalBench-DF: A Benchmark for Evaluating Speech LLM Robustness to Disfluency
- Qwen2.5 Technical Report
- Qwen2.5-Omni Technical Report
- Qwen3-Omni Technical Report
- Qwen3 Technical Report
- GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
- InteractiveOmni: A Unified Omni-modal Model for Audio-Visual Multi-turn Dialogue
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering