Proactive for Uncertainty: Cause-Aware Error Diagnosis and Interactive Clarification for Spoken Dialogue Systems
cs.CL, eess.AS
Submitted: 2026-05-25
Updated: 2026-09-24
Code: https://github.com/tatsu-lab/stanford_alpaca
Terminology
Sources
- Asking Clarifying Questions in Open-Domain Information-Seeking Conversations
- FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs
- GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
- Qwen2-Audio Technical Report
- Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models
- Moshi: a speech-text foundation model for real-time dialogue
- Resolving Intent Ambiguities by Retrieving Discriminative Clarifying Questions
- CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training
- Towards ASR Robust Spoken Language Understanding Through In-Context Learning With Word Confusion Networks
- SPGISpeech 2.0: Transcribed multi-speaker financial audio for speaker-tagged transcription
- Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction
- X-Talk: On the Underestimated Potential of Modular Speech-to-Speech Dialogue System
- ASR Error Correction using Large Language Models
- BLSTM-Based Confidence Estimation for End-to-End Speech Recognition
- MUSAN: A Music, Speech, and Noise Corpus
- SALMONN: Towards Generic Hearing Abilities for Large Language Models
- Word-level confidence estimation for RNN transducers
- Step-Audio 2 Technical Report
- Qwen3-Omni Technical Report
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering