Cocktail-Talker: Multi-Speaker Dialog Modeling in Noisy Social Environments with Turn Action GRPO
Xilin Jiang, Riki Shimizu, Sukru Samet Dindar, Junkai Wu, Zhongweiyang Xu, Nima Mesgarani
cs.SD, cs.CL, cs.MM, eess.AS
Submitted: 2026-07-30
Code: https://github.com/xi-j/Cocktail-Talker
License: http://creativecommons.org/licenses/by/4.0/
The gist: Spoken dialog systems are typically designed for clean, dyadic interactions in which a single user and an assistant take turns speaking.
Terminology
Abstract
Spoken dialog systems are typically designed for clean, dyadic interactions in which a single user and an assistant take turns speaking. Real-world social conversations, however, are often more ambiguous: multiple speakers may participate in the same conversation amid irrelevant speech and background noise. Each utterance may be directed to the assistant, addressed to another speaker, or completely irrelevant. In such settings, the assistant must decide not only what to say, but also whether to speak at all. In this paper, we introduce Cocktail-Talker, a speech LLM framework for multi-speaker spoken dialog modeling in noisy social environments. We model the assistant's behavior with three action tokens: <respond>, <listen>, and <ignore>, placed before a response or silence. Cocktail-Talker is trained via supervised finetuning and reinforcement learning to generate the appropriate action token and, only in <respond> mode, a speech response. To prepare the training data, we develop Cocktail-DialogGen, an LLM-based data pipeline that simulates realistic multi-speaker dialogs with speaker roles across diverse social settings. Together, these components take a step toward spoken dialog systems that interact more naturally and selectively in complex social environments.
Sources
- Moshi: a speech-text foundation model for real-time dialogue
- PersonaPlex: Voice and Role Control for Full Duplex Conversational Speech Models
- Qwen2.5-Omni Technical Report
- Qwen3-Omni Technical Report
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming
- LLaMA-Omni: Seamless Speech Interaction with Large Language Models
- Qwen2-Audio Technical Report
- Mini-Omni2: Towards Open-source GPT-4o with Vision, Speech and Duplex Capabilities
- AVMeme Exam: A Multimodal Multilingual Multicultural Benchmark for LLMs' Contextual and Cultural Knowledge and Thinking
- MSU-Bench: Towards Understanding the Conversational Multi-talker Scenarios
- M3-SLU: Evaluating Speaker-Attributed Reasoning in Multimodal Large Language Models
- Full-Duplex-Bench: A Benchmark to Evaluate Full-duplex Spoken Dialogue Models on Turn-taking Capabilities
- MeanFlow-TSE: One-Step Generative Target Speaker Extraction with Mean Flow
- Focus Then Listen: An Empirical Study of Plug-and-Play Audio Enhancer for Noise-Robust Large Audio Language Models
- LibriMix: An Open-Source Dataset for Generalizable Speech Separation
- Speak or Stay Silent: Context-Aware Turn-Taking in Multi-Party Dialogue
- Qwen3-TTS Technical Report
- Step-Audio 2 Technical Report
- Kimi-Audio Technical Report
Related papers
- Few-Shot Open-Set Audio Classification via Transductive Prototype Refinement and Class Logit Enhancement
- Spectral Masking and Interpolation Attack (SMIA): A Black-box Adversarial Attack against Voice Authentication and Anti-Spoofing Systems
- AVMeme Exam: A Multimodal Multilingual Multicultural Benchmark for LLMs' Contextual and Cultural Knowledge and Thinking
- SoundWeaver: Compositional Warm-Starting for Text-to-Audio Diffusion Serving
- WASIL: In-the-Wild Arabic Spoken Interactions with LLMs
- Efficient Audiovisual Speech Processing via MUTUD: Multimodal Training and Unimodal Deployment