Learning Natural Conversational Behavior in Tandem Speech-to-Speech Models with Randomized Guidance
cs.CL, eess.AS
Submitted: 2026-09-25
Updated: 2026-09-25
Terminology
Sources
- Moshi: a speech-text foundation model for real-time dialogue
- MoshiRAG: Asynchronous Knowledge Retrieval for Full-Duplex Speech Language Models
- Thinking While Speaking: Inference-Time Knowledge Transfer for Responsive and Intelligent Conversational Voice Agents
- Multi-Faceted Interactivity Alignment in Full-Duplex Speech Models
- ConversationalVoice: Full-Duplex Speech Data from Real Conversations through Source-Faithful Reconstruction and Conversation-Grounded Expansion
- DuplexChat: Constructing Speaker-Separated Full-Duplex Dialogue Speech at Scale for Spoken Dialogue Language Modeling
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering