RelayS2S: A Dual-Path Speculative Generation for Real-Time Dialogue
cs.AI
Submitted: 2026-03-24
Updated: 2026-09-09
Comments: EMNLP 2026 Findings
Code: https://github.com/mailong25/relays2s
License: http://creativecommons.org/licenses/by/4.0/
The gist: Real-time spoken dialogue systems face a fundamental tension between latency and response quality.
Terminology
Abstract
Real-time spoken dialogue systems face a fundamental tension between latency and response quality. End-to-end speech-to-speech (S2S) models respond immediately and naturally handle turn-taking, backchanneling, and interruption, but produce semantically weaker outputs. Cascaded pipelines (ASR -> LLM) deliver stronger responses at the cost of latency that grows with model size. We present RelayS2S, a hybrid architecture that runs two paths in parallel upon turn detection. The fast path - a duplex S2S model - speculatively drafts a short response prefix that is streamed immediately to TTS for low-latency response onset, while continuing to monitor live audio events. The slow path - a cascaded ASR -> LLM pipeline - generates a higher-quality continuation conditioned on the committed prefix, producing an uninterrupted utterance. A lightweight learned verifier gates the handoff, committing the prefix when appropriate or falling back gracefully to the cascaded pipeline. With GPT-4.1 as the back-end, RelayS2S substantially reduces response latency while preserving nearly all of the cascaded pipeline's textual quality. On synthetic voice dialogues, it achieves a P90 first-chunk latency of 81 ms, excluding TTS and network latency, compared with 1,006 ms for the cascaded baseline. On real voice dialogues, RelayS2S reduces average first-chunk latency by 479 ms while retaining 99% of the cascaded pipeline's textual quality. These benefits become larger as the slow-path model scales. Because the prefix handoff requires no architectural modification to either component, RelayS2S serves as a lightweight, drop-in addition to existing cascaded pipelines. Our code is publicly available at: https://github.com/mailong25/relays2s
Sources
- SpeakStream: Streaming Text-to-Speech with Interleaved Data
- FireRedChat: A Pluggable, Full-Duplex Voice Interaction System with Cascaded and Semi-Cascaded Implementations
- Moshi: a speech-text foundation model for real-time dialogue
- CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models
- Topical-Chat: Towards Knowledge-Grounded Open-Domain Conversations
- SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model
- GPT-4o System Card
- KAME: Tandem Architecture for Enhancing Knowledge in Real-Time Speech-to-Speech Conversational AI
- ChipChat: Low-Latency Cascaded Conversational Agent in MLX
- Discourse-Aware Dual-Track Streaming Response for Low-Latency Spoken Dialogue Systems
- X-Talk: On the Underestimated Potential of Modular Speech-to-Speech Dialogue System
- VoxCeleb: a large-scale speaker identification dataset
- PersonaPlex: Voice and Role Control for Full Duplex Conversational Speech Models
- AudioPaLM: A Large Language Model That Can Speak and Listen
- Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM
- Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming
- LTS-VoiceAgent: A Listen-Think-Speak Framework for Efficient Streaming Voice Interaction via Semantic Triggering and Incremental Reasoning
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection