Full-Duplex Speech Models Take the Floor When Asked, Not When Needed
cs.CL, cs.HC
Submitted: 2026-09-17
Updated: 2026-09-17
Comments: 5 pages
Code: https://github.com/vocaliodmiku/take-the-floor
License: http://creativecommons.org/licenses/by/4.0/
The gist: Full-duplex speech models listen and speak at once, promising always-on assistants.
Terminology
Abstract
Full-duplex speech models listen and speak at once, promising always-on assistants. Yet they must also decide when they should speak. Human listeners speak when addressed or when the speaker stops, but also self-select to correct a false claim, supply a missing word, or warn of danger. We ask whether full-duplex models do the same. To separate the reason to speak from the opportunity, we construct context-matched English monologues in which only the trigger utterance varies within a topic, define 10 conditions from turn-allocation rules, and compress inter-word pauses to limit opportunities created by silence. Across five model families, being addressed and silence are far more reliable triggers than false facts or hazards. Frame-level text-token probabilities in Moshi and PersonaPlex are lower for false facts than for Neutral when averaged over the first 2,s after trigger end. Pauses or permission to interrupt do not close this gap either. Given the floor, Moshi and PersonaPlex answer most direct questions, yet the proportion of non-empty false-fact replies that challenge the claim is only.14--.15, and the proportion of hazard replies that warn of danger is.04--.07. This paper thus identifies a gap in both speech initiation and response content. Closing it requires genuine content understanding and intervention decisions grounded in it.
Sources
- Moshi: a speech-text foundation model for real-time dialogue
- PersonaPlex: Voice and Role Control for Full Duplex Conversational Speech Models
- Multi-Faceted Interactivity Alignment in Full-Duplex Speech Models
- Full-Duplex-Bench: A Benchmark to Evaluate Full-duplex Spoken Dialogue Models on Turn-taking Capabilities
- FD-Bench: A Full-Duplex Benchmarking Pipeline Designed for Full Duplex Spoken Dialogue Systems
- Full-Duplex Interaction in Spoken Dialogue Systems: A Comprehensive Study from the ICASSP 2026 HumDial Challenge
- FLEXI: Benchmarking Full-duplex Human-LLM Speech Interaction
- Instruct-FD: Can Your Full-Duplex Speech System Follow Turn-Taking Instructions?
- Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM
- MiniCPM-o 4.5: Towards Real-Time Full-Duplex Omni-Modal Interaction
- Raon-Speech Technical Report
- DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering