SteerablePlex: Can We Steer Full-Duplex Models?
eess.AS, cs.CL
Submitted: 2026-10-08
Updated: 2026-10-08
Code: https://github.com/snakers4/silero-vad
Terminology
Sources
- Moshi: a speech-text foundation model for real-time dialogue
- Qwen3 Technical Report
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Optimizing Conversational Quality in Spoken Dialogue Systems with Reinforcement Learning from AI Feedback
- Decoupling Conversational Dynamics in Full-Duplex Spoken Models through Reinforcement Learning
- Raon-Speech Technical Report
Related papers
- X-VC: Zero-shot Streaming Voice Conversion in Codec Space
- Autoregressive Guidance of Deep Spatially Selective Filters using Bayesian Tracking for Efficient Extraction of Moving Speakers
- Anonymization, Not Elimination: Utility-Preserved Speech Anonymization
- Towards Audio Token Compression in Large Audio Language Models
- WaveScat: Wavelet Scattering Front-Ends with Self-Supervised Features for Speech Deepfake Detection
- ProPS: Prompted Profile Synthesis for Natural Language-Conditioned Speaker Embedding Distributions