Liberating LLM Capabilities in Full-Duplex Speech Models
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Liberating LLM Capabilities in Full-Duplex Speech Models".
Jane: The paper was written by Luoyuan Zhang, Bokai Xu, Junbo Cui, Weiyue Sun, Yingjing Xu et al. from.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: We are starting our look at 'Liberating LLM Capabilities in Full-Duplex Speech Models' by Luoyuan Zhang and his colleagues. This paper addresses a massive gap in how we talk to AI, where models are often stuck only being able to speak back to us.
Jane: It sounds like they want to break those chains, Tom, because right now, if you ask a voice assistant for code or a table, it just reads it out loud like a robot.
Tom: That is exactly the frustration I feel every time I try to use these tools for anything productive.
Lu: The authors are quite clever to use the word 'liberating' in the title because they see text-native intelligence as something currently being suppressed by the speech modality. They want to give these models back their ability to show us things visually while they talk.
Jane: So, instead of just a voice, we get a full collaborator?
Lu: Precisely, Jane, and it is more than just adding a screen; it is about making sure the model uses its full brain. Most current systems treat text as a hidden thought process that never reaches the user.
Meng: I am looking at the technical approach they mention for this liberation, and I wonder if this requires a total overhaul of our current hardware setups.
Tom: That is a fair concern, Meng, but it seems they might have found a way to do this without making everything more expensive or complicated.
Meng: If they can keep the underlying architecture standard, then engineers like me can actually deploy this in real products without needing a supercomputer in every pocket.
Lalam: This shift toward multi-channel interaction will fundamentally change our relationship with digital entities. We are moving away from simple tools and toward partners that can communicate through both sound and sight simultaneously.
Jane: It feels like we are finally getting a glimpse of what true multimodal intelligence looks like in practice.
Tom: We will explore the specific mechanics of how they actually achieve this in our next segment.
Summary: Jane: Now that we know the goal, let's look at the methodology which they call Listen-Write-Speak, or LWS. It is a tri-channel system where the model listens, writes, and speaks all at once.
Tom: That sounds like a lot for one model to juggle without getting confused or lagging behind.
Jane: It does sound intense, but they have organized it into these one-second units to keep everything synchronized.
Lu: The most brilliant part is how the writing channel stays active even while you are still talking to the AI. You can see the model's thoughts or notes appearing on your screen before it even says a single word in response.
Tom: Wait, so I could be mid-sentence and see the model's summary of my request pop up?
Lu: Yes, and that incremental writing helps build a shared understanding between you and the machine. It acts like an external working memory that we can all see.
Meng: I was reading about how they actually implement this, and it is surprisingly elegant because they use a Token Schema.
Jane: Does that mean they do not have to redesign the actual brain of the model?
Meng: They really don't, which is a huge win for scalability. They just use special tokens in a standard Transformer to tell the model when to listen, write, or speak. It is all handled within one shared causal context.
Lalam: This transparency through visible writing will bridge the gap between human intent and machine execution. When we can see the reasoning process unfolding, we can trust the interaction much more deeply.
Tom: It seems like they have found a way to let the AI show its work while it talks, which is something we have always wanted.
Jane: We should probably check if this actually works as well as it sounds in their tests.
Improvements: Tom: Moving into the results, we need to see if this Listen-Write-Speak method actually lives up to the hype. The researchers tested it on several benchmarks, and the performance is quite striking.
Jane: They reached a score of four point seven two on the VoiceBench AlpacaEval, which is incredibly high for these types of speech-to-text evaluations.
Tom: That puts them right up there with the most advanced models we have seen lately.
Meng: I was particularly interested in the latency and how they handled real-world interruptions. They achieved a turn-taking latency of about zero point four eight seconds, which is very snappy for a system doing this much work.
Jane: And they also managed to keep the written and spoken parts consistent, didn't they?
Meng: They did, with a consistency rate of ninety-two point six percent between what is written and what is said. For an engineer, that level of alignment is vital because you don't want the AI telling you one thing while the screen shows something completely different.
Lu: Their performance on URO-Bench was also very impressive, especially regarding complex reasoning tasks in both Chinese and English. It shows that adding the writing channel doesn't hurt the model's ability to think deeply.
Tom: What happens if I suddenly interrupt the model while it is responding?
Lu: The system handles those interruptions naturally because the listening channel never stops receiving audio, even during a reply. It can pivot almost instantly without losing its place.
Lalam: This ability to handle dynamic, unpredictable human behavior makes the AI feel much more like a natural participant in a conversation. It moves us away from rigid commands and toward fluid social interaction.
Jane: It sounds like they have successfully turned a one-way street into a two-way highway.
Tom: Let's wrap this all up and think about what it means for the future.
Conclusion: Tom: We have covered so much ground today regarding 'Liberating LLM Capabilities in Full-Duplex Speech Models' by Luoyuan Zhang and his team. It is clear that this tri-channel approach changes the game for how we interact with machines.
Jane: I am still thinking about that idea of seeing the AI's thoughts on screen while it talks to us. It makes the whole experience feel so much more collaborative and less like a black box.
Tom: It really does, Jane, and it opens up doors for everything from coding to scientific research.
Lu: I can see this being used in classrooms or even in high-stakes professional environments where you need to see the math or the code as it is being discussed. The possibilities for remote collaboration are endless.
Meng: From my side, the fact that this works within standard architectures makes it feel very ready for real-world deployment. We can start building these types of interfaces much sooner than I thought.
Lalam: This advance will help integrate AI into our culture as a tool that supports both our voices and our visual thoughts. It allows digital intelligence to exist in the same multi-sensory space that humans do.
Jane: That is a beautiful way to put it, Lalam, and it really summarizes why this research is so important.
Tom: We have certainly learned a lot about how visible writing can change the landscape of speech models.
Jane: It has been such a blast discussing this with all of you, but we have to wrap up the show now.
Tom: Thank you all for joining us to break down this incredible research on 'Liberating LLM Capabilities in Full-Duplex Speech Models'. We will see you next time!
Jane: Goodbye everyone!
cs.CL, cs.AI, cs.SD
Submitted: 2026-05-04
Updated: 2026-09-15
Code: https://github.com/tatsu-lab/alpaca
Project page: https://royalzhang.com/project/lws-page
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 84/100
The gist: This paper presents Listen-Write-Speak (LWS), a "text-first tri-channel paradigm" designed to overcome the inherent limitations of speech-based large language models.
Key concepts
- Listen-Write-Speak (LWS)
- A tri-channel methodology where the model listens to audio, writes text or visual data to a screen, and speaks simultaneously. By organizing these actions into one-second units, the system maintains synchronization and allows users to see the model's reasoning or notes in real-time.
- Token Schema
- A method using special tokens within a standard Transformer architecture to direct the model's actions. This allows the AI to manage listening, writing, and speaking within one shared context, making the technology scalable and easy to deploy on existing hardware without a total overhaul.
- Full-Duplex Interaction
- A communication style that allows for simultaneous two-way interaction. Because the listening channel remains active even while the model is speaking, the system can handle natural human interruptions and pivot instantly, making conversations feel more fluid and less like rigid command-and-response exchanges.
Terminology
Summary
This paper presents Listen-Write-Speak (LWS), a text-first tri-channel paradigm
designed to overcome the inherent limitations of speech-based large language models. By enabling an autoregressive LLM to simultaneously listen, write visible text, and speak in real-time, LWS liberates text-native capabilities such as code generation, structured analysis, and multi-step reasoning
that are typically suppressed when models are restricted to spoken outputs.
The speech interaction bottleneck
Current speech-based LLMs are often constrained to spoken replies, which limits their user-facing outputs to what can be verbalized.
This forces text-native capabilities into a much narrower subspace of speakable language,
making it difficult for models to provide precise, structured, or spatially organized information. While existing work has improved reasoning or turn-taking, most approaches still treat text as a hidden intermediate state
or a subordinate modality
rather than a first-class output channel. Consequently, when a user requests something like executable code or a Markdown table, the model is forced to dictate it aloud rather than providing an editable, visible representation.
The LWS architecture and token schema
LWS implements its tri-channel behavior within a standard autoregressive Transformer using a Token Schema
that requires no architectural modifications.
The interaction is partitioned into discrete temporal Units,
where three channels operate concurrently under a shared causal attention context:
-
The Listening (Perception) Channel: Continuously ingests the user’s audio stream to enable
full-duplex interaction.
-
The Visible Writing (Cognition) Channel: Always active, externalizing the model’s text-native cognition as
unconstrained free-form text
during both listening and speaking. -
The Speaking (Expression) Channel: Generates natural spoken language as a real-time oral response.
To prevent information leakage and simplify next-token prediction, the authors separate the cognitive phases into Listening Units (which build incremental understanding) and Speaking Units (which produce richer structured artifacts). This distinction ensures that speaking generation cannot leak information from a listening-phase writing segment
that has not yet concluded.
A two-stage data pipeline
Because public corpora lack per-second cognitive annotations aligned to an audio timeline,
the authors developed a specialized two-stage pipeline:
-
Stage 1 (Offline cognitive synthesis): A strong teacher LLM synthesizes three distinct text streams—a streaming reasoning chain, a concise voice response, and the original structured response—while adhering to a
strict causal constraint
to ensure reasoning is consistent with the progressively revealed input timeline. -
Stage 2 (Online timeline construction): This stage combines these synthesized streams with real audio recordings and CTC-based character-level alignments. It also incorporates
interruption augmentation
to train the model on how to handle overlapping speech and take-over behavior.
Empirical performance and consistency
LWS demonstrates strong results across several benchmarks, achieving a 4.72 score on VoiceBench AlpacaEval and outperforming open-source baselines. The model achieves 92.6% writing–speaking consistency,
indicating that the spoken response remains well-aligned with the visible written output. On URO-Bench, LWS consistently outperforms its internal ablations, proving that visible writing can serve as a first-class output channel
without sacrificing responsiveness. In full-duplex evaluations on Full-Duplex-Bench, the system manages interruptions and turn-taking effectively, maintaining a 4.02 GPT-4o quality
under user interruption while sustaining real-time interaction with minimal overhead.
Improvements for AI systems
Improvement 1: Transition from single-modality output to a Tri-Channel Listen-Write-Speak (LWS) Token Schema.
- What the improved system can do: Instead of forcing all reasoning into a linear spoken stream, the AI can simultaneously ingest audio, generate visible free-form text (such as executable Python code, Markdown tables, or mathematical derivations), and emit natural speech. This allows for
text-native
capabilities—like providing a structured summary in a table while verbally explaining it—without requiring complex new architectural decoders or cross-modal fusion modules.
Improvement 2: Integration of Dual-Phase Cognitive Writing (ls cogn and reply cogn) within a shared causal context.
- What the improved system can do: The system can perform
cognition while listening
by displaying incremental reasoning or partial understandings on-screen while the user is still speaking. During the response phase, it can performcognition while speaking,
where it generates high-fidelity written artifacts (like a complex code block) in parallel with its oral explanation, ensuring the visible output is more detailed and structured than the spoken version.
Improvement 3: Implementation of a Two-Stage Causal Data Pipeline for per-second temporal alignment.
- What the improved system can do: By training on synthesized, second-by-second annotations that are strictly causally constrained to the input timeline, the system can achieve high-performance full-duplex interaction. This enables natural interruption handling, seamless turn-taking, and low-latency response (sub-0.5s) by allowing the model to build a response plan incrementally as audio tokens arrive, rather than waiting for a complete turn to finish.
Abstract
Speech-based large language models are typically constrained to spoken replies, which limits their user-facing outputs to what can be verbalized and suppresses text-native capabilities such as code generation, structured analysis, and multi-step reasoning in realtime interaction, for tasks that require persistent, structured, and inspectable intermediate outputs. Existing work improves spoken reasoning or full-duplex turn-taking, but still treats text as a hidden intermediate state or a subordinate modality rather than a first-class output channel. We propose Listen-Write-Speak (LWS), a text-first tri-channel paradigm in which a single autoregressive LLM continuously listens to user audio, writes visible free-form text as its primary output, and speaks a realtime oral response in parallel under a shared causal attention context. This behavior is implemented entirely through a Token Schema, requiring no architectural modifications, and learned via a two-stage data pipeline that synthesizes per-second cognitive annotations consistent with the revealed input timeline. Empirically, LWS demonstrates strong full-duplex interaction on Full-Duplex-Bench, reaches 4.72 on VoiceBench AlpacaEval, achieves 92.6% writing-speaking consistency, and consistently outperforms its internal ablations on URO-Bench. These results suggest that visible writing can serve as a first-class output channel for speech interaction without sacrificing realtime responsiveness. The code and dataset are available on the project page: https://royalzhang.com/project/lws-page/.
Sources
- VoiceBench: Benchmarking LLM-Based Voice Assistants
- SHANKS: Simultaneous Hearing and Thinking for Spoken Language Models
- STITCH: Simultaneous Thinking and Talking with Chunked Reasoning for Spoken Language Models
- Moshi: a speech-text foundation model for real-time dialogue
- Kimi-Audio Technical Report
- CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models
- FlexDuo: A Pluggable System for Enabling Full-Duplex Capabilities in Speech Dialogue Systems
- Can Speech LLMs Think while Listening?
- DrVoice: Parallel Speech-Text Voice Conversation Model via Dual-Resolution Speech Representations
- Step-Audio-R1 Technical Report
- Step-Audio 2 Technical Report
- Chronological Thinking in Full-Duplex Spoken Dialogue Language Models
- Mind-Paced Speaking: A Dual-Brain Approach to Real-Time Reasoning in Spoken Language Models
- Mini-Omni-Reasoner: Token-Level Thinking-in-Speaking in Large Speech Models
- Qwen3-Omni Technical Report
- URO-Bench: Towards Comprehensive Evaluation for End-to-End Spoken Dialogue Models
- Qwen3 Technical Report
- MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering