DuplexAct-Bench: Broadening Full-Duplex Speech Evaluation toward Proactive Interaction across Diverse Behavioral Requirements
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "DuplexAct-Bench: Broadening Full-Duplex Speech Evaluation toward Proactive Interaction across Diverse Behavioral Requirements".
Tom: Existing full-duplex speech benchmarks cover only subsets of realtime interaction behaviors, often under limited contextual conditions.
Jane: First, who's behind it and why it matters.
Paper summary: Tom: So, what are the main points from this "DuplexAct-Bench: Broadening Full-Duplex Speech Evaluation toward Proactive Interaction across Diverse Behavioral Requirements" paper? Essentially, the thesis is that existing full-duplex speech benchmarks only cover limited subsets of real-time interaction behaviors under constrained conditions. The authors introduce DuplexAct-Bench to systematically cover six complementary behaviors—interruption, yielding, proactive initiation, active silence, and backchanneling—across three contexts: Pre-session, In-session, and Noexplicit.
Jane: Exactly. What they claim is that this new benchmark allows for a much more complete evaluation of these systems by testing them on both Timing and Content metrics across all these scenarios. They are looking at how well the systems perform when the behavior needs to be established beforehand, introduced during the actual interaction, or if no instruction is given at all.
Lu: The significance they point out is that their results show substantial variation across those behaviors, conditions, and systems when evaluated on this new benchmark. They specifically highlight frequent mismatches between how high the semantic quality of the behavior is versus when it actually happens in time during the interaction.
Meng: That variation sounds important for practical application; if a system can score well on content but consistently misses the required timing for proactive initiation, that means it might fail in a real-world scenario where spontaneity matters. How does this affect deployment decisions?
Lalam: I see it as moving evaluation away from simple pass/fail and toward understanding the nuances of conversational flow. It suggests that we need to look at the entire interaction tapestry rather than just single isolated actions, which is a much richer way to design AI for communication.
Conclusion: Tom: Looking at the title, "DuplexAct-Bench: Broadening Full-Duplex Speech Evaluation toward Proactive Interaction across Diverse Behavioral Requirements," it really tells us where the field needs to go next in testing these systems. The authors are pushing past just checking if a system can react correctly during a conversation and demanding that they handle the whole spectrum of participation, from being interrupted to initiating things on their own.
Jane: They're showing that simply having good content quality isn't enough if the timing is off for those proactive behaviors, which is something they emphasize in their findings across all one thousand two hundred ninety trials. The implication for us is that we can no longer rely on one metric to define success; we have to look at both what the AI says and precisely when it says it.
Lu: I think the impact on future development lies in designing training methods that specifically target these complex behavioral requirements across all three conditions—Pre-session, In-session, and Noexplicit—so that the system learns not just what to say but also how and when to participate appropriately.
Meng: For my team, this suggests we need a testing framework that mirrors this complexity so we can actually pinpoint exactly where our current models are falling short when they encounter ambiguous or spontaneous interaction situations. It gives us a clearer target for improvement in robustness.
Lalam: For me, the bigger picture is that if we can build systems that are robust across all these behavioral requirements, it could lead to AI assistants that feel much more natural and responsive in extended conversations, moving them from just answering questions to being genuine conversational partners.
Keyue Xing, Wentao Ding, Mengmeng Wang, Wenming Tu, Zilong Zheng, Yipeng Kang
State Key Laboratory of General Artificial Intelligence, BIGAI, China
cs.CL, cs.HC
Submitted: 2026-09-30
Updated: 2026-09-30
Code: https://github.com/viitor-ai/viitor-voice
Project page: https://alitaxky.icu/DuplexAct-Bench
Importance score: 83/100
The gist: Existing full-duplex speech benchmarks cover only subsets of realtime interaction behaviors, often under limited contextual conditions.
Key concepts
- Agent Interruption (AI)
- This behavior measures when an agent should interrupt the user's speech. Success depends on timing: the interruption must happen after a specific valid point and before the user finishes speaking. It tests if the system knows when it is appropriate to jump into a conversation.
- Contextual Conditions
- These are three ways to set up an interaction test: Pre-session (behavior is set beforehand), In-session (behavior is stated during the talk), and Noexplicit (the behavior must be guessed from the surrounding words or sounds). Testing across these conditions reveals how robust a system's understanding of context truly is.
- Content Quality vs. Timing
- Evaluation looks at two things: Content Quality, which judges if the intended participation behavior makes sense semantically and prosodically; and Timing Appropriateness, which checks if the action happens at the right moment relative to other speech. The study finds that strong scores in one area don't guarantee overall success.
- Proactive Initiation (PI)
- This behavior assesses the agent's ability to start an interaction on its own without being prompted. It is found to be particularly challenging, with low success rates observed, suggesting current systems are weak at independently driving the conversation forward.
Terminology
Summary
Existing full-duplex speech benchmarks cover only subsets of realtime interaction behaviors, often under limited contextual conditions. DuplexAct-Bench introduces a bilingual benchmark that systematically covers six complementary behaviors—from interruption and yielding to proactive initiation—across Pre-session, In-session, and Noexplicit conditions. Across 1,290 English and Chinese streaming trials evaluating 12 full-duplex speech systems on Timing and Content, the results reveal substantial variation across behaviors, conditions, and systems. These findings show that current systems remain far from robustly managing when, whether, and how to participate as real-time interaction unfolds.
The gist
DuplexAct-Bench is a bilingual benchmark that systematically covers six complementary interaction behaviors—from interruption and yielding to proactive initiation—across Pre-session, In-session, and Noexplicit conditions.
Behavioral Coverage and Contexts
The benchmark systematically evaluates six complementary interaction behaviors: Agent Interruption (AI), User Interruption (UI), Interruption Resistance (IR), Active Silence (AS), Proactive Initiation (PI), and Agent Backchannel (AB). These behaviors are evaluated under three contextual conditions that differ in how the intended behavior is specified or implied: Pre-session, where the behavioral requirement is established through a persistent profile before interaction; In-session, where it is explicitly introduced during the streamed interaction; and Noexplicit, where no behavioral instruction is given and the appropriate behavior must be inferred from the semantic or acoustic context.
Evaluation Dimensions
The evaluation protocol assesses full-duplex speech agents based on two dimensions: content quality and timing appropriateness. Content evaluates how appropriately the intended participation behavior is realized in the given interaction situation, scored from 0 to 5 across four dimensions: Semantic Fulfillment (0–2), Constraint Adherence (0–1), Contextual Consistency (0–1), and Prosodic Appropriateness (0–1). Timing requirements vary across behaviors and interaction contexts, with specific success criteria defined for each behavior. For example, Agent Interruption success requires the first semantically valid interruption to occur after the annotated earliest valid interruption point and before the end of user audio.
Data Construction
The data construction involves drafting user-side utterances, expected behavior, content requirements, and target speaking style using GPT5.6 [16] for each trial construction. English and Chinese trials follow the same construction protocol. For Agent Backchannel, its Pre-session subset is based on dialogue texts from real Chinese crosstalk (xiangsheng) performances consolidated by GPT-5.6, while its No-explicit subset is selected from channel-separated otoSpeech conversations [17] in which one speaker channel contains only backchannels. Trials are rendered with ViiTorVoice [18], mixing white noise into all user audio as background noise and aligning all audio events on a shared timeline using manually verified VAD boundaries.
Results Summary
Across 1,290 bilingual streaming trials spanning 30 interaction scenarios, the results reveal substantial variation across behaviors and contextual conditions,
showing that current systems remain far from consistently handling the full repertoire of real-time interaction behaviors.
For instance, while some systems achieve high scores in specific metrics—such as Freeze-Omni achieving 88.8% BCR for Pre-session Active Silence—other key behaviors show difficulty, with No-explicit Proactive Initiation having a best BCR of only 8.8%. The study concludes that strong individual metrics do not necessarily imply joint success,
and that Content or Timing alone does not characterize successful participation.
Case studies illustrate the importance of evaluating interaction behavior against its contextual requirements, showing that full-duplex behavior depends not only on when an agent speaks, but also on what it says and whether it should speak at all.
Conclusion
The evaluation reveals substantial variation across interaction behaviors and conditions, with no system performing consistently well across the full range of settings. Strong semantic quality or favorable timing alone often fails to translate into successful interaction, while proactive behaviors remain particularly challenging. These results highlight a substantial gap between current full-duplex speech capabilities and robustly managing when, whether, and how to participate as real-time interaction unfolds.
References
[1] Alexandre Defossez et al., “Moshi: A speech-text foundation model for real-time dialogue,” arXiv preprint arXiv:2410.00037, 2024.
[2] Xiong Wang et al., “Freeze-omni: A smart and low latency speech-to-speech dialogue model with frozen LLM,” in Proc. 42nd Int. Conf. Mach. Learn. (ICML), 2025, pp. 63345–63354.
[16] OpenAI, “GPT-5.6: Frontier intelligence that scales with your ambition,” 2026, [Online]. Available: https://openai.
Improvements for AI systems
Based on the DuplexAct-Bench paper, here are specific improvements for existing full-duplex speech AI systems:
-
Acknowledge and systematically evaluate proactive interaction behaviors: Current systems heavily favor turn-taking and user-triggered responses. They must be explicitly trained to autonomously determine whether, when, and how to participate based on evolving context (Pre-session, In-session, No-explicit).
-
Implement robust behavioral switching across contexts: Systems should demonstrate consistent performance across the three contextual conditions (Pre-session profile establishment, In-session explicit instruction/inference, and No-explicit inference) for all six behaviors (Interruption, Yielding, Resistance to Interruption, Active Silence, Proactive Initiation).
-
Enhance semantic fulfillment and constraint adherence: Models must be refined to ensure that when they participate (e.g., during an interruption), the response not only meets the required information or correction but also strictly adheres to any given profile, instruction, role, language constraints, or output requirements.
-
Improve prosodic appropriateness: Systems need fine-tuning on AnyAudio-Judge rubrics to ensure that the realization of intended emotion or speaking style is aligned with the context and dialogue state.
-
Develop
Timing Awareness
for every behavior: Implement behavior-specific timing metrics (e.g., calculating offsets like L for Interruption or Lresp) to ensure actions are taken at the correct temporal points relative to reference events, rather than just reacting immediately. -
Master complex
Interruption Resistance
: Systems must be able to maintain the intended ongoing activity or execute scenario-specific recovery requirements when faced with non-disruptive user input without prematurely stopping their task or yielding unnecessarily. -
Develop specialized
Active Silence
logic: Implement a mechanism to correctly identify and maintain silence during annotated evaluation intervals, ensuring speech is withheld when contextually appropriate, rather than generating irrelevant filler speech.
These improvements will enable the improved AI system to move beyond basic turn-taking and achieve robust, natural-sounding participation in complex, real-time human conversations by mastering the nuanced when, whether, and how
of interaction.
Sources
- Moshi: a speech-text foundation model for real-time dialogue
- Qwen2.5-Omni Technical Report
- BayLing-Duplex: Native Full-Duplex Speech Dialogue with a Single Autoregressive LLM
- Full-Duplex-Bench: A Benchmark to Evaluate Full-duplex Spoken Dialogue Models on Turn-taking Capabilities
- FD-Bench: A Full-Duplex Benchmarking Pipeline Designed for Full Duplex Spoken Dialogue Systems
- FLEXI: Benchmarking Full-duplex Human-LLM Speech Interaction
- Full-Duplex Interaction in Spoken Dialogue Systems: A Comprehensive Study from the ICASSP 2026 HumDial Challenge
- Full-Duplex-Bench-v3: Benchmarking Tool Use for Full-Duplex Voice Agents Under Real-World Disfluency
- DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action
- DuplexSpeechBench-IFEval: Evaluating Implicit Instruction Following in Full-Duplex Voice Agents
- GPT-4o System Card
- AnyAudio-Judge: A Dynamic Rubric-Based Benchmark and Evaluator for Audio Instruction Following
- MiniCPM-o 4.5: Towards Real-Time Full-Duplex Omni-Modal Interaction
- VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction
- Raon-Speech Technical Report
- DuplexCascade: Full-Duplex Speech-to-Speech Dialogue with VAD-Free Cascaded ASR-LLM-TTS Pipeline and Micro-Turn Optimization
- Qwen3.5-Omni Technical Report
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering