X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction

arXiv:2608.10878 · cs.CL, eess.AS · Submitted 2026-08-11 · Read on arXiv

X Square Robot

cs.CL, eess.AS

Submitted: 2026-08-11

Updated: 2026-09-08

Code: https://github.com/X-Square-Robot/X2-Turn

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 75/100

The gist: X2-Turn presents a frame-synchronous turn state prediction method via delayed-stream modeling, extending a pretrained Voxtral Realtime streaming ASR model with a parallel turn state prediction head.

Terminology

Summary

X2-Turn presents a frame-synchronous turn state prediction method via delayed-stream modeling, extending a pretrained Voxtral Realtime streaming ASR model with a parallel turn state prediction head. The architecture builds on delayed-stream modeling, where a causal audio encoder maps 16-kHz waveforms to frame-level features, a temporal adapter downsamples to 12.5 Hz, and a decoder-only language model emits one token per 80-ms step. The proposed dual-head design preserves the original ASR head while adding a parallel turn state head, with both heads operating on shared hidden states, enabling joint prediction of ASR tokens and fine-grained turn states at the frame level within a single streaming forward pass. The joint objective combines ASR token-level cross-entropy with turn state loss, weighted by λ = 0.1. At inference, autoregressive decoding is driven solely by the ASR head, and turn state predictions do not feed back into the decoding loop, so turn state errors do not affect subsequent ASR decoding.

The method defines five turn state tokens: for user silence, for active speech without semantic content, for active speech with partial semantic content, for active speech with complete semantic content, and for backchannel signals or filler words. Unlike prior work, refers specifically to early active speech before sufficient semantic content is observed, not non-speech or background noise. The paper introduces ASR-anchored supervision, which projects word-level turn annotations onto the frame-level positions of corresponding ASR tokens. Each word is represented by a single word-boundary token [W] followed by its subword tokens, with [W] anchored to word onset rather than offset. The target position is computed as pi = round(si/∆) + nτ, where ∆ = 80 ms and nτ = τ/∆ is the target delay in frames. Turn state labels are assigned to all positions occupied by [W] and subword tokens, while all unoccupied positions are labeled, including pre-speech silence, inter-word pauses, mid-utterance pauses, and trailing silence.

Training uses two stages. Stage 1 performs streaming ASR adaptation on approximately 26k hours of Chinese-English ASR data (14k hours Chinese, 12k hours English) from corpora including AISHELL 1-4, AliMeeting, WenetSpeech, KeSpeech, LibriSpeech, GigaSpeech, TED-LIUM, and VoxPopuli. Stage 2 performs joint ASR and turn state fine-tuning on turn-taking data: approximately 126 hours of Chinese EasyTurn training subset and approximately 249 hours of English Fisher telephone conversations. Qwen3-ForceAligner provides word-level timestamps, and Qwen3.5-Plus serves as an LLM annotator for word-level semantic turn state labeling. The streaming delay τ is sampled per batch between 1 and 30 frames (80-2400 ms) in both stages. The turn state head is initialized as a copy of the ASR head. The backbone is Voxtral-Mini-4B-Realtime, with full fine-tuning of both the causal audio encoder and language decoder.

Latency is measured relative to word end: Li = τ - (ei - si), where [si, ei] is the word span. For cascaded baselines, latency includes inference time plus front-end VAD delay latencyvad. On the EasyTurn Chinese and English test sets, X2-Turn with τ = 480 ms achieves 91.00% ACCcomp, 93.00% ACCincomp, and 96.00% ACCbc on Chinese, with 288 ms latency; on English, it achieves 92.10% ACCcomp and 84.60% ACCincomp with 225 ms latency. Compared to the fully streaming SoulX-Duplug baseline, X2-Turn consistently outperforms it while remaining competitive in latency, and requires no auxiliary separately optimized ASR model at inference. Compared to cascaded VAD-dependent baselines, X2-Turn matches the accuracy of the strongest cascaded system while operating fully streaming, yielding a better accuracy-latency trade-off across both languages.

Ablation on streaming delay τ (320, 400, 480 ms) shows a consistent latency-accuracy trade-off. As τ decreases from 480 ms to 320 ms, latency drops substantially (from 288 ms to 120 ms on Chinese; from 225 ms to 65 ms on English), while turn state accuracy degrades only mildly (Chinese average accuracy from 92.00% to 90.67%; English from 88.49% to 85.09%). This indicates robustness under tight latency constraints, allowing suitable operating points for different latency requirements.

For ASR performance, Stage1-ASR at τ = 480 ms outperforms chunk-based streaming baselines Uni-ASR and Freeze-Omni across nearly all evaluation sets (AISHELL-1, test-meeting, test-net, GigaSpeech, LS-clean, LS-other), demonstrating that frame-wise prediction with short lookahead is competitive with or superior to fixed-chunk streaming. After Stage2-Turn joint training, turn-state supervision introduces set-dependent degradation relative to Stage1-ASR at the same delay, but the model still matches or exceeds Freeze-Omni, indicating the frame-synchronous backbone retains most of its recognition advantage under multi-task optimization.

The paper concludes that X2-Turn achieves accurate turn-taking detection while maintaining low latency on bilingual EasyTurn test sets, with the streaming delay τ offering a controllable trade-off between turn-taking accuracy, response latency, and ASR quality. Future work will further balance ASR and turn state objectives and improve robustness in more challenging conversational settings.

Improvements for AI systems

Improvements to AI Systems:

  1. Unified Streaming ASR + Turn-Taking Head
  • Integrate a parallel turn-state prediction head into any streaming ASR model (e.g., Whisper Realtime, Conformer) using shared hidden states, with a weighted joint loss (λ=0.1) to avoid ASR degradation.

  • The improved system can detect user turn boundaries (idle, incomplete, complete, backchannel) in real time without a separate VAD or cascaded pipeline.

  1. ASR-Anchored Supervision for Frame-Level Turn Labels
  • Replace heuristic VAD-based labels with word-boundary-anchored turn states, projecting word-level annotations onto ASR token positions using a fixed delay τ.

  • The improved system can learn precise turn states even during mid-word pauses or trailing silence, reducing false complete triggers and improving natural conversation flow.

  1. Delayed-Stream Modeling with Controllable Latency
  • Use a causal encoder + temporal adapter (12.5 Hz) with a decoder emitting one token per 80 ms, and sample τ (1–30 frames) per batch during training.

  • The improved system can be tuned for latency vs. accuracy trade-offs (e.g., 120 ms at 90.7% accuracy vs. 288 ms at 92.0% on Chinese), enabling deployment in both low-latency (voice assistants) and high-accuracy (meeting transcription) scenarios.

  1. Dual-Head Decoding with Error Isolation
  • Keep ASR head as the sole autoregressive driver; turn-state head predictions are non-feedback, preventing turn errors from corrupting ASR output.

  • The improved system can maintain ASR quality (matching or exceeding chunk-based baselines) while adding turn detection, even under noisy or ambiguous speech.

  1. Two-Stage Training for Domain Adaptation
  • First adapt the backbone on large-scale ASR data (26k hours), then fine-tune jointly on smaller turn-taking data (126–249 hours) with LLM-annotated semantic labels.

  • The improved system can generalize to bilingual (Chinese-English) conversations with minimal turn-taking data, reducing annotation cost.

  1. Five-State Turn Tokenization
  • Use distinct tokens for idle, noidle (early speech), incomplete, complete, and backchannel, where noidle is specifically early active speech before semantic content.

  • The improved system can distinguish filler words and backchannels from true turn completions, enabling more natural barge-in handling and smoother multi-party dialogue.

  1. Frame-Synchronous Inference with No Chunking
  • Replace fixed-chunk streaming with frame-wise prediction, enabling sub-100 ms response to user speech onset.

  • The improved system can start processing speech as soon as acoustic features arrive, reducing perceived latency in interactive applications like real-time translation or voice-controlled robots.

Capabilities of the Improved AI System:

  • Real-time, fully streaming turn-taking detection in bilingual conversations with accuracy >90% for complete/incomplete states and >84% for backchannels.

  • Adjustable latency from 65 ms to 288 ms without retraining, allowing dynamic trade-offs based on device or network constraints.

  • Robust ASR performance that matches or exceeds fixed-chunk streaming models, even after multi-task fine-tuning.

  • No reliance on external VAD or cascaded ASR at inference, simplifying deployment and reducing error propagation.

  • Ability to handle overlapping speech, inter-word pauses, and trailing silence correctly via ASR-anchored labels.

Sources

Related papers