Streaming Translation and Transcription Through Speech-to-Text Causal Alignment
cs.CL
Submitted: 2026-03-12
Updated: 2026-09-10
Comments: 15 pages, 9 figures
Code: https://github.com/SYSTRAN/faster-whisper
License: http://creativecommons.org/licenses/by/4.0/
The gist: Simultaneous machine translation (SiMT) has traditionally relied on offline machine translation models coupled with human-engineered heuristics or learned policies.
Terminology
Abstract
Simultaneous machine translation (SiMT) has traditionally relied on offline machine translation models coupled with human-engineered heuristics or learned policies. We propose Hikari, a policy-free, end-to-end model for simultaneous speech-to-text translation and streaming transcription. We also introduce Decoder Time Dilation, a mechanism that counteracts the overrepresentation of WAIT tokens in training. We present a supervised fine-tuning strategy that trains the model to recover from delays, significantly improving the quality-latency trade-off. Despite its modest size, Hikari delivers competitive translation quality at consistently low latency, comparing favorably with published IWSLT 2026 submissions up to 38x larger and with proprietary API systems across en-ja, en-de, and en-ru. We release our model weights and code to facilitate further research.
Sources
- MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
- Can neural machine translation do simultaneous translation?
- Moshi: a speech-text foundation model for real-time dialogue
- Qwen2.5 Technical Report
- Gemma 3 Technical Report
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- Streaming Sequence-to-Sequence Learning with Delayed Streams Modeling
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering