Efficient One-to-Many Translation with Joint Multi-Stream Diffusion
cs.CL, cs.AI
Submitted: 2026-09-14
Updated: 2026-09-14
License: http://creativecommons.org/licenses/by/4.0/
The gist: One-to-many machine translation (MT) is computationally expensive for autoregressive (AR) systems, which suffer from linear latency scaling with both sequence length and the number of target
Terminology
Abstract
One-to-many machine translation (MT) is computationally expensive for autoregressive (AR) systems, which suffer from linear latency scaling with both sequence length and the number of target languages. We explore how diffusion can enable multilingual translation with a discrete diffusion framework that refines all target languages in parallel, achieving sublinear latency scaling with the number of targets, and supports deployment as a single unified model to replace multiple independent systems. Conditioned on a continuous semantic anchor rather than source tokens, our framework supports zero-shot transfer to unseen source languages without retraining, maintaining approximately 75% of its supervised translation quality on zero-shot sources. We investigate the quality-latency frontier and find that with accelerated sampling, it achieves comparable supervised quality to AR baselines with a 2 times speedup and 11.9% better zero-shot BLEU. These results highlight the potential of joint multi-stream diffusion as a practical and flexible alternative for efficient one-to-many translation.
Sources
- A Framework for Hierarchical Multilingual Machine Translation
- Forward-Learned Discrete Diffusion: Learning how to noise to denoise faster
- LLaDA2.0: Scaling Up Diffusion Language Models to 100B
- XDLM: Cross-lingual Diffusion Language Model for Machine Translation
- DPad: Efficient Diffusion Language Models with Suffix Dropout
- No Language Left Behind: Scaling Human-Centered Machine Translation
- Transformer-Encoder Trees for Efficient Multilingual Machine Translation and Speech Translation
- Deep Encoder, Shallow Decoder: Reevaluating Non-autoregressive Machine Translation
- dLLM-Cache: Accelerating Diffusion Large Language Models with Adaptive Caching
- Fast-dLLM: Training-free Acceleration of Diffusion LLM by Enabling KV Cache and Parallel Decoding
- Dream 7B: Diffusion Large Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering