Survival-Guided Length Control for Efficient Diffusion Language Models
cs.CL
Submitted: 2026-08-26
Updated: 2026-08-26
Comments: EMNLP 2026 (Main Conference)
Code: https://github.com/DreamLM/Dream
License: http://creativecommons.org/licenses/by/4.0/
The gist: Diffusion language models (DLMs) generate text by iteratively denoising masked sequences, but standard decoding either fixes the sequence length or relies on ad hoc stopping rules, often leading to
Terminology
Abstract
Diffusion language models (DLMs) generate text by iteratively denoising masked sequences, but standard decoding either fixes the sequence length or relies on ad hoc stopping rules, often leading to unnecessary denoising steps. We recast length selection as a discrete-time survival problem over the end-of-sequence token and propose a plug-in, training-free length predictor that can be added to any existing DLM. Across reasoning and code-generation benchmarks, survival-guided length decoding speeds up inference by up to 7 times while preserving task accuracy. We further find that predicted lengths vary widely even within the same dataset, making model performance sensitive to the chosen length.
Sources
- Evaluating Large Language Models Trained on Code
- Training Verifiers to Solve Math Word Problems
- Diffusion Language Models Know the Answer Before Decoding
- dLLM-Cache: Accelerating Diffusion Large Language Models with Adaptive Caching
- Program Synthesis with Large Language Models
- Accelerated Sampling from Masked Diffusion Models via Entropy Bounded Unmasking
- Your Absorbing Discrete Diffusion Secretly Models the Conditional Distributions of Clean Data
- Fast-dLLM: Training-free Acceleration of Diffusion LLM by Enabling KV Cache and Parallel Decoding
- Dream 7B: Diffusion Large Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering