Masked Diffusion Decoding as x-Prediction Flow
Weitian Wang, Lianlei Shan, Shubham Rai, Cecilia De La Parra, Akash Kumar
cs.CL
Submitted: 2026-08-17
Updated: 2026-08-18
Comments: under review
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: Masked diffusion language models (MDLMs) generate text by iteratively unmasking tokens, but their standard decoder reduces each step to a binary action: a position is either committed to a single
Terminology
Abstract
Masked diffusion language models (MDLMs) generate text by iteratively unmasking tokens, but their standard decoder reduces each step to a binary action: a position is either committed to a single token or left fully masked, with no representation of partial belief in between. This all-or-nothing regime discards rich predictive information and forces premature, irrevocable commitments, leading to poor performance under a limited decoding budget. In this paper, we reinterpret mask prediction as clean-state prediction (x-prediction) and show that it can be used to induce a continuous flow in input embedding space. Building on this view, we propose a continuous decoding framework for MDLMs where tokens can accumulate partial progress at each diffusion step and remain revisable. To match the uneven contextual constraints across positions in language, we replace the globally synchronous schedule in image diffusion with a confidence-based asynchronous update in which the diffusion progress is token-wise accumulated. Additionally, we introduce a lightweight policy network and formulate its training as a reinforcement learning problem. Applied to pretrained LLaDA, our continuous decoder reaches 97% of its performance on the HumanEval dataset with 25% of decoding budget.
Sources
- GPT-4 Technical Report
- The Llama 3 Herd of Models
- Large Language Diffusion Models
- LLaDA2.0: Scaling Up Diffusion Language Models to 100B
- Flow Matching for Generative Modeling
- Back to Basics: Let Denoising Generative Models Denoise
- Your Absorbing Discrete Diffusion Secretly Models the Conditional Distributions of Clean Data
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Tulu 3: Pushing Frontiers in Open Language Model Post-Training
- Program Synthesis with Large Language Models
- Evaluating Large Language Models Trained on Code
- DiffuSeq: Sequence to Sequence Text Generation with Diffusion Models
- Soft-Masked Diffusion Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering