DA-DLM: Explicitly Modeling Token Dependencies in Diffusion Language Models
cs.CL
Submitted: 2026-09-14
Updated: 2026-09-14
Code: https://github.com/jipy0222/DA-DLM
Project page: http://skylion007.github.io/OpenWebTextCorpus
License: http://creativecommons.org/licenses/by/4.0/
The gist: Diffusion Language Models (DLMs) generate text by iteratively denoising a masked sequence, independently predicting multiple tokens at each step.
Terminology
Abstract
Diffusion Language Models (DLMs) generate text by iteratively denoising a masked sequence, independently predicting multiple tokens at each step. This conditional independence discards inter-token dependencies and degrades coherence-an issue that parallels the multi-modality problem in Non-Autoregressive Translation (NAT). Drawing on the Directed Acyclic Transformer (DAT), which tackles this problem in NAT via a Directed Acyclic Graph (DAG), we propose DA-DLM, a model that adapts DAG-based dependency modeling to DLMs' iterative setting through a position-oriented DAG design. The position-oriented DAG binds node groups to fixed output positions so that tokens fixed in earlier steps anchor neighboring predictions via learned transitions, and evolves with denoising to focus on remaining uncertainty as anchors accumulate. On language modeling, open-ended generation, and summarization, DA-DLM consistently outperforms Block Diffusion, especially under fewer denoising steps, and matches autoregressive models while preserving the parallel generation advantage. Our code is publicly available at https://github.com/jipy0222/DA-DLM.
Sources
- Understanding and Accelerating the Training of Masked Diffusion Language Models
- Optimizing Non-Autoregressive Transformers with Contrastive Learning
- DAPD: Dependency-Aware Parallel Decoding via Attention for Diffusion LLMs
- LLaDA2.1: Speeding Up Text Diffusion via Token Editing
- LLaDA2.0: Scaling Up Diffusion Language Models to 100B
- Breaking the Factorization Barrier in Diffusion Language Models
- One Billion Word Benchmark for Measuring Progress in Statistical Language Modeling
- Why mask diffusion does not work
- Fast-dLLM: Training-free Acceleration of Diffusion LLM by Enabling KV Cache and Parallel Decoding
- Dependency-Guided Parallel Decoding in Discrete Diffusion Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering