Attention-Discounted Adaptive Sampler for Masked Diffusion Language Models
cs.CL, cs.AI
Submitted: 2026-06-09
Updated: 2026-09-25
License: http://creativecommons.org/licenses/by/4.0/
The gist: Masked diffusion language models can reduce inference steps by revealing multiple tokens per denoising iteration, but this parallelism is fragile: positions that are individually confident may be
Terminology
Abstract
Masked diffusion language models can reduce inference steps by revealing multiple tokens per denoising iteration, but this parallelism is fragile: positions that are individually confident may be unsafe to commit together when their predictions are coupled. Existing training-free samplers such as Top- k, Fast-dLLM, and EB-Sampler mainly control how many tokens to reveal, while often ranking candidates by token-wise scores that ignore interactions within the selected set. We propose ADAS, a training-free reranking rule that leaves the base sampler's stopping rule unchanged and greedily discounts each token-wise confidence score according to its attention to already selected positions, weighted by their prediction uncertainty. Across LLaDA-8B-Base and Dream-7B-Base on the reasoning benchmarks GSM8K and MATH500 and the code benchmarks HumanEval and MBPP, plugging ADAS into all three samplers improves low-NFE performance at matched denoiser evaluations by 9.11 and 10.46 percentage points on average, respectively, with 3.1% per-forward runtime overhead.
Sources
- Where-to-Unmask: Ground-Truth-Guided Unmasking Order Learning for Masked Diffusion Language Models
- Program Synthesis with Large Language Models
- Learning to Parallel: Accelerating Diffusion Large Language Models via Learnable Parallel Decoding
- Evaluating Large Language Models Trained on Code
- Measuring Mathematical Problem Solving With the MATH Dataset
- Improving Discrete Diffusion Unmasking Policies Beyond Explicit Reference Policies
- Accelerating Diffusion LLMs via Adaptive Parallel Decoding
- Learning Unmasking Policies for Diffusion Language Models
- DAPD: Dependency-Aware Parallel Decoding via Attention for Diffusion LLMs
- KLASS: KL-Guided Fast Inference in Masked Diffusion Models
- Plan for Speed: Dilated Scheduling for Masked Diffusion Language Models
- Path Planning for Masked Diffusion Model Sampling
- Simple and Effective Masked Diffusion Language Models
- Dream 7B: Diffusion Large Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering