From Truncation to Commitment: Persistent Context in Uniform Discrete Diffusion
cs.LG, cs.AI, math.PR, stat.ML
Submitted: 2026-09-01
Updated: 2026-09-01
Comments: 23 pages
Project page: http://skylion007.github.io/OpenWebTextCorpus
License: http://creativecommons.org/licenses/by/4.0/
The gist: Uniform-state discrete diffusion models update all tokens in parallel while keeping every position revisable.
Terminology
Abstract
Uniform-state discrete diffusion models update all tokens in parallel while keeping every position revisable. Even when the commonly used top- p rule leaves only one candidate at a position, that choice affects only the current reverse step and can be revised at the next sampling step. We ask what changes when selected hypotheses instead become persistent context for later predictions. We therefore propose committed reveal sampling (CRS), a training-free sampler that stores selected argmax tokens and inserts them into subsequent model inputs. Our analysis gives a rationale for selecting later and for keeping selected tokens visible. Under the exact forward process, the Bayes error of selecting a clean token cannot increase as noise decreases, while in a simple latent-mode model, keeping the selected token visible helps later parallel predictions agree on the same sequence-level choice. Empirically, paired experiments on Duo-distilled then separate this persistent effect from single-step top- p restriction and scalar temperature scaling. Under the same finalization rule, CRS without top- p truncation reaches lower generative perplexity (GenPPL) than fixed p=0.95 and p=0.9 baselines across budgets of 8--64 function evaluations (NFE). At 64 NFE, the comparison at matched unigram entropy also gives lower GenPPL for CRS, yielding a more favorable GenPPL--entropy tradeoff. Base Duo shows the same direction in a descriptive comparison, while other diversity and continuation metrics can rank these operating points differently. These results identify support restriction and persistent context as distinct controls of that tradeoff.
Sources
- Uniform Diffusion Models Revisited: Leave-One-Out Denoiser and Absorbing State Reformulation
- Simplex Relaxation for Discrete Diffusion
- Deferred Commitment Decoding for Diffusion Language Models
- The Path Matters: Learning a Token-Commitment Policy for Diffusion Language Models
- TACG: Trajectory-Aware Commit Gating for Diffusion Language Model Decoding
- When to Commit? Towards Variable-Size Self-Contained Blocks for Discrete Diffusion Language Models
- Don't Commit Alone: Joint Token Commitment in Diffusion Language Models
- Dream 7B: Diffusion Large Language Models
- Sumi: Open Uniform Diffusion Language Model from Scratch
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks