Vocabulary-size-independent Convergence of Discrete Diffusion Models: adjoint equations induce the right space
cs.LG, math.ST, stat.ML, stat.TH
Submitted: 2026-05-17
Updated: 2026-09-07
License: http://creativecommons.org/licenses/by/4.0/
The gist: Discrete diffusion has become a leading framework for generative modeling in various applications including language, vision, and biology.
Terminology
Abstract
Discrete diffusion has become a leading framework for generative modeling in various applications including language, vision, and biology. Existing convergence theory, however, exhibits fundamental limitations. KL-based analyses diverge under singular priors such as the masked distribution, while bounds in total variation (TV) depend on the vocabulary size S and become vacuous for modern language tasks, where vocabularies contain hundreds of thousands of tokens. We develop a unified adjoint-equation-based framework that establishes vocabulary-size-independent convergence guarantees in any integral probability metric (IPM). To the best of our knowledge, our bounds are the first to be entirely free of S and applicable to both masked and uniform priors. Importantly, our results can extend existing step complexity guarantees to any IPM. Also, our theory relies only on a single standard rate-matrix regularity assumption and applies to general priors. Five novel techniques drive our improvements: 1. working in the space of observables via adjoint equations rather than directly with probability measures; 2. a regularity analysis that yields bounds on any IPM; 3. a coupling argument that removes S-dependence under uniform transitions; and 4. score-marginal cancellation and 5. exit-routing techniques that remove S-dependence under masked transitions. Our framework thus sharply departs from prior analyses and avoids the shortcomings of pathspace-KL and existing TV-based approaches. Beyond convergence bounds, our framework provides a versatile toolkit for further theoretical study of discrete diffusion models, including principled choices of loss functions and vocabulary-size-independent step complexity.
Sources
- Muse: Text-To-Image Generation via Masked Generative Transformers
- Optimal Inference Schedules for Masked Diffusion Models
- Non-Asymptotic Convergence of Discrete Diffusion Models: Masked and Random Walk dynamics
- Efficient Sampling with Discrete Diffusion Models: Sharp and Adaptive Guarantees
- Neural Continuous-Time Markov Chain: Discrete Diffusion via Decoupled Jump Timing and Direction
- From Scores to Gibbs Correctors: Accelerating Uniform-Rate Discrete Diffusion Models
- Sharp Convergence Rates for Masked Diffusion Models
- On integral probability metrics, \phi-divergences and binary classification
- Scheduling Thoughts: Learning the Order of Thought in Diffusion Language Models
- Masked Diffusion Modeling for Anomaly Detection
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks