From Position Risks to Block Survival: Faster Generation for Diffusion Language Models
cs.CL, cs.AI
Submitted: 2026-09-27
Updated: 2026-09-27
Terminology
Sources
- Block Diffusion: Interpolating Between Autoregressive and Diffusion Language Models
- LLaDA2.0: Scaling Up Diffusion Language Models to 100B
- LLaDA2.1: Speeding Up Text Diffusion via Token Editing
- Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads
- Accelerating Large Language Model Decoding with Speculative Sampling
- DFlash: Block Diffusion for Flash Speculative Decoding
- SDAR: A Synergistic Diffusion-AutoRegression Paradigm for Scalable Sequence Generation
- DSpark: Confidence-Scheduled Speculative Decoding with Semi-Autoregressive Generation
- Beyond Autoregression: Fast LLMs via Self-Distillation Through Time
- Efficient-DLM: From Autoregressive to Diffusion Language Models, and Beyond in Speed
- Better & Faster Large Language Models via Multi-token Prediction
- Scaling Diffusion Language Models via Adaptation from Autoregressive Models
- S2D2: Fast Decoding for Diffusion LLMs via Training-Free Self-Speculation
- GRIFFIN: Effective Token Alignment for Faster Speculative Decoding
- Domino: Decoupling Causal Modeling from Autoregressive Drafting in Speculative Decoding
- EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty
- EAGLE-3: Scaling up Inference Acceleration of Large Language Models via Training-Time Test
- Discrete Diffusion Modeling by Estimating the Ratios of the Data Distribution
- SpecInfer: Accelerating Generative Large Language Model Serving with Tree-based Speculative Inference and Verification
- Large Language Diffusion Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering