Back on Track: Aligning Rewards and States for Reasoning in Diffusion Large Language Models
cs.CL
Submitted: 2026-06-07
Updated: 2026-09-06
Comments: Accepted to EMNLP 2026
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Training Verifiers to Solve Math Word Problems
- IIB-LPO: Latent Policy Optimization via Iterative Information Bottleneck
- Boosting the Generalization and Reasoning of Vision Language Models with Curriculum Reinforcement Learning
- Scaling Diffusion Language Models via Adaptation from Autoregressive Models
- DiffuCoder: Understanding and Improving Masked Diffusion Models for Code Generation
- Block Diffusion: Interpolating Between Autoregressive and Diffusion Language Models
- SSD-LM: Semi-autoregressive Simplex-based Diffusion Language Model for Text Generation and Modular Control
- SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization
- Wide-In, Narrow-Out: Revokable Decoding for Efficient and Effective DLLMs
- SDAR: A Synergistic Diffusion-AutoRegression Paradigm for Scalable Sequence Generation
- Omni-Diffusion: Unified Multimodal Understanding and Generation with Masked Discrete Diffusion
- Score-Based Generative Modeling through Stochastic Differential Equations
- Discrete Diffusion Modeling by Estimating the Ratios of the Data Distribution
- Dirichlet Flow Matching with Applications to DNA Sequence Design
- Consolidating Reinforcement Learning for Multimodal Discrete Diffusion Models
- wd1: Weighted Policy Optimization for Reasoning in Diffusion Language Models
- s1: Simple test-time scaling
- Wan: Open and Advanced Large-Scale Video Generative Models
- Scaling up Masked Diffusion Models on Text
- MRO: Enhancing Reasoning in Diffusion Language Models via Multi-Reward Optimization
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering