AdaFlash: Adaptive Speculative Decoding via On-Policy Distilled Diffusion Drafters
cs.LG, cs.CL
Submitted: 2026-07-21
Updated: 2026-09-26
Comments: COLM'26 Workshop (Spotlight)
Code: https://github.com/ZinYY/AdaFlash
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- OpenCodeInstruct: A Large-scale Instruction Tuning Dataset for Code LLMs
- LLaDA2.0: Scaling Up Diffusion Language Models to 100B
- Accelerating Large Language Model Decoding with Speculative Sampling
- Training Verifiers to Solve Math Word Problems
- Self Speculative Decoding for Diffusion Large Language Models
- Distilling the Knowledge in a Neural Network
- Fast and Accurate Causal Parallel Decoding using Jacobi Forcing
- ReFusion: A Diffusion Large Language Model with Parallel Autoregressive Decoding
- EAGLE-3: Scaling up Inference Acceleration of Large Language Models via Training-Time Test
- Let's Verify Step by Step
- Not-a-Bandit: Provably No-Regret Drafter Selection in Speculative Decoding for LLMs
- AngelSlim: A more accessible, comprehensive, and efficient toolkit for large model compression
- d3LLM: Ultra-Fast Diffusion LLM using Pseudo-Trajectory Distillation
- When RL Meets Adaptive Speculative Training: A Unified Training-Serving System
- Revolutionizing Reinforcement Learning Framework for Diffusion Large Language Models
- Fast-dLLM: Training-free Acceleration of Diffusion LLM by Enabling KV Cache and Parallel Decoding
- Qwen3 Technical Report
- Taming Masked Diffusion Language Models via Consistency Trajectory Reinforcement Learning with Fewer Decoding Step
- Multi-Candidate Speculative Decoding
- Dream 7B: Diffusion Large Language Models
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks