DIP: Dynamic In-Context Planner For Diffusion Language Models
cs.CL, cs.AI
Submitted: 2026-01-06
Updated: 2026-08-29
Comments: EMNLP Findings 2026
Code: https://github.com/wmd3i/DIP
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
The gist: Diffusion language models (DLMs) have shown strong potential for general natural language tasks with in-context examples.
Terminology
Abstract
Diffusion language models (DLMs) have shown strong potential for general natural language tasks with in-context examples. Existing In-Context Learning (ICL) approaches largely inherit the practice of autoregressive language models (ARLMs), incorporating all examples into a fixed prompt. However, applying this rigid, static-prompt paradigm to DLMs incurs substantial computational overhead, as the model must evaluate the maximum context length at every step. We address this inefficiency with a key discovery: the block-wise KV-cache mechanism inherent to DLM inference enables the low-cost dynamic adjustment of the context. Following this intuition, our core idea is to start generation with a minimal prompt and progressively insert additional examples on the fly only when the generated tokens are of low confidence. Through rigorous empirical evaluations, we observe that average verified token confidence correlates strongly with generation accuracy, making it a reliable and computationally efficient signal of token quality. Formally, we propose Dynamic In-Context Planner (DIP), a context-optimization algorithm based on average verified confidence that dynamically ranks and inserts in-context examples during generation, rather than providing all examples up front. Experimental results on math and coding benchmarks with LLaDA-1.5 and LLaDA-8B-Instruct show that DIP achieves up to 1.59 times and 1.36 times speedups, respectively, while largely preserving the generation quality of the fixed-prompt baseline. Code: https://github.com/wmd3i/DIP
Sources
- Retrieval-Augmented Generation for Large Language Models: A Survey
- OPTAGENT: Optimizing Multi-Agent LLM Interactions Through Verbal Reinforcement Learning for Enhanced Reasoning
- dParallel: Learnable Parallel Decoding for dLLMs
- Training Verifiers to Solve Math Word Problems
- d$^2$Cache: Accelerating Diffusion-Based LLMs via Dual Adaptive Caching
- Thinking Inside the Mask: In-Place Prompting in Diffusion LLMs
- dLLM-Cache: Accelerating Diffusion Large Language Models with Adaptive Caching
- Better RAG using Relevant Information Gain
- Sparse-dLLM: Accelerating Diffusion LLMs with Dynamic Cache Eviction
- Fast-dLLM: Training-free Acceleration of Diffusion LLM by Enabling KV Cache and Parallel Decoding
- Dream 7B: Diffusion Large Language Models
- A Survey on Parallel Text Generation: From Parallel Decoding to Diffusion Language Models
- LLaDA 1.5: Variance-Reduced Preference Optimization for Large Language Diffusion Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering