Enabling KV Caching of Shared Prefix for Diffusion Language Models
cs.LG, cs.AI
Submitted: 2026-05-26
Updated: 2026-09-02
Comments: Accepted to EMNLP 2026 Main Conference. Code: https://github.com/OSSS-KU/BiCache
Code: https://github.com/OSSS-KU/BiCache
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
The gist: Key-value (KV) caching for shared prefixes is essential for high-throughput large language model (LLM) serving, but it faces critical challenges in emerging diffusion language models (DLMs).
Terminology
Abstract
Key-value (KV) caching for shared prefixes is essential for high-throughput large language model (LLM) serving, but it faces critical challenges in emerging diffusion language models (DLMs). In DLMs, bidirectional attention means that updating any token dynamically alters the entire context and its corresponding KVs. Thus, existing caching techniques developed for LLMs, which assume that KVs remain invariant once computed, corrupt the shared prefix KVs. Our experiments show that applying these techniques to DLMs causes model accuracy to collapse to near zero. To unlock high-throughput DLM serving, we propose bidirectional prefix caching, BiCache, the first KV caching technique for shared prefixes in DLMs. BiCache is designed based on key observations from our comprehensive analysis: shared prefix KVs remain stable and reusable in shallow layers, while the depth of shallow layers depends on the fraction of shared prefix tokens in each request. Thus, BiCache dynamically identifies a safe layer depth for reusing shared prefix KVs and eliminates redundant computation. Evaluations demonstrate that BiCache significantly improves serving throughput by 36.3%-98.3% compared to existing techniques without accuracy collapse (only 0-1.8% difference).
Sources
- Training Verifiers to Solve Math Word Problems
- Prompt Cache: Modular Attention Reuse for Low-Latency Inference
- FlashDLM: Accelerating Diffusion Language Model Inference via Efficient KV Caching and Guided Diffusion
- d$^2$Cache: Accelerating Diffusion-Based LLMs via Dual Adaptive Caching
- Beyond Next-Token Prediction: A Performance Characterization of Diffusion versus Autoregressive Language Models
- Large Language Diffusion Models
- Reinforcement Learning is all You Need
- d3LLM: Ultra-Fast Diffusion LLM using Pseudo-Trajectory Distillation
- CD4LM: Consistency Distillation and aDaptive Decoding for Diffusion Language Models
- WeDLM: Reconciling Diffusion Language Models with Standard Causal Attention for Fast Inference
- Diffusion LLMs Can Do Faster-Than-AR Inference via Discrete Diffusion Forcing
- Fast-dLLM v2: Efficient Block-Diffusion LLM
- TiDAR: Think in Diffusion, Talk in Autoregression
- Fast-dLLM: Training-free Acceleration of Diffusion LLM by Enabling KV Cache and Parallel Decoding
- TurboSpec: Closed-loop Speculation Control System for Optimizing LLM Serving Goodput
- dLLM-Cache: Accelerating Diffusion Large Language Models with Adaptive Caching
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks