WaveFront Decoding: Parallelized Self-Speculative Decoding for Looped Language Models
cs.LG
Submitted: 2026-09-19
Updated: 2026-09-26
Comments: 17 pages, 7 figures, 8 tables
License: http://creativecommons.org/licenses/by/4.0/
The gist: Looped language models repeatedly apply a weight-shared block to increase effective depth without increasing parameter count, but the resulting T sequential recurrent-block calls per generated token
Terminology
Abstract
Looped language models repeatedly apply a weight-shared block to increase effective depth without increasing parameter count, but the resulting T sequential recurrent-block calls per generated token substantially increase decoding latency. To address the issue, we introduce Wavefront Decoding (WFD), a training-free self-speculative decoding framework designed for looped language models. WFD exploits two properties of these architectures: intermediate recurrence outputs provide effective draft predictions, and weight sharing allows token states at different positions and recurrence depths to be processed in one batched recurrent-block call. WFD organizes these mixed-depth states into a diagonal wavefront, continuously drafting new positions at shallow depth while advancing earlier positions toward full-depth verification. Unlike the phase-separated draft-then-verify schedule, WFD therefore co-batches drafting and verification within the same recurrent calls, while rejected drafts are corrected using full-depth predictions. Across six Spec-Bench task categories, WFD achieves 2.42x speedup on Ouro-2.6B and 3.54x on Huginn-3.5B over autoregressive decoding, consistently outperforming draft-then-verify. Cross-recurrence KV sharing further reduces wavefront KV traffic and increases WFD's speedup to 4.81x on Huginn-3.5B.
Sources
- Accelerating Large Language Model Decoding with Speculative Sampling
- Training Verifiers to Solve Math Word Problems
- Universal Transformers
- Efficient Parallel Samplers for Recurrent-Depth Models and Their Connection to Diffusion Language Models
- Measuring Mathematical Problem Solving With the MATH Dataset
- LoopFormer: Elastic-Depth Looped Transformers for Latent Reasoning via Shortcut Modulation
- Teaching Pretrained Language Models to Think Deeper with Retrofitted Recurrence
- Looped Latent Attention: Cross-Loop KV Compression for Looped Transformers
- LoopUS: Recasting Pretrained LLMs into Looped Latent Refinement Models
- Parcae: Scaling Laws For Stable Looped Language Models
- Fast Transformer Decoding: One Write-Head is All You Need
- Memory-Efficient Looped Transformer: Decoupling Compute from Memory in Looped Language Models
- Parallel Loop Transformer for Efficient Test-Time Computation Scaling
- LoopCoder-v2: Only Loop Once for Efficient Test-Time Computation Scaling
- Scaling Latent Reasoning via Looped Language Models
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks