DARTree: Speculative Diffusion Decoding with Autoregressive Draft Trees

arXiv:2608.13524 · cs.LG · Submitted 2026-08-13 · Read on arXiv

Tianyi Li, Yaxin Luo, Xinyi Shang, Zhiqiang Shen

Mohamed bin Zayed University of Artificial Intelligence

cs.LG

Submitted: 2026-08-13

Updated: 2026-08-14

Code: https://github.com/VILA-Lab/DARTree

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 95/100

The gist: DARTree is a training-free speculative decoding method that extends a pretrained causally corrected block-parallel drafter from a single chain to a speculative tree.

Terminology

Summary

DARTree is a training-free speculative decoding method that extends a pretrained causally corrected block-parallel drafter from a single chain to a speculative tree. It is designed to accelerate autoregressive language models by verifying multiple draft tokens in parallel.

The paper identifies a key challenge: "Because the corrected scores of a node’s children depend on that node’s path-specific state, each heap pop must be followed by correction-head inference and subsequent heap pushes before the next node can be expanded. This interleaving forces correction-head inference to proceed one node at a time, making tree construction a substantial source of latency." DARTree addresses this by decoupling causal correction from sequential best-first search.

The method works as follows: "DARTree expands a fixed number of nodes at each depth and evaluates their corrected distributions in one batch, avoiding the node-wise correction interleaved with heap operations. Since wider layer expansion adds little latency, DARTree first constructs a wider supertree and then applies best-first pruning later to obtain a compact verification tree, decoupling correction-head inference from sequential heap operations. The approach uses depth-wise batched AR expansion, deferred best-first pruning, and a single standard target-model tree-verification pass."

The contributions are threefold: "We identify the sequential bottleneck caused by coupling path-conditioned causal correction with node-wise best-first tree construction. We introduce a training-free, depth-wise batched tree-construction method that applies causal correction across multiple branches in parallel, achieving lossless speedup. We propose a deferred and simplified best-first pruning strategy for selecting the final verification tree."

In experiments across seven math, code, and chat benchmarks with Qwen3-4B and Qwen3-8B at temperatures 0 and 1, DARTree achieves the highest overall average acceptance length and speedup in all four model–temperature configurations. Specifically, On GSM8K with Qwen3-4B at T = 0, DARTree accepts 12.97 tokens per verification round and reaches a 9.73× speedup. Its acceptance length is 98.6% higher than DFlash and 27.9% higher than Domino in the same setting.

The paper also notes that DARTree (pruned) achieves the highest overall average τ and speedup in all four configurations, outperforming DDTree and Domino by up to 28.9% and 46.8% in τ, and by up to 22.7% and 40.1% in speedup, respectively. The only exception is AIME25 with Qwen3-8B at T = 1, where DDTree is marginally faster (4.56× versus 4.50×), while DARTree still attains a higher acceptance length.

Ablation studies show that "Sequential Correction w. Heap... obtains acceptance lengths close to DARTree, confirming the quality of fully conditional node-wise search, but requires roughly 70 ms per round, more than twice DARTree’s latency on all three datasets. DARTree also transfers to other correction heads: DARTree improves both metrics in all six dataset–temperature settings: acceptance length increases by 14.6–40.6%, while speedup improves by up to 34.3%" when applied to DSpark's Markov head.

The paper acknowledges limitations: "DARTree relies on a pretrained diffusion drafter equipped with a causal correction head. Consequently, its training-free property does not directly apply to naive diffusion drafters, which require additional training to incorporate such a head. It also notes that as a speculative decoding method, DARTree does not reduce the total number of FLOPs; instead, it uses additional computation to reduce inference latency."

Improvements for AI systems

Improvements to AI Systems Based on DARTree:

  1. Batch-Parallel Speculative Decoding for Latency-Critical Deployment
  • Improvement: Replace sequential, node-by-node tree construction in speculative decoding with depth-wise batched expansion and deferred pruning.

  • Capability: The AI system can generate multiple draft branches simultaneously, verify them in a single target-model pass, and reduce per-token latency by up to 9.73× (e.g., on GSM8K) without any additional training or fine-tuning. This is ideal for real-time applications like chatbots, code completion, and interactive reasoning where low latency is critical.

  1. Lossless Acceleration for Autoregressive Models with Correction Heads
  • Improvement: Decouple path-conditioned causal correction from sequential heap operations, enabling parallel correction-head inference across all nodes at a given depth.

  • Capability: The system can maintain the exact same output distribution (lossless speedup) while achieving higher acceptance lengths (e.g., 98.6% higher than DFlash and 27.9% higher than Domino on GSM8K). This makes it suitable for production systems where output quality must remain identical to the original model.

  1. Adaptive Tree Pruning for Memory and Compute Efficiency
  • Improvement: Construct a wider supertree first, then apply best-first pruning to select a compact verification tree, avoiding the need to store or process all branches.

  • Capability: The AI system can dynamically balance exploration (wider search) and exploitation (focused verification) based on available memory and compute, leading to up to 28.9% higher token acceptance rate and 22.7% faster speedup compared to prior tree-based methods like DDTree, while using fewer resources.

  1. Model-Agnostic Integration with Existing Diffusion Drafters
  • Improvement: Apply DARTree’s batched construction and pruning strategy to any pretrained drafter with a causal correction head (e.g., DSpark’s Markov head).

  • Capability: The system can be retrofitted onto existing speculative decoding pipelines, improving acceptance length by 14.6–40.6% and speedup by up to 34.3% across diverse datasets and temperatures, without retraining the drafter or target model.

  1. Temperature-Robust Acceleration for Diverse Reasoning Tasks
  • Improvement: Use DARTree’s depth-wise batched expansion to handle both greedy (T=0) and stochastic (T=1) decoding settings uniformly.

  • Capability: The AI system can maintain high speedups and acceptance lengths across math, code, and chat benchmarks (e.g., GSM8K, AIME25, HumanEval) at both temperatures, making it reliable for varied user prompts and creative vs. deterministic tasks.

  1. Reduced Wall-Clock Time for Multi-Branch Reasoning
  • Improvement: Eliminate the sequential bottleneck where each heap pop triggers correction-head inference; instead, batch all corrections per depth.

  • Capability: The system can reduce per-round latency from 70 ms (sequential) to under 35 ms (batched) on the same hardware, enabling faster multi-step reasoning, chain-of-thought generation, and self-consistency sampling in interactive AI assistants.

  1. Deployable on Existing Hardware without Specialized Kernels
  • Improvement: Use only standard target-model tree-verification passes and batched inference, requiring no custom CUDA kernels or hardware modifications.

  • Capability: The improved AI system can be deployed on commodity GPUs and CPUs, achieving significant speedups in edge or cloud environments where specialized acceleration hardware is unavailable.

Abstract

Speculative decoding losslessly accelerates autoregressive language models by verifying multiple draft tokens in parallel. Diffusion-based drafters further reduce proposal latency by predicting an entire token block in parallel, but their position-wise distributions are marginal rather than conditioned on tokens selected along each draft path. Existing recurrent correction incorporates causal information along a single draft chain, whereas diffusion-based tree construction broadens candidate coverage without carrying this correction along individual branches. We introduce DARTree, a training-free speculative decoding method that extends a pretrained AR correction head from chains to trees. DARTree first constructs a fixed-width candidate tree by expanding and scoring all nodes at each depth in a single batch, and then only applies best-first pruning to select the verification tree, decoupling AR-head inference from sequential heap operations. Across seven math, code, and chat benchmarks, DARTree achieves the highest average acceptance length and speedup in all four model--temperature configurations, accepting up to 12.97 tokens per verification round, 98.6% more than DFlash and 27.9% more than Domino in the same setting, and reaching up to 9.73 times lossless speedup over locally measured autoregressive decoding.

Sources

Related papers