Test-Time Scaling with Diffusion Language Models via Reward-Guided Stitching

arXiv:2602.22871 · cs.CL, cs.AI · Submitted 2026-02-26 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Test-Time Scaling with Diffusion Language Models via Reward-Guided Stitching".

Jane: The gist The proposed framework introduces Stitching Noisy Diffusion Thoughts,

Tom: First, who's behind it and why it matters.

Paper summary: Tom: So we've covered the basics of "Test-Time Scaling with Diffusion Language Models via Reward-Guided Stitching," which is this framework that turns cheap diffusion sampling into a reusable pool of step-level candidates.

Jane: To sum up, it’s a three-part system: exploring with diffusion, evaluating every intermediate step with a PRM, and stitching the best steps into a rationale for an AR solver to do the final heavy lifting.

Lu: What this really means is that we can use diffusion models for broad exploration without paying the high computational cost of running long derivations repeatedly during testing.

Meng: The implication for practical application is clear: we get much better accuracy and much lower latency when we need to solve complex tasks in real-time, because that final AR solver only runs once.

Lalam: It really shifts the focus from trying to find one perfect long path to building a high-quality evidence set that guides a fast final computation.

Tom: The authors are Roy Miles, Aysim Toker, Andreea-Maria Oncescu, Songcen Xu, and Jiankang Deng. They've shown that this method can push accuracy up to +thirty point six percent over vanilla diffusion decoding and offer latency reductions of up to one point eight times.

Jane: It’s a practical way to strengthen the accuracy-latency trade-off for large language models in complex reasoning tasks, making them much more usable in real-world scenarios.

Conclusion: Tom: So we’re wrapping up on "Test-Time Scaling with Diffusion Language Models via Reward-Guided Stitching." Basically, they took those cheap diffusion thoughts and turned them into a reusable pool of high-quality steps using a stitching mechanism to speed up testing for large language models.

Jane: It’s about taking the messy exploration phase and turning it into something structured that an autoregressive solver can actually use to find the final answer faster.

Lu: The authors, Miles, Toker, Oncescu, Xu, and Deng. They designed this modular pipeline where you explore broadly with diffusion first then score those steps using a process reward model before stitching them back together for the final recomputation.

Meng: From an engineering standpoint, it’s neat because the exploration part is parallel and cheap while the heavy synthesis only happens once at the very end when you really need to compute.

Lalam: It means we don't have to run those massive, slow derivations repeatedly just to test a model; we pre-process a high-quality evidence set that’s ready for fast final checks.

Tom: That’s the core idea—using quality scores from an external reward model to filter and combine the noisy diffusion outputs into something more reliable than any single path.

Jane: It really changes how we think about scaling these models; it moves us away from just generating one long chain of thought toward building a robust, evidence-based reasoning toolkit that’s ready for deployment.

Lu: I think the real power here is model agnosticism—it doesn't matter what kind of diffusion model you use as long as it can generate those initial candidate trajectories and you have a way to score the intermediate steps.

Meng: That flexibility is huge because we don't have to redesign the whole system every time we switch diffusion architectures; you just swap out the components.

Tom: Exactly, so it’s not tied down to one specific architecture or one type of reasoning path; it’s a framework for reusing step-level quality regardless of how that initial exploration happened.

Jane: It shifts the focus from just making the initial thinking process better to making the final answer generation process much more efficient and accurate on demand.

Tom: So, we’ve seen how they use this to boost accuracy by over thirty percent while cutting latency by nearly two times; that’s a big win for real-time AI.

Jane: It shows that combining cheap exploration with targeted verification can actually lead to significant gains without needing a massive increase in training data or computational resources.

Tom: Next up, we’re going to look at what the authors did when they tested this against other popular decoding methods and saw how it stacked up in terms of raw numbers.

Huawei London Research Center

cs.CL, cs.AI

Submitted: 2026-02-26

Updated: 2026-10-08

Comments: NeurIPS 2026. Code available at https://github.com/roymiles/diffusion-stitching

Code: https://github.com/roymiles/diffusion-stitching

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 91/100

The gist: The gist The proposed framework introduces Stitching Noisy Diffusion Thoughts, a self-consistency mechanism that converts cheap diffusion-sampled reasoning into a reusable pool of step-level

Key concepts

Masked Diffusion Language Models
These models are used for broad exploration by iteratively denoising tokens in parallel rather than sequentially. They are well-suited for generating diverse, low-cost reasoning chains of thought from a partially masked starting sequence.
Process Reward Model (PRM)
An off-the-shelf model that scores every intermediate step in a reasoning trajectory. It outputs a scalar confidence score indicating the quality of that specific step based on the problem and history, allowing for fine-grained evaluation.
Stitching Mechanism
The process of combining high-quality steps from multiple diverse reasoning paths into one coherent rationale. Instead of picking the single best path, it collects all steps above a certain confidence threshold and uses them to condition a final autoregressive solver.
Autoregressive (AR) Solver
A lightweight model invoked only once at the end to recompute the final answer. It takes the original problem and the stitched, evidence-conditioned rationale as input, prioritizing high-confidence steps from the pool.

Terminology

Summary

The gist The proposed framework introduces Stitching Noisy Diffusion Thoughts, a self-consistency mechanism that converts cheap diffusion-sampled reasoning into a reusable pool of step-level candidates to improve test-time scaling for large language models.

How it works

The method proposes a modular pipeline that separates exploration from synthesis: "we first generate a diverse pool of low-cost thoughts via a diffusion model (Section 3.2), evaluate their utility at the step level using an off-the-shelf process reward model (Section 3.3), and finally employ a stitching mechanism (Section 3.4) to recombine high-confidence steps into a coherent rationale that guides an AR solver to recompute the final answer" The diffusion component supplies breadth, the PRM provides a fine-grained notion of step quality, and the AR solver converts a partially redundant set of high-quality steps into a more accurate final response.

The pipeline consists of three main components:

  1. Explore: given an input problem, we sample N diverse chains-of-thought using a masked diffusion language model with a confidence-based sampling procedure that encourages diversity while preserving high-confidence content.

  2. Evaluate: we score every intermediate step in every trajectory using an off-the-shelf process reward model (PRM), yielding a global pool of candidate steps with quality estimates. This allows for retaining useful intermediate results even when an overall trajectory later derails.

  3. Stitch + Recompute: rather than selecting a single “best” trajectory (e.g., by maximizing the geometric mean of step scores), we collect the highest-quality steps across paths and concatenate them into a composite rationale. A lightweight autoregressive (AR) model then conditions on the original problem and this stitched rationale and recomputes only the final answer.

Key Components and Mechanisms

The paper details specific mechanisms for each stage of the pipeline:

Masked diffusion language models are well suited to broad exploration: starting from a partially-masked sequence, they iteratively denoise tokens in parallel, rather than generating tokens in a strict chronological order

For step evaluation, the quality score is determined by an off-the-shelf process reward model (PRM): the PRM outputs a scalar confidence: r(n)t = PRMϕx, s(n)1:t ∈ [0, 1], where the score is conditioned on the problem x and the partial solution history s(n)1:t. This step-level view is crucial because it allows for retaining useful intermediate derivations even when a reasoning path later derails.

The stitching process involves several choices, including:

we retain all steps in P with a PRM score at least δ, and we also include every step from the best full trace (to serve as an anchor)

The final answer is generated by an AR solver using evidence-conditioned decoding: "We construct a solver prompt by concatenating the problem x with the confidence-annotated evidence list E. We explicitly instruct the solver to treat these steps as evidence, prioritizing high-confidence entries while ignoring conflicts".

Empirical Results and Analysis

The framework demonstrates significant performance gains across various benchmarks:

Stitching yields large gains over vanilla diffusion decoding by converting diverse but noisy reasoning paths into a higher-quality evidence set

The authors report that the training-free pipeline outperform[s] the LLaDA baseline decoded with a high confidence threshold (γ = 0.9), despite using a much lower threshold (γ = 0.7), by an average of 23.9% average points.

The results show that "step-level reuse can deliver state-of-the-art accuracy among diffusion-based approaches while remaining latency-friendly: exploration is cheap and parallel in diffusion, and the AR solver is invoked only once to produce the final answer".

The method achieves up to +30.6% absolute accuracy points over vanilla diffusion decoding. Furthermore, it achieves up to a 1.8× latency reduction relative to both traditional diffusion models (e.g., Dream, LLaDA) and unified architectures (e.g., TiDAR).

Ablation and Extensions

The study investigates the impact of various parameters:

We ablate both how we aggregate diffusion traces and when an AR solver is needed in Table 2

They show that selective stitching yields the best accuracy. The results indicate that a moderate-to-high diversity is important for strong performance: too low temperatures leads to poor exploration and diversity, while too high degrades the quality of each trace.

The framework is also model-agnostic, as it only assumes (i) a generator that can produce candidate reasoning trajectories (here, masked diffusion LMs) and (ii) a PRM that can score intermediate steps.

Conclusion

The paper concludes that the framework successfully introduces a training-free reasoning framework that combines inexpensive diffusion-based exploration with steplevel verification and selective aggregation. This approach "substantially strengthens accuracy–latency trade-offs: it improves accuracy by up to +30.6% over vanilla diffusion decoding and remains competitive with strong autoregressive baselines (e.g., +4.3% on average), while reducing sequential solver compute to a single AR solve and yielding up to 3.2× fewer forward passes in latency-critical settings".

--- Page 1 ---

The gist The proposed framework introduces Stitching Noisy Diffusion Thoughts, a self-consistency mechanism that converts cheap diffusion-sampled reasoning into a reusable pool of step-level candidates to improve test-time scaling for large language models.

How it works

The method proposes a modular pipeline that separates exploration from synthesis: "we first generate a diverse pool of low-cost thoughts via a diffusion model (Section 3.2), evaluate their utility at the step level using an off-the-shelf process reward model (Section 3.3), and finally employ a stitching mechanism (Section 3.4) to recombine high-confidence steps into a coherent rationale that guides an AR solver to recompute the final answer".

Improvements for AI systems

  1. Bold header: Step-level self-consistency framework

This framework replaces trajectory-level selection with PRM-scored step selection and stitching over diffusion-sampled chains of thought, allowing for step-level recombination to recover useful intermediate work from partial attempts.

  1. Bold header: Modular pipeline for reasoning

The system separates roles into exploration (diffusion), evaluation (PRM scoring), and solution synthesis (AR solver), which enables a modular pipeline separates exploration from evaluation and solution synthesis, avoiding monolithic unified hybrids while preserving broad search.

  1. Bold header: Enhanced test-time scaling

The approach improves accuracy by up to 30.6% over vanilla diffusion decoding while achieving up to a 1.8× latency reduction relative to both traditional diffusion models (e.g., Dream, LLaDA) and unified architectures (e.g., TiDAR).

  1. Bold header: Robust low-confidence sampling

The pipeline remains accurate even when the diffusion sampler is run with a low sampling confidence, because it scores each step and filter out those of low quality, allowing for operation with a lower confidence threshold to substantially reduce latency.

  1. Bold header: Code generation specialization

The framework extends to code by using different step definitions, such as treating each line in the code as a step for MBPP, resulting in concise solutions with 3.21× fewer forward passes than the baseline.

Abstract

Reasoning with large language models often benefits from generating multiple chains-of-thought, but existing aggregation strategies are typically trajectory-level (e.g., selecting the best trace or voting on the final answer), discarding useful intermediate work from partial or "nearly correct" attempts. We propose Stitching Noisy Diffusion Thoughts, a self-consistency framework that turns cheap diffusion-sampled reasoning into a reusable pool of step-level candidates. Given a problem, we (i) sample many diverse, low-cost reasoning trajectories using a masked diffusion language model, (ii) score every intermediate step with an off-the-shelf process reward model (PRM), and (iii) stitch these highest-quality steps across trajectories into a composite rationale. This rationale is then used to recompute only the final answer. This modular pipeline separates exploration (diffusion) from evaluation and solution synthesis, avoiding monolithic unified hybrids while preserving broad search. Across math reasoning benchmarks, we find that step-level recombination is most beneficial on harder problems, and ablations highlight the importance of the final solver in converting stitched but imperfect rationales into accurate answers. Using low-confidence diffusion sampling with parallel, independent rollouts, our training-free framework improves average accuracy by up to 23.8% across six math and coding tasks. At the same time, it achieves up to a 1.8x latency reduction relative to both traditional diffusion models (e.g., Dream, LLaDA) and unified architectures (e.g., TiDAR). The code is available https://github.com/roymiles/diffusion-stitching

Related papers