RFG: Self-Improving Diffusion Large Language Models with Reward-Free Guidance
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "RFG: Self-Improving Diffusion Large Language Models with Reward-Free Guidance".
Tom: The gist: Reward-free guidance (RFG) is a principled method for guiding the reasoning trajectory of diffusion large language models (dLLMs) without explicit process reward,
Jane: First, who's behind it and why it matters.
Title and authors: Tom: So we’re looking at this paper, "RFG: Self-Improving Diffusion Large Language Models with Reward-Free Guidance," and it’s tackling something big in how we guide these diffusion large language models for reasoning.
Jane: Right, Tom, the authors are proposing a way to guide the thinking process of these models without needing an explicit reward model for every single step. It's a principled method they call reward-free guidance, or RFG.
Lu: What’s really interesting is that they bypass the need for those dense annotations you usually have to build a process reward model for each intermediate step, which is tough because diffusion models generate text in an any-order fashion and those intermediate states are just partially masked sentences.
Meng: So, if I'm hearing right, they’re using something else to figure out what a good reasoning step looks like? Something that doesn't require training a separate reward model specifically for the reasoning path?
Tom: Exactly. They parameterize this process reward by using the log-likelihood ratios of two models: an enhanced version of the language model and a reference version, which is just another language model.
Jane: And they set up this specific formulation where the enhanced model can be any off-the-shelf diffusion large language model that’s been post-trained with reinforcement learning or supervised fine-tuning.
Lu: That means you don't need to retrain anything new just to get this guidance, because you leverage existing models that already have some form of training applied to them.
Meng: So it’s a kind of test-time scaling technique, then? Using the difference between those two model likelihoods to steer the generation during sampling.
Tom: That’s right. They formulate this as a guidance strength, w, which they control by balancing the policy model and the reference model log probabilities.
Jane: The core mathematical idea is that you can get a guided log probability by taking the original log probability and adding a term involving that guidance strength w.
Lu: And they show how this step reward can be written as the difference between two quantities, Q t and Q t+one which allows them to do reward-guided sampling by reweighting those denoising transitions <ref:2509.25604#pg1>.
Meng: That sounds complex for what it is, but the practical implication I see is that it lets us get better reasoning from models we already have without spending time and resources training an entirely new reward mechanism.
Tom: The results they show are pretty strong across four benchmarks—GSM8K, MATH-five hundred for math, and HumanEval and MBPP for code generation. They report accuracy gains of up to nine point two percent on those tasks.
Jane: And what’s really compelling is that the paper confirms these improvements aren't just from making the models bigger or faster; they come from this principled guidance framework instead of just scaling up computation.
Lu: It shows that this approach works well even when you look at different post-training methods on the base diffusion model, confirming its robustness across various training setups.
Title and authors: Meng: I noticed they mention that it performs better than a naive ensemble baseline with the same amount of compute, which is a good comparison for showing why this method is useful.
Tom: It’s definitely not just about getting a few extra points on a leaderboard; they’re showing that this RFG framework can produce more coherent multi-step derivations in math and code that are actually more robust.
Jane: And the qualitative analysis suggests the generated code isn't just syntactically correct, but it actually handles things like missing edge conditions better.
Lu: I think what’s really exciting here is that this method is fully model-agnostic and training-free, meaning you can apply it to any diffusion large language model that already exists.
Meng: That’s a huge practical thing for deployment. If we can use it without retraining, the path to improving reasoning capabilities becomes much faster.
Tom: So we’re moving from needing a reward model trained on the whole trajectory to just using those log-likelihood ratios as a shortcut for guidance during sampling.
Jane: The interpretation they give is that this formulation, pi RFG = pi theta + w(pi theta - pi ref), essentially steers the enhanced model's distribution along a direction defined by the difference between the two models with strength w.
Lu: And they suggest that choosing a negative value for w, specifically between minus one and zero, effectively takes a step back from the over-optimized model to correct for that potential over-optimization.
Meng: That makes sense, it’s like fine-tuning the guidance itself to find a sweet spot where you don't just blindly follow the policy model too closely.
Tom: So this RFG paper really establishes itself as a general test-time scaling framework for diffusion large language models that doesn't rely on external reward models.
Jane: It sets up a foundation for how we can potentially improve reasoning in generative models, and they see it as a scalable alternative to the cost of retraining.
Lu: And they’re pointing toward this being a potential foundation for broader alignment improvements, even into multimodal diffusion and agentic reasoning systems down the road.
Meng: I just wonder where we can apply this most immediately in production; is it something that fits into our existing inference pipeline right now?
Tom: The results strongly suggest that the gain comes from this principled guidance framework itself, not just brute-force computational scaling of the model.
Jane: So to wrap up, RFG gives us a way to guide diffusion large language models during testing using existing model likelihoods without needing a separate reward training process.
Lu: It’s a powerful test-time scaling technique that shows the effectiveness and robustness of this guidance method across different post-training methods.
Meng: The main thing for me is the model agnosticism; it means we can use whatever instruction-tuned or RL-enhanced model we have right now to get better performance without any extra training work.
Tom: That’s the takeaway: RFG provides a practical, training-free way to scale test-time reasoning, and that’s where we should focus our attention next.
The paper's summary: Tom: So, we’re looking at this paper now, RFG: Self-Improving Diffusion Large Language Models with Reward-Free Guidance. Basically, they’ve figured out a way to guide these massive language models during the generation process without having to train a separate reward model for every single step.
Jane: That's right, Tom. They’re using the existing models themselves—specifically an enhanced version and a reference version—to create this guidance signal. It lets you steer the model toward better reasoning just by comparing how those two different versions of the AI perform on a given task.
Lu: The real clever part is how they take that comparison, which is already there, and they break it down into these smaller pieces we call PRMs for each denoising step. It’s like taking one big idea about what a good final answer looks like and distributing that instruction across every single tiny move the AI makes while it's writing.
Meng: So, if I'm hearing you right, they’re not building a brand new reward system from scratch for math or code; they’re repurposing the existing likelihood ratios of the models we already have to create a guide for test time. It sounds like a huge win for practical application because it doesn't require massive new training runs just to get better results.
Tom: Exactly, Meng. It’s fully model-agnostic, which means you can plug in almost any instruction-tuned or RL-enhanced diffusion model you have and get this guidance without any extra training on your side. The paper shows that this method consistently boosts performance across math and code benchmarks by a solid nine point two percent.
Jane: And what’s important here is the interpretation of the guidance formula itself, pi RFG = pi theta + w(pi theta - pi ref). This just means you’re taking that original model's probability and nudging it in a specific direction defined by the difference between the enhanced and reference models, controlled by that strength w.
Lu: And they even found a sweet spot for that guidance strength, w, where performance stays strong across a wide range of values. That means you don't have to get bogged down trying to perfectly tune that parameter; you can find a good balance without intensive hyperparameter hunting.
Tom: That’s what the authors emphasized, and it’s huge for deployment because it makes scaling test-time reasoning much more accessible than we thought. It moves the focus from just brute-force computation to this principled framework for steering the AI's thinking.
Jane: So, what does this actually change for someone who just listens to you? It means that if you’re using a diffusion model right now for complex reasoning, you could potentially improve its output quality during testing without having to go through the whole costly cycle of retraining and fine-tuning a dedicated reward mechanism.
Lu: And the vision here goes beyond just simple accuracy numbers. They showed that this guidance actually leads to more coherent multi-step derivations in math and code, meaning the AI isn't just getting lucky on a single answer, it’s building a better logical path for itself.
Meng: That coherence is what matters in production environments; we need reliable reasoning, not just random correct outputs. If you can make the AI's internal thought process more robust through this method, that could significantly reduce the kinds of subtle errors we see in complex code generation or multi-step logic problems.
Tom: It really sets up a foundation for how we think about alignment and reasoning improvements across all generative models, not just diffusion ones. This test-time guidance feels like a scalable alternative to the expensive retraining cycles we’ve been doing.
Jane: And looking ahead, they suggest this could be a starting point for even bigger things, maybe extending this concept into multimodal diffusion or even agentic systems where the AI needs to plan and reason across different types of data.
Lu: It opens up a whole new avenue for how we can build systems that are better at complex tasks because we’re giving them a principled way to self-correct during the generation itself.
The paper's improvements: Tom: So, we’re talking about how this method improves itself—the self-improving part of RFG—and what that actually means for making these models smarter over time. Basically, they show how you can use the guidance mechanism to actually adjust the policy model itself during sampling.
Jane: Right, Tom. The paper suggests that by using this reward-free guidance formula, you’re not just getting a better answer in one go; you're steering the model's internal trajectory toward a higher quality path for every single token it generates.
Lu: They show that this can be done without needing to re-train the whole policy model from scratch, because the guidance mechanism is built on top of existing RL or SFT models. It’s like giving the model a very smart steering wheel during its journey instead of just setting a fixed speed.
Meng: So, if I were running this in production, what does that mean for my engineering pipeline? Does it mean we can run inference faster because the guidance is built into the sampling process rather than being an extra computational layer we have to add on top?
Tom: That’s a fair question, Meng. The paper points out that this formulation is designed to be applied during reweighted denoising transitions, which means it integrates into the existing diffusion sampling steps without adding huge overhead to the inference process itself.
Jane: It also addresses the issue of over-optimization that some models can get stuck in where they just keep repeating what they already know, by using that guidance strength w to take a step back from being too aggressively tuned.
Lu: The vision here is really about creating a more dynamic, adaptive reasoning system. Instead of a static model that gives the same output every time, RFG suggests we can create models that actively optimize their own thinking based on how they are performing against our reference standard.
Tom: And the results confirm this isn't just theoretical tweaking; when they tested it on those challenging reasoning tasks like GSM8K and MATH-five hundred the improvements were substantial—up to nine point two percent across the board.
Jane: So, for someone listening who just wants to know what’s new: this paper shows a training-free way to improve AI reasoning by using existing models as a guide during testing, leading to more logical and robust outputs in math and code.
Meng: That robustness is key for me. When an AI writes code or solves a complex problem, I don't want it to fail on subtle edge cases; this method seems designed specifically to push the model toward those more complete solutions.
Lu: It feels like we’re moving toward a system where the AI isn't just memorizing patterns but is actually learning *how* to reason better by following this reward-free path. Imagine that for creative tasks, or even agentic systems where planning is crucial.
Tom: Exactly, Lu. This framework establishes RFG as a general way to scale reasoning capabilities at test time without having to build a whole new alignment training pipeline every single time we want an upgrade.
Jane: So, it’s about making the path to better reasoning much more scalable and less dependent on massive retraining efforts for every small gain.
Lu: And that scalability is what makes this interesting for the wider AI community; it offers a general framework that can apply across different types of diffusion models and even hint at future applications in agentic reasoning.
Conclusion: Tom: So, we’re wrapping up this deep dive on RFG: Self-Improving Diffusion Large Language Models with Reward-Free Guidance. It really boils down to this principled method for guiding diffusion models during testing without needing a separate reward model trained on the whole process.
Jane: That’s right, Tom. The core idea is using those likelihood ratios between an enhanced and a reference model to steer the generation based on that guidance strength w we talked about earlier. It’s about test-time reasoning improvement without the heavy cost of retraining for alignment.
Lu: It opens up a whole new space for how we think about aligning these systems; it suggests that test-time guidance could be a scalable alternative to costly retraining cycles when improving complex reasoning abilities.
Meng: I think the practical implication is that we can deploy models faster and get better quality outputs without having to run massive, multi-stage reinforcement learning training just for every minor performance bump. It makes the process much more efficient for building robust systems.
Lalam: From my perspective, this means our ability to create more coherent and reliable content is going up because we’re giving the AI a better internal compass during its work; it helps us build a culture where the AI's outputs are consistently high quality across all tasks.
Tom: It really does. The authors show that even with this test-time guidance, you get solid gains on benchmarks like GSM8K and HumanEval, which is pretty impressive for a training-free method.
Jane: And remember, the results also showed that this method doesn't rely on intensive hyperparameter tuning to get good performance; there’s a wide plateau of strong results across many values of w.
Lu: That robustness is what makes it compelling; it suggests the framework itself is solid, not just a lucky result from one specific training setup. The future work they hint at points toward applying this to even more complex things like multimodal diffusion or agentic reasoning systems down the road.
Meng: I agree, Lu. If we can generalize this guidance mechanism beyond text generation into visual tasks or planning tasks, that would be a huge win for practical engineering applications right now.
Lalam: For me, it means the models we build are going to be more trustworthy because their reasoning steps are inherently guided toward better logical outcomes rather than just random guessing.
Tom: That’s the big picture, folks. RFG really establishes itself as a solid foundation for scalable reasoning improvements in diffusion large language models using existing likelihood ratios during testing.
Jane: It’s a lot to take in, but at its heart it’s about finding a clever way to guide the AI's thinking without needing that explicit reward training process we usually have to do.
Tianlang Chen, Minkai Xu, Jure Leskovec, Stefano Ermon
Stanford University
cs.CL, cs.LG
Submitted: 2025-09-29
Updated: 2026-10-04
Importance score: 86/100
The gist: The gist: Reward-free guidance (RFG) is a principled method for guiding the reasoning trajectory of diffusion large language models (dLLMs) without explicit process reward, achieving significant
Key concepts
- Reward-free guidance (RFG)
- RFG is a technique that guides diffusion models during sampling by using log-likelihood ratios of different models instead of training an explicit reward model. It reparameterizes the process reward to steer the generation trajectory towards desired outcomes without needing external feedback or retraining.
- Log-likelihood ratio
- This concept involves calculating the difference in log probabilities between two models, such as an enhanced policy model and a reference model. This ratio is used to define a reward signal that indicates how much better one version of the model performs compared to another.
- Classifier-free guidance (CFG) analogy
- RFG draws inspiration from classifier-free guidance, which reparameterizes a hypothetical classifier as the likelihood ratio between conditional and unconditional diffusion models. This idea encourages the model to sample from a desired conditional distribution over an unconditional one by reweighting denoising steps.
- Guidance strength (w)
- The hyperparameter 'w' controls how strongly the guidance steers the sampling process. It is derived from the ratio of beta and gamma and determines how much the model leans on the log-likelihood difference to correct its trajectory, allowing users to tune performance without extensive retraining.
Terminology
Summary
The gist: Reward-free guidance (RFG) is a principled method for guiding the reasoning trajectory of diffusion large language models (dLLMs) without explicit process reward, achieving significant performance gains across mathematical reasoning and code generation benchmarks.
How it works
-
RFG parameterizes the process reward by using the
log-likelihood ratios of the enhanced and reference dLLMs
wherethe enhanced model can be easily obtained by any off-the-shelf dLLM that has been post-trained with reinforcement learning (RL) or supervised fine-tuning (SFT)
in page 1. -
The reward model is formulated as the
log-likelihood ratio of the policy and reference models rθ(x) = β log pθ(x) pref(x)
where both are implemented as dLLMs, which is a reparameterization widely adopted in various RL and preference optimization literature in page 3. -
The PRMs for each denoising step follow the same log-likelihood ratio form but are calculated on partially observed responses rθ(xt−1xt) = β log pθ(xt−1xt) pref(xt−1xt) where
intermediate generations are typically partial sentences with tokens masked in random positions
in page 3. -
The step reward is then written as rtθ(xt−1xt) = Qtθ − Qt+1θ = β log pθ(xt−1xt) pref(xt−1xt) where
with such PRMs, we can then conduct reward-guided sampling from dLLMs by reweighted denoising transitions
in page 4. -
The guided log probability is written as log p∗(xt−1xt) = log pθ(xt−1xt)e1/γ rtθ(xt−1xt) + C = (1 + w) log pθ(xt−1xt) − w log pref(xt−1xt) + C where
w = β/γ is a new hyperparameter to control the guidance strength
in page 4.
Connections to diffusion guidance method
"A widely adopted technique for guiding diffusion sampling is called classifier-free guidance (CFG) (Ho & Salimans, 2022), which pushes samples towards high class-confidence regions by reweighted denoising steps"
"The key idea of CFG is to reparameterize a hypothetical classifier as the likelihood ratio of conditional and unconditional diffusion models, which encourages the model to draw samples from density pconditional over punconditional"
"The key idea of RFG is that instead of training an explicit PRM on incomplete intermediate generations, we specially parameterize an outcome reward model (ORM) with pretrained dLLMs and freely decompose it to PRMs for each diffusion step"
Implementation and Validation
-
The practical implementation involves computing logits from both the policy model pθ and the reference model pref, then combining them using the guidance strength w towards the guided logit distribution as detailed in Algorithm 1 in page 5.
-
The framework is
fully model-agnostic and training-free
because pθ can be any off-the-shelf RL-enhanced or instruction-tuned model, allowing performance improvementwithout any additional training
in page 5. -
Experiments on four challenging benchmarks—GSM8K, MATH-500 for mathematical reasoning, and HumanEval, MBPP for code generation—show that RFG consistently yields
significant improvements across all tasks and model types,
achieving accuracy gains of up to 9.2% in page 1. -
Qualitative analysis shows that RFG produces
more coherent multi-step derivations
on mathematical reasoning tasks and generates code that isnot only syntactically correct but also more robust, reducing common errors such as missing edge conditions or incomplete logic
in page 7. -
Sensitivity analysis reveals the framework's robustness, showing a
wide plateau of strong performance across a broad range of w values,
indicating that substantial gains can be achievedwithout intensive hyperparameter tuning
in page 8.
Guidance Interpretation
"The RFG formulation, log πRFG = log πθ+w(log πθ − log πref), can be interpreted as steering the enhanced model’s distribution, log πθ, along an optimization direction defined by the difference (log πθ−log πref) with strength w"
"By selecting −1 < w < 0, the guidance effectively takes a step back from the over-optimized model, correcting for this over-optimization"
This framework establishes RFG as a general training-free framework that scales test-time reasoning without reliance on external reward models
in page 1.
The paper concludes by envisioning this framework as a "foundation for broader alignment and reasoning improvements in generative models, including multimodal diffusion and agentic reasoning systems, where test-time guidance offers a scalable and general alternative to costly retraining in page 9. The results confirm that the performance gain stems from the
principled guidance framework rather than mere computational scaling" in page 6.
<ref:2509.
Improvements for AI systems
-
textbfReweighting Denoising Transitions via RFG Formulation Instantly: The improved system can perform reward-guided sampling by utilizing a formulation that shows
by simply parameterizing a reward model on the whole diffusion sampling trajectory, we can freely obtain PRMs for each denoising step without any additional training.
This allows for test-time guidance usinglog πRFG = log πθ+w(log πθ − log πref).
-
textbfModel-Agnostic Test-Time Scaling: The system gains the capability to
significantly improve performance over the policy model itself without any additional training
by leveragingany off-the-shelf RL or instruction-tuned models
as the policy model pθ, demonstrating that RFG isfully model-agnostic and training-free.
-
textbfAdaptive Guidance Strength Tuning: The improved system can dynamically adjust its reasoning trajectory by utilizing the observation that it exhibits a
wide plateau of strong performance across a broad range of w values,
allowing it to find an optimal balance between the policy and reference models, even when the enhanced model isnot perfectly tuned for every target task.
Sources
- Block Diffusion: Interpolating Between Autoregressive and Diffusion Language Models
- Program Synthesis with Large Language Models
- Large Language Monkeys: Scaling Inference Compute with Repeated Sampling
- Evaluating Large Language Models Trained on Code
- Training Verifiers to Solve Math Word Problems
- Inference-Time Scaling of Diffusion Language Models via Trajectory Refinement
- Continuous diffusion for categorical data
- DiffuCoder: Understanding and Improving Masked Diffusion Models for Code Generation
- Bayesian Flow Networks
- DiffusionBERT: Improving Generative Masked Language Models with Diffusion Models
- Classifier-Free Diffusion Guidance
- Mercury: Ultra-Fast Language Models Based on Diffusion
- LaViDa: A Large Diffusion Language Model for Multimodal Understanding
- s1: Simple test-time scaling
- Large Language Diffusion Models
- Diffusion Beats Autoregressive in Data-Constrained Settings
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- A General Framework for Inference-time Scaling and Steering of Diffusion Models
- Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
- wd1: Weighted Policy Optimization for Reasoning in Diffusion Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering