How to Guide Your Language Flow
cs.LG, cs.AI
Submitted: 2026-09-16
Updated: 2026-09-16
License: http://creativecommons.org/licenses/by/4.0/
The gist: We introduce a new method to guide flow matching models.
Terminology
Abstract
We introduce a new method to guide flow matching models. Our approach, which we call probe guidance, uses the frozen internal states of an existing diffusion model to construct a guidance signal. This works using a similar principle as autoguidance, but eliminates the need for an additional forward pass at inference time and provides a reliable path to ensure that the weak and strong model share similar dynamics. We apply and benchmark this method on continuous diffusion language models, where probe guidance sets a new state-of-the-art performance on unconditional generation. When applied to a 1.7B diffusion language model, probe guidance consistently improves on multiple choice question answering benchmarks. Using our probes, we study the traditional autoguidance setting where the strong model is a weak checkpoint, and find that the weak model must come from a low-entropy region of training. These findings both provide a practical way to improve diffusion language models and shed light on the actual mechanism behind autoguidance, which is currently poorly understood.
Sources
- Classifier-Free Guidance is a Predictor-Corrector
- Analog Bits: Generating Discrete Data using Diffusion Models with Self-Conditioning
- LangFlow: Continuous Diffusion Rivals Discrete in Language Modeling
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- Scaling Categorical Flow Maps
- Continuous diffusion for categorical data
- A Frozen Pixel-Space Diffusion Model Can Guide Itself with Its Own Samples
- Classifier-Free Diffusion Guidance
- ELF: Embedded Language Flows
- The Principles of Diffusion Models
- Flow Map Language Models: One-step Language Modeling via Continuous Denoising
- A Survey on Diffusion Language Models
- Flow Matching for Generative Modeling
- Generative Frontiers: Why Evaluation Matters for Diffusion Language Models
- Categorical Flow Maps
- Scaling Beyond Masked Diffusion Language Models
- Score-Based Generative Modeling through Stochastic Differential Equations
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks