Amortizing intractable inference in diffusion models for vision, language, and control
cs.LG, cs.CV
Submitted: 2024-05-31
Updated: 2026-08-28
Comments: NeurIPS 2024; code: https://github.com/GFNOrg/diffusion-finetuning
DOI: 10.52202/079017-2422
Code: https://github.com/GFNOrg/diffusion-finetuning
License: http://creativecommons.org/licenses/by/4.0/
The gist: Diffusion models have emerged as effective distribution estimators in vision, language, and reinforcement learning, but their use as priors in downstream tasks poses an intractable posterior
Terminology
Abstract
Diffusion models have emerged as effective distribution estimators in vision, language, and reinforcement learning, but their use as priors in downstream tasks poses an intractable posterior inference problem. This paper studies amortized sampling of the posterior over data, x about p post(x) proportional to p(x)r(x), in a model that consists of a diffusion generative model prior p(x) and a black-box constraint or likelihood function r(x). We state and prove the asymptotic correctness of a data-free learning objective, relative trajectory balance, for training a diffusion model that samples from this posterior, a problem that existing methods solve only approximately or in restricted cases. Relative trajectory balance arises from the generative flow network perspective on diffusion models, which allows the use of deep reinforcement learning techniques to improve mode coverage. Experiments illustrate the broad potential of unbiased inference of arbitrary posteriors under diffusion priors: in vision (classifier guidance), language (infilling under a discrete diffusion LLM), and multimodal data (text-to-image generation). Beyond generative modeling, we apply relative trajectory balance to the problem of continuous control with a score-based behavior prior, achieving state-of-the-art results on benchmarks in offline reinforcement learning.
Sources
- Posterior samples of source galaxies in strong gravitational lenses with score-based priors
- From discrete-time policies to continuous-time diffusion samplers: Asymptotic equivalences and faster training
- Continuous diffusion for categorical data
- D4RL: Datasets for Deep Data-Driven Reinforcement Learning
- Neural Processes
- IDQL: Implicit Q-Learning as an Actor-Critic Method with Diffusion Policies
- Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems
- G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment
- Discrete Diffusion Modeling by Estimating the Ratios of the Data Distribution
- Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning
- Improved off-policy training of diffusion samplers
- Transition Path Sampling with Improved Off-Policy Training of Diffusion Path Samplers
- Fine-Tuning of Continuous-Time Diffusion Models as Entropy-Regularized Control
- Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation
- Text Infilling
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks