Policy-based Tuning of Autoregressive Image Models with Instance- and Distribution-Level Rewards
cs.LG, cs.CV
Submitted: 2026-03-24
Updated: 2026-07-09
Comments: Accepted at the European Conference on Computer Vision (ECCV), 2026. This version includes the appendix
Journal ref: Computer Vision - ECCV 2026, Lecture Notes in Computer Science, vol. 17070, pp. 320-339, Springer, 2026
DOI: 10.1007/978-3-032-37013-6_18
Code: https://github.com/bugrabaran/ar-
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Avoiding mode collapse in diffusion models fine-tuned with reinforcement learning
- Training Diffusion Models with Reinforcement Learning
- Directly Fine-Tuning Diffusion Models on Differentiable Rewards
- Fluid: Scaling Autoregressive Text-to-image Generative Models with Continuous Tokens
- DreamSim: Learning New Dimensions of Human Visual Similarity using Synthetic Data
- VA-$\pi$: Variational Policy Alignment for Pixel-Aware Autoregressive Generation
- STAGE: Stable and Generalizable GRPO for Autoregressive Image Generation
- Aligning Text-to-Image Diffusion Models with Reward Backpropagation
- Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding
- Proximal Policy Optimization Algorithms
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- ImageReFL: Balancing Quality and Diversity in Human-Aligned Diffusion Models
- Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation
- SimpleAR: Pushing the Frontier of Autoregressive Visual Generation through Pretraining, SFT, and RL
- Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis
- AR-GRPO: Training Autoregressive Image Generation Models via Reinforcement Learning
- Group Critical-token Policy Optimization for Autoregressive Image Generation
- Fine-Tuning Language Models from Human Preferences
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks