Multi-Axis Max@K Reinforcement Learning for Representative Diversity in Text-to-Image Generation
cs.LG, cs.AI, cs.CV
Submitted: 2026-07-16
Updated: 2026-09-17
Comments: Accepted at WACV 2027
Code: https://github.com/KuOnoda/multi-axis-maxk
License: http://creativecommons.org/licenses/by/4.0/
The gist: Text-to-image (T2I) models can synthesize realistic, prompt-aligned images, yet samples generated for the same prompt often cover only a small subset of visually distinct modes.
Terminology
Abstract
Text-to-image (T2I) models can synthesize realistic, prompt-aligned images, yet samples generated for the same prompt often cover only a small subset of visually distinct modes. This limits diversity and, for person-centric prompts, can reflect or amplify demographic skew. We formalize this problem as target-mode coverage, the coverage of a predefined set of semantically specified modes, and propose multi-axis max@K, a group-based reinforcement learning objective for improving it in diffusion-based T2I models. Given a group of samples and one score per target mode, multi-axis max@K first takes the maximum score across samples for each mode and then sums these per-mode maxima. The resulting credit assignment gives a sample positive weight on a mode only when it raises that mode's group maximum, so different samples can contribute to different modes. We validate the credit-assignment mechanism on a synthetic mixture and on SD3.5-M with deterministic pixel-based color rewards, and then apply the same objective to perceived-appearance fairness. On held-out prompts, multi-axis max@K improves the Fairness Score by 0.23-0.36 over the base model under three automatic evaluators, while maintaining image quality and text alignment. Code is available at https://github.com/KuOnoda/multi-axis-maxk.
Sources
- Consistency-diversity-realism Pareto fronts of conditional image generative models
- The Best of N Worlds: Aligning Reinforcement Learning with Best-of-N Sampling via max@k Optimisation
- Vector Policy Optimization: Training for Diversity Improves Test-Time Search
- Qwen2.5-VL Technical Report
- HoloFair: Unified T2I Fairness Evaluation and Fair-GRPO Debiasing
- Inference-Aware Fine-Tuning for Best-of-N Sampling in Large Language Models
- Debiasing Vision-Language Models via Biased Prompts
- Fair Diffusion: Instructing Text-to-Image Generation Models on Fairness
- Improving Dynamic Object Interactions in Text-to-Video Generation with AI Feedback
- Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning
- TempFlow-GRPO: When Timing Matters for GRPO in Flow Models
- Debiasing Diffusion Model: Enhancing Fairness through Latent Representation Learning in Stable Diffusion Model
- FairGen: Enhancing Fairness in Text-to-Image Diffusion Models via Self-Discovering Latent Directions
- DiverseGRPO: Mitigating Mode Collapse in Image Generation via Diversity-Aware GRPO
- Beyond the Dirac Delta: Mitigating Diversity Collapse in Reinforcement Fine-Tuning for Versatile Image Generation
- Retry Policy Gradients in Continuous Action Spaces
- Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying
- OrderGrad: Optimizing Beyond the Mean with Order-Statistic Policy Gradient Estimation
- On Advantage Estimates for Max@K Policy Gradients
- Finite-Time Regret Analysis of Retry-Aware Bandits
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks