CARINOX: Inference-time Scaling with Category-Aware Reward-based Initial Noise Optimization and Exploration
cs.CV, cs.CL
Submitted: 2025-09-22
Updated: 2026-09-05
Comments: Accepted to TMLR (2026)
Project page: https://amirkasaei.com/carinox
License: http://creativecommons.org/licenses/by/4.0/
The gist: Text-to-image diffusion models, such as Stable Diffusion, can produce high-quality and diverse images but often fail to achieve compositional alignment, particularly when prompts describe complex
Terminology
Abstract
Text-to-image diffusion models, such as Stable Diffusion, can produce high-quality and diverse images but often fail to achieve compositional alignment, particularly when prompts describe complex object relationships, attributes, or spatial arrangements. Recent inference-time approaches address this by optimizing or exploring the initial noise under the guidance of reward functions that score text-image alignment without requiring model fine-tuning. While promising, each strategy has intrinsic limitations when used alone: optimization can stall due to poor initialization or unfavorable search trajectories, whereas exploration may require a prohibitively large number of samples to locate a satisfactory output. Our analysis further shows that neither single reward metrics nor ad-hoc combinations reliably capture all aspects of compositionality, leading to weak or inconsistent guidance. To overcome these challenges, we present Category-Aware Reward-based Initial Noise Optimization and EXploration (CARINOX), a unified framework that combines noise optimization and exploration with a principled reward selection procedure grounded in correlation with human judgments. Evaluations on two complementary benchmarks covering diverse compositional challenges show that CARINOX raises average alignment scores by approximately 19% on both T2I-CompBench++ and HRS, consistently outperforming state-of-the-art optimization and exploration-based methods across all major categories, while preserving image quality and diversity. The project page is available at https://amirkasaei.com/carinox/.
Sources
- D-Flow: Differentiating through Flows for Controlled Generation
- Make It Count: Text-to-Image Generation with an Accurate Number of Objects
- Davidsonian Scene Graph: Improving Reliability in Fine-grained Evaluation for Text-to-Image Generation
- OIL-AD: An Anomaly Detection Framework for Sequential Decision Sequences
- Benchmarking Spatial Relationships in Text-to-Image Generation
- VersaT2I: Improving Text-to-Image Models with Versatile Reward
- CLIPScore: A Reference-free Evaluation Metric for Image Captioning
- Diffusion Model-Based Image Editing: A Survey
- CoMat: Aligning Text-to-Image Diffusion Model with Image-to-Text Concept Matching
- Counting Guidance for High Fidelity Text-to-Image Synthesis
- If at First You Don't Succeed, Try, Try Again: Faithful Diffusion-based Text-to-Image Generation by Selection
- Evaluating the Evaluators: Metrics for Compositional Text-to-Image Generation
- Inference-Time Scaling for Flow Models via Stochastic Generation and Rollover Budget Forcing
- Test-time Alignment of Diffusion Models without Reward Over-optimization
- All Seeds Are Not Equal: Enhancing Compositional Text-to-Image Generation with Reliable Random Seeds
- Divide & Bind Your Attention for Improved Generative Semantic Nursing
- ORIGEN: Zero-Shot 3D Orientation Grounding in Text-to-Image Generation
- T2I-Adapter: Learning Adapters to Dig out More Controllable Ability for Text-to-Image Diffusion Models
- GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models
- SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models