Test-time Prompt Refinement for Text-to-Image Models
cs.LG
Submitted: 2025-07-22
Updated: 2026-09-08
Comments: Accepted to ICCV 2025, MARS2 Workshop. Total 14 pages, 12 figures and 3 tables
Code: https://github.com/hafeezkhan909/Test-time-Image-Refinement
License: http://creativecommons.org/licenses/by/4.0/
The gist: Text-to-image (T2I) generation models have made significant strides but still struggle with prompt sensitivity: even minor changes in prompt wording can yield inconsistent or inaccurate outputs.
Terminology
Abstract
Text-to-image (T2I) generation models have made significant strides but still struggle with prompt sensitivity: even minor changes in prompt wording can yield inconsistent or inaccurate outputs. To address this challenge, we introduce a closed-loop, test-time prompt refinement framework that requires no additional training of the underlying T2I model, termed TIR. In our approach, each generation step is followed by a refinement step, where a pretrained multimodal large language model (MLLM) analyzes the output image and the user's prompt. The MLLM detects misalignments (e.g., missing objects, incorrect attributes) and produces a refined and physically grounded prompt for the next round of image generation. By iteratively refining the prompt and verifying alignment between the prompt and the image, TIR corrects errors, mirroring the iterative refinement process of human artists. We demonstrate that this closed-loop strategy improves alignment and visual coherence across multiple benchmark datasets, all while maintaining plug-and-play integration with black-box T2I models. Code is available at https://github.com/hafeezkhan909/Test-time-Image-Refinement.
Sources
- eDiff-I: Text-to-Image Diffusion Models with an Ensemble of Expert Denoisers
- LLM Blueprint: Enabling Text-to-Image Generation with Complex and Detailed Prompts
- Prompt-to-Prompt Image Editing with Cross Attention Control
- GPT-4o System Card
- Zero-shot Text-guided Infinite Image Synthesis with LLM guidance
- Aligning Text-to-Image Models using Human Feedback
- LLM-grounded Diffusion: Enhancing Prompt Understanding of Text-to-Image Diffusion Models with Large Language Models
- VideoDirectorGPT: Consistent Multi-scene Video Generation via LLM-Guided Planning
- Inference-Time Scaling for Diffusion Models beyond Scaling Denoising Steps
- PhyBench: A Physical Commonsense Benchmark for Evaluating Text-to-Image Models
- GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models
- SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis
- DiffusionAgent: Navigating Expert Models for Agentic Image Generation
- Hierarchical Text-Conditional Image Generation with CLIP Latents
- Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
- Qwen2 Technical Report
- Controllable Text-to-Image Generation with GPT-4
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks