OSPO: Object-Centric Self-Improving Preference Optimization for Text-to-Image Generation
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "OSPO: Object-Centric Self-Improving Preference Optimization for Text-to-Image Generation".
Tom: Object-centric Self-improving Preference Optimization (OSPO) is a self-improving framework designed to enhance object-level text–image alignment for Text-to-Image (T2I) generation, specifically targeting object hallucination and fine-grained alignment failures.
Jane: First, who's behind it and why it matters.
Title and authors: Tom: So we're diving into the details of this paper, "OSPO: Object-Centric Self-Improving Preference Optimization for Text-to-Image Generation," and looking at what it actually proposes in terms of its title. It sounds pretty intense, doesn't it?
Jane: It definitely does sound intense, Tom; focusing on object centric optimization suggests they’re trying to fix a really specific problem in image generation where things often get fuzzy around what objects look like.
Lu: The title itself tells us that the core idea is to move beyond just general text-image alignment and focus directly on making sure the model understands individual objects better during the training process.
Meng: I see that, but I wonder how they manage to make this object centric approach work without needing massive amounts of external preference data, which is a big question for practical deployment.
Lalam: It’s smart they put "Self-Improving" in there; that suggests the system isn't just taking instructions and doing math on them, but it has a mechanism to actually get better on its own over time.
The paper's summary: Tom: Now, let's talk about what OSPO actually does inside this framework. Essentially, the paper outlines a five-stage process where the AI autonomously builds its own high-quality preference data without needing any outside human help to do the heavy lifting.
Jane: That’s wild; so it generates prompts, creates variations, makes images, and then uses self-generated questions to check which image is better—all in one loop. It’s a closed system for generating its own supervision.
Lu: The key thing they do in the middle stages is take a base prompt and intentionally mess with it using strategies like replacing or swapping objects, then densifying those pairs so that both prompts share the same general context but are different at the object level.
Meng: That perturbation step sounds incredibly complex computationally; generating multiple enriched prompt pairs for every initial idea must be quite demanding on the underlying MLLM's processing power during that second stage.
Lalam: But that’s the point, Meng; it forces the AI to learn exactly what those fine-grained details are, which is how we want our models to improve their precision over time.
The paper's improvements: Tom: Focusing on the actual results and methodology of "OSPO: Object-Centric Self-Improving Preference Optimization for Text-to-Image Generation," the authors highlight a few key methodological improvements that set it apart from previous methods.
Jane: One major improvement seems to be how they use attention masks extracted from intermediate layers to create an object mask, which is a really clever way for the AI to visually pinpoint where specific objects are in an image.
Lu: That object mask generation is crucial because it allows the framework to guide the optimization directly at the level of individual visual tokens, which seems much more effective than just looking at global image scores.
Meng: From an engineering standpoint, managing that attention weight extraction and binarization process must be quite delicate; if the mask isn't accurate, the subsequent loss function won't work well, so stability is a major concern.
Lalam: I think this whole approach could fundamentally change our culture in AI because it means we can build systems that are intrinsically more reliable and less prone to making up objects when creating complex visuals.
Conclusion: Tom: So, wrapping up our discussion on "OSPO: Object-Centric Self-Improving Preference Optimization for Text-to-Image Generation," the authors show that this five-stage pipeline successfully creates its own preference data and uses a specialized loss function for object optimization.
Jane: It really is a fantastic paper because it demonstrates a way to build high-quality training supervision autonomously, which simplifies so much of the current workflow for those of us trying to improve generative models.
Lu: The real takeaway here is that we’re moving toward generative models that truly grasp spatial relationships and compositional details rather than just generating visually pleasing, but ultimately shallow, pictures.
Meng: I'm still impressed by how they manage the engineering complexity of generating those object masks during training without needing massive pre-labeled datasets to start with, which speaks to scalability.
Lalam: This whole development could redefine generative AI by making it inherently more reliable and detailed, pushing us toward systems that possess genuine object-level understanding instead of just superficial visual mimicry.
Yoonjin Oh, Yongjin Kim, Hyomin Kim, Donghwan Chi, Sungwoong Kim
Korea University
cs.CV
Submitted: 2025-05-28
Updated: 2026-09-25
Importance score: 83/100
The gist: Object-centric Self-improving Preference Optimization (OSPO) is a self-improving framework designed to enhance object-level text–image alignment for Text-to-Image (T2I) generation, specifically
Key concepts
- OSPO
- Object-Centric Self-Improving Preference Optimization is a self-improving framework designed to enhance object-level text–image alignment for Text-to-Image generation. It targets problems like object hallucination and fine-grained alignment failures by allowing the system to improve on its own.
- Self-Improving Framework
- This suggests the AI system has a mechanism to get better over time without needing massive amounts of external human preference data. It autonomously builds its own supervision through a closed loop where it generates prompts, images, and uses self-generated questions to check quality.
- Object Mask Generation
- The authors use attention masks extracted from intermediate layers to create an object mask. This is a clever method that allows the framework to visually pinpoint specific objects within an image, guiding optimization directly at the level of individual visual tokens.
- Fine-Grained Alignment Failures
- This refers to issues in text-to-image generation where the model struggles with precise details regarding objects. OSPO specifically targets these failures by focusing on making sure the model understands individual objects better during training.
Terminology
Summary
Object-centric Self-improving Preference Optimization (OSPO) is a self-improving framework designed to enhance object-level text–image alignment for Text-to-Image (T2I) generation, specifically targeting object hallucination and fine-grained alignment failures. OSPO is a five-stage framework that autonomously constructs fine-grained, object-focused T2I preference pairs without relying on any external data or auxiliary models.
The five stages of the OSPO framework are as follows:
-
First, it generates initial text prompts categorized into four semantic types: Attribute, Layout, Non-spatial Relationship, and Complex Composition.
-
Second, it perturbs each text prompt into multiple variations using three strategies: Replace (substituting an object or attribute with one not originally present), Swap (exchanging the positions of objects or attributes), and Drop (removing an object or attribute). Each original–perturbed pair is jointly densified by the MLLM.
-
Third, it generates candidate preferred and non-preferred images from each densified prompt pair. During image generation, a binary token-level mask, referred to as the Object Mask, is extracted from attention weights to indicate whether each visual token belongs to the object region of the image. This mask is obtained by averaging attention distributions across heads and layers and binarizing it using OTSU’s adaptive thresholding method. This process is repeated for all objects described in the original prompt, and a single object mask M is obtained for each generated image by taking the union of these masks.
-
Fourth, it performs Self-VQA using self-generated decompositional questions to evaluate each candidate image's prompt fidelity. The alignment score S(y) for an image y is computed as the average probability margin between the Yes and No responses across a set of atomic semantic elements:
sk (y) = p(yes y, qk) − p(no qk), S(y) = PK / k=1 sk (y).
The framework filters out noisy supervision by discarding pairs if the aggregated alignment score for the preferred image falls below a threshold, and similarly discards non-preferred images if they satisfy sk (yl) > 0 for all k. -
Finally, it performs object-centric preference optimization using an object-weighted loss derived from the predicted object masks. The final training objective is defined as:
LOSPO = LObj-SimPO + λ LSFT. Here, λ is a loss weight for SFT loss, which we set λ = 2 by default.
The framework explicitly constructs object-centric preference data without relying on external data or external models. It introduces a new approach that leverages attention-based object masks together with an object-weighted SimPO loss to enhance object-specific fidelity. The Object-weighted SimPO loss is defined as: w(yw) LObj-SimPO = − E(x,yw,yl)∼D log σ log πθ (yw x) yw ! w(yl) − log πθ (yl x) − γ, yl
where the weighting scheme strengthens gradients on object-relevant tokens by defining wt = β · (1 + α mt), mt ∈ [0, 1], where t denotes token position and α and β control the emphasis on object-relevant visual tokens.
The main contributions of OSPO include:
We present OSPO, a five-stage self-improving framework that mitigates object hallucination in T2I generation without relying on either external datasets or auxiliary models.
"We propose a pipeline that constructs high-quality, object-centric preference data with shared global semantics but fine-grained local differences, and incorporates object-aware supervision using attention-based object masks with an object-weighted SimPO loss."
We demonstrate the empirical effectiveness of OSPO with substantial object-level alignment improvements across multiple T2I benchmarks, surpassing prior self-improving and specialized diffusion models.
Extensive experiments on three compositional image generation benchmarks demonstrate that OSPO significantly improves fine-grained alignment and reduces object hallucination, outperforming prior self-improving methods and even specialized diffusion-based text-to-image models.
The framework's performance is analyzed across various factors, including sample size (showing consistent performance improvement as the total number of training samples increases), candidate image pair size (where sampling multiple candidate pairs per prompt leads to consistent performance improvements), and the impact of prompt densification, which was shown to consistently yield higher performance on both benchmarks.
Furthermore, OSPO achieves comparable performance to other methods like T2I-R1 and FocusDiff but with substantially lower computational cost, demonstrating its efficiency.
The results show substantial improvements in the Attribute category across benchmarks.
Improvements for AI systems
Here are the specific improvements that can be made to AI systems by implementing the Object-centric Self-improving Preference Optimization (OSPO) framework, and what these improved systems will be capable of:
The implementation of OSPO transforms existing Text-to-Image (T2I) generation models by moving beyond general visual coherence to achieve high levels of semantic fidelity at the object level. The resulting AI systems will possess the following capabilities:
- Enhanced Fine-Grained Alignment and Attribute Precision:
Large language models (MLLMs) will no longer hallucinate or distort specific object attributes (color, shape, texture) because OSPO explicitly enforces object-centric preference optimization using an object-weighted SimPO loss.
- Mitigation of Object Hallucination and Omission Errors:
The system will significantly reduce the generation of non-existent objects or the omission/distortion of objects described in complex prompts by constructing highly reliable, object-specific training data from scratch, eliminating reliance on external preference datasets.
- Superior Spatial Relationship Modeling:
By decomposing prompts into categories like Layout
(2D/3D spatial relationships) and optimizing based on object masks derived from attention weights, the model will demonstrate a substantially improved ability to render accurate spatial configurations of objects within an image.
- Robustness to Compositional Complexity:
The framework is specifically designed to handle Complex Composition
prompts by utilizing perturbation strategies (Replace, Swap, Drop) and densification techniques that ensure generated pairs share global context while differing only in fine-grained object details, leading to better handling of long and intricate text inputs.
- Self-Improving Capability without External Data Dependency:
The system will operate in a self-improving loop by autonomously generating its own high-quality preference signals (using Self-VQA) and optimizing its generation process, drastically reducing the cost, complexity, and scalability issues associated with traditional human/AI preference data collection.
- Improved Training Efficiency through Targeted Loss Functions:
The combination of an Object-weighted SimPO loss (emphasizing object-relevant tokens via spatial weighting) and a Supervised Fine-Tuning (SFT) loss will provide complementary guidance, ensuring that the model learns both fine-grained object details and globally coherent visual structure simultaneously.
In summary, the improved AI system will be capable of generating text-to-image outputs that are not only aesthetically pleasing but also semantically faithful to highly detailed textual descriptions—a critical leap from current MLLMs which often struggle with precise object attributes and spatial accuracy.
Sources
- T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT
- UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning
- FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL
- SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation
- Microsoft COCO Captions: Data Collection and Evaluation Server
- Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling
- Davidsonian Scene Graph: Improving Reliability in Fine-grained Evaluation for Text-to-Image Generation
- Making LLaMA SEE and Draw with SEED Tokenizer
- CLIPScore: A Reference-free Evaluation Metric for Image Captioning
- ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment
- Show-o: One Single Transformer to Unify Multimodal Understanding and Generation
- Show-o2: Improved Native Unified Multimodal Models
- Self-Rewarding Large Vision-Language Models for Optimizing Prompts in Text-to-Image Generation
- Idea2Img: Iterative Self-Refinement with GPT-4V(ision) for Automatic Image Design and Generation
- mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
- RLAIF-V: Open-Source AI Feedback Leads to Super GPT-4V Trustworthiness
- Proximal Policy Optimization Algorithms
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Emu: Generative Pretraining in Multimodality
- Chameleon: Mixed-Modal Early-Fusion Foundation Models
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models