TPSO: Training-Free Diverse Image Generation via Semantic Prompt Embedding Optimization
cs.CV, cs.CL, cs.LG
Submitted: 2025-11-25
Updated: 2026-08-31
Comments: Accepted at the 2026 International Joint Conference on Neural Networks (IJCNN 2026). 8 pages, 5 figures
Code: https://github.com/Open-Debin/TPSO
License: http://creativecommons.org/licenses/by/4.0/
The gist: Image diversity remains a fundamental challenge for text-to-image diffusion models.
Terminology
Abstract
Image diversity remains a fundamental challenge for text-to-image diffusion models. Low-diversity generation often leads to repetitive outputs, increasing sampling redundancy and hindering both creative exploration and downstream applications. A key factor is the tendency of diffusion models to collapse toward strong modes in the learned distribution. Existing attempts to improve diversity, such as steering-based guidance, often introduce distortions that degrade image quality. To address this issue, we propose Token-Prompt Embedding Space Optimization (TPSO), a training-free and model-agnostic module. TPSO introduces learnable parameters to explore underrepresented regions of the token embedding space, reducing the tendency to repeatedly sample from strong modes of the distribution. Meanwhile, a prompt-level semantic constraint regulates distribution shifts, preventing quality degradation while preserving semantic fidelity. Extensive experiments on MS-COCO across three representative diffusion backbones demonstrate that TPSO substantially improves diversity, boosting performance from 1.10 to 4.18, while maintaining image quality with only a modest inference-time overhead of 3.6% to 8.9%. Code is available at: https://github.com/Open-Debin/TPSO.
Sources
- Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow
- MM2Latent: Text-to-facial image generation and editing in GANs with multimodal assistance
- Hierarchical Text-Conditional Image Generation with CLIP Latents
- CADS: Unleashing the Diversity of Diffusion Models through Condition-Annealed Sampling
- Denoising Diffusion Implicit Models
- Discovering Latent Graphs with GFlowNets for Diverse Conditional Image Generation
- Consistency-diversity-realism Pareto fronts of conditional image generative models
- The Crystal Ball Hypothesis in diffusion models: Anticipating object positions from initial noise
- Particle Guidance: non-I.I.D. Diverse Sampling with Diffusion Models
- The Vendi Score: A Diversity Evaluation Metric for Machine Learning
- Classifier-Free Diffusion Guidance
- Flow Matching for Generative Modeling
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models