CompArt: Operationalizing Aesthetic Alignment in Text-to-Image Generation via Principles of Art
cs.CV, cs.AI
Submitted: 2025-03-15
Updated: 2026-09-16
Journal ref: Proceedings of the 2026 International Conference on Multimedia Retrieval (ICMR '26), pp. 2200-2209, Association for Computing Machinery (ACM), 2026
Code: https://github.com/jin-zhe/ArtDapter
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
The gist: Text-to-Image (T2I) diffusion models have made rapid progress on semantic alignment (generating what is described in the prompt), yet users still lack reliable control over aesthetic composition (how
Terminology
Abstract
Text-to-Image (T2I) diffusion models have made rapid progress on semantic alignment (generating what is described in the prompt), yet users still lack reliable control over aesthetic composition (how visual elements are put together). Prior work often treats aesthetics as a single, preference-driven notion (e.g., "high quality", "detailed", "breathtaking"), which does not map cleanly to compositional intent. We propose Aesthetic Alignment: aligning generated images to explicit, user-specified compositional constraints. We operationalize these constraints using the Principles of Art (PoA)-e.g., Balance, Rhythm, and Emphasis-commonly used in art education to describe composition. To support this task, we introduce CompArt, a dataset of 80,032 WikiArt images augmented with captions and PoA analyses produced by a multimodal LLM under structured prompting. We further propose ArtDapter, a lightweight and disentangled adapter that enables steering a pretrained T2I model along 10 PoA dimensions while retaining the base model's semantic capability. Experiments on CompArt show improved adherence to PoA controls over strong baselines under a dual evaluation protocol.
Sources
- Microsoft COCO Captions: Data Collection and Evaluation Server
- Scaling Instruction-Finetuned Language Models
- eDiff-I: Text-to-Image Diffusion Models with an Ensemble of Expert Denoisers
- Training-Free Structured Diffusion Guidance for Compositional Text-to-Image Synthesis
- Personalizing Text-to-Image Generation via Aesthetic Gradients
- Divide & Bind Your Attention for Improved Generative Semantic Nursing
- M6: A Chinese Multimodal Pretrainer
- Compositional Visual Generation with Composable Diffusion Models
- ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment
- T2I-CompBench++: An Enhanced and Comprehensive Benchmark for Compositional Text-to-image Generation
- GPT-4 Technical Report
- SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis
- Zero-Shot Text-to-Image Generation
- Hierarchical Text-Conditional Image Generation with CLIP Latents
- Linguistic Binding in Diffusion Models: Enhancing Attribute Correspondence through Attention Map Alignment
- DreamSync: Aligning Text-to-Image Generation with Image Understanding Feedback
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- U-Net: Convolutional Networks for Biomedical Image Segmentation
- Large-scale Classification of Fine-Art Paintings: Learning The Right Metric on The Right Feature
- LAION-5B: An open large-scale dataset for training next generation image-text models
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models