CompArt: Operationalizing Aesthetic Alignment in Text-to-Image Generation via Principles of Art

arXiv:2503.12018 · cs.CV, cs.AI · Submitted 2025-03-15 · Read on arXiv

cs.CV, cs.AI

Submitted: 2025-03-15

Updated: 2026-09-16

Journal ref: Proceedings of the 2026 International Conference on Multimedia Retrieval (ICMR '26), pp. 2200-2209, Association for Computing Machinery (ACM), 2026

DOI: 10.1145/3805622.3810794

Code: https://github.com/jin-zhe/ArtDapter

License: http://creativecommons.org/licenses/by-nc-nd/4.0/

The gist: Text-to-Image (T2I) diffusion models have made rapid progress on semantic alignment (generating what is described in the prompt), yet users still lack reliable control over aesthetic composition (how

Terminology

Abstract

Text-to-Image (T2I) diffusion models have made rapid progress on semantic alignment (generating what is described in the prompt), yet users still lack reliable control over aesthetic composition (how visual elements are put together). Prior work often treats aesthetics as a single, preference-driven notion (e.g., "high quality", "detailed", "breathtaking"), which does not map cleanly to compositional intent. We propose Aesthetic Alignment: aligning generated images to explicit, user-specified compositional constraints. We operationalize these constraints using the Principles of Art (PoA)-e.g., Balance, Rhythm, and Emphasis-commonly used in art education to describe composition. To support this task, we introduce CompArt, a dataset of 80,032 WikiArt images augmented with captions and PoA analyses produced by a multimodal LLM under structured prompting. We further propose ArtDapter, a lightweight and disentangled adapter that enables steering a pretrained T2I model along 10 PoA dimensions while retaining the base model's semantic capability. Experiments on CompArt show improved adherence to PoA controls over strong baselines under a dual evaluation protocol.

Sources

Related papers