TASTE: A Designer-Annotated Multi-Dimensional Preference Dataset for AI-Generated Graphic Design
cs.CV, cs.AI, stat.AP
Submitted: 2026-05-20
Updated: 2026-09-21
Code: https://github.com/purvanshi-lica/taste
License: http://creativecommons.org/licenses/by/4.0/
The gist: Text-to-image models now generate graphic design at production scale, yet their supervision still comes primarily from photo-style preference datasets with a single overall verdict per comparison.
Terminology
Abstract
Text-to-image models now generate graphic design at production scale, yet their supervision still comes primarily from photo-style preference datasets with a single overall verdict per comparison. Designers evaluate designs along several distinct axes (e.g., typography, layout, color harmony) that a single preference label collapses. We release TASTE(Typography, Aesthetics, Spatial, Tone, Etc.), a multi-dimensional preference dataset in which two disjoint cohorts of five professional designers each ranked outputs from four current text-to-image models across nine criteria along with per-image hallucination flags. We pair the dataset with two contributions. First, a criterion-agnostic signal-validation framework based on Kendall's τ, majority-vote probability, and Condorcet cycles against exact iid-uniform nulls; the analysis reveals significant but moderate designer agreement, with every TASTE criterion rejecting the random-rater null. Second, we benchmark preference models on TASTE and find that off-the-shelf VLM judges and dedicated T2I scorers fail to reach majority agreement with the designer panel, while a small MLP head trained directly on TASTE substantially narrows the gap to the single-rater ceiling, setting a baseline for future TASTE-trained preference models.
Sources
- E-comIQ-ZH: A Human-Aligned Dataset and Benchmark for Fine-Grained Evaluation of E-commerce Posters with Chain-of-Thought
- Graphic-Design-Bench: A Comprehensive Benchmark for Evaluating AI on Graphic Design Tasks
- Predicting Visual Importance Across Graphic Design Types
- DistortBench: Benchmarking Vision Language Models on Image Distortion Identification
- LICA: Layered Image Composition Annotations for Graphic Design Research
- A Survey of Multimodal Hallucination Evaluation and Detection
- DesignPref: Capturing Personal Preferences in Visual Design Generation
- ImagenWorld: Stress-Testing Image Generation Models with Explainable Human Evaluation on Open-ended Real-World Tasks
- Identifying and Mitigating Position Bias of Multi-image Vision-Language Models
- Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis
- VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness
- From Fragment to One Piece: A Survey on AI-Driven Graphic Design
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models