Can MLLMs Critique Like Humans? Evaluating Open-Ended Aesthetic Reasoning in Multimodal Large Language Models
cs.CL
Submitted: 2026-06-29
Updated: 2026-08-31
License: http://creativecommons.org/licenses/by/4.0/
The gist: Open-ended aesthetic critique is a challenge for multimodal large language models (MLLMs): it has no single correct answer, and most aesthetic evaluation measures models against numeric scores rather
Terminology
Abstract
Open-ended aesthetic critique is a challenge for multimodal large language models (MLLMs): it has no single correct answer, and most aesthetic evaluation measures models against numeric scores rather than the written critiques people actually give. We ask whether MLLM critiques are close to human ones, scoring eight open-weight MLLMs from 7 B to 397 B, plus GPT-5.5, against multiple ranked human critiques for each of 1, 227 r/photocritique posts under eight prompt conditions. Reference-based similarity gives a misleading picture. In absolute terms the stricter lexical and learned metrics align only weakly with human critiques while a coarse embedding cosine reports broad topical overlap, yet requesting shorter critiques raises those scores and withholding the image barely changes them: the similarity reflects length, the post text, and a stable critiquing style more than image-specific observation. An LLM judge sharpens the question rather than settling it: in the primary condition all four judges prefer the frontier models' critiques to the human ones, but on the 7 -- 8 B models they diverge wildly, from 9% to 81% preference on identical pairs. Asked instead how similar each pair is in substance, those judges and two human annotators agree, rating every model between 1.81 and 2.59 on a 1 -- 5 scale, close to ``mostly different''. Behaviorally, the models diverge in ways the scores do not surface: they cover nearly every aesthetic aspect where humans are selective and repeat themselves across critiques of one photo, even when prompted to write at human length. We argue that reference-based similarity rewards a fluent, comprehensive critique style rather than the selectivity and specificity of human critique.
Sources
- VISTA: Vision-Language Inference for Training-Free Stock Time-Series Analysis
- LLaVA-OneVision: Easy Visual Task Transfer
- RoBERTa: A Robustly Optimized BERT Pretraining Approach
- Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators
- pFedMMA: Personalized Federated Fine-Tuning with Multi-Modal Adapter for Vision-Language Models
- Few-Shot Adversarial Low-Rank Fine-Tuning of Vision-Language Models
- MMLoP: Multi-Modal Low-Rank Prompting for Efficient Vision-Language Adaptation
- The Llama 3 Herd of Models
- AesBench: An Expert Benchmark for Multimodal Large Language Models on Image Aesthetics Perception
- Q-Align: Teaching LMMs for Visual Scoring via Discrete Text-Defined Levels
- BERTScore: Evaluating Text Generation with BERT
- UNIAA: A Unified Multi-modal Image Aesthetic Assessment Baseline and Benchmark
- InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
- Training-Free Adversarial Robustness in Computational MRI
- NeuroMambaLLM: Dynamic Graph Learning of fMRI Functional Connectivity in Autistic Brains Using Mamba and Language Model Reasoning
- Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering