AesCanvas: A Large-Scale Dataset and Benchmark for Aesthetic Critique and Contextual Suitability
cs.CV, cs.AI
Submitted: 2026-08-27
Updated: 2026-08-27
Terminology
Sources
- LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
- Qwen3-VL Technical Report
- UniPercept: Towards Unified Perceptual-Level Image Understanding across Aesthetics, Quality, Structure, and Texture
- ArtiMuse: Fine-Grained Image Aesthetics Assessment with Joint Scoring and Expert-Level Understanding
- GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning
- Seeing Is Believing? A Benchmark for Multimodal Large Language Models on Visual Illusions and Anomalies
- AesExpert: Towards Multi-modality Foundation Model for Image Aesthetics Perception
- AesBench: An Expert Benchmark for Multimodal Large Language Models on Image Aesthetics Perception
- Bridging Cognitive Gap: Hierarchical Description Learning for Artistic Image Aesthetics Assessment
- Benchmarking Deflection and Hallucination in Large Vision-Language Models
- MINERVA-Cultural: A Benchmark for Cultural and Multilingual Long Video Reasoning
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
- WorldBench: A Challenging and Visually Diverse Multimodal Reasoning Benchmark
- BERTScore: Evaluating Text Generation with BERT
- Teaching LMMs for Image Quality Scoring and Interpreting
- ReactBench: A Cause-Driven Benchmark for Multimodal Hallucination via Systematic Evaluation
- UNIAA: A Unified Multi-modal Image Aesthetic Assessment Baseline and Benchmark
- InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
- MLLM-as-a-Judge: Assessing Multimodal LLM-as-a-Judge with Vision-Language Benchmark
- The Photographer Eye: Teaching Multimodal Large Language Models to Understand Image Aesthetics like Photographers
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models