Evaluating the Evaluators: Diagnosing Large Multimodal Models for AI-Generated Image Assessment
cs.CV
Submitted: 2026-09-29
Updated: 2026-09-29
Code: https://github.com/Kwai-Kolors/Kolors
Terminology
Sources
- LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
- Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling
- Microsoft COCO Captions: Data Collection and Evaluation Server
- Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
- Exploring the Naturalness of AI-Generated Images
- Emerging Properties in Unified Multimodal Pretraining
- FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark
- Seedream 3.0 Technical Report
- The Llama 3 Herd of Models
- AesBench: An Expert Benchmark for Multimodal Large Language Models on Image Aesthetics Perception
- Zebra-CoT: A Dataset for Interleaved Vision Language Reasoning
- Evaluating Text-to-Visual Generation with Image-to-Text Generation
- Improving Video Generation with Human Feedback
- Step1X-Edit: A Practical Framework for General Image Editing
- DeepSeek-VL: Towards Real-World Vision-Language Understanding
- Ovis2.5 Technical Report
- Lego: Learning to Disentangle and Invert Personalized Concepts Beyond Object Appearance in Text-to-Image Diffusion Models
- GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models
- WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation
- SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models