VGA-BenchV2: An Expanded Unified Benchmark and Multi-Model Framework for Evaluating Video Aesthetics and Generation Quality
cs.CV, cs.AI
Submitted: 2026-08-26
Updated: 2026-08-26
Comments: IJCAI 2026
Code: https://github.com/genmoai/models
License: http://creativecommons.org/licenses/by/4.0/
The gist: We introduce VGA-BenchV2, an extended human-aligned benchmark and optimization framework for jointly evaluating and improving video generation quality and aesthetic value.
Terminology
Abstract
We introduce VGA-BenchV2, an extended human-aligned benchmark and optimization framework for jointly evaluating and improving video generation quality and aesthetic value. Built upon VGA-Bench, VGA-BenchV2 preserves the original fine-grained taxonomy with two primary dimensions-Aesthetic and Generation-and 52 sub-dimensions. Guided by this taxonomy, we curate 1,016 diverse prompts and collect over 60,000 videos generated by 12 mainstream video generation models. More importantly, VGA-BenchV2 substantially expands human-labeled supervision by adding 36,000 task-level annotations, including 16,200 for aesthetic quality, 13,200 for aesthetic tagging, and 6,600 for generation quality, corresponding to 13.46x, 11.15x, and 1.55x scale-ups over VGA-Bench, respectively. Leveraging this enlarged annotation corpus, we develop a hybrid evaluator architecture consisting of VAQA-Net for continuous aesthetic scoring and two Qwen-based Large Vision-Language Model evaluators, VTag-Net and VGQA-Net, for aesthetic tagging and generation quality assessment. Extensive experiments demonstrate strong alignment with human judgments across diverse generation models. Beyond evaluation, VGA-BenchV2 further introduces an evaluation-to-optimization pipeline, where the learned aesthetic evaluator serves as a reward model for reinforcement learning-based generator fine-tuning. This closes the loop from benchmark construction and human supervision to automated evaluation and model optimization, enabling video generators to improve not only in realism but also in aesthetic quality and human preference alignment. Resources are available at https://huggingface.co/datasets/BestiVictoryLab/VGA-Bench.
Sources
- Qwen3-VL Technical Report
- Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets
- AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning
- LTX-Video: Realtime Video Latent Diffusion
- VGA-Bench: A Unified Benchmark and Multi-Model Framework for Video Aesthetics and Generation Quality Evaluation
- HunyuanVideo: A Systematic Framework For Large Video Generative Models
- Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models
- Flow-GRPO: Training Flow Matching Models via Online RL
- VideoFusion: Decomposed Diffusion Models for High-Quality Video Generation
- Latte: Latent Diffusion Transformer for Video Generation
- VADB: A Large-Scale Video Aesthetic Database with Professional and Multi-Dimensional Annotations
- Score-Based Generative Modeling through Stochastic Differential Equations
- Human-Centric Foundation Models: Perception, Generation and Agentic Modeling
- Wan: Open and Advanced Large-Scale Video Generative Models
- ModelScope Text-to-Video Technical Report
- CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer
- VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models