VidOmni-Bench: A Benchmark for Fine-Grained Video Understanding via Spatio-Temporal Event Verification across Complexity and Duration
cs.CV, cs.AI
Submitted: 2026-09-18
Updated: 2026-10-03
Code: https://github.com/HumanSignal/label-studio
License: http://creativecommons.org/licenses/by/4.0/
The gist: While Video Large Language Models (Video-LLMs) have recently demonstrated strong performance, reliably evaluating their fine-grained video understanding remains challenging.
Terminology
Abstract
While Video Large Language Models (Video-LLMs) have recently demonstrated strong performance, reliably evaluating their fine-grained video understanding remains challenging. Existing benchmarks often rely on question answering or ground-truth caption matching, where models may succeed through superficial cues and incomplete annotations. To this end, we introduce VidOmni-Bench, a benchmark that requires models to verify whether each event in dense video captions is supported by the video. VidOmni-Bench consists of 500 videos spanning five complexity types and diverse durations from 4 seconds to 90 minutes. After collecting videos along these axes, we use diverse Video-LLMs to generate dense captions and obtain human-verified sentence-level labels, where sentences containing incorrect events serve as hard negatives for evaluation. Our experiments on VidOmni-Bench reveal three key findings: (i) Video-LLMs frequently generate hallucinated descriptions in dense video captioning; (ii) they also struggle as verifiers, failing to reliably detect plausible but incorrect event descriptions; and (iii) model weaknesses vary across video complexity and duration, revealing diverse, model-specific bottlenecks in current Video-LLMs.
Sources
- Qwen3-VL Technical Report
- TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models
- Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding?
- Exploring Hallucination of Large Multimodal Models in Video Understanding: Benchmark, Analysis and Mitigation
- The Llama 3 Herd of Models
- NarrativeTrack: Evaluating Entity-Centric Reasoning for Narrative Understanding
- NOAH: Benchmarking Narrative Prior driven Hallucination and Omission in Video Large Language Models
- TimeBlind: A Spatio-Temporal Compositionality Benchmark for Video LLMs
- LLaVA-OneVision: Easy Visual Task Transfer
- VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation
- HalDec-Bench: Benchmarking Hallucination Detector in Image Captioning
- OpenAI GPT-5 System Card
- video-SALMONN 2: Caption-Enhanced Audio-Visual Large Language Models
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
- Qwen3 Technical Report
- Kwai Keye-VL 1.5 Technical Report
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models