Beneath the Scores: Rethinking Hallucination Evaluation for Video Understanding Models
cs.CV, cs.AI
Submitted: 2026-09-24
Updated: 2026-09-24
Terminology
Sources
- When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
- Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding?
- Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding
- Distorted or Fabricated? A Survey on Hallucination in Video LLMs
- LongVideoAgent: Multi-Agent Reasoning with Long Videos
- ELV-Halluc: Benchmarking Semantic Aggregation Hallucinations in Long Video Understanding
- VideoHallucer: Evaluating Intrinsic and Extrinsic Hallucinations in Large Video-Language Models
- EventHallusion: Diagnosing Event Hallucinations in Video LLMs
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models