Video-Index: A Curated Meta-Benchmark for Video Understanding
cs.CV
Submitted: 2026-10-01
Updated: 2026-10-01
Terminology
Sources
- Lost in Time: A New Temporal Benchmark for VideoLLMs
- LVBench: An Extreme Long Video Understanding Benchmark
- VideoHallucer: Evaluating Intrinsic and Extrinsic Hallucinations in Large Video-Language Models
- StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding
- TOMATO: Assessing Visual Temporal Reasoning Capabilities in Multimodal Foundation Models
- Benchmark Designers Should "Train on the Test Set" to Expose Exploitable Non-Visual Shortcuts
- BenchAgents: Multi-Agent Systems for Structured Benchmark Creation
- Harbor Adapters and Harbor-Index: Infrastructure and a Curated Meta-Dataset for Large-Scale Agentic Evaluation
- Generalization through Memorization: Nearest Neighbor Language Models
- From Entropy to Epiplexity: Rethinking Information for Computationally Bounded Intelligence
- TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models
- MMWorld: Towards Multi-discipline Multi-faceted World Model Evaluation in Videos
- Qwen3-VL Technical Report
- Learning to Reason in 4D: Dynamic Spatial Understanding for Vision Language Models
- Qwen3 Technical Report
- The Vendi Score: A Diversity Evaluation Metric for Machine Learning
- SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features
- VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding
- InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
- VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models