TimeBlind: A Spatio-Temporal Compositionality Benchmark for Video LLMs
cs.CV, cs.AI
Submitted: 2026-01-30
Updated: 2026-09-08
Comments: For code and data, see https://baiqi-li.github.io/timeblind_project/
Project page: https://baiqi-li.github.io/timeblind_project
License: http://creativecommons.org/licenses/by/4.0/
The gist: Fine-grained spatio-temporal understanding is essential for video reasoning and embodied AI.
Terminology
Abstract
Fine-grained spatio-temporal understanding is essential for video reasoning and embodied AI. Yet, while Multimodal Large Language Models (MLLMs) master static semantics, their grasp of temporal dynamics remains brittle. We present TimeBlind, a diagnostic benchmark for compositional spatio-temporal understanding. Inspired by cognitive science, TimeBlind categorizes fine-grained temporal understanding into three levels: recognizing atomic events, characterizing event properties, and reasoning about event interdependencies. Unlike benchmarks that conflate recognition with temporal reasoning, TimeBlind leverages a minimal-pairs paradigm: video pairs share identical static visual content but differ solely in temporal structure, utilizing complementary questions to neutralize language priors. Evaluating over 20 state-of-the-art MLLMs (e.g., GPT-5, Gemini 3 Pro) on 600 curated instances (2400 video-question pairs), reveals that the Instance Accuracy (correctly distinguishing both videos in a pair) of the best performing MLLM is only 48.2%, far below the human performance (98.2%). These results demonstrate that even frontier models rely heavily on static visual shortcuts rather than genuine temporal logic, positioning TimeBlind as a vital diagnostic tool for next-generation video understanding. Dataset and code are available at https://baiqi-li.github.io/timeblind project/.
Sources
- Qwen3-VL Technical Report
- TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models
- Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning?
- PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding
- GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning
- Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos
- VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling
- TempCompass: Do Video LLMs Really Understand Videos?
- VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation
- A Shortcut-aware Video-QA Benchmark for Physical Understanding via Minimal Video Pairs
- MINERVA: Evaluating Complex Video Reasoning
- GenAI-Bench: Evaluating and Improving Compositional Text-to-Visual Generation
- TOMATO: Assessing Visual Temporal Reasoning Capabilities in Multimodal Foundation Models
- Seeing the Arrow of Time in Large Multimodal Models
- EgoLife: Towards Egocentric Life Assistant
- Video-MMLU: A Massive Multi-Discipline Lecture Understanding Benchmark
- Cambrian-S: Towards Spatial Supersensing in Video
- ExAct: A Video-Language Benchmark for Expert Action Analysis
- Kwai Keye-VL Technical Report
- MiniCPM-V 4.5: Cooking Efficient MLLMs via Architecture, Data, and Training Recipe
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models