TempCloze: Can Video-LLMs Identify the Missing Middle?
cs.CV, cs.AI
Submitted: 2026-09-01
Updated: 2026-09-01
Comments: EMNLP 2026 Findings
Code: https://github.com/CedricPei/Temporal-Cloze
License: http://creativecommons.org/licenses/by/4.0/
The gist: Temporal reasoning benchmarks for Video-LLMs are often mediated by language, leaving room for linguistic shortcuts from option wording, answer correlations, or language priors.
Terminology
Abstract
Temporal reasoning benchmarks for Video-LLMs are often mediated by language, leaving room for linguistic shortcuts from option wording, answer correlations, or language priors. To reduce such shortcuts, we introduce TempCloze, a video cloze benchmark for evaluating visual temporal reasoning in Video-LLMs. Given the beginning and ending clips of a video, models must identify the true missing middle from four candidates. TempCloze contains 1,521 carefully filtered videos from seven sources, mainly long-take and egocentric videos. We construct same-source distractors along three dimensions: Semantic asks what event should happen, Alignment probes when it should occur, and Progression tests how it should unfold, while shared scenes and objects reduce appearance cues. Our evaluation of 10 proprietary and 21 open-source Video-LLMs reveals Alignment as the primary bottleneck: models often recognize plausible semantic content and local event progression but struggle with temporal alignment. We further conduct error pattern and behavioral sensitivity analyses on TempCloze-Mixed and TempCloze-Hard with four representative models to examine where errors arise and how candidate order, context direction, visible span, frame density, and test-time scaling influence model choices.
Sources
- Qwen3-VL Technical Report
- Qwen2.5-VL Technical Report
- Artifacts or Abduction: How Do LLMs Answer Multiple-Choice Questions Without the Question?
- TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models
- FIBER: Fill-in-the-Blanks as a Challenging Video Understanding Evaluation Framework
- Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- Lost in Time: A New Temporal Benchmark for VideoLLMs
- Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
- TGIF-QA: Toward Spatio-Temporal Reasoning in Visual Question Answering
- MiraData: A Large-Scale Video Dataset with Long Durations and Structured Captions
- Kimi-VL Technical Report
- Efficient Memory Management for Large Language Model Serving with PagedAttention
- TVQA: Localized, Compositional Video Question Answering
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
- MVBench: A Comprehensive Multi-modal Video Understanding Benchmark
- TempCompass: Do Video LLMs Really Understand Videos?
- Video Cloze Procedure for Self-Supervised Spatio-Temporal Learning
- Robust Visual Question Answering: Datasets, Methods, and Future Challenges
- Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models