TEMPURA: Temporal Event Masked Prediction and Understanding for Reasoning in Action
cs.CV, cs.AI
Submitted: 2025-05-02
Updated: 2026-09-22
Comments: CoLM 2026
License: http://creativecommons.org/licenses/by/4.0/
The gist: Understanding causal event relationships and achieving fine-grained temporal grounding in videos remain challenging for vision-language models (VLMs).
Terminology
Abstract
Understanding causal event relationships and achieving fine-grained temporal grounding in videos remain challenging for vision-language models (VLMs). We propose TEMPURA (Temporal Event Masked Prediction and Understanding for Reasoning in Action), a two-stage training framework that enhances the video temporal understanding of VLMs. Inspired by infilling techniques in language modeling, TEMPURA first performs masked event prediction, learning to reconstruct missing events and generate step-by-step causal explanations from dense event annotations. It then learns video segmentation and dense captioning, decomposing videos into non-overlapping events with detailed, timestamp-aligned descriptions. We train TEMPURA on VER, our large-scale dataset of 500K videos annotated with temporally aligned event descriptions and structured reasoning steps. Experiments on video temporal grounding and highlight detection benchmarks show that TEMPURA substantially improves strong base VLMs across model families and scales, confirming that combining event-level reasoning with fine-grained temporal segmentation is an effective recipe for video temporal understanding.
Sources
- Qwen2.5-VL Technical Report
- Efficient Training of Language Models to Fill in the Middle
- AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark
- TimeMarker: A Versatile Video-LLM for Long and Short Video Understanding with Superior Temporal Localization Ability
- VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
- The Llama 3 Herd of Models
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Liger Kernel: Efficient Triton Kernels for LLM Training
- GPT-4o System Card
- LLaVA-OneVision: Easy Visual Task Transfer
- LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
- Temporal Preference Optimization for Long-Form Video Understanding
- GroundingGPT:Language Enhanced Multi-modal Grounding Model
- Visual CoT: Advancing Multi-Modal Language Models with a Comprehensive Dataset and Benchmark for Chain-of-Thought Reasoning
- FiLM: Fill-in Language Models for Any-Order Generation
- MovieChat+: Question-aware Sparse Memory for Long Video Question Answering
- SketchAgent: Language-Driven Sequential Sketch Generation
- Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models
- HawkEye: Training Video-Text LLMs for Grounding Text in Videos
- VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models