Beyond Frame Selection: Rethinking Long-Video Understanding with MLLMs
cs.CV
Submitted: 2026-08-06
Updated: 2026-09-26
Terminology
Sources
- PaLM 2 Technical Report
- Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames
- MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens
- Qwen3-VL Technical Report
- Qwen2.5-VL Technical Report
- LongVILA: Scaling Long-Context Visual Language Models for Long Videos
- VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
- Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
- GPT-4o System Card
- LLaVA-OneVision: Easy Visual Task Transfer
- VideoMind: A Chain-of-LoRA Agent for Temporal-Grounded Video Reasoning
- Video-RAG: Visually-aligned Retrieval-Augmented Long Video Comprehension
- Video-MTR: Reinforced Multi-Turn Reasoning for Long Video Understanding
- Qwen2 Technical Report
- mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
- Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning
- Long Context Transfer from Language to Vision
- LLaVA-Video: Video Instruction Tuning With Synthetic Data
- FOCUS: Efficient Keyframe Selection for Long Video Understanding
- VideoLucy: Deep Memory Backtracking for Long Video Understanding
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models