Where to Focus: Query-Modulated Multimodal Keyframe Selection for Long Video Understanding
cs.CV, cs.MM
Submitted: 2026-04-19
Updated: 2026-09-25
Terminology
Sources
- LLaMA: Open and Efficient Foundation Language Models
- Less is More: Token-Efficient Video-QA via Adaptive Frame-Pruning and Semantic Graph Integration
- Qwen3 Technical Report
- Frame-Voyager: Learning to Query Frames for Video Large Language Models
- FOCUS: Efficient Keyframe Selection for Long Video Understanding
- Longformer: The Long-Document Transformer
- Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning?
- Logic-in-Frames: Dynamic Keyframe Search via Visual Semantic-Logical Verification for Long Video Understanding
- GPT-4o System Card
- LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
- Ring Attention with Blockwise Transformers for Near-Infinite Context
- Too Many Frames, Not All Useful: Efficient Strategies for Long-Form Video QA
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models