CueKFS: Agentic Cue-Driven Keyframe Selection for Long Video Understanding
cs.CV, cs.AI, cs.CL
Submitted: 2026-09-25
Updated: 2026-09-25
Terminology
Sources
- Qwen2.5-VL Technical Report
- HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering
- Wavelet-based Frame Selection by Detecting Semantic Boundary for Long Video Understanding
- FrameOracle: Learning What to See and How Much to See in Videos
- Divide, then Ground: Adapting Frame Selection to Query Types for Long-Form Video Understanding
- KeyVideoLLM: Towards Large-scale Video Keyframe Selection
- BOLT: Boost Large Vision-Language Model Without Training for Long-form Video Understanding
- CLIP-It! Language-Guided Video Summarization
- Decomposing Queries into Tool Calls for Long-Video Keyframe Retrieval
- MDP3: A Training-free Approach for List-wise Frame Selection in Video-LLMs
- MSJoE: Jointly Evolving MLLM and Sampler for Efficient Long-Form Video Understanding
- Adaptive Keyframe Sampling for Long Video Understanding
- SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features
- Where to Focus: Query-Modulated Multimodal Keyframe Selection for Long Video Understanding
- Query-Conditioned Evidential Keyframe Sampling for MLLM-Based Long-Form Video Understanding
- LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding
- Generative Frame Sampler for Long Video Understanding
- T*: Re-thinking Temporal Search for Long-Form Video Understanding
- Frame-Voyager: Learning to Query Frames for Video Large Language Models
- Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models