Query-aligned video frame selection for long video understanding
cs.CV, cs.AI
Submitted: 2026-09-15
Updated: 2026-09-15
Terminology
Sources
- VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
- Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
- KeyVideoLLM: Towards Large-scale Video Keyframe Selection
- SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models
- Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
- VaQuitA: Enhancing Alignment in LLM-Assisted Video Understanding
- LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token
- LLaVA-Video: Video Instruction Tuning With Synthetic Data
- MLVU: Benchmarking Multi-task Long Video Understanding
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models