MarKey: Marginal Utility Guided Greedy Keyframe Selection for Long Video Understanding
cs.CV, cs.CL
Submitted: 2026-09-14
Updated: 2026-09-14
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Qwen3-VL Technical Report
- Qwen2.5-VL Technical Report
- TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models
- LongVILA: Scaling Long-Context Visual Language Models for Long Videos
- EmoBench-M: Benchmarking Emotional Intelligence for Multimodal Large Language Models
- MMHBench: A Multi-Perspective Benchmark for Mental Health Understanding in Long-Form Videos
- VidCompress: Memory-Enhanced Temporal Compression for Video Understanding in Large Language Models
- ReFoCUS: Reinforcement-guided Frame Optimization for Contextual Understanding
- LLaVA-OneVision: Easy Visual Task Transfer
- Aria: An Open Multimodal Native Mixture-of-Experts Model
- LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
- Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input
- PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning
- mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models
- Frame-Voyager: Learning to Query Frames for Video Large Language Models
- Zelda: Video Analytics using Vision-Language Models
- LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding
- Gemini: A Family of Highly Capable Multimodal Models
- Script-a-Video: Deep Structured Audio-visual Captions via Factorized Streams and Relational Grounding
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models