Select, Compress, Reinvest: A Controlled Study of Visual-Token Allocation in Long-Video MLLMs
cs.CV, cs.CL
Submitted: 2026-09-03
Updated: 2026-09-03
Code: https://github.com/codeprakhar25/omp-keyframe-sampling
Project page: https://vision.cs.utexas.edu/projects/adaalloc
Terminology
Sources
- LDDR: Linear-DPP-Based Dynamic-Resolution Frame Sampling for Video MLLMs
- Event-Anchored Frame Selection for Effective Long-Video Understanding
- LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
- AdaCodec: A Predictive Visual Code for Video MLLMs
- Adaptive Greedy Frame Selection for Long Video Understanding
- ReQuest: Rethinking-based Question-Aware Frame Selection for Long-Form Video QA
- Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs
- QCA: Query- and Content-Aware Keyframe Selection for Long Video Understanding
- AdaRD-key: Adaptive Relevance-Diversity Keyframe Sampling for Long-form Video understanding
- AdaptToken: Entropy-based Adaptive Token Selection for MLLM Long Video Understanding
- InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
- Qwen3-VL Technical Report
- FOCUS: Efficient Keyframe Selection for Long Video Understanding
- MDP3: A Training-free Approach for List-wise Frame Selection in Video-LLMs
- Adaptive Keyframe Sampling for Long Video Understanding
- LVBench: An Extreme Long Video Understanding Benchmark
- Query-Conditioned Evidential Keyframe Sampling for MLLM-Based Long-Form Video Understanding
- Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors
- LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models