VideoMM: Adaptive Macro-Micro Inference for Efficient Video MLLMs
cs.AI, cs.CL, cs.CV, cs.MM
Submitted: 2026-09-15
Updated: 2026-09-15
Code: https://github.com/adfh917k/VideoMM
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Qwen2.5-VL Technical Report
- PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling
- LVAgent: Long Video Understanding by Multi-Round Dynamical Collaboration of MLLM Agents
- Accelerating Large Language Model Decoding with Speculative Sampling
- An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models
- Taming the Fragility of KV Cache Eviction in LLM Inference
- Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
- CriticalKV: Optimizing KV Cache Eviction from an Output Perturbation Perspective
- Not All Heads Matter: A Head-Level KV Cache Compression Method with Integrated Retrieval and Reasoning
- Fast Inference from Transformers via Speculative Decoding
- MMInference: Accelerating Pre-filling for Long-Context VLMs via Modality-Aware Permutation Sparse Attention
- Qwen3 Technical Report
- GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning
- SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs
- LVBench: An Extreme Long Video Understanding Benchmark
- Self-Consistency Improves Chain of Thought Reasoning in Language Models
- VideoChat-A1: Thinking with Long Videos by Chain-of-Shot Reasoning
- Think, Then Verify: A Hypothesis-Verification Multi-Agent Framework for Long Video Understanding
- LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding
- Retrieval Head Mechanistically Explains Long-Context Factuality
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection