AdaVSkip: Adaptive Visual Token Skipping Across Layers For Efficient MLLMs Inference
cs.CV, cs.AI
Submitted: 2026-09-14
Updated: 2026-09-14
License: http://creativecommons.org/licenses/by/4.0/
The gist: Multimodal large language models (MLLMs) require substantial computation to process numerous visual tokens across all transformer layers.
Terminology
Abstract
Multimodal large language models (MLLMs) require substantial computation to process numerous visual tokens across all transformer layers. Most methods for efficient MLLM inference exploit horizontal redundancy by compressing visual tokens. Beyond token reduction, recent studies exploit vertical redundancy through early exit or fixed-layer skipping. However, we find that the extent and distribution of this redundancy vary across inputs and differ between self-attention and MLP modules. Motivated by these observations, we propose AdaVSkip, which equips each layer with two lightweight routers that independently determine whether visual tokens pass through by or skip the self-attention and MLP modules. These decisions collectively define an input-specific visual-computation path, but their discrete and non-differentiable nature makes learning effective paths challenging. To address this challenge, we develop a progressive two-stage training framework that updates only the routers while keeping the backbone frozen. Stage I establishes an initial routing policy through supervised training with input-specific targets derived from module-wise necessity scores. To further align the routing policy with task performance, Stage II uses reinforcement learning to optimize routing decisions with direct feedback from generated answers. It combines an answer correctness reward with a skip-consistency reward that discourages excessive retention of visual-token computation. Across three MLLM backbones, AdaVSkip maintains strong task performance with substantially less computation. On LLaVA-NeXT-7B, AdaVSkip reduces FLOPs by 53.2% while preserving the original model's average performance. Combining it with visual token compression increases this reduction to 91.2%, while retaining 97.2% of the original performance on average.
Sources
- Qwen3-VL Technical Report
- Qwen2.5-VL Technical Report
- How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites
- SkipDecode: Autoregressive Skip Decoding with Batching and Caching for Efficient LLM Inference
- MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
- Seed1.5-VL Technical Report
- LLaVA-OneVision: Easy Visual Task Transfer
- SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension
- TokenPacker: Efficient Visual Projector for Multimodal LLM
- Evaluating Object Hallucination in Large Vision-Language Models
- InternVL-X: Advancing and Accelerating InternVL Series with Efficient Visual Token Compression
- ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning
- Mixture-of-Depths: Dynamically allocating compute in transformer-based language models
- Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings
- Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs
- LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token
- SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference
- SkipGPT: Dynamic Layer Pruning Reinvented with Token Awareness and Module Decoupling
- InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models