Who Speaks for the Pruned? Visual Token Pruning as Coverage Optimization
cs.CV, cs.CL, cs.LG
Submitted: 2026-09-02
Updated: 2026-09-02
Comments: Accepted to EMNLP 2026 main
License: http://creativecommons.org/licenses/by/4.0/
The gist: Visual token pruning reduces the inference cost of vision-language models (VLMs), but most methods only ask which tokens to keep.
Terminology
Abstract
Visual token pruning reduces the inference cost of vision-language models (VLMs), but most methods only ask which tokens to keep. This retained-token view can keep redundant high-scoring tokens while leaving discarded evidence without a close representative. We propose CoverPruner, a training-free pruner that asks the complementary demand-side question: after a token is removed, which surviving original token represents it for the target VLM? CoverPruner formulates pruning as Representational Coverage Maximization (RCM), covering the full projected visual-token set with query-weighted demand. It instantiates RCM with projector-space coverage and a lightweight first-layer attention probe. Across multiple VLM architectures and compression rates, CoverPruner achieves the best average accuracy among all compared methods, with the largest gains usually appearing under aggressive compression.
Sources
- Qwen2.5-VL Technical Report
- Matryoshka Multimodal Models
- Similarity-Aware Token Pruning: Your VLM but Faster
- LLaVA-OneVision: Easy Visual Task Transfer
- LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models
- Less is More: A Simple yet Effective Token Reduction Method for Efficient Multi-modal LLMs
- LLaVA-Video: Video Instruction Tuning With Synthetic Data
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models