ToDRE: Effective Visual Token Pruning via Token Diversity and Task Relevance

arXiv:2505.18757 · cs.CV · Submitted 2025-05-24 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "ToDRE: Effective Visual Token Pruning via Token Diversity and Task Relevance".

Jane: Visual token pruning aims to compress and prune redundant visual tokens which play a critical role in efficient inference with large vision-language models (LVLMs).

Tom: First, who's behind it and why it matters.

Title and authors: Jane: The paper is titled "ToDRE: Effective Visual Token Pruning via Token Diversity and Task Relevance," and the authors are Duo Li, Zuhao Yang, Xiaoqin Zhang, Ling Shao, and Shijian Lu.

Tom: It’s about this two-stage framework they designed to compress visual tokens effectively for large vision-language models.

Jane: They showed that these two factors—token diversity and task relevance—are crucial and should be handled separately for better pruning results.

The paper's summary: Tom: So, the main idea is they introduce ToDRE, a training-free method that uses a greedy max-sum diversification algorithm to select a subset of diverse and representative visual tokens first.

Jane: That’s the diversity part. They start by picking an initial pivot token and then keep adding tokens that are least similar to what they already have, ensuring the final set covers a broad range of visual information.

Lu: And then they move to the second stage, which is relevance-driven reduction. This involves looking at how attention shifts across different layers of the large language model.

Meng: They adaptively drop visual tokens in specific decoder layers if both cross-modal attention ratios, which are the text-to-visual and visual-to-text ones, fall below a certain threshold called tau.

The paper's improvements: Tom: The core improvement is this two-stage approach, which handles token selection based on diversity first and then removes irrelevant tokens based on where the model’s attention fades.

Jane: That second stage is guided by observing an "information migration" phenomenon, where cross-modal attention gets strong in early layers but then drops off in deeper layers.

Lu: The theoretical justification they give is that intra-modal redundancy and cross-modal redundancy are statistically independent in the embedding space, which validates why this two-stage compression paradigm works.

Meng: Experimentally, they show that ToDRE prunes up to ninety percent of visual tokens after the vision encoder and all visual tokens in certain LLM decoder layers.

Conclusion: Tom: So, to wrap up, ToDRE is a training-free method that first selects a maximally diverse subset of visual tokens using that greedy max-sum diversification algorithm, then removes all remaining visual tokens once cross-modal attention fades.

Jane: This combination allows them to achieve a two point six times speed-up in total inference time while keeping ninety-five point zero percent model performance and excellent compatibility across twelve benchmarks.

Lu: From my perspective, the potential here is that by keeping only the most diverse and relevant tokens, we might be able to build models that are significantly more efficient for real-world deployment.

Meng: I’m thinking about how this impacts the practical side of things; reducing inference time by that much means we could run these complex vision tasks on much less computational power.

Lalam: For me, as a language model, it means the resulting representations are cleaner because they don't carry unnecessary visual noise from redundant tokens.

Tom: So, ToDRE is a training-free framework that tackles token pruning by looking at diversity and relevance separately to get significant speed gains while keeping performance high.

Duo Li, Zuhao Yang, Xiaoqin Zhang, Ling Shao, Shijian Lu

CCDS, NTU · ZJUT, China · Terminus AI Lab, UCAS, China

cs.CV

Submitted: 2025-05-24

Updated: 2026-10-04

Importance score: 92/100

The gist: Visual token pruning aims to compress and prune redundant visual tokens which play a critical role in efficient inference with large vision-language models (LVLMs).

Key concepts

Diversity-Driven Token Selection
This stage uses a 'greedy max-sum diversification algorithm' to choose a subset of visual tokens that are as different from each other as possible. It starts with one key token and iteratively adds the next token that is least similar to the ones already chosen, ensuring the final set covers a broad spectrum of visual information.
Relevance-Driven Token Reduction
This stage targets task-irrelevant tokens by analyzing how attention shifts within the LLM. It identifies specific decoder layers where both visual-to-text and text-to-visual attention ratios drop below a threshold, allowing the framework to drop all visual tokens in those layers, effectively removing noise.
Orthogonality of Redundancy
The paper proves that visual redundancy (tokens that are similar to each other) and cross-modal redundancy (tokens that don't relate well to the text) are statistically independent. This mathematical property justifies using a two-stage process because the diversity selection and relevance reduction tackle these two distinct types of redundancy separately.
Information Migration
This phenomenon describes how cross-modal attention strength changes as you move deeper into the LLM's decoder layers. The framework exploits this by dropping visual tokens in later layers where this attention has significantly diminished, indicating those tokens are no longer useful for the specific task.

Terminology

Summary

Visual token pruning aims to compress and prune redundant visual tokens which play a critical role in efficient inference with large vision-language models (LVLMs). The gist: ToDRE incorporates token diversity and task relevance, two largely neglected yet critical factors that help preserve indispensable and informative visual cues and improve pruning robustness and answer accuracy as illustrated in the coffee cup localization task.<ref:2505.18757#pg2>

How it works

The proposed framework is a two-stage, training-free, and plug-and-play token compression technique that incorporates both visual token diversity and task-specific relevance for effective token pruning and efficient LVLM inference<ref:2505.18757#pg3>. This approach circumvents the limitations of existing methods that rely on a single metric by treating inter-token diversity and token-task relevance as two orthogonal factors, which are crucial yet complementary in conveying useful information<ref:2505.18757#pg4>.

The framework operates in two distinct stages:

  1. Diversity-Driven Token Selection: This first stage utilizes a greedy max-sum diversification algorithm to select a maximally diverse subset of visual tokens prior to LLM input<ref:2505.18757#pg5>. The process involves two steps: (1) initializing a retention set by selecting the initial pivot token, and (2) iteratively adding the token that minimizes its cumulative similarity to the current set<ref:2505.18757#pg7>.

  2. Relevance-Driven Token Reduction: This second stage leverages an “information migration” mechanism within certain decoder layers of the large language model (LLM) to eliminate task-irrelevant visual tokens<ref:2505.18757#pg10>. The layer selection is adaptive, dropping all visual tokens within a selected layer if both cross-modal attention ratios, visual-to-text and text-to-visual attention ratios both fall below a predefined threshold τ, are lower than that threshold.

Key Mechanisms

The core of the diversity selection is the greedy max-sum diversification algorithm, which starts from a designated pivot token—selected based on [CLS] attention—and iteratively picks a new token by minimizing its cumulative similarity to the already selected set using cosine similarity. This ensures that the retained set is maximally diverse, thereby preserving a broad spectrum of visual information and enhancing the token representativeness at high pruning ratios.

The relevance-driven reduction stage is guided by observing an “information migration” phenomenon where cross-modal attention is strong in early layers but fades in deeper layers. The framework adaptively selects one layer in the latter half of the LLM decoder (where crossmodal attention has significantly diminished) and drops all visual tokens within that layer, which removes visual tokens irrelevant to the given question and thus further eliminates redundant computation during inference.

Theoretical Justification

The two-stage paradigm is theoretically justified by the orthogonality of intra-modal redundancy and cross-modal redundancy in the embedding space. By mapping visual tokens (V) and text tokens (T) onto mutually orthogonal sub-spaces such that W⊤ V WT = 0 (V ⊥ T), the paper proves that intra-modal redundancy and cross-modal redundancy are statistically independent in the embedding space, validating the effectiveness of this two-stage compression paradigm.

Experimental Validation

Extensive experiments over four widely adopted LVLMs and twelve multimodal benchmarks demonstrate the superior and consistent effectiveness of ToDRE. Results show that ToDRE prunes 90% of visual tokens after the vision encoder as well as all visual tokens in certain LLM decoder layers, leading to a 2.6× speed-up in total inference time while maintaining 95.0% model performance plus excellent model compatibility. Ablation studies confirm that applying Stage 1 only yields substantial time savings, while the combination of Stage 1 and Stage 2 (ToDRE) reduces inference time by 42.5% and 61.4% at the 25% and 10% token retention ratios, respectively, while even improving performance (up to +0.2%).

Conclusion

ToDRE is a training-free, architecture-agnostic framework that first selects a maximally diverse subset of visual tokens via a greedy max-sum diversification algorithm, then removes all remaining visual tokens once cross-modal attention fades. Experiments on twelve image- and videolanguage benchmarks show that ToDRE prunes up to 90% of visual tokens while preserving 95.0% of the original performance, achieving 2.6× faster inference and 14.5% lower memory usage than uncompressed baselines.

REFERENCES

[1] Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. Gpt-4 technical report. arXiv preprint arXiv:2303.08774](https://arxiv.org/abs/2303.08774)

[2] Saeed Ranjbar Alvar, Gursimran Singh, Mohammad Akbari, and Yong Zhang. Divprune: Diversity-based visual token pruning for large multimodal models. arXiv preprint arXiv:2503.02175](https://arxiv.org/abs/2503.02175)

[3] Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, et al. Qwen technical report. arXiv preprint arXiv:2309.16609](https://arxiv.org/abs/2309.16609)

[4] Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou. Qwen-vl: A frontier large vision-language model with versatile abilities. arXiv preprint arXiv:2308.12966](https://arxiv.org/abs/2308.12966)

[5] Shuai Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Sibo Song, Kai Dang, Peng Wang, Shijie Wang, Jun Tang, et al. Qwen2. 5-vl technical report. arXiv preprint arXiv:2502.13923](https://arxiv.org/abs/2502.13923)

[6] Daniel Bolya, Cheng-Yang Fu, Xiaoliang Dai, Peizhao Zhang, Christoph Feichtenhofer, and Judy Hoffman. Token merging: Your vit but faster. In Proceedings of the International Conference on Learning Representations, 2023. 1](https://arxiv.org/abs/2308.14577)

[7] Mu Cai, Jianwei Yang, Jianfeng Gao, and Yong Jae Lee. Matryoshka multimodal models. In Workshop on VideoLanguage Models@ NeurIPS 2024, 2024. 3](https://arxiv.org/abs/2411.17686)

[8] Yuxuan Cai, Jiangning Zhang, Haoyang He, Xinwei He, Ao Tong, Zhenye Gan, Chengjie Wang, and Xiang Bai. Llava-kd: A framework of distilling multimodal large language models. arXiv preprint arXiv:2410.16236](https://arxiv.org/abs/2410.16236)

[9] Wenhao Chai, Enxin Song, Yilun Du, Chenlin Meng, Vashisht Madhavan, Omer Bar-Tal, Jenq-Neng Hwang, Saining Xie, and Christopher D. Manning. Auroracap: Efficient, performant video detailed captioning and a new benchmark, 2025. 5](https://arxiv.org/abs/2504.07491)

[10] Liang Chen, Haozhe Zhao, Tianyu Liu, Shuai Bai, Junyang Lin, Chang Zhou, and Baobao Chang. An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models.

Improvements for AI systems

  1. Improved visual token selection via greedy max-sum diversification will select a maximally diverse subset of visual tokens by iteratively selecting tokens that minimize its cumulative similarity to the current set, which preserves a broad spectrum of visual information.

  2. Enhanced task relevance through the Relevance-driven Token Reduction mechanism will dynamically identify a pivot decoder layer where both cross-modal attention ratios fall below a threshold τ, allowing all visual tokens are removed at this point, thereby eliminating redundant computation during inference.

  3. The system can achieve a 2.6× speed-up in total inference time while maintaining 95.0% model performance plus excellent model compatibility by effectively pruning tokens both in the embedding space and the LLM decoder, as shown by the result: ToDRE prunes 90% of visual tokens after the vision encoder as well as all visual tokens in certain LLM decoder layers.

Sources

Related papers