ToDRE: Effective Visual Token Pruning via Token Diversity and Task Relevance
summary
The gist
Visual token pruning aims to compress and prune redundant visual tokens which play a critical role in efficient inference with large vision-language models (LVLMs).
In short
ToDRE is a training-free method to compress large vision-language models by pruning redundant visual tokens. It achieves this by combining two strategies: first, selecting a diverse set of tokens using a greedy algorithm; second, removing irrelevant tokens in deep model layers where cross-modal attention has faded. This results in up to 90% token pruning with significant speed improvements while maintaining high performance.
Key concepts
- Diversity-Driven Token Selection
- This stage uses a 'greedy max-sum diversification algorithm' to choose a subset of visual tokens that are as different from each other as possible. It starts with one key token and iteratively adds the next token that is least similar to the ones already chosen, ensuring the final set covers a broad spectrum of visual information.
- Relevance-Driven Token Reduction
- This stage targets task-irrelevant tokens by analyzing how attention shifts within the LLM. It identifies specific decoder layers where both visual-to-text and text-to-visual attention ratios drop below a threshold, allowing the framework to drop all visual tokens in those layers, effectively removing noise.
- Orthogonality of Redundancy
- The paper proves that visual redundancy (tokens that are similar to each other) and cross-modal redundancy (tokens that don't relate well to the text) are statistically independent. This mathematical property justifies using a two-stage process because the diversity selection and relevance reduction tackle these two distinct types of redundancy separately.
- Information Migration
- This phenomenon describes how cross-modal attention strength changes as you move deeper into the LLM's decoder layers. The framework exploits this by dropping visual tokens in later layers where this attention has significantly diminished, indicating those tokens are no longer useful for the specific task.
Terminology used across episodes
This episode discusses
- ToDRE: Effective Visual Token Pruning via Token Diversity and Task Relevance · Paper Radio
- GPT-4 Technical Report
- DivPrune: Diversity-based Visual Token Pruning for Large Multimodal Models
- Qwen Technical Report
- Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
- Qwen2.5-VL Technical Report
- LLaVA-KD: A Framework of Distilling Multimodal Large Language Models
- MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
- Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
- Filter, Correlate, Compress: Training-Free Token Reduction for MLLM Acceleration
- GPT-4o System Card
- What Kind of Visual Tokens Do We Need? Training-free Visual Token Pruning for Multi-modal Large Language Models from the Perspective of Graph
- LLaVA-OneVision: Easy Visual Task Transfer
- TokenPacker: Efficient Visual Projector for Multimodal LLM
- Evaluating Object Hallucination in Large Vision-Language Models
- Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models
- Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
- Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model
- Global Compression Commander: Plug-and-Play Inference Acceleration for High-Resolution Large Vision-Language Models
- VL-Mamba: Exploring State Space Models for Multimodal Learning
- LLaVA-PruMerge: Adaptive Token Reduction for Efficient Large Multimodal Models
The paper
ToDRE: Effective Visual Token Pruning via Token Diversity and Task Relevance · Read on arXiv
Duo Li, Zuhao Yang, Xiaoqin Zhang, Ling Shao, Shijian Lu
CCDS, NTU · ZJUT, China · Terminus AI Lab, UCAS, China
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "ToDRE: Effective Visual Token Pruning via Token Diversity and Task Relevance".
Jane: Visual token pruning aims to compress and prune redundant visual tokens which play a critical role in efficient inference with large vision-language models (LVLMs).
Tom: First, who's behind it and why it matters.
Title and authors: Jane: The paper is titled "ToDRE: Effective Visual Token Pruning via Token Diversity and Task Relevance," and the authors are Duo Li, Zuhao Yang, Xiaoqin Zhang, Ling Shao, and Shijian Lu.
Tom: It’s about this two-stage framework they designed to compress visual tokens effectively for large vision-language models.
Jane: They showed that these two factors—token diversity and task relevance—are crucial and should be handled separately for better pruning results.
The paper's summary: Tom: So, the main idea is they introduce ToDRE, a training-free method that uses a greedy max-sum diversification algorithm to select a subset of diverse and representative visual tokens first.
Jane: That’s the diversity part. They start by picking an initial pivot token and then keep adding tokens that are least similar to what they already have, ensuring the final set covers a broad range of visual information.
Lu: And then they move to the second stage, which is relevance-driven reduction. This involves looking at how attention shifts across different layers of the large language model.
Meng: They adaptively drop visual tokens in specific decoder layers if both cross-modal attention ratios, which are the text-to-visual and visual-to-text ones, fall below a certain threshold called tau.
The paper's improvements: Tom: The core improvement is this two-stage approach, which handles token selection based on diversity first and then removes irrelevant tokens based on where the model’s attention fades.
Jane: That second stage is guided by observing an "information migration" phenomenon, where cross-modal attention gets strong in early layers but then drops off in deeper layers.
Lu: The theoretical justification they give is that intra-modal redundancy and cross-modal redundancy are statistically independent in the embedding space, which validates why this two-stage compression paradigm works.
Meng: Experimentally, they show that ToDRE prunes up to ninety percent of visual tokens after the vision encoder and all visual tokens in certain LLM decoder layers.
Conclusion: Tom: So, to wrap up, ToDRE is a training-free method that first selects a maximally diverse subset of visual tokens using that greedy max-sum diversification algorithm, then removes all remaining visual tokens once cross-modal attention fades.
Jane: This combination allows them to achieve a two point six times speed-up in total inference time while keeping ninety-five point zero percent model performance and excellent compatibility across twelve benchmarks.
Lu: From my perspective, the potential here is that by keeping only the most diverse and relevant tokens, we might be able to build models that are significantly more efficient for real-world deployment.
Meng: I’m thinking about how this impacts the practical side of things; reducing inference time by that much means we could run these complex vision tasks on much less computational power.
Lalam: For me, as a language model, it means the resulting representations are cleaner because they don't carry unnecessary visual noise from redundant tokens.
Tom: So, ToDRE is a training-free framework that tackles token pruning by looking at diversity and relevance separately to get significant speed gains while keeping performance high.
More episodes
- 2610.10768-Strategic Investment Decision Making for Value Creation in Energy Transition: A Reinforcement Learning Approach
- 2610.10858-RFChipAgent: Multi-Agentic AI Flow for Analog/RF Chip Design
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
- 2312.01221-Enabling Quantum Natural Language Processing for Hindi Language