AdaptInfer: Adaptive Token Pruning for Vision-Language Model Inference via Dynamical Text Guidance

arXiv:2508.06084 · cs.CV · Submitted 2025-08-08 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "AdaptInfer: Adaptive Token Pruning for Vision-Language Model Inference via Dynamical Text Guidance".

Tom: Vision–language models (VLMs) have achieved impressive performance on multimodal reasoning tasks such as visual question answering, image captioning and so on,

Jane: First, who's behind it and why it matters.

Title and authors: Tom: So we're talking about AdaptInfer today, and the title is "AdaptInfer: Adaptive Token Pruning for Vision-Language Model Inference via Dynamical Text Guidance". It sounds pretty technical, but really, what it means is this new way to cut down the massive number of vision tokens that these large models process when they are just doing inference.

Jane: Exactly. The authors are focusing on how to make that pruning process smarter by letting the model change its strategy dynamically while it's running. It's moving away from just using static rules for deciding which parts of the vision input are important during that prefill stage.

Lu: What’s interesting about this is they aren't just looking at attention patterns once; they are trying to capture those dynamic signals happening inside the model while it’s actually running its inference pass. They want to use that movement to make a better decision about which vision tokens to keep.

Meng: So, instead of applying a fixed filter, the system tries to figure out token importance in real time based on what the model is currently doing layer by layer. That sounds like it could save a lot of processing time for deployment.

Lalam: I think this is important because when we think about making these models useful for everyday applications, like quick visual answers, cutting down that huge token count during runtime really helps with speed and cost.

The paper's summary: Tom: Okay, so the core idea here in AdaptInfer is to propose a plug-and-play framework for adaptive vision token pruning. They introduce a fine-grained, dynamic text-guided pruning mechanism that reuses existing layer-wise attention maps.

Jane: That means they are taking those attention maps they already have from the model's layers and using them to create a prior—a kind of educated guess—about which vision tokens are most valuable at each step.

Lu: They take that prior and then reweight the text-to-vision attention matrix based on it, calculating an importance score for every single visual token by averaging across all the attention heads.

Meng: And they use those scores to rank all the visual tokens and just keep the top k for that specific layer during inference. It’s a way to prune based on what's happening right then and there.

Lalam: So, it’s like giving each layer of the model a personalized importance list for the vision part, which is much more flexible than just pruning everything based on one single text prompt before we even start.

Tom: It sounds like they are trying to solve the problem where static methods fail because they don't account for how the importance of text tokens changes as the model refines its internal understanding across different layers.

The paper's improvements: Jane: One big improvement is that instead of relying on just text prompt guidance, AdaptInfer uses dynamic text guidance during inference to better capture the evolving importance of those tokens.

Tom: And they don't stop there; they use offline analysis of cross-modal attention shifts to identify consistent inflection points in the model’s layers where pruning will be most effective.

Lu: They found specific layers that act as these inflection points, like layer one ten and twenty in LLaVA-one point five-7B, or layer zero nine and nineteen in Qwen2-VL-2B.

Meng: That offline analysis is key because it gives you a principled schedule instead of just guessing where to cut things off based on trial and error during testing. It’s data-driven timing for the pruning.

Lalam: And the computational complexity analysis shows that even with all this extra work, the additional FLOPs for pruning are really minimal compared to the main transformer computations during prefill. It’s designed to be lightweight and plug-and-play.

Tom: So we're looking at a system that doesn't just prune tokens blindly; it uses dynamic guidance informed by where the model naturally shifts its focus, creating a schedule that is both principled and efficient.

Conclusion: Jane: So to wrap up, AdaptInfer gives us this plug-and-play solution for adaptive vision token pruning in VLMs that uses dynamic text guidance to prune vision tokens based on layer-wise attention scores.

Tom: It achieves a significant reduction in CUDA latency, showing a sixty-one point three percent drop while keeping an average accuracy of ninety-three point one percent on the LLaVA-one point five-7B model, which is pretty solid for this type of method.

Lu: The paper shows that this approach is generalizable across different models and even different architectures because you can do a simple offline attention shift analysis to tailor the schedule to other models with different layer counts or parameter scales.

Meng: From an engineering standpoint, it’s good that the overhead is small, but the real win here for someone building something practical is that it delivers high accuracy under all the token budgets tested.

Lalam: It’s really interesting because they showed how you can maintain high performance while drastically cutting down on vision tokens, which opens up possibilities for much faster multimodal reasoning applications in the near future.

Weichen Zhang, Zhui Zhu, Ningbo Li, Shilong Tao, Kebin Liu, *Yunhao Liu

Global Innovation Exchange, Tsinghua University

cs.CV

Submitted: 2025-08-08

Updated: 2026-10-04

Code: https://github.com/fartashf/vsepp

Importance score: 92/100

The gist: Vision–language models (VLMs) have achieved impressive performance on multimodal reasoning tasks such as visual question answering, image captioning and so on, but their inference cost remains a

Key concepts

Dynamic Text Guidance
Instead of using static prompts for guidance, AdaptInfer dynamically determines the importance of text tokens at each inference layer. This allows the method to capture the 'dynamic fluidity of information' across layers, providing a more informed way to score and prune vision tokens in real-time during processing.
Text-to-Text Attention Matrix ($A_{t2t}$)
This matrix captures how text tokens attend to other text tokens within the model at each layer. By aggregating these scores, AdaptInfer estimates a prior importance weight for each text token, which is then used to reweight the attention paid to vision tokens in the next step.
Cross-Modal Attention Shifts
This involves analyzing how attention patterns change between different layers of the VLM during inference. Identifying consistent inflection points in these shifts helps authors design a principled and efficient pruning schedule, suggesting when aggressive token removal will yield the best computational savings without sacrificing performance.

Terminology

Summary

Vision–language models (VLMs) have achieved impressive performance on multimodal reasoning tasks such as visual question answering, image captioning and so on, but their inference cost remains a significant challenge due to the large number of vision tokens processed during the prefill stage. Existing pruning methods often rely on directly using the attention patterns or static text prompt guidance, failing to exploit the dynamic internal signals generated during inference. AdaptInfer addresses this by proposing a plug-and-play framework for adaptive vision token pruning in VLMs, which leverages dynamic text guidance and an offline analysis of cross-modal attention shifts to create a principled and efficient pruning schedule.

The gist: AdaptInfer is a plug-and-play framework for adaptive vision token pruning in VLMs that reuses layer-wise text-to-text attention maps to construct soft priors over text-token importance, allowing more informed scoring of vision tokens at each stage, and identifies consistent attention inflection points to propose a more principled and efficient pruning schedule. <ref:2508.06084#pg2>

How it works

AdaptInfer is a plug-and-play framework for adaptive vision token sparsification in which VLM dynamically determines text token guidance during inference, allowing the method to be seamlessly integrated as a plugin into existing VLMs <ref:2508.06084#pg4>. The mechanism involves three main steps:

  1. Extracting the text-to-text attention matrix Ah t2t ∈ R T ×T for each pruning layer, where T is the number of text tokens, and aggregating attention scores across the query dimension to estimate importance: w(h) = X Σ i=1 A(h)t2t [i,:] ∈ R T <ref:2508.06084#pg8>.

  2. Using this prior to reweight the text-to-vision attention matrix A(h)t2v ∈ R T ×V, where V is the number of visual tokens remaining, and computing the importance score of each visual token on average of all attention heads: s = 1 Σ H Σ h=1 w(h)⊤ · A(h)t2v ∈ R V <ref:2508.06084#pg10>.

  3. Ranking all visual tokens and retaining the top-k for the current layer based on these scores, denoted as Ik = TopK(s, k).

How it works (Schedule Design)

The method proposes a principled pruning schedule inspired by offline analysis of cross-modal attention shifts to identify consistent inflection points in inference <ref:2508.06084#pg4>. The authors identified consistent attention inflection points at layer 1, 10, and 20 in LLava-1.5-7B (Liu et al., 2023b) and at layer 0, 9, and 19 on Qwen2-VL-2B (Wang et al., 2024a), suggesting that aggressive pruning immediately after these layers is a more effective and computationally efficient strategy on LLava and Qwen <ref:2508.06084#pg8>. The proposed schedule involves selecting to prune vision tokens after layer 1 and after layer 20 on LLaVA-1.5-7B, with an additional pruning location at layer 10 to balance caution and efficiency <ref:2508.06084#pg10>.

How it works (Computational Complexity)

The computational complexity analysis shows that the additional FLOPs for pruning are minimal relative to the main transformer computations during the prefill stage, estimated as F LOP sprune = T2 + 2TV. The method is lightweight and plug-and-play, introducing little additional computational overhead because both attention matrices are already computed during the forward pass <ref:2508.06084#pg10>.

How it works (Results and Evaluation)

Experimental results verified the effectiveness of AdaptInfer, showing that it reduces CUDA latency by 61.3% while maintaining an average accuracy of 93.1% on vanilla LLaVA-1.5-7B <ref:2508.06084#pg10>. In Table 1, AdaptInfer achieves the highest overall accuracy scores under both 128 and 64 vision token budgets on LLava-1.5-7B. Furthermore, in latency tests on LLaVA-1.5-7B, AdaptInfer reaches a lower average cuda latency of 33.0 ms per sample than PDrop and SparseVLM <ref:2508.06084#pg10>.

How it works (Performance Analysis)

The dynamic text guidance mechanism outperforms static text guidance because it dynamically infers the relative importance of text tokens during inference at each layer, whereas static methods fail to capture the dynamic fluidity of information across layers <ref:2508.06084#pg10>. In Qwen2-VL-2B, AdaptInfer consistently matches or outperforms SparseVLM across both image-based QA and multi-frame video QA benchmarks. The visualization study shows that ADAPTINFER rapidly suppresses tokens associated with semantically irrelevant background areas while consistently preserving those aligned with question-critical cues.

How it works (Pruning Tolerance)

The pruning-tolerance evaluation on MME showed that code reasoning, count and color-based tasks have a lower pruning tolerance relatively, while calculation and OCR tasks outperform the original Qwen due to higher pruning tolerance and the noise reduction effect by token pruning. Instance-level decision analysis revealed that the vast majority of predictions remain unchanged after pruning, with the fraction of Correct to Error cases being small and comparable to Error to Correct cases <ref:2508.06084#pg10>.

How it works (Generalizability)

The proposed adaptive pruning schedule is generalizable across multimodal tasks and different models by performing a simple, offline attention shift analysis, which allows the schedule to be tailored to other models with different parameter scales or architectures <ref:2508.06084#pg10>. The authors tested AdaptInfer on LLava-1.5-13B and Qwen2-VL-2B, confirming that the token-sparsification strategy scales gracefully with model size <ref:2508.06084#pg10>.

How it works (Conclusion)

AdaptInfer proposes a novel plug-and-play solution for VLM acceleration via dynamic text-guided pruning, which introduces minimal additional computational overhead while maintaining high accuracy and achieving SOTA accuracy under all the token budgets chosen in the experiments. Specially, AdaptInfer achieves SOTA accuracy under all the token budgets chosen in the experiments. For instance, AdaptInfer reduces CUDA latency by 61.3% and retains an average accuracy of 93.1% on LLaVA-1.5-7B, with only 64 vision tokens preserved per layer on average <ref:2508.06084#pg10>.

REFERENCES

Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katie Millicah, Malcolm Reynolds, Roman Ring, Eliza Rutherford, Serkan Cabihiimannarilalajaiyiannaa. Flamingo: a visual language model for few-shot learning. In Proceedings of the 36th International Conference on Neural Information Processing Systems, NIPS ’22, Red Hook, NY, USA, 2022. Curran Associates Inc. ISBN 9781713871088.<ref:2508.06084#pg2>

Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C. Lawrence Zitnick, and Devi Parikh. Vqa: Visual question answering. In 2015 IEEE International Conference on Computer Vision (ICCV), pp. 2425–2433, 2015. doi: 10.1109/ICCV.2015.2798339.<ref:2508.06084#pg8>

Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou. Qwen-vl: A versatile vision-language model for understanding localization text reading and beyond. arXiv preprint arXiv:2308.12966, 2023.<ref:2508.06084#pg8>

Daniel Bolya, Cheng-Yang Fu, Xiaoliang Dai, Peizhao Zhang, Christoph Feichtenhofer, and Judy Hoffman. Token merging: Your ViT but faster. In International Conference on Learning Representations, 2023.<ref:2508.06084#pg2>

Rishi Bommasani, Drew A. Hudson, Ehsan Adeli, Russ B. Altman, Simran Arora, Sydney von Arx, Michael S.

Improvements for AI systems

  1. textbf Dynamic Text Guidance Mechanism for Vision Token Pruning (AdaptInfer): This mechanism allows VLMs to dynamically infer their relative importance during inference at each predefined pruning layer, meaning the model can adapt its vision token selection based on evolving internal representations, overcoming the limitation where any static text-prompt–guided pruning approach is likely to fail in capturing this evolving semantic alignment.

  2. textbf Principled Pruning Schedule Based on Attention Shift Analysis: The method utilizes offline analysis of cross-modal attention shifts and identify consistent inflection locations in inference, which inspires a more principled and efficient pruning schedule by identifying stable regions (e.g., layers 2–9 and 20+) where importance rankings are reliable for pruning, instead of relying on empirical rule-of-thumb or extensive hyperparameter optimization experiments.

  3. textbf Enhanced Inference Efficiency and Accuracy Trade-off: The system achieves significant computational savings, such as reducing CUDA latency by 61.3% while maintaining high accuracy (an average accuracy of 93.1% on vanilla LLaVA-1.5-7B), demonstrating that the method is both lightweight and plug-and-play and offers a data-driven and architecture-aware solution compared to heuristic schemes.

  4. textbf Robust Performance Across Model Scales: The framework is shown to be generalizable across different VLM architectures, as it can be easily transferred by performing an offline attention shift analysis, confirming its robustness when applied to models like LLaVA-1.5-13B and Qwen2-VL-2B with different layer counts.

Sources

Related papers