AdaptInfer: Adaptive Token Pruning for Vision-Language Model Inference via Dynamical Text Guidance

summary

Video file (mp4)

The gist

Vision–language models (VLMs) have achieved impressive performance on multimodal reasoning tasks such as visual question answering, image captioning and so on, but their inference cost remains a

In short

AdaptInfer is a plug-and-play framework for pruning vision tokens in Vision-Language Models (VLMs) during inference. It uses dynamic text guidance, leveraging layer-wise text attention maps to estimate token importance and create an efficient pruning schedule based on cross-modal attention shifts, significantly reducing latency while maintaining high accuracy.

Key concepts

Dynamic Text Guidance
Instead of using static prompts for guidance, AdaptInfer dynamically determines the importance of text tokens at each inference layer. This allows the method to capture the 'dynamic fluidity of information' across layers, providing a more informed way to score and prune vision tokens in real-time during processing.
Text-to-Text Attention Matrix ($A_{t2t}$)
This matrix captures how text tokens attend to other text tokens within the model at each layer. By aggregating these scores, AdaptInfer estimates a prior importance weight for each text token, which is then used to reweight the attention paid to vision tokens in the next step.
Cross-Modal Attention Shifts
This involves analyzing how attention patterns change between different layers of the VLM during inference. Identifying consistent inflection points in these shifts helps authors design a principled and efficient pruning schedule, suggesting when aggressive token removal will yield the best computational savings without sacrificing performance.

Terminology used across episodes

This episode discusses

The paper

AdaptInfer: Adaptive Token Pruning for Vision-Language Model Inference via Dynamical Text Guidance · Read on arXiv

Weichen Zhang, Zhui Zhu, Ningbo Li, Shilong Tao, Kebin Liu, *Yunhao Liu

Global Innovation Exchange, Tsinghua University

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "AdaptInfer: Adaptive Token Pruning for Vision-Language Model Inference via Dynamical Text Guidance".

Tom: Vision–language models (VLMs) have achieved impressive performance on multimodal reasoning tasks such as visual question answering, image captioning and so on,

Jane: First, who's behind it and why it matters.

Title and authors: Tom: So we're talking about AdaptInfer today, and the title is "AdaptInfer: Adaptive Token Pruning for Vision-Language Model Inference via Dynamical Text Guidance". It sounds pretty technical, but really, what it means is this new way to cut down the massive number of vision tokens that these large models process when they are just doing inference.

Jane: Exactly. The authors are focusing on how to make that pruning process smarter by letting the model change its strategy dynamically while it's running. It's moving away from just using static rules for deciding which parts of the vision input are important during that prefill stage.

Lu: What’s interesting about this is they aren't just looking at attention patterns once; they are trying to capture those dynamic signals happening inside the model while it’s actually running its inference pass. They want to use that movement to make a better decision about which vision tokens to keep.

Meng: So, instead of applying a fixed filter, the system tries to figure out token importance in real time based on what the model is currently doing layer by layer. That sounds like it could save a lot of processing time for deployment.

Lalam: I think this is important because when we think about making these models useful for everyday applications, like quick visual answers, cutting down that huge token count during runtime really helps with speed and cost.

The paper's summary: Tom: Okay, so the core idea here in AdaptInfer is to propose a plug-and-play framework for adaptive vision token pruning. They introduce a fine-grained, dynamic text-guided pruning mechanism that reuses existing layer-wise attention maps.

Jane: That means they are taking those attention maps they already have from the model's layers and using them to create a prior—a kind of educated guess—about which vision tokens are most valuable at each step.

Lu: They take that prior and then reweight the text-to-vision attention matrix based on it, calculating an importance score for every single visual token by averaging across all the attention heads.

Meng: And they use those scores to rank all the visual tokens and just keep the top k for that specific layer during inference. It’s a way to prune based on what's happening right then and there.

Lalam: So, it’s like giving each layer of the model a personalized importance list for the vision part, which is much more flexible than just pruning everything based on one single text prompt before we even start.

Tom: It sounds like they are trying to solve the problem where static methods fail because they don't account for how the importance of text tokens changes as the model refines its internal understanding across different layers.

The paper's improvements: Jane: One big improvement is that instead of relying on just text prompt guidance, AdaptInfer uses dynamic text guidance during inference to better capture the evolving importance of those tokens.

Tom: And they don't stop there; they use offline analysis of cross-modal attention shifts to identify consistent inflection points in the model’s layers where pruning will be most effective.

Lu: They found specific layers that act as these inflection points, like layer one ten and twenty in LLaVA-one point five-7B, or layer zero nine and nineteen in Qwen2-VL-2B.

Meng: That offline analysis is key because it gives you a principled schedule instead of just guessing where to cut things off based on trial and error during testing. It’s data-driven timing for the pruning.

Lalam: And the computational complexity analysis shows that even with all this extra work, the additional FLOPs for pruning are really minimal compared to the main transformer computations during prefill. It’s designed to be lightweight and plug-and-play.

Tom: So we're looking at a system that doesn't just prune tokens blindly; it uses dynamic guidance informed by where the model naturally shifts its focus, creating a schedule that is both principled and efficient.

Conclusion: Jane: So to wrap up, AdaptInfer gives us this plug-and-play solution for adaptive vision token pruning in VLMs that uses dynamic text guidance to prune vision tokens based on layer-wise attention scores.

Tom: It achieves a significant reduction in CUDA latency, showing a sixty-one point three percent drop while keeping an average accuracy of ninety-three point one percent on the LLaVA-one point five-7B model, which is pretty solid for this type of method.

Lu: The paper shows that this approach is generalizable across different models and even different architectures because you can do a simple offline attention shift analysis to tailor the schedule to other models with different layer counts or parameter scales.

Meng: From an engineering standpoint, it’s good that the overhead is small, but the real win here for someone building something practical is that it delivers high accuracy under all the token budgets tested.

Lalam: It’s really interesting because they showed how you can maintain high performance while drastically cutting down on vision tokens, which opens up possibilities for much faster multimodal reasoning applications in the near future.

More episodes

← Home