ERASE: Eliminating Redundant Visual Tokens via Adaptive Two-Stage Token Pruning

arXiv:2605.09982 · cs.CV · Submitted 2026-05-11 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "ERASE: Eliminating Redundant Visual Tokens via Adaptive Two-Stage Token Pruning".

Jane: The gist The ERASE framework is an adaptive two-stage vision token pruning framework that identifies and retains salient tokens through pruning strategies adaptive to image complexity,

Tom: First, who's behind it and why it matters.

Paper summary: Jane: So to wrap up this paper on ERASE: Eliminating Redundant Visual Tokens via Adaptive Two-Stage Token Pruning, the authors are showing how to prune visual tokens by making the pruning strategy change depending on image complexity. It jointly addresses both visual redundancy and contextual relevance in a dynamic way.

Lu: The implication is that we can move toward vision processing that isn't just brute-force compression but something intelligently guided by the input data itself, using entropy to guide the initial cut and text relevance to guide the deeper pruning layers.

Meng: It means for those building these systems, they don't have to pick one fixed way to prune; they can let the system decide whether it needs early-layer pruning or deeper layer processing based on how complex the visual scene is.

Tom: They’re proving that this adaptive approach keeps accuracy high even when you apply aggressive pruning ratios, which is a tough hurdle for any token reduction method. ERASE consistently achieves a superior efficiency-accuracy trade-off compared to previous vision token pruning methods across diverse architectures and benchmarks.

Lalam: It suggests that future advancements in VLM efficiency won't just come from bigger models, but from smarter ways to handle the sheer volume of visual data they process efficiently.

Jane: That’s the main point—it’s about designing a more flexible system where the pruning isn't static, but responsive to what it’s actually looking at. We need this kind of adaptive mechanism in multimodal AI.

Conclusion: Tom: So we're wrapping up on ERASE, this paper titled "ERASE: Eliminating Redundant Visual Tokens via Adaptive Two-Stage Token Pruning." Basically, what they’ve built is a two-stage system that prunes the vision tokens in large models based on how complex the picture actually is.

Jane: It’s smart because it doesn't use one fixed pruning rule for every image; instead, it adapts its strategy depending on the visual input.

Lu: The authors used entropy scores from the raw image to decide how much of the visual data to keep in that first stage, and then they look at text relevance for the second stage.

Meng: It sounds like they’re trying to save a ton of computation time during generation by cutting out what’s visually redundant before it even gets into the main language model.

Lalam: And Lalam sees this as really important because if you can compress those vision tokens efficiently, it makes running huge models for tasks way more practical for everyday use.

Tom: Exactly. The core idea is that complex images need different pruning treatments than simple ones, and ERASE figures out which treatment to apply on the fly.

Jane: It’s about balancing keeping enough visual information to be accurate with cutting down the massive computational overhead that comes from huge image inputs in modern AI.

Lu: Think about it this way—for a really messy, detailed photo, you need to keep more tokens than for a simple landscape. ERASE handles that complexity adjustment directly.

Tom: And the results they showed are pretty telling; they’re seeing big speedups on high-resolution images while still keeping accuracy quite high compared to previous methods.

Meng: I’m interested in how practical this is for real deployment; does this framework run smoothly across different types of AI models, or is it picky?

Jane: The paper tested it on several architectures, and the authors showed that ERASE works well even when you switch between different model designs.

Lu: That transferability they demonstrated, applying configurations optimized for one model to a completely different one without losing much performance, that’s pretty significant for the research community.

Tom: Yeah, it points toward a more flexible way of building these multimodal systems where the compression happens naturally within the framework itself.

Jane: It suggests that future work could focus on how this adaptive mechanism interacts with other parts of the model's processing pipeline. (Sound of music swells slightly)

Tom: We’ve got a lot to think about here, and next up, we’re going to look at how this adaptive pruning compares directly against some of the older methods that tried to tackle this problem before ERASE came along.

Department of Electrical and Computer Engineering, Sungkyunkwan University · Department of Semiconductor Systems Engineering, Sungkyunkwan University

cs.CV

Submitted: 2026-05-11

Updated: 2026-10-08

Comments: 19 pages, 11 figures

Code: https://github.com/Tuna-Luna/ERASE

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 84/100

The gist: The gist The ERASE framework is an adaptive two-stage vision token pruning framework that identifies and retains salient tokens through pruning strategies adaptive to image complexity, demonstrating

Key concepts

Image Entropy
This measures the information density within an image by quantifying how much information is present. High entropy indicates areas with rapid changes, like edges, which are visually rich. Low entropy suggests smooth, redundant regions that can be safely pruned without losing important visual detail.
Stage 1: Image-level Pruning
This initial stage removes redundant tokens directly from the raw image based on its complexity. ERASE uses Bayesian Optimization to set pruning ratios for different complexity levels, ensuring that visually simple images are pruned differently than complex ones.
Context-aware Pruning
The second stage prunes tokens based on their relevance to the input text. It dynamically selects the appropriate decoder layer—early for simple images and mid-to-late layers for complex ones—to maintain accuracy while targeting text-relevant information.

Terminology

Summary

The gist The ERASE framework is an adaptive two-stage vision token pruning framework that identifies and retains salient tokens through pruning strategies adaptive to image complexity, demonstrating superior efficiency-performance trade-off over existing methods

Motivation

Recent advancements in Vision-Language Models (VLMs) enable large language models (LLMs) to process high-resolution images, significantly improving realworld multimodal understanding. However, this capability introduces a large number of vision tokens, resulting in substantial computational overhead. Existing approaches predominantly rely on learned semantic features within the model to capture visual redundancy, and they lack adaptive mechanisms to adjust pruning strategies according to the complexity of the input image. Vision token scaling in modern VLMs is quadratic with image resolution, leading to a significant increase in prefill latency and KV cache memory usage. For example, a 4K image produces around 16K tokens. Therefore, efficient VLM inference requires effective vision token compression.

Proposed ERASE Framework

ERASE is a hierarchical vision token pruning framework that compresses vision tokens through two complementary pruning stages. Specifically, ERASE consists of (1) image-level pruning, which removes visually redundant regions using entropy scores computed directly from the raw image, and (2) context-aware pruning, which further prunes vision tokens at adaptive decoder layers based on input-specific contextual redundancy.

Stage 1: Image-level vision token pruning

Stage 1 of ERASE removes redundant vision tokens directly from the input image while preserving visual information. Based on this intuition, ERASE uses image entropy to estimate image complexity and apply Bayesian Optimization to determine the optimal pruning ratio for each complexity level. Entropy quantifies the amount of information in a discrete random variable X as defined by H(X) = −Σ P(xi) ln P(xi). Redundancy mainly arises from spatially continuous regions, resulting in low entropy, whereas high-information regions typically occur near edges where pixel values change rapidly, producing high entropy. Global raw image complexity is defined with the median of all patch-level entropy scores across the image. To determine the optimal pruning ratio for each image, ERASE partitions image complexity into discrete levels and assigns a target Stage 1 pruning ratio to each level based on predefined entropy thresholds Θ.

Stage 2: Context-aware vision token pruning

Stage 2 adopts an instruction-aware strategy that preserves tokens relevant to the input text. ERASE uses attention scores to measure text-vision relevance, but introduces a dynamic mechanism for selecting the decoder layer used for relevance estimation and pruning. The hypothesis is that simple images require only early-layer processing to identify text-relevant regions, whereas complex images require deeper layers to capture fine-grained details relevant to the prompt. This is validated by comparing accuracy retained after pruning tokens at early versus mid-to-late decoder layers, showing that pruning at mid-to-late layers maintains stable accuracy across different global entropy values. ERASE introduces an image-complexity-aware dual-layer selection mechanism where images with overall entropy H¯ ≤ Θ⌈Θ/2⌉ are classified as simple and undergo early-layer pruning, while images exceeding this threshold use mid-to-late layer pruning.

Experiment and Results

Extensive experiments demonstrate that ERASE consistently achieves a superior efficiency-performance trade-off over existing vision token pruning methods across diverse VLM architectures and benchmarks. For Qwen2.5-VL-7B, at a token pruning ratio of 85%, ERASE retains 89.46% of the original model accuracy, whereas the best prior method retains only 78.19%. ERASE consistently outperforms previous methods across all pruning ratios on both models. Furthermore, ERASE generalizes across diverse architectures because it uses raw images for Stage 1 pruning and attention scores for Stage 2 pruning. In terms of efficiency, ERASE achieves the highest overall task speedup among all evaluated methods, reaching up to 1.56× end-to-end speedup on high-resolution images with around 16K vision tokens.

Conclusion

In this paper, ERASE is an adaptive two-stage vision token pruning framework for efficient VLM inference. ERASE jointly addresses visual redundancy and contextual relevance through entropy-guided image-level pruning and adaptive text-conditioned token pruning. In particular, ERASE dynamically adjusts both the Stage 1 pruning ratio and the Stage 2 pruning layer according to image complexity. Extensive experiments across multiple VLM architectures and challenging benchmarks demonstrate that ERASE consistently achieves superior efficiency-accuracy trade-offs compared to previous methods, preserving higher accuracy even under aggressive pruning ratios while significantly reducing end-to-end latency in high-resolution settings.

Algorithm of ERASE

Algorithm 1 outlines the ERASE framework, utilizing a complexity configuration set T pre-optimized via Bayesian Optimization. In Stage 1, the global image entropy (H¯) is computed as the median of M patch-level local entropies. Based on H¯, the algorithm retrieves a retention ratio (r1) and a target pruning layer (k) from T, preserving only the top ⌊M · r1⌋ high-entropy tokens. In Stage 2 (layer k), lightweight text-to-vision cross-attention identifies and retains the top Kfinal text-relevant tokens. Finally, ERASE retrospectively evicts KV cache entries corresponding to pruned vision tokens for all layers up to k, significantly reducing memory usage during autoregressive generation.

D Details on Bayesian Optimization

The objective of the Bayesian Optimization is to maximize a joint reward that balances task performance and computational efficiency, defined by F = α · Accuracy + (1 − α) · Σ ci · pi. The dataset for Bayesian Optimization is meticulously sampled from diverse benchmarks, comprising solely instances where the target VLM originally predicted the correct answers.

Transferability of Bayesian Optimization

To demonstrate the transferability of ERASE, ERASE† demonstrates remarkable robustness by retaining 87.00% of the base performance when applied to InternVL3-8B using configurations optimized for Qwen2.5-VL-7B. Crucially, ERASE† still significantly outperforms the best competing baseline (IVC-Prune, 75.40%).

Efficiency Analysis

ERASE successfully strikes an optimal balance between computational efficiency and task accuracy through its adaptive two-stage pruning mechanism. While IVC-Prune achieves the highest accuracy among prior methods, it obtains minimal speedup in overall task latency because its prefill latency reduction is bottlenecked by its late-layer pruning strategy. ERASE employs a lightweight non-iterative entropy calculation at the input stage, removing a substantial portion of redundant tokens before they enter the LLM backbone.

Entropy Analysis

Following the same settings as in Fig. 5, patches with a local entropy below 3.1 (indicating low information density) are visualized in the middle column, while the remaining highly informative patches are shown on the right. As the global entropy increases, the proportion of low-entropy regions visibly decreases, substantiating our hypothesis.

C Algorithm of ERASE

Algorithm 1 outlines the ERASE framework, utilizing a complexity configuration set T pre-optimized via Bayesian Optimization. In Stage 1, the global image entropy (H¯) is computed as the median of M patch-level local entropies. Based on H¯, the algorithm retrieves a retention ratio (r1) and a target pruning layer (k) from T, preserving only the top ⌊M · r1⌋ high-entropy tokens. The median of the optimized thresholds (Θ) acts as a definitive boundary: H¯ below this median categorizes the image as simple (triggering Stage 2 early), while exceeding it delays Stage 2 to a mid-to-late layer. In Stage 2 (layer k), lightweight text-to-vision cross-attention identifies and retains the top Kfinal text-relevant tokens. Notably, if Stage 1 already yields ≤ Kfinal tokens, Stage 2 is entirely bypassed, directly pruning to Kfinal tokens. Finally, we retrospectively prune the KV cache of preceding layers (1... k) for the evicted tokens to optimize memory efficiency.

D.3 Obtained parameter values

Table 13 details the optimized entropy thresholds and pruning ratios for each model. Within the Qwen2.5-VL series, both models exhibit similar optimal configurations due to their shared architectural design, with minor deviations likely stemming from differences in model capacity. Conversely, the Qwen3-VL series adopts a more conservative approach to classifying simple images using lower entropy thresholds yet applies significantly higher pruning ratios to them.

**E.

Improvements for AI systems

  1. The ERASE framework enables efficient VLM inference by compressing vision tokens through two complementary pruning stages that jointly exploit redundancy at both image and contextual levels. This allows AI systems to process high-resolution images with significantly reduced computational overhead, as demonstrated by achieving up to a 1.56× end-to-end speedup for Qwen2.5-VL on 4K images compared to the base model.

  2. The adaptive mechanism dynamically adjusts pruning strategies based on image complexity, specifically using image entropy to estimate image complexity and assigning a target Stage 1 pruning ratio derived from Bayesian Optimization. This means the system can aggressively prune simple scenes (simple images require only early-layer processing) while using deeper layers for complex ones, leading to preserved accuracy even under extreme compression.

  3. The context-aware pruning stage in ERASE uses a dynamic mechanism for selecting the decoder layer used for relevance estimation and pruning, which is selected based on image complexity: Images with overall entropy H¯ ≤ Θ⌈Θ/2⌉ are classified as simple and undergo early-layer pruning, whereas images exceeding this threshold are classified as complex and use mid-to-late layer pruning. This ensures that tokens relevant to the input text are retained while minimizing processing through unnecessary layers.

  4. The system can achieve superior performance on fine-grained visual reasoning tasks by preserving critical information: ERASE consistently outperforms previous methods across all pruning ratios on both models, and it is shown to perform comparable to the base model on standard visual grounding tasks while successfully preserving and capturing the fine-grained details within the image.

Sources

Related papers