Confidence under Visual Token Pruning: Removed Evidence and Risk-Controlled Token Budgets for MLLMs

arXiv:2604.12035 · cs.CV · Submitted 2026-04-13 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Confidence under Visual Token Pruning".

Jane: Visual token pruning affects how multimodal large language models (MLLMs) believe their answers, and this effect can be quantified by calibration metrics.

Tom: First, who's behind it and why it matters.

Title and authors: Tom: So, to recap, this paper examines the effect of removing visual tokens on model calibration in MLLMs, specifically looking at how different selection methods—like coverage versus attention—handle that trade-off against accuracy. The authors highlight a key finding regarding which selection rule matters most.

Jane: That's right; they are comparing various ways to prune those tokens and measuring the resulting agreement between what the model predicts and what it actually gets right, which is calibration. It’s not just about getting more tokens out of there, it’s about keeping things reliable.

Lu: The authors also introduce a specific evaluation pitfall where simply zeroing out the hidden states of pruned positions can actually destroy accuracy by making the model behave randomly because those zeroed states still absorb attention mass (<ref:2604.12035#pg2>).

Meng: That's a huge practical warning, Lu; if we implement pruning naively by just zeroing out parts of the model's state, we risk collapsing the performance entirely, so that correction they suggest is vital for real-world deployment.

Lalam: I think this whole investigation into how selection rules dictate calibration is super important because it tells us precisely what kind of visual information we should be prioritizing when we compress these models for use.

The paper's summary: Tom: Moving on to the core findings, the paper demonstrates that moderate coverage-based pruning improves calibration even when there is no statistically significant drop in accuracy on benchmarks like POPE, which is pretty interesting because it suggests a way to get better confidence without sacrificing correctness.

Jane: That’s what they found: using coverage-based pruning with a specific setting, like moving from five hundred seventy-six down to one hundred twenty-eight tokens on LLaVA-one point five for the POPE benchmark, reduced the expected calibration error from zero point zero four one down to zero point zero one six without changing how many correct answers the model gives <ref:2604.12035#pg0,expected calibration error from 0.041>.

Lu: They explain this through their evidence-coverage account, showing that accuracy is strongly associated with kept-set coverage with a Spearman correlation of +zero point eight nine, while mean confidence isn't really related to it at all (rho = -zero point zero three) (<ref:2604.12035#pg1>).

Meng: So, the mechanism they are pointing to is that the accuracy is tied directly to which visual evidence we retain, but the confidence score doesn't necessarily follow that same trend, which is a crucial distinction for us when designing compression pipelines.

Lalam: That means we don't have to worry about sacrificing accuracy just because we are trying to make the model leaner; instead, by focusing on keeping relevant evidence, we can maintain better reliability in the confidence scores.

The paper's improvements: Tom: The authors suggest several ways to improve how we use this information, highlighting that the selection rule itself is a major calibration knob that operates before any other post-hoc adjustments are made.

Jane: They show that coverage-based selection really dominates on calibration across different budgets, and they even found that less saliency weight helps at every budget compared to using only saliency, which is a big hint for tuning our methods.

Lu: The paper points out that the ordering of selectors stays stable across different architectures like GQA and LLaVA-NeXT, confirming that this selection mechanism is fairly universal in its effect on calibration (<ref:2604.12035#pg2>).

Meng: For practical implementation, this suggests we should stop treating saliency as the automatic default for pruning; instead, using a hybrid selector that balances coverage and saliency seems like a much safer approach to get that improved calibration they observed.

Lalam: It’s really empowering to see how this research gives us actionable levers—like tuning the alpha parameter in SCOPE—to actively manage our model's confidence quality rather than just accepting whatever the standard pruning routine throws at us.

Conclusion: Tom: So, wrapping up "Confidence under Visual Token Pruning: Removed Evidence and Risk-Controlled Token Budgets for MLLMs," the main message is that we need to report calibration alongside accuracy because two selectors with equal accuracy can differ by a factor of three in their expected calibration error.

Jane: That means for anyone working on compressing multimodal models, focusing solely on the raw accuracy number misses a big piece of the picture regarding how trustworthy those confidence scores really are.

Lu: The paper confirms that moderate coverage-based pruning is a solid strategy for achieving better calibration at unchanged accuracy, and this evidence-coverage account gives us a solid framework for understanding performance trade-offs.

Meng: For us in the engineering side, the guidance here is pretty clear: use hybrid selectors to expose saliency weight as a free calibration knob while always running that no-drop equivalence sanity check to prevent accidental accuracy collapses.

Lalam: I’m really excited because this work provides a concrete standard for evaluating compression; it tells us that confidence quality should be treated as a standard axis right alongside accuracy and the computational cost of these token compressions.

Kaizhen Tan, Yang Feng, Heqing Du, Hanzhe Hong, Siru Tao, Xin Xu

Carnegie Mellon University · Columbia University

cs.CV

Submitted: 2026-04-13

Updated: 2026-10-04

Importance score: 86/100

The gist: Visual token pruning affects how multimodal large language models (MLLMs) believe their answers, and this effect can be quantified by calibration metrics.

Key concepts

Calibration
Calibration measures how well a model's predicted confidence matches its actual correctness. A calibrated model's high confidence predictions are usually correct, and low confidence predictions are usually incorrect. The study shows that pruning methods can improve this relationship.
Coverage-based Pruning
This selection rule keeps tokens based on how much of the visual information they represent across the entire image. It aims to select a set of tokens that collectively cover the most important visual evidence, rather than just picking individually salient pixels.
Attention-based Selection
This method selects tokens based on their importance within the model's attention mechanism. It prioritizes visual information that is highly relevant to the model's internal processing, focusing on where the model 'looks' most closely during inference.

Terminology

Summary

Visual token pruning affects how multimodal large language models (MLLMs) believe their answers, and this effect can be quantified by calibration metrics. The core finding is that selection rules for pruning matter more than the token budget alone, with coverage-based pruning improving calibration at unchanged accuracy on POPE while attention-based selection preserves confidence as accuracy deteriorates.

How it works

The study systematically investigates the interaction between visual token pruning and model calibration, defined as the agreement between confidence and correctness. The research compares various selection signals—such as saliency, coverage, diversity, optimal transport—against a baseline random selection across different MLLMs (LLaVA-1.5-7B, LLaVA-NeXT-Vicuna-7B, Qwen2-VL) and benchmarks (POPE, ScienceQA-IMG). A controlled testbed called SCOPE is used to interpolate between pure coverage and saliency-weighted coverage by sweeping the saliency exponent α.

Key Findings on Calibration

The empirical results show that moderate coveragebased pruning improves calibration at statistically unchanged accuracy on the POPE benchmark (ECE 0.041→0.016 at K=128). This suggests that the selection rule is more critical than the token budget itself. Specifically, less saliency weight helps at every budget, and saliency-only selection is significantly worse calibrated than random at aggressive budgets. The authors identify an evidence-coverage account, where accuracy tracks kept-set coverage (Spearman ρ=+0.89) while confidence is independent of it (ρ=−0.03). This implies that overconfidence is the gap that opens as coverage falls (ρ=−0.92).

The Role of Selection Rules

The selection rule dictates the calibration outcome, as evidenced by how different selectors rank against each other. On POPE, coverage-based selection dominates on calibration at both budgets. The ordering across various conditions is stable on GQA and LLaVA-NeXT, where coverage beats random at both ratios. Conversely, attention-based selection preserves confidence while accuracy deteriorates. The paper notes that the cross-selector ordering holds on GQA, on LLaVA-NeXT, and on Qwen2-VL, demonstrating that the selector itself is a calibration knob operating before post-hoc methods.

Evaluation Pitfalls and Robustness Checks

The study identifies an important evaluation pitfall: zeroing rather than removing pruned tokens can reduce accuracy to chance. The authors present a corrected implementation of FastV which physically removes tokens, restoring sensible behavior and revealing the honest failure mode: flat, unshakeable confidence over collapsing accuracy. Furthermore, they confirm that confidence definitions matter more in principle, less in practice, noting that margin and entropy are monotone transforms of max-probability confidence. Verbalized self-confidence is found to be nearly degenerate at this scale, leaving first-token probabilities as the most informative channel for analysis.

Task and Model Dependence

The effect of pruning is task and model-dependent, with a clear boundary where selection rule ordering ceases to matter. On ScienceQA-IMG, coverage stops ordering calibration at all because performance is driven by language priors. In contrast, on POPE, the coverage–saliency axis separates methods by up to 3× in ECE. The account generalizes across different backbones (LLaVA-1.5-7B, LLaVA-NeXT-Vicuna-7B, Qwen2-VL), showing that the calibration free lunch and the α-trend transfer when budgets are matched by keep ratio. However, on Qwen2-VL at 10% retention, pruning costs accuracy (2.4 points) to reduce ECE (from.064 to.040), indicating a trade rather than a free lunch in some regimes.

Conclusion and Guidance

The paper concludes that Confidence quality should join accuracy and FLOPs as a standard axis for evaluating token compression in multimodal models. For practical deployment, moderate coverage-based pruning buys cheaper inference and better-ranked confidence at little or no accuracy cost. The primary guidance is to report calibration alongside accuracy, as two selectors with equal accuracy can differ by 3× in ECE, and to use hybrid selectors to expose saliency weight as a free calibration knob. The study emphasizes using a no-drop-equivalence sanity check for pruned baselines.

The gist: Moderate coverage-based pruning improves calibration at unchanged accuracy on POPE, while attention-based selection preserves confidence as accuracy falls away. The selection rule matters more than the token budget alone. This is explained by an evidence-coverage account where accuracy tracks kept-set coverage, and overconfidence is the gap between that coverage and mean confidence.

Improvements for AI systems

Based on the research presented in this paper, here are specific, actionable improvements for AI systems:

  1. Acknowledge Calibration alongside Accuracy: When deploying or evaluating any visual token pruning method (e.g., FastV, SCOPE), always report the Expected Calibration Error (ECE) and Brier scores alongside accuracy. This moves beyond simple accuracy wins metrics to assess the reliability of the model's confidence scores, which is critical for downstream decision-making systems.

  2. Adopt Coverage-Based Pruning as a Default Strategy: For tasks where accuracy gains are marginal or non-existent (especially when language priors dominate, like ScienceQA), use coverage-based pruning (like SCOPE with α=0) as the primary method.

  3. Use Saliency Weighting for Confidence Tuning: When deploying a pruning system, use saliency-only selection only if you specifically want to preserve the model's pre-existing high-attention anchors, acknowledging that this often leads to exacerbated overconfidence gaps (as seen in FastV).

  4. Implement No-Drop Equivalence Checks for Pruning: When developing or deploying pruning pipelines (especially those involving hidden state zeroing), implement a rigorous check to ensure the pruned model behaves identically to the unpruned model when no tokens are actually dropped. This prevents catastrophic accuracy loss due to implementation shortcuts (as demonstrated by the FastV zeroing artifact).

  5. Employ Hybrid Selectors for Robustness: Use hybrid selection methods (like SCOPE's default setting, α=1) as a robust baseline. These selectors tend to balance coverage and saliency well, providing better calibration than pure saliency or pure coverage on many benchmarks.

  6. Contextualize Pruning with Task Type: Understand that the optimal selector depends on the task structure. For tasks dominated by language priors (e.g., ScienceQA), token selection rules have little impact on accuracy, suggesting that a simpler, high-coverage rule is sufficient to maintain reasonable calibration without risking accuracy degradation.

  7. Mitigate Overconfidence Gaps: Recognize that pruning often concentrates confidence loss on incorrect answers more than correct ones. System designers should treat the gap between confidence and correctness (overconfidence) as a primary metric to monitor during model compression, rather than focusing solely on accuracy gains.

  8. Avoid Query-Conditioned Selection for General Pruning: Be cautious with selection rules that are conditioned directly by the LLM's internal attention mechanism (query-conditioned). These methods can lead to significantly worse calibration than coverage-based methods under aggressive pruning budgets.

This paper shows that what tokens are kept (coverage) matters more for confidence quality than how important those tokens seem to be based on internal attention alone.

Sources

Related papers