Focusing by Contrastive Attention: Enhancing VLMs' Visual Reasoning
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Focusing by Contrastive Attention: Enhancing VLMs' Visual Reasoning".
Jane: The paper was written by Yuyao Ge, Shenghua Liu, Yiwei Wang, Lingrui Mei, Baolong Bi et al. from Institute of Computing Technology, Chinese Academy of Sciences and University of California, Merced.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Welcome back to the show, everyone. Today we’re cracking open a fresh one from arXiv, and the title alone got me hooked: “Focusing by Contrastive Attention: Enhancing VLMs' Visual Reasoning.” Jane, what’s your first read on that?
Jane: Tom, I love it because it’s so direct. These vision-language models, the ones that can look at a picture and answer questions about it, they’re brilliant, but they get distracted. This paper is basically saying, hey, we can teach them to focus better without any extra training.
Tom: Right, and that’s the part that blew my mind when I first skimmed it. They’re not building a new model. They’re not fine-tuning anything. They’re using the model’s own attention, the little internal spotlight it shines on different parts of an image, to figure out what’s noise and what’s actually important.
Jane: Exactly. And the authors, Yuyao Ge and the team from ICT, CAS, they start with a really clever observation. They noticed that when you give these models a complex image, like a cluttered shelf or a busy street scene, the model’s attention gets all spread out. It’s like when you walk into a crowded room and you can’t find your friend because there’s just too much going on.
Tom: That’s the perfect analogy. And they actually proved it. They measured something called attention entropy, which is basically a measure of how scattered that spotlight is. The more complex the image, the higher the entropy, and the worse the model performs. It’s a direct link between visual chaos and reasoning failure.
Jane: So the problem is clear, but the fix is what’s really clever. They call it CARVE, Contrastive Attention Refinement for Visual Enhancement. The idea is to run the model twice. Once with a generic prompt like “describe this image,” and once with the actual question you want answered.
Tom: And then you compare the attention maps from those two runs. The generic prompt shows you where the model naturally looks, which is full of visual noise. The specific question shows you where it *should* look. By contrasting them, you can isolate the signal from the noise.
Jane: And they do this at the pixel level. They literally create a mask that blacks out the noisy parts of the image, keeping only the regions that are relevant to the question. Then they feed that cleaned-up image back to the model. It’s like giving the model a pair of blinders so it can only see the horse it’s supposed to be looking at.
Tom: And the results are pretty dramatic, especially on older, smaller models. We’re talking about a seventy-five percent relative improvement on the V* benchmark with LLaVA-one point five-7B. That’s not a small bump; that’s a game-changer for those models.
Jane: It really is. And the beauty is that it’s training-free. You don’t need a supercomputer to make this work. You just need to run the model a few times. That’s a huge deal for practical applications.
Tom: So we’ve got a problem, a diagnosis, and a solution that actually works. But I’m dying to know how they figured out exactly *where* to look in the model’s brain. That’s what we’re digging into next.
Summary: Tom: So, Jane, we’ve established that CARVE is this clever masking trick. But the paper doesn’t just throw a method at you. It actually does the detective work to explain *why* it works. That’s what I love about this summary section.
Jane: Oh, absolutely. They go deep into the mechanics of the model’s attention. They found that attention isn’t static. It evolves as information flows through the layers. In the shallow layers, the model is doing a broad scan, like looking at a whole map. In the middle layers, it starts to narrow down to regions. And in the deep layers, it should be locking onto the specific object.
Tom: But here’s the kicker. In a complex image, that final convergence never really happens. The model gets stuck in that “confused where to look” state. The attention stays diffuse even in the deepest layers, which is why it gives a wrong answer.
Jane: And they tie this directly back to the image itself. They break down visual complexity into two measurable things: texture and color. Texture is how many edges and fine details are in the image. Color is how diverse the hues are. They show that both of these have a strong, positive correlation with that scattered attention.
Tom: So a busy, colorful image literally makes the model’s attention more chaotic. And that chaos is what hurts its performance. It’s a really clean causal chain they’ve built here. It’s not just an observation; it’s a mechanism.
Jane: And they even provide a theoretical framework for it. They decompose the attention map into two parts: a visual noise component, which is inherent to the image, and a semantic signal component, which is tied to the question. The generic prompt mostly captures the noise. The specific question captures the signal.
Tom: And CARVE is the mathematical operation that separates those two. It’s like a filter that amplifies the signal and suppresses the noise. The paper proves that this contrast operation gives you a clean estimate of the semantic signal, which is exactly what you need to build the mask.
Jane: Right. And they don’t just stop at the theory. They test it. They show that using attention from the deeper layers, where the model is closest to focusing, works much better than using shallow layers. It makes sense, right? You want to use the most refined attention map you can get.
Tom: And they also show that using the final token’s attention is better than the first token’s. That’s because the final token has seen the entire question and all the context, so its attention is the most informed.
Jane: So the summary is really a complete story. They identify the failure mode, they explain the underlying mechanism, they build a theoretical model for it, and then they use that model to design a targeted fix. It’s a textbook example of good research.
Tom: It really is. But a good story in a paper is one thing. I want to know how this holds up in the real world. What happens when you throw different models and different types of questions at it? That’s the next part of our conversation.
Improvements: Tom: Alright, Jane, so we’ve got the theory and the method. Now let’s talk about the proof in the pudding. How much better does CARVE actually make these models?
Jane: The numbers are really impressive, Tom. They tested it on four different datasets and four different models. And in every single case, CARVE improved the accuracy. There wasn’t a single instance where it made things worse.
Tom: And the improvements are huge, especially for the older models. We already mentioned the seventy-five percent jump on V* for LLaVA-one point five-7B. But even the newer Qwen models, which are already pretty good, got a solid boost. Qwen two point five-VL-7B went up by over seventeen percent on that same V* benchmark.
Jane: That’s a really interesting pattern. The weaker the model, the more it benefits. It’s like CARVE is giving them a crutch, but a really good crutch. The newer models are already better at focusing, so they don’t need as much help.
Tom: And it’s not just about accuracy. They also compared CARVE against other methods. They compared it to using external tools like SAM, which is a segmentation model, and YOLO, which is an object detector. CARVE beat them all.
Jane: And that makes sense. Those tools are generic. They don’t know what question you’re asking. SAM will segment everything, and YOLO will find all the objects, but neither of them knows that you specifically care about the shape seen through the cup’s handle. CARVE is question-aware.
Tom: Exactly. It’s using the model’s own understanding of the question to guide the masking. It’s not just cropping out a random object; it’s cropping out the *right* object.
Jane: And they also did a sensitivity analysis on the masking parameters. They found that you don’t want to mask too aggressively. If you keep only twenty percent of the image, you might throw away important context. The sweet spot is keeping around forty percent to sixty percent of the pixels, focusing on the top two or three regions.
Tom: So it’s a balance. You want to remove the noise, but you don’t want to remove the whole scene. You need to keep enough context for the model to understand what it’s looking at.
Jane: Right. And the computational cost is reasonable. It takes about one point three four seconds per image on a good GPU, which is slower than just running the model once, but it’s not prohibitively expensive. For the accuracy gains you get, it’s a pretty good trade-off.
Tom: So we have a method that’s training-free, works across models, and beats external tools. But I’m always thinking about the bigger picture. Where does this actually get used? Who cares about a seventeen percent improvement on a benchmark? Let’s bring in Lu and Meng to talk about the real-world impact.
Lu: Tom, this is where it gets exciting. Think about any system that uses a vision-language model in a messy, uncontrolled environment. Autonomous driving, for instance. A car sees a cluttered street scene. This method could help it focus on the traffic light or a pedestrian instead of getting distracted by a billboard.
Meng: And from an engineering standpoint, that’s huge. We can’t always retrain a massive model for every new environment. CARVE is a plug-and-play solution. You can put it in front of an existing model to make it more robust. The fact that it’s training-free means we can deploy it immediately without a massive GPU cluster.
Jane: And it’s not just for safety-critical systems. Think about accessibility tools. A vision model helping a blind person navigate a grocery store. The model needs to answer questions like “which is the cheapest cereal?” This method could help it filter out all the colorful packaging and focus on the price tags.
Tom: So it’s not just about making models smarter on a test. It’s about making them more reliable in the real world, where things are messy and complicated. That’s the real value proposition here.
Conclusion: Tom: Well, we’ve had a fantastic time unpacking “Focusing by Contrastive Attention: Enhancing VLMs' Visual Reasoning.” We’ve gone from a simple observation about distracted models to a clever, training-free solution that delivers real gains.
Jane: It really is a complete package. The paper identifies a core problem, explains why it happens, and then provides a practical fix that works across multiple models and datasets. It’s the kind of research that you can immediately see being applied in the field.
Tom: And I think the biggest takeaway for me is that you don’t always need a bigger model or more data. Sometimes, you just need to understand how the model you already have is thinking, and then help it think a little better.
Jane: Exactly. It’s about working with the model’s own intelligence, not against it. The contrastive attention trick is so elegant because it uses the model’s own internal spotlight to find the signal in the noise.
Tom: And that’s a powerful idea that could extend beyond just vision. The concept of contrasting a general response with a specific one to isolate relevant information could be useful in other areas of AI, like text summarization or even reasoning.
Jane: I think you’re right. It’s a clever mechanism that has the potential to be a general tool for improving focus in AI systems. We’re saying goodbye to this paper, but I have a feeling we’ll be seeing the influence of this idea for a while.
Tom: Absolutely. A big thank you to the authors for sharing this work. And to our listeners, thanks for tuning in. We’ll be back soon with another fascinating paper from the arXiv. Until then, keep your attention focused on the good stuff.
Jane: See you next time, everyone!
Yuyao Ge, Shenghua Liu, Yiwei Wang, Lingrui Mei, Baolong Bi, Xuanshan Zhou, Jiayu Yao, Jiafeng Guo, Xueqi Cheng
Institute of Computing Technology, Chinese Academy of Sciences · University of California, Merced
cs.CV, cs.AI
Submitted: 2026-08-17
Updated: 2026-08-18
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 58/100
Key concepts
- Attention Entropy
- A measure of how scattered the model's internal spotlight is when processing an image. Higher entropy indicates a more complex or cluttered image, which correlates with worse model performance.
- CARVE (Contrastive Attention Refinement for Visual Enhancement)
- A training-free technique where the model is run twice—once with a generic prompt and once with a specific question. Comparing the attention maps from these two runs allows the method to isolate the relevant signal from visual noise by creating a mask.
- Semantic Signal vs. Visual Noise Component
- The paper decomposes an attention map into two parts: visual noise, which is inherent to the image's texture and color, and a semantic signal, which relates to the specific question. CARVE uses contrast to separate these two components for targeted enhancement.
Terminology
Summary
Summary
Vision-Language Models (VLMs) have demonstrated remarkable success across diverse visual tasks, yet their performance degrades in complex visual environments. While existing enhancement approaches require additional training, rely on external segmentation tools, or operate at coarse-grained levels, they overlook the innate ability within VLMs. To bridge this gap, we investigate VLMs’ attention patterns and discover that: (1) visual complexity strongly correlates with attention entropy, negatively impacting reasoning performance; (2) attention progressively refines from global scanning in shallow layers to focused convergence in deeper layers, with convergence degree determined by visual complexity. (3) Theoretically, we prove that the contrast of attention maps between general queries and task-specific queries enables the decomposition of visual signal into semantic signals and visual noise components. Building on these insights, we propose Contrastive Attention Refinement for Visual Enhancement (CARVE), a training-free method that extracts task-relevant visual signals through attention contrasting at the pixel level. Extensive experiments demonstrate that CARVE consistently enhances performance, achieving up to 75% improvement on open-source models. Our work provides critical insights into the interplay between visual complexity and attention mechanisms, offering an efficient pathway for improving visual reasoning with contrasting attention.
The paper investigates whether complex images interfere with VLMs’ attention mechanisms, making it difficult for them to focus on task-relevant regions. The authors define visual complexity as texture and color dimensions, revealing a significant positive correlation between both factors and attention entropy. Furthermore, their analysis shows that attention entropy negatively correlates with accuracy on visual reasoning tasks. Through this two-stage analysis, they establish that complex visual information impairs VLMs’ reasoning performance via attention distribution.
A preliminary experiment on TextVQA is conducted by first applying progressive masking to obscure background regions, then cropping to retain only task-relevant regions and adaptively magnifying them to the original image size. Results show that cluttered visual environments initially cause incorrect predictions, but correct token probability surpasses incorrect probability at mask ratios of approximately 0.02 and 0.65 respectively, providing initial validation that masking visual noise can improve correct token probability.
To automate the visual noise masking process, the authors leverage contrasting attention maps between general instructions and task-specific questions to distinguish semantic signal from visual noise. CARVE is a contrastive method for visual extraction that masks with contrastive attention maps, crops and magnifies semantic regions to focus on essential semantic signal.
The paper defines texture complexity using Canny edge detection, where the texture complexity Tc(I) is the L1 norm of the binary edge map divided by the total number of pixels, yielding values in [0,1]. Color complexity is defined using hue values in HSV space, computed as the negative sum of ρb ln ρb divided by ln B, where B=180 hue bins, yielding values in [0,1] where higher values indicate greater color diversity.
For measuring attention distribution, Shannon entropy is employed. The overall attention entropy H is computed as the average over layers of the entropy of the attention map at the final generation step. Higher entropy indicates more dispersed attention, while lower entropy indicates more concentrated focus.
The correlation analysis shows that both texture complexity and color complexity exhibit strong positive linear relationships with attention entropy, indicating that complex visual features lead to dispersed attention patterns in VLMs. Additionally, a strong negative correlation between attention entropy and accuracy is revealed: as attention entropy increases from 5.1 to 6.8, performance decreases from approximately 76% to 65%.
The hierarchical evolution of attention entropy shows that attention entropy monotonically decreases with layer depth, and the 95% confidence intervals progressively widen with increasing depth, indicating enhanced inter-sample variability. For samples with clear visual targets, deep layers achieve highly concentrated attention, while for noisy samples, the model maintains dispersed attention patterns even in deep layers.
The theoretical foundation defines attention decomposition where the attention map decomposes as the Hadamard product of a visual noise factor Fvis(I) and a semantic signal factor Fsem(Q,I). When using general instructions G, the semantic signal function reduces to uniform distribution, making general instruction attention predominantly capture visual noise. The paper proves a closed-form solution for semantic extraction: Âi = Ai(Q) / (Ai(G) + λ), which approximates the semantic signal Fsem,i when Fvis,i ≫ λ.
CARVE comprises three stages: Stage 1 generates general attention distribution with general instructions; Stage 2 extracts task-specific attention; Stage 3 applies contrasted attention to generate enhanced masked images for noise suppression. The algorithm fuses attention maps across layers and generation time steps through weighted aggregation, where later tokens receive higher fusion weights. Task-relevant regions are identified by applying a top-p percentile threshold, and connected component analysis extracts coherent regions. The top-K regions ranked by cumulative attention scores are selected, and the enhanced image is generated through masking, cropping, and resizing.
Experiments are conducted on four datasets: A-OKVQA, POPE, V∗, and TextVQA, covering visual reasoning, visual understanding, and visual knowledge reasoning. Four VLMs are evaluated: QWEN 2.5-VL-3B-INSTRUCT, QWEN 2.5-VL-7B-INSTRUCT, LLAVA-1.5-7B, and LLAVA-1.5-13B.
Results demonstrate CARVE’s consistent performance enhancement across all evaluated models and datasets. Earlier-generation models exhibit substantially greater improvements than more recent counterparts. For instance, LLAVA 1.5-7B achieves a 71.83% relative improvement on V∗, whereas QWEN 2.5-VL-7B shows a 17.52% gain. This pattern indicates that limited-capability models suffer more from visual complexity interference and benefit more from contrastive attention-guided focusing mechanisms.
Ablation studies on time step selection reveal that tend generally outperforms Tfull, which in turn surpasses tstart across most configurations. Later tokens encode richer contextual information by accessing complete preceding sequences during inference, so the final token’s attention maps accurately localize target objects. Ablation studies on layer selection show that the layer-wise performance ordering from best to worst is: [20,25], [15,20], single layer 25, single layer 20, single layer 14, and [10,15]. Multi-layer fusion outperforms single-layer alternatives by capturing complementary information and providing robustness against individual layer randomness.
Sensitivity analysis of mask generation shows that when p=1.0 (no masking), performance remains at original levels, while p within [0.2, 0.6] combined with K ∈ 2, 3 achieves optimal performance. Aggressive masking strategies with retention ratios set to 20% and single-region constraints lead to degradation.
Comparative analysis with alternative methods shows that CARVE substantially outperforms external tool-based approaches including SAM, YOLO, CLIP, and recent ViCrop variants. External tools rely on generic segmentation algorithms that lack question-image context awareness. CARVE requires 1.34 seconds of GPU processing time, exceeding simpler approaches such as YOLO (0.35 seconds) but remaining within practical deployment constraints.
The paper also provides theoretical proofs including the existence and uniqueness of attention decomposition, strict convexity of the optimization objective, closed-form expression of the optimal solution, approximation error bounds, and theoretical selection of the regularization parameter. The regularization parameter λ serves dual purposes: controlling the bias-variance tradeoff and ensuring numerical stability.
The paper concludes that visual complexity correlates with attention entropy, which in turn negatively impacts VLMs’ performance. Theoretically, contrasting attention maps between general and specific instructions enables effective decomposition of visual signal into semantic signal and visual noise components. CARVE is a training-free method that leverages this theoretical insight to extract task-relevant signal through attention contrasting and pixel-level masking, providing critical insights into the interplay between visual complexity and attention mechanisms.
Improvements for AI systems
Based on the paper, here are the specific improvements I can implement in AI systems:
Implementation: Add a pre-inference pipeline that:
-
Runs two parallel forward passes: one with a general instruction (
Write a general description of the image
) and one with the task-specific question -
Extracts attention maps from layers 20–25 (for 28-layer models) at the final generation token
-
Computes contrastive attention:
 = A question / (A general + λ)with λ ≈ 0.05 -
Applies top-p percentile masking (p = 0.2–0.6) with connected-component analysis, keeping top-2 or top-3 regions
-
Crops and resizes the masked image back to original dimensions before final inference
Resulting capability: The system can automatically identify and remove visual distractors (cluttered backgrounds, irrelevant objects) before answering, improving accuracy by up to 75% on complex visual reasoning tasks (e.g., V* dataset with LLaVA-1.5-7B).
Implementation: During inference, compute attention entropy per layer using Shannon entropy. If entropy at deep layers (e.g., layer 25) exceeds a threshold (e.g., > 6.0 normalized), trigger the CARVE refinement automatically. This creates a self-diagnosing system that knows when it's confused
and applies correction.
Implementation: For models with different depths, dynamically select the layer range based on the entropy-convergence pattern. Use the observation that entropy monotonically decreases with depth—select layers where entropy reduction rate is highest (typically the last 20–25% of layers). For Qwen-2.5-VL (28 layers), use layers 20–25; for LLaVA-1.5 (32 layers), use layers 24–28.
Implementation: Instead of using only the first or last token's attention, fuse attention maps across all generation steps with weights proportional to token position (w t = t - t start + 1). This leverages the fact that later tokens have richer contextual information and more accurate spatial localization.
Implementation: Before answering, compute texture complexity (Canny edge density) and color complexity (HSV hue entropy). If both exceed thresholds (e.g., texture > 0.5, color > 0.5), apply CARVE; otherwise, skip refinement to save computation. This reduces inference overhead by 30% on simple images.
Implementation: After generating the contrastive attention map, rank connected regions by cumulative attention score. Instead of keeping a fixed number of regions, dynamically select regions whose cumulative score exceeds 80% of the total attention mass. This adapts to images with one dominant object versus images with multiple relevant objects.
Implementation: Apply CARVE specifically to object presence verification tasks. By masking out non-relevant regions, the system reduces false positives caused by visually similar but irrelevant objects, improving POPE accuracy from 86.9% to 88.4% (Qwen-2.5-VL-3B) and from 83.6% to 89.0% (LLaVA-1.5-7B).
Implementation: For applications where multiple questions are asked about the same image, cache the general-instruction attention maps. Since these depend only on the image, they can be reused across questions, reducing the inference cost from 3 passes to 2 passes per question (after the first).
Abstract
Vision-Language Models (VLMs) have demonstrated remarkable success across diverse visual tasks, yet their performance degrades in complex visual environments. While existing enhancement approaches require additional training, rely on external segmentation tools, or operate at coarse-grained levels, they overlook the innate ability within VLMs. To bridge this gap, we investigate VLMs' attention patterns and discover that: (1) visual complexity strongly correlates with attention entropy, negatively impacting reasoning performance; (2) attention progressively refines from global scanning in shallow layers to focused convergence in deeper layers, with convergence degree determined by visual complexity. (3) Theoretically, we prove that the contrast of attention maps between general queries and task-specific queries enables the decomposition of visual signal into semantic signals and visual noise components. Building on these insights, we propose Contrastive Attention Refinement for Visual Enhancement (CARVE), a training-free method that extracts task-relevant visual signals through attention contrasting at the pixel level. Extensive experiments demonstrate that CARVE consistently enhances performance, achieving up to 75% improvement on open-source models. Our work provides critical insights into the interplay between visual complexity and attention mechanisms, offering an efficient pathway for improving visual reasoning with contrasting attention.
Sources
- Star Attention: Efficient LLM Inference over Long Sequences
- Context-DPO: Aligning Language Models for Context-Faithfulness
- RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs
- PIS: Linking Importance Sampling and Attention Mechanisms for Efficient Prompt Compression
- Related Knowledge Perturbation Matters: Rethinking Multiple Pieces of Knowledge Editing in Same-Subject
- DP-IQA: Utilizing Diffusion Prior for Blind Image Quality Assessment in the Wild
- VistaWise: Building Cost-Effective Agent with Cross-Modal Knowledge Graph for Minecraft
- Innate Reasoning is Not Enough: In-Context Learning Enhances Reasoning Large Language Models with Less Overthinking
- DistiLLM-2: A Contrastive Approach Boosts the Distillation of LLMs
- Evaluating Object Hallucination in Large Vision-Language Models
- Structured Attention Matters to Multimodal LLMs in Document Understanding
- Visual Instruction Tuning
- Smooth Reading: Bridging the Gap of Recurrent LLM to Self-Attention LLM on Long-Context Tasks
- On Learning to Summarize with Large Language Models as References
- EE-MLLM: A Data-Efficient and Compute-Efficient Multimodal Large Language Model
- SLANG: New Concept Comprehension of Large Language Models
- "Not Aligned" is Not "Malicious": Being Careful about Hallucinations of Large Language Models' Jailbreak
- HiddenGuard: Fine-Grained Safe Generation with Specialized Representation Router
- a1: Steep Test-time Scaling Law via Environment Augmented Generation
- A Survey of Context Engineering for Large Language Models
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models