Attend to Evidence: Evidence-Anchored Spatial Attention Supervision for Multimodal RLVR

arXiv:2605.30912 · cs.CV, cs.CL · Submitted 2026-05-29 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Attend to Evidence: Evidence-Anchored Spatial Attention Supervision for Multimodal RLVR".

Jane: The paper was written by the authors from Harbin Institute of Technology and Zhongguancun Academy and Zhongguancun Institute of Artificial Intelligence and Nankai University and Shanghai Jiaotong University and Zhejiang University.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

The Mechanism: Tom: We’ve covered how the core problem with current RLVR methods is that they only reward the final correct answer, but we're curious about what EASE actually does to fix that gap.

Jane: The paper "Attend to Evidence: Evidence-Anchored Spatial Attention Supervision for Multimodal RLVR" solves this by mapping those specific evidence regions—the parts of the image that justify the answer—into a smoothed target distribution over all visual tokens.

Lu: It’s not a simple hard switch, but a probability map that guides the response-to-image attention during training, which is a mathematically sophisticated way to ensure the model learns from multiple supporting regions simultaneously.

Meng: And their approach is highly efficient because they apply this guidance only to trajectories where the model already achieved a high score, making it an extremely targeted and resource-efficient training signal that minimizes wasted compute.

Lalam: This targeted learning ensures that we are strengthening the visual connection for successful reasoning, not just adding noise or random patterns to every single response in the entire training set.

Tom: It's clear they have designed this mechanism with both theoretical elegance and practical efficiency in mind, which is impressive considering the complexity of multimodal data.

Jane: This soft approach is brilliant because it allows the model to smoothly transition its focus, rather than being abruptly forced into one single box, which is much more realistic for complex reasoning tasks.

Lu: The mathematical framework they use—the Gaussian mixture target—is designed to be robust, ensuring that the attention mechanism doesn't just clump on a single point but covers all the necessary supporting regions.

Meng: And their approach is highly efficient because they apply this guidance only to trajectories where the model already achieved a high score, making it an extremely targeted and resource-efficient training signal.

Lalam: This targeted learning ensures that we are strengthening the visual connection for successful reasoning, not just adding noise to every single response in the training set.

Tom: It’s clear they have designed this mechanism with both theoretical elegance and practical efficiency in mind. But how do we know if these gains hold up across different levels of AI complexity?

Jane: We need to see if this works on small, medium, and large models, right?

Results and Gains: Tom: Moving into the results, the paper "Attend to Evidence: Evidence-Anchored Spatial Attention Supervision for Multimodal RLVR" shows that EASE performs very consistently across different model sizes.

Jane: The empirical results show a consistent improvement across different model sizes—specifically Qwen2 point 5-VL-7B, Qwen3-VL-4B, and Qwen3-VL-8B—proving the scalability of this approach.

Lu: This suggests that the "process supervision" signal is extremely robust across different architectures, which is a major theoretical win for consistency in how we develop AI intelligence.

Meng: From an engineering standpoint, seeing these gains consistently across various backbones means we can implement this as a powerful module regardless of the specific LLM platform we are using today.

Lalam: It’s truly exciting to see the model becoming more aligned with visual truth, making it less prone to hallucination and much more dependable when facing complex reasoning challenges.

Tom: These gains are significant too, ranging from two point five to three point one points on benchmarks like visual math and perception tests, which is a huge jump in model reliability for those fine-grained tasks.

Jane: The data shows that EASE is not just improving one type of performance, but boosting the average score across multiple diverse benchmarks, which speaks volumes about its versatility.

Lu: This suggests that the "process supervision" signal is extremely robust across different architectures, which is a major theoretical win for consistency in how we develop AI intelligence.

Meng: From an engineering standpoint, seeing these gains consistently across various backbones means we can implement this as a powerful module regardless of the specific LLM platform we are using today.

Lalam: It’s truly exciting to see the model becoming more aligned with visual truth, making it less prone to hallucination and much more dependable when facing complex reasoning challenges.

Tom: It's clear that this improvement in alignment is a repeatable, measurable outcome across model generations. But what does this mean for the practical deployment of these large systems?

Jane: We need to know if we can trust these gains when it translates into real-world tasks like medical or industrial applications.

Conclusion: Tom: So, we've seen how "Attend to Evidence: Evidence-Anchored Spatial Attention Supervision for Multimodal RLVR" provides a structured way to improve multimodal AI by focusing on the internal process rather than just the final outcome score.

Jane: It’s about making sure that when an answer is right, it’s because the AI was actually looking at the right evidence, not just some lucky coincidence or a language prior bias.

Lu: This paper successfully operationalizes visual evidence acquisition as a measurable training signal for visual grounding in reinforcement learning, which is a huge step toward formalizing visual reasoning.

Meng: The practical impact of this work is that we're getting more reliable tools for complex tasks like medical diagnosis or autonomous inspection because they are truly grounded in the image data.

Lalam: I hope this framework helps us build an AI culture where "seeing" and "knowing" are always aligned, ensuring that our machines truly understand the world they interact with.

Tom: We've covered a lot of ground today, but as we wrap up, let's give one last thought on the future possibilities.

Jane: It seems like a great time to look ahead at how this could feed into the next generation of model training.

Lu: I can already see creative applications where the AI needs to verify its own hypotheses against visual evidence, not just predict an answer.

Meng: I'm looking forward to seeing how we can optimize this EASE mechanism further in production environments to maximize inference speed.

Lalam: It’s a beautiful foundation for building a trustworthy relationship between the AI and the human users who depend on its reliable outputs.

Final Thoughts: Tom: Well, we’ve spent a lot of time today breaking down this paper, and it's pretty clear that the core of "Attend to Evidence: Evidence-Anchored Spatial Attention Supervision for Multimodal RLVR" is about fundamentally improving how AI understands the relationship between its output and the visual world.

Jane: Exactly. It’s about moving past just being right on a test—it’s about ensuring that when we get an answer, we know *why* it' by making sure the AI was properly looking at the evidence in the image.

Lu: From a theoretical perspective, I think this opens up incredible possibilities for complex reasoning where visual input is critical; it allows us to formalize exactly how much of our AI's "thought process" is actually grounded in reality.

Meng: And I see huge practical benefits here too, because as an engineer, I want to know that we can deploy systems—say in automated inspection or medical imaging—that aren't just guessing but are rigorously tethered to the visual data.

Lalam: I think this framework will help build a much higher level of trust in AI applications; it ensures that the machines are not just generating plausible-sounding answers, but truly understanding the world they operate within.

Tom: That is such an important distinction, Lalam, and it brings us right back to that idea of moving away from blind luck toward genuine process.

Jane: It's reassuring to see this kind of deliberate design in reinforcement learning—it’s like adding a safety checklist to the AI's internal reasoning process.

Lu: This capability really suggests we can push past current limitations in how visual and linguistic information is synthesized within a single model architecture.

Meng: We need to start thinking about how to scale this EASE concept across different deployment environments, ensuring that its reliability doesn' doesn't degrade when we move beyond the lab settings.

Lalam: It’s a powerful combination of engineering rigor and ethical responsibility, ensuring that the AI’s "eyes" are as reliable as its "mind."

Tom: Absolutely, it’s a remarkable advancement in multimodal RLVR.

Jane: We really hope this work paves the way for future solutions to complex visual reasoning tasks.

Harbin Institute of Technology · Zhongguancun Academy · Zhongguancun Institute of Artificial Intelligence · Nankai University · Shanghai Jiaotong University · Zhejiang University

cs.CV, cs.CL

Submitted: 2026-05-29

Updated: 2026-09-03

Comments: Accepted to EMNLP 2026

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 87/100

The gist: " Reinforcement learning with verifiable rewards (RLVR) has become a standard method for improving vision-language models (VLMs).

Key concepts

EASE
Evidence-Anchored Spatial Attention Supervision is the core mechanism of the paper. It maps specific evidence regions—the parts of an image that justify an answer—into a smoothed target distribution over all visual tokens. This guides response-to-image attention during training, allowing the model to learn from multiple supporting regions simultaneously rather than just a single point.
Process Supervision
This signal is used in the paper to supervise the internal process of reasoning. It ensures that when an AI achieves a high score, it was actually looking at the correct visual evidence. This soft approach strengthens the visual connection for successful reasoning, making it more realistic and robust than hard switching.
Scalability across Model Sizes
The results show that EASE performs consistently across different model sizes, including Qwen2 point 5-VL-7B, Qwen3-VL-4B, and Qwen3-VL-8B. This suggests the process supervision signal is robust across various architectures and can be implemented as a powerful module regardless of the specific LLM platform used.
Visual Grounding
This refers to the AI's ability to connect its linguistic output with visual input accurately. The paper operationalizes visual evidence acquisition as a measurable training signal for visual grounding in reinforcement learning, helping models become more aligned with visual truth and less prone to hallucination.

Terminology

Summary

"

Reinforcement learning with verifiable rewards (RLVR) has become a standard method for improving vision-language models (VLMs). However, traditional RLVR relies solely on outcome rewards, which are final answers. The authors argue that these outcome-only signals are insufficient because they do not tell the model which image regions justify an answer.

This lack of process supervision creates a perception credit assignment problem. A correct answer may result from various factors—such as language priors, dataset bias, a label-correlated shortcut, or guessing—rather than genuine visual reasoning. The inability to track the visual evidence acquisition step means that models can produce plausible text while remaining weakly grounded and vulnerable to hallucination.

The authors introduce EASE (Evidence-Anchored Spatial Attention), a framework designed to augment multimodal RLVR with visual-evidence process supervision. EASE converts annotated evidence regions into a smoothed visual-token target and uses it to guide response-to-image attention during RL training.

Crucial design choices regarding its application include:

  1. Training vs. Inference: The annotations are used solely as privileged training labels, while the model's inference phase remains unchanged, requiring only the original image and question.

  2. Focus on Success: The guidance is applied only to high-reward trajectories that receive a high verifier reward, ensuring that low-reward responses do can have unreliable attention patterns.

EASE leverages visual evidence by mapping annotated regions to a soft target distribution over the model's visual tokens.

1. Evidence Annotation Pipeline:

The process involves several steps:

  • Step 1 (Extraction): GPT-4.1-mini is prompted to identify the smallest set of visible and localizable objects or regions that a human would inspect to verify the answer.

  • Step 2 (Localization): Grounding DINO is used to localize these phrases, and results from a locally deployed Qwen3.5-27B are merged if their Intersection over Union (IoU) is high.

  • Step 3 (Validation): Gemini 2.5 Flash-Lite validates the proposed boxes against the image, ensuring that the box contains visible evidence matching [the query] and directly helps verify a.

The resulting annotated pool is split into single-evidence examples (for local grounding) and multi-evidence examples (for cross-region reasoning).

2. Evidence Target Construction:

  • ** Single Box:** A Gaussian component is defined for each box B k = (x k, y k, w k, h k centered at mu k with diagonal covariance k).

  • ** Multi-Evidence:** For multiple supporting regions are represented by an equal-weight mixture of these Gaussian components.

  • ** Smoothing:** To prevent zero probability for background tokens, the target is smoothed: P target(v) = (1 - alpha)P evid(v) + alpha U(v), where alpha controls background smoothing.

3. Response-to-Vision Attention:

EASE measures how response tokens attend to visual tokens during the actor update. It extracts attention from a semantic middle layer, specifically* = 2L/3. The model's conditional spatial attention is averaged over H heads and t valid response tokens, yielding P model(v).

4. Reward-Aware RL Objective:

The total objective is defined as:

L total = L RL + lambda attn L EASE

Where L EASE is the attention loss, defined by the KL divergence between the model's attention and the evidence-anchored target: L g,t = D KL(P target P model). This loss penalizes visual-token normalized attention assigned outside the evidence-anchored target. This auxiliary loss is gated by a threshold tau, meaning it is only applied to trajectories where the raw verifier reward r(g) tau.

EASE was tested across three VLM backbones: Qwen2.5-VL-7B, Qwen3-VL4B, and Qwen3-VL8B.

  • Performance: EASE consistently improves outcome-only RL baselines compared to the strong baseline DAPO. The gains are significant: on Qwen2.5-VL-7B, EASE raises the average score from 70.5 to 73.4; on Qwen3-VL4B, it goes from 75.7 to 78.8; and on Qwen3-VL8B, it improves from 78.4 to 80.9 (Table 1).

  • Benchmark Gains: EASE showed gains across perception-heavy reasoning, hallucination robustness, visual math, and logic benchmarks.

  • Alignment Diagnostics: The auxiliary objective successfully improved the model's spatial alignment with evidence. Metrics like Evidence Attention Mass (EAM), Pointing Accuracy (PoA), and Multi-evidence Coverage demonstrated that EASE encouraged attention to align with annotated evidence regions rather than just improving final answer accuracy.

The authors conclude that EASE successfully provides a practical complement to final-answer rewards by strengthening the model's measured acquisition of visual evidence.

The limitations acknowledged are:

  1. EASE relies on high-quality annotations during training, which introduces data construction costs.

  2. It is most suitable for tasks where supporting evidence can be localized to one or more image regions; it is less suitable for holistic, stylistic, affective, or globally distributed cues.

  3. The method supervises where attention is allocated but does not directly optimize the degree of visual reliance on the input.

Improvements for AI systems

Based on the provided research, here are the specific improvements that can be made to existing AI systems and what those improved systems can achieve.


Improvement: Integrate EASE as an auxiliary process supervision term into the Reinforcement Learning with Verifiable Rewards (RLVR) training pipeline, moving beyond reliance on outcome-only rewards. The system will now incorporate a set of pre-annotated evidence regions (B) derived from the QA corpus into the training signal.

Specific Mechanism:

  • Target Generation: For each annotated evidence box B, generate a soft visual-token target (P target) by parameterizing it as a Gaussian mixture distribution (Section 3.3). This replaces the hard binary mask approach with a continuous, probabilistic target that allows for tolerance to noise.

  • Layer Selection: The system will extract and regularize response-to-vision attention specifically from the semantic middle layer,* = 2L/3, where L is the total transformer layers. This avoids both shallow representations and output-adjacent lexical dominance (Section 3.4).

  • Attention Regularization: The system will minimize the KL divergence between the model’s internal response-to-vision attention distribution (P model) and this evidence target (P target), as defined by L g,t = D(P model P target).

By implementing these changes, the improved AI system gains several critical capabilities:

  1. Enhanced Visual Grounding: The system will consistently align its generated textual response with the specific image regions that justify its claims, making it significantly harder for the model to generate plausible text based on language priors or dataset bias (addressing the perception credit assignment problem).

  2. Increased Hallucination Robustness: By forcing alignment between internal attention and verifiable evidence, the system is highly resistant to hallucinations derived from visual illusions or lack of direct evidence within the image.

  3. Superior Multimodal Reasoning: The system will demonstrate improved performance in tasks requiring complex reasoning (e.g., MathVerse-V, MMK12), as it is explicitly trained to cover multiple supporting regions simultaneously rather than focusing on a single salient cue.

  4. Verifiable Process Trace: During the RL training phase, the system provides quantifiable evidence that it has acquired the necessary visual information before generating an answer, allowing for process-aware debugging and a higher degree of confidence in its final output.

Abstract

Reinforcement learning with verifiable rewards (RLVR) improves vision-language models (VLMs) by optimizing outcome rewards derived from final answers. However, such outcome-only rewards do not tell the model which image regions justify an answer. For questions that require visual grounding, these rewards cannot distinguish responses supported by relevant visual evidence from those produced by language-prior shortcuts or lucky guesses. We introduce EASE (Evidence-Anchored Spatial Attention), which augments multimodal RLVR with visual-evidence process supervision. EASE converts annotated evidence regions into a smoothed visual-token target and uses it to guide response-to-image attention during RL training, but only on high-reward trajectories. The annotations are used solely as privileged training labels, while inference requires only the original image and question. Across Qwen2.5-VL-7B, Qwen3-VL-4B, and Qwen3-VL-8B, EASE raises average scores over DAPO by 2.5 to 3.1 points on perception, hallucination, visual math, and multimodal reasoning benchmarks. Diagnostics and ablations show that EASE better aligns visual attention with annotated evidence regions.

Sources

Related papers