EviRover: Reinforcing Agentic Perception Beyond a Glance

arXiv:2609.40230 · cs.CV, cs.AI · Submitted 2026-09-30 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "EviRover: Reinforcing Agentic Perception Beyond a Glance".

Tom: Visual perception is conventionally formulated as a one-shot prediction from a single glance at an image, an assumption that often fails in real-world scenarios requiring fine-grained details or external knowledge.

Jane: First, who's behind it and why it matters.

Title and authors: Jane: To summarize, "EviRover" reformulates perception under insufficient evidence as an active process rather than a direct prediction from the first look. It acknowledges that when fine-grained details or external knowledge are needed, a single glance simply isn't enough to determine the final output.

Lu: They treat perception as this stricter test of evidence seeking, where the agent’s actual goal becomes acquiring and correctly linking the missing evidence to the visual content. This fundamentally changes how we view visual understanding in AI systems.

Meng: So, if we think about a complex scene, instead of the AI freezing on an ambiguous area, it’s trained to initiate a search sequence—maybe it tries cropping first, then maybe it queries an external search engine if that fails. That’s much more dynamic than just one fixed prediction path.

Lalam: The summary emphasizes that this method requires the model to translate perceptual needs into a series of actionable steps—a location, a boundary, or a count—which demands fine-grained visual understanding. It’s about operationalizing the need for more information.

Tom: So, when we look at the results presented in "EviRover: Reinforcing Agentic Perception Beyond a Glance," they show that this agent successfully performs complex tasks like finding details occupying less than zero point one percent of an image area. That level of precision is really impressive when compared to standard approaches.

Jane: And the authors demonstrate that by training this agent using these specialized datasets, EviRover improves over its backbone model by thirty points on the EviLens benchmark across five different perception categories.

Lu: That comparison shows that even a relatively smaller backbone model can get close to performance levels comparable with advanced proprietary models when it learns this new way of searching.

Meng: I'm thinking about the practical side here, Lu. If we can reliably handle those tiny targets or complex group photo queries, that means we could deploy vision systems in fields that currently require expert human intervention for fine detail spotting.

Lalam: That capability to handle both internal visual ambiguity and external knowledge synthesis makes the resulting AI much more versatile for real-world deployment across various domains.

The paper's summary: Tom: The paper suggests several key improvements to the perception pipeline itself, focusing on enabling the agent to search effectively beyond a single glance. This involves learning when and how to search for missing evidence before producing an output.

Jane: One major improvement is moving from direct prediction to an evidence-seeking process, which means the system actively decides if it needs more information and selects the right tool—whether that’s a fine-grained visual inspection or reaching out to external knowledge sources.

Lu: They also detail a comprehensive set of tools the agent uses, including four for external evidence acquisition like text search, and three specifically for fine-grained visual inspection tasks such as cropping or verifying a specific part.

Meng: Having that explicit set of tools is crucial because it gives the model defined actions to take when it hits a wall, instead of just crashing or hallucinating an answer based on incomplete input. It makes the agent's behavior much more controllable.

Lalam: The training method itself is also an improvement; they use Agentic Reinforcement Learning to teach the model precisely when and how to search to push perception beyond that initial single glance. This teaches a sophisticated policy for information gathering.

Tom: And the results confirm this is effective, showing that EviRover substantially improves over its backbone by thirty points on EviLens across all five categories, which is a solid lift for any system trying to achieve high accuracy.

Jane: Furthermore, they showed that this new skill generalizes well, with a fifteen-point gain on the BrowseComp-VL benchmark, meaning the capability learned under insufficient evidence transfers to other types of multimodal reasoning tasks.

Lu: That generalization is what excites me; it suggests we aren't just training a specialized detector, but developing a more fundamentally intelligent agent capable of handling novel visual ambiguities.

The paper's improvements: Tom: So, to wrap up the discussion on "EviRover: Reinforcing Agentic Perception Beyond a Glance," we’ve seen how this agent moves perception from a single guess to an iterative, evidence-seeking process. It uses specific tools and specialized training to handle things that are too small or require outside context.

Jane: Essentially, the paper shows that when we frame perception as acquiring and relating missing evidence, the AI gets much better at tasks requiring fine-grained visual understanding. This method yields significant performance gains across benchmarks like EviLens.

Meng: From an engineering standpoint, the development of these two data pipelines and the subsequent RL training shows a reliable path for building more robust agents that don't just rely on the initial input quality but actively compensate for its limitations. It’s about designing systems that are resilient to real-world data imperfections.

Lalam: My analysis is that this work contributes a powerful mechanism to the underlying AI architecture, teaching it a sophisticated policy for information gathering when faced with ambiguity, which should improve the general reliability of vision models. It gives our models a better internal way to handle uncertainty.

Lu: I think the most exciting implication is that this approach validates agentic perception as a viable path for tackling complex visual reasoning tasks that are currently stalled by the single-glance assumption. It opens up possibilities we hadn't fully mapped out yet.

Tom: That’s the big picture, Lu. We’re looking at a paper called "EviRover: Reinforcing Agentic Perception Beyond a Glance," and it clearly demonstrates that training an AI to search for evidence is a very effective way to enhance its perceptual abilities.

Jane: It’s clear that this research has some serious implications for how we design vision systems that need to be accurate in detail and handle ambiguity effectively. We've had a great discussion about how this moves beyond simple prediction into actual investigation.

Meng: It’s definitely something to keep an eye on as we look at how these agentic search strategies can be integrated into practical applications where precision is critical. We need to see how this translates into production systems soon.

Lalam: This paper, "EviRover: Reinforcing Agentic Perception Beyond a Glance," shows that enhancing the model’s ability to search for evidence is a strong way to build more reliable and versatile vision models.

Lu: It’s an interesting direction for research, showing how structured agentic behavior can lead to tangible improvements in complex visual reasoning capabilities. We should definitely follow the work on this path.

Conclusion: Tom: So we’ve covered how EviRover shifts perception from a single prediction to an active evidence-seeking process using tools for both external search and fine visual inspection, right?

Jane: That's right, Tom, it really makes sense when you think about how we teach AI to handle things that aren't immediately obvious in one look.

Lu: I think the way they formalized perception under insufficient evidence as a stricter test of evidence seeking is really interesting from a theoretical standpoint; it frames ambiguity as a solvable problem requiring targeted acquisition.

Meng: From an engineering standpoint, having that structured toolset and the staged training process—SFT followed by Agentic RL—shows how you can actually build reliable systems that know when to stop looking and when to keep searching, which is something we need for real deployment.

Lalam: I think the biggest cultural implication here is how it redefines what we expect from visual intelligence; instead of just a predictor, we start expecting an AI that knows how to investigate and ask for more information when it doesn't have enough data.

Tom: Exactly! The performance gains they showed on EviLens, especially in localization and recognition, really prove that this evidence-seeking approach leads to tangible improvements over previous methods.

Jane: It’s impressive how they managed to get the model performing so well just by teaching it *how* to search instead of just teaching it *what* to predict initially.

Lu: The generalization results on BrowseComp-VL are what really suggest this isn't just a fix for one dataset; it points toward a more fundamentally capable reasoning mechanism that applies broadly across different modalities.

Meng: I see the practical impact in areas where precision is key, like medical imaging or detailed inspection tasks, where getting those small details right is essential for accurate diagnosis or manufacturing.

Lalam: For me, this advances the culture because it shows we can build AI that doesn't just parrot information; it can actively probe for truth when the initial signal is weak.

Tom: Absolutely! So to wrap up, EviRover proves that training an AI to search for evidence is a very effective way to enhance its perceptual abilities and general reasoning skills.

Jane: It’s a solid piece of work, Tom; understanding how the AI decides when it needs more data is a really valuable concept for us all.

Lu: And I think this paper sets a new precedent for how we approach ambiguous visual tasks by focusing on the process of evidence acquisition rather than just the final output.

Meng: For me, it’s about moving past models that just guess and toward systems that can intelligently plan their next move when they hit a wall.

Lalam: This EviRover work really solidifies the idea that structured agentic behavior is a powerful way to build more robust and versatile vision intelligence for everyone.

Kaixuan Fan, Kaituo Feng, Tianshuo Peng, Yilei Jiang

MMLab, The Chinese University of Hong Kong · Fudan University

cs.CV, cs.AI

Submitted: 2026-09-30

Updated: 2026-09-30

Code: https://github.com/kxfan2002/evirover

Importance score: 82/100

The gist: Visual perception is conventionally formulated as a one-shot prediction from a single glance at an image, an assumption that often fails in real-world scenarios requiring fine-grained details or

Key concepts

Evidence-seeking process
This redefines perception as a search task. Instead of just predicting from one look, the agent must determine if more evidence is needed and then actively search through finer visual details or external knowledge sources to resolve the task. It treats perception as a test of acquiring and relating missing information.
Data Generation Pipelines
Since real data for this setting was scarce, two pipelines were created. One collects high-resolution images of very small targets that need fine inspection. The second generates external evidence by verifying identities or synthesizing group photos using image generation models followed by quality filtering.
Agentic Reinforcement Learning (RL)
The RL stage trains the model to learn the strategy of searching. It uses a set of ten tools, including external search and fine-grained visual inspection tools, to decide when and how to use them. The goal is to optimize the agent's policy for finding missing evidence before making a final perception output.
EviLens Benchmark
This human-verified benchmark tests perception across five categories: localization, recognition, spot-the-difference, segmentation, and counting. It uses specific metrics tailored to each task to measure performance rigorously. EviRover achieved top results in localization and recognition on this challenging set.

Terminology

Summary

Visual perception is conventionally formulated as a one-shot prediction from a single glance at an image, an assumption that often fails in real-world scenarios requiring fine-grained details or external knowledge. This paper introduces EviRover, the first agent explicitly trained to resolve perceptual queries through interaction by learning when and how to search for missing evidence beyond a single glance.

The gist

EviRover is the first perception agent explicitly trained to search both within and beyond an image to resolve perceptual queries, moving beyond a single glance by deciding whether further evidence is needed and which action can provide it before producing a location, a boundary, or a count.

Perception Under Insufficient Evidence

The paper reformulates perception under insufficient evidence as an evidence-seeking process rather than direct prediction from the initial observation. This formulation acknowledges that in settings where fine-grained visual details or knowledge beyond the image are required, a single glance is insufficient to determine the output. The model must then search through finergrained visual inspection or external knowledge sources to resolve the perceptual task. Perception is treated as a stricter test of evidence seeking, where the agent's objective is not just prediction but acquiring and correctly relating missing evidence to the visual content.

Data Generation Pipelines

To address the absence of data for this setting, two dedicated data generation pipelines were designed:

  1. For evidence present in the image but not discernible at a glance, they collect high-resolution images containing targets so small that they cannot be resolved without inspecting the image at a finer scale, such as I-spy and spot-the-difference puzzles.

  2. For evidence beyond the image, they construct queries over group photographs of public figures and anime characters through two procedures: (1) starting from group photographs and verifying identities, or (2) starting from a reference image of a known individual or character and synthesizing a group photograph using GPT-Image-2 followed by filtering with Seed-2.0-Pro for visual quality and identity preservation.

EviRover Architecture and Training

EviRover is trained in two stages: Supervised Fine-Tuning (SFT) followed by Agentic Reinforcement Learning (RL). The SFT stage initializes tool use and task-specific prediction on EviRover-SFT-5K. The RL stage optimizes the model using EviRover-RL-12K, teaching it to learn when and how to search to push perception beyond a single glance. The agent operates with a set of ten tools, including four for external evidence acquisition (text search, image search, etc.) and three for fine-grained visual inspection (cropping, verify part), alongside task-specific prediction tools like SAM3-based mask generation.

Performance and Generalization

Experiments demonstrate that EviRover substantially improves over its backbone by 30 points on EviLens across the five categories, reaching performance comparable to advanced proprietary models. These gains extend beyond EviLens to WebEyes and conventional perception benchmarks like ReasonSeg and RefCOCOg. Furthermore, the results show that learning to search under insufficient evidence generalizes well, with a 15-point gain on BrowseComp-VL (general multimodal benchmark). This suggests that learning to search under insufficient evidence not only enhances perception but also yields capabilities that generalize well beyond the training distribution.

Evaluation Benchmark: EviLens

EviLens is a human-verified benchmark comprising 688 instances across five categories: localization, recognition, spot-the-difference, segmentation, and counting. The evaluation metrics are tailored to each task: mean box IoU and R@0.5 for localization/recognition; gIoU and cIoU for segmentation; exact-match accuracy for counting; and micro/macro F1 under greedy one-to-one matching for spot-the-difference. EviRover achieves the best localization and recognition results of all evaluated models on EviLens, with a clear margin on localization (0.444 IoU versus 0.373 for Gemini-3.5-Flash).

Agentic Reinforcement Learning Details

The RL stage uses EMA-GRPO to optimize the SFT model, utilizing task-specific outcome rewards in [0, 1]. For segmentation, the reward depends on IoU between predicted and ground truth masks generated by SAM3. For spot-the-difference, a soft set-level reward is used where candidate pairs are weighted by their overlap: STP = X (i,j)∈M IoU(ˆpi, gj) α, with the threshold τ deliberately set below 0.5 to allow for partial credit and a dense learning signal during training. The policy loss aggregation uses sequence-mean-token-mean reduction to avoid length bias induced by token-mean reduction.

Tool Use Policy

The agent operates under a strict system prompt requiring each assistant turn to be exactly one of three formats: (1)

Improvements for AI systems

Based on the provided scientific paper, here are specific improvements for AI systems derived from EviRover, categorized by capability:


)AI System Improvements Derived from EviRover

The core improvement lies in shifting perception from a single-glance prediction to an evidence-seeking process. This allows AI agents to handle complex queries where the answer requires knowledge acquisition or fine-grained visual inspection beyond what is immediately visible.

Here are the specific improvements and capabilities:

  1. Agentic Evidence Acquisition and Search Strategy:

AI systems can now be trained to dynamically decide whether information is present in the image or if external search/inspection is required.

  • The AI can perform a multi-step reasoning process: identify what evidence is missing (e.g., fine-grained detail, external knowledge) and then select the appropriate tool from its toolkit (e.g., cropping, web search, reverse image search).

  • Unlike previous agents that only gather evidence for a textual answer, EviRover treats the perceptual output itself as the objective of evidence seeking (location/boundary/count).

  1. Handling Perception Under Insufficient Evidence:

AI systems can reliably resolve perception tasks where visual information is ambiguous or incomplete.

  • The AI can handle perception under insufficient evidence, such as locating extremely small targets (<0.1% image area) or identifying individuals based on indirect references (recognition).

  • The system moves beyond simple detection to perform complex reasoning, such as navigating multi-hop queries (e.g., Find the lead actor in the film that won Award C) by iteratively replacing entities with indirect descriptions and performing targeted searches.

  1. Fine-Grained Visual Inspection and Verification:

AI systems can perform detailed visual analysis necessary for high-precision tasks.

  • The AI can use tools like 'crop' to zoom into specific regions when a target is hard to see or densely surrounded by clutter.

  • It can use 'verify part' to confirm if a candidate box tightly frames the target, ensuring high precision in localization and segmentation outputs.

  1. Advanced Spot-the-Difference Resolution:

AI systems can accurately detect subtle visual changes between two panels (left/right).

  • The system uses a specialized comparison tool ('compare lr') to side-by-side view regions, enabling it to identify subtle differences like color variations or missing elements that require cross-panel comparison.
  1. Generalization and Robustness Across Benchmarks:

AI systems trained with this evidence-seeking approach exhibit superior performance across diverse tasks and benchmarks.

  • The improved perception capability transfers from specialized perceptual datasets (EviLens) to conventional vision benchmarks (ReasonSeg, RefCOCOg).

  • The learned behavior generalizes to broader multimodal reasoning tasks like MMMU, MathVerse, and agentic visual search (BrowseComp-VL), indicating a fundamental enhancement in the model's general intelligence.

)What the Improved AI System Can Do (Specific Examples)

An AI system improved by EviRover can perform the following:

  1. High-Precision Object Localization:

Instead of failing when a target is tiny, it will use 'crop' and 'verify part' sequentially to isolate and precisely locate objects occupying less than 0.1% of the image area, achieving high IoU scores even in cluttered scenes (e.g., finding a specific small logo on a complex blueprint).

  1. Knowledge-Intensive Entity Resolution:

It can answer queries requiring external knowledge, such as identifying an actor based on a vague description or tracing relationships across multiple entities (e.g., Find the director of the film starring Actor A who won the Oscar in 2025). This is achieved through iterative, multi-hop reasoning combined with web search.

  1. Subtle Difference Detection:

It can reliably spot minute visual discrepancies between two images presented side-by-side, even if those differences are subtle (e.g., identifying a misplaced artifact or a minor color shift in a technical diagram).

  1. Complex Counting and Segmentation:

It can perform accurate counting tasks based on shared attributes (e.g., Count all people wearing blue hats) and generate highly precise segmentation masks using SAM3, even when the target is partially occluded, by leveraging targeted visual inspection tools.

  1. Adaptive Search Behavior:

When faced with an unknown query, the system won't guess; it will first perform internal visual inspection ('crop') to reduce uncertainty. If that fails, it will intelligently decide whether to use 'text search' for external facts or 'image search' to find reference images before attempting a final localization or count.

Sources

Related papers