Beacon: Knowing When and How to Perform Agentic Visual Reasoning

arXiv:2607.28595 · cs.CV · Submitted 2026-07-30 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Beacon: Knowing When and How to Perform Agentic Visual Reasoning".

Jane: Beacon introduces a novel agentic visual reasoning model designed to improve performance on complex tasks by focusing on two critical dimensions of tool use: Mode Adaptiveness and Tool Effect.

Tom: First, who's behind it and why it matters.

Title and authors: Tom: So, as we look at the title and authors of "Beacon: Knowing When and How to Perform Agentic Visual Reasoning," it immediately tells us this work is concerned with the decision-making aspect of agentic systems. It’s not just about having a tool available; it’s about knowing precisely when that tool is necessary for a given visual task.

Jane: Exactly, Tom; and the team behind it—Qixun Wang, Yang Shi, Letian Cheng, Zhuoran Zhang, Yan He, Yuqi Tang, Qi Zhang, Xinlei Yu, Ruizhe Chen—they bring a strong group of expertise to this kind of multimodal research. It shows a collaborative effort from different institutions working on visual reasoning challenges.

Lu: The authors are clearly tackling the core limitation we see in many current models: the tendency toward redundant or unhelpful tool calls, which the preprint summary points out as an issue with existing agentic visual reasoning models.

Meng: I noticed they’re testing this across thirteen diverse benchmarks, which suggests they aren't just looking at one specific type of problem, but trying to prove this adaptiveness holds up broadly.

Lalam: It’s exciting that the authors are systematically analyzing these properties across such a wide variety of tasks; it gives us a much clearer picture of where the current limitations lie in complex visual problem-solving.

The paper's summary: Tom: Okay, let’s talk about what they actually found in terms of their main findings regarding the summary of "Beacon: Knowing When and How to Perform Agentic Visual Reasoning." Essentially, the paper shows that existing models struggle with Mode Adaptiveness, meaning they don't always recognize when a tool is truly needed.

Jane: That’s right; and they also look at Tool Effect, which measures whether using the tool actually helps on hard problems while not causing trouble on easier ones. The paper demonstrates that current models often show limited Mode Adaptiveness, and the gains from using tools on difficult examples are frequently canceled out by errors made on problems that could be solved with just text.

Lu: What I find particularly interesting is their finding that existing agentic visual reasoning models exhibit limited Mode Adaptiveness, while the gains produced by tool use on hard examples are largely offset by the harm introduced on easy examples.

Meng: That’s a critical point for me; it means we have to be very careful how we integrate these tools so they don't introduce more problems than they solve on simpler instances.

Lalam: This analysis clearly shows that the authors didn't just focus on tool frequency, but on whether the tool use genuinely extends the model’s capability beyond what text alone can manage; that distinction is key to understanding this work.

The paper's improvements: Tom: Moving on to the actual proposed improvements in "Beacon: Knowing When and How to Perform Agentic Visual Reasoning," the authors introduce a new RL framework built around two main mechanisms to fix these shortcomings. They propose a supervised fine-tuning pipeline alongside this novel reinforcement learning approach.

Jane: The core of their suggested improvement involves developing this high-quality SFT data synthesis pipeline, which is designed to give the model better fundamental code-use capabilities from the start. Then, they build on that with a specific RL framework to guide the model toward better decision-making during execution.

Lu: They propose Necessity-Aware Adaptive Reward, or NAAR, which prioritizes text-only reasoning when it’s enough and only gives a full reward for correct code-based solutions when no text is available. That’s a very direct way to encourage the model to be selective about its tool use.

Meng: And then there's Hint-Guided Capability Expansion, or HCE, which uses expert hints during the rollout process to help address those 'wasted' hard examples by recovering useful reasoning signals.

Lalam: These mechanisms sound really powerful because they directly target the core issues: encouraging necessity in NAAR and strengthening capability through HCE, which is a smart way to ensure the model learns how to use tools effectively for genuine extension.

Conclusion: Tom: So, wrapping up with the conclusion of "Beacon: Knowing When and How to Perform Agentic Visual Reasoning," the paper really emphasizes that we need to measure agentic visual reasoning by whether models know when tools are needed and how they use them to get real capability gains. It boils down to adaptive tool-invocation behavior rather than just how often they call a tool.

Jane: That’s the main message, Tom; the authors show that Beacon achieves better average performance across thirteen benchmarks and proves it has stronger tool-invocation adaptiveness and the largest gap between tool gain and harm, showing that its use effectively extends the model’s capabilities beyond text-only reasoning.

Lu: I think what they really highlighted is that these improvements stem from increasing accuracy in reasoning-mode selection, not just a simple switch to either pure text or code-assisted reasoning.

Meng: From an engineering standpoint, seeing that gap between tool gain and harm of plus three point one four percent is a strong metric because it means the tools are adding measurable value when they are used correctly in the right context.

Lalam: It’s super inspiring to see this level of structured analysis on tool use; it gives us a blueprint for training future AI to be genuinely resourceful rather than just reactive.

Tom: That really brings us full circle, and I think it sets a new standard for how we evaluate these complex agentic systems. We’ve seen how Beacon achieves adaptive tool-invocation behavior that is largely absent from existing agentic visual reasoning models, invoking tools more frequently on samples that are difficult to solve through tool-free reasoning.

Jane: Absolutely; the whole point of "Beacon: Knowing When and How to Perform Agentic Visual Reasoning" is showing us that our goal shouldn't just be high tool frequency, but rather using those tools in a way that produces actual capability gains. It’s a big step forward in understanding how to build more useful AI agents.

Lu: We should keep looking at these kinds of RL frameworks because they show how we can design reward systems that explicitly encourage the right kind of behavior for complex reasoning tasks.

Meng: I'm curious to see if this approach scales well when we move from visual tasks to, say, more abstract planning problems where the tool set is much larger.

Lalam: It’s exciting because it validates the idea that careful design in the training and reward stages can lead to models that are genuinely more capable of complex visual tasks.

Qixun Wang, Yang Shi, Letian Cheng Zhuoran Zhang, Yan He Yuqi Tang, Qi Zhang Xinlei Yu Ruizhe Chen Tianrun Xu Yuanxing Zhang Pengfei Wan Haotian Wang Xianghua Ying

Peking University · Kling Team HKUST(GZ) · CUHK ZJU · THU

cs.CV

Submitted: 2026-07-30

Updated: 2026-09-29

Importance score: 91/100

The gist: Beacon introduces a novel agentic visual reasoning model designed to improve performance on complex tasks by focusing on two critical dimensions of tool use: Mode Adaptiveness and Tool Effect.

Key concepts

Mode Adaptiveness
This refers to a model's ability to recognize precisely when a specific tool is truly needed for a visual task. Existing models often lack this, leading to redundant or unhelpful tool calls.
Tool Effect
This measures whether using a tool actually helps solve hard problems without causing errors on easier ones. The paper shows current models often have gains from tools on hard examples canceled out by errors on simpler ones.
Necessity-Aware Adaptive Reward (NAAR)
This proposed reward mechanism prioritizes text-only reasoning when sufficient and only rewards full code solutions when no text is available, encouraging selective tool use.
Hint-Guided Capability Expansion (HCE)
This framework uses expert hints during the rollout process to help recover useful reasoning signals, addressing 'wasted' hard examples by strengthening capability.

Terminology

Summary

Beacon introduces a novel agentic visual reasoning model designed to improve performance on complex tasks by focusing on two critical dimensions of tool use: Mode Adaptiveness and Tool Effect. This research addresses the limitations of existing models, which often exhibit redundant or unhelpful tool calls, by proposing a reinforcement learning framework that encourages adaptive invocation based on task necessity and strengthens the model's capability to genuinely extend its reasoning beyond text-only methods. The paper systematically analyzes these properties across 13 diverse benchmarks, demonstrating that Beacon achieves superior overall performance by achieving stronger Mode Adaptiveness and the largest gap between tool-induced gain and harm compared to prior agentic visual reasoning models.

Mode Adaptiveness and Tool Effect Analysis

The core contribution of this work is the comprehensive analysis of two key dimensions: Mode Adaptiveness (MA), which characterizes whether an MLLM can recognize when tools are truly necessary, and Tool Effect (TE), which characterizes the actual impact of tool use—whether it extends capabilities on hard problems while avoiding errors on easy ones. The authors conduct a "comprehensive analysis to quantify these two properties and empirically reveal that existing agentic visual reasoning models exhibit limited MA, while the gains produced by tool use on hard examples are largely offset by the harm introduced on easy examples. Beacon is proposed as a solution, achieving stronger overall performance, improved Mode Adaptiveness, and genuine tool-induced performance gains" through its core mechanisms.

Beacon's Training Paradigm

Beacon employs an SFT-then-RL training paradigm to achieve its goals. During the Supervised Fine-Tuning (SFT) stage, the model is equipped with fundamental code-use capabilities using a highquality data synthesis pipeline that equips the model with fundamental code-use capabilities. This pipeline involves three stages:

  1. Sampling five responses from the base model and retaining those answered correctly at most twice.

  2. Using a powerful model (Gemini 3.1 Pro) to generate code-assisted reasoning trajectories for these hard examples, retaining only those that yield correct answers.

  3. Introducing an additional refinement procedure using Gemini 3.1 Pro to remove redundant or repetitive tool calls and ensure insufficient reliance on tool outputs are eliminated.

Reinforcement Learning Framework

The Reinforcement Learning stage introduces two mechanisms to improve the model's reasoning-mode adaptiveness and extend its capabilities:

  1. Necessity-Aware Adaptive Reward (NAAR): This reward prioritizes text-only reasoning when sufficient, assigning a reduced reward to correct code-based responses when text is present, while assigning a full reward to correct code-based solutions when no text is available, thereby encouraging the model to use tools only when necessary.

  2. Hint-Guided Capability Expansion (HCE): This mechanism addresses the issue of hard examples being 'wasted' by injecting expert-generated hints into the rollout process. These hints are constructed by having an expert model solve a hard example and then extracting the crucial reasoning steps (text or code-use steps) and the expected subgoal of each step to create an answer-free hint, which is then used to guide the policy exploration.

Evaluation and Results

Beacon was evaluated on 13 benchmarks spanning diverse visual reasoning tasks, including high-resolution visual search (V, HRBench), spatial and perceptual reasoning (RealWorldQA, BLINK), and compositional reasoning (VisualPuzzles). The results demonstrate that Beacon achieves the best average performance among all open-source models, ranking first on 11 of the 13 benchmarks. Furthermore, Beacon exhibits stronger tool-invocation adaptiveness and the largest gap between tool-induced gain and harm, with a Tool Effect delta of +3.14%. Ablation studies confirm that both components contribute to performance: "The Necessity-Aware Adaptive Reward primarily improves Mode Adaptiveness by encouraging the model to select the reasoning mode according to task necessity, while Hint-Guided Capability Expansion strengthens Tool Effect by recovering useful learning signals from all-wrong rollout groups."

Conclusion and Takeaways

The study concludes that agentic visual reasoning should be measured not merely by tool frequency but by whether models know when tools are necessary and how to use them to produce genuine capability gains. The key takeaways include:

  1. Beacon exhibits adaptive tool-invocation behavior, invoking tools more frequently on samples that are difficult to solve through tool-free reasoning.

  2. Beacon achieves the best average performance and the largest gap between tool gain and tool harm, indicating that its tool use more effectively extends the model’s capabilities beyond text-only reasoning.

  3. The training dynamics confirm these improvements arise from increasingly accurate reasoning-mode selection rather than a simple shift toward either pure-text or code-assisted reasoning.

Figure 1: Agentic visual reasoning models should use tools adaptively and effectively.

**Figure 2: (a) The tool-call ratio at different levels of text-only accuracy.

Improvements for AI systems

Here are specific improvements for existing AI systems based on the Beacon agentic visual reasoning model, categorized by capability:


)1. Enhanced Reasoning Mode Adaptiveness (MA) and Efficiency:

Beacon's core strength is its ability to distinguish between text-only solvability and tool necessity. AI systems should be improved by integrating a Necessity-Aware Adaptive Reward mechanism into their reinforcement learning (RL) phases, similar to Beacon’s NAAR.

  • Improvement: Implement a reward function that gives the highest credit for correct text-only answers when sufficient, and only assigns high rewards for tool use when text-only reasoning fails or is insufficient.

  • Capability: The system will drastically reduce unnecessary computational overhead (avoiding redundant tool calls) while ensuring that when tools are invoked, they are strictly necessary to solve hard problems. This leads to significantly more efficient agentic workflows.

)2. Genuine Tool Effect and Capability Extension (TE):

Beacon demonstrates a large gap between Tool-Gain and Tool-Harm, meaning its tool use genuinely extends the model's capabilities beyond what text alone can achieve.

  • Improvement: Develop RL frameworks that explicitly optimize for maximizing Tool Gain (successful problem solving via tools) while minimizing Tool Harm (performance degradation on easy problems). This requires sophisticated reward shaping and hint-guided rollout strategies to expose the model to high-value, tool-assisted trajectories.

  • Capability: The system will solve complex, multi-step visual reasoning tasks that are fundamentally unsolvable through text alone (e.g., detailed image manipulation or spatial analysis), leading to genuine performance gains on challenging benchmarks rather than just marginal improvements over text-only reasoning.

)3. Robust and Interpretable Tool Execution:

Beacon's training pipeline emphasizes structured trajectory formats and meticulous observation logging after every tool call, even when the result is not immediately obvious.

  • Improvement: Mandate a strict "tool-call -> tool-response -> observation" sequence for all agentic models. Furthermore, integrate a refinement stage (like Beacon’s Prompt Box 3) that forces the model to discard redundant calls and synthesize meaningful reasoning from tool outputs before generating the final answer.

  • Capability: The system will produce more reliable and verifiable results by explicitly grounding its final answer in concrete visual evidence obtained through code execution, making errors traceable and debugging significantly easier.

)4. Advanced Knowledge Recovery via Hint-Guided Exploration:

The Hint-Guided Capability Expansion (HCE) mechanism shows that conventional RL often fails to find hard solutions because it gets stuck in low-reward loops.

  • Improvement: Implement an HCE module that uses expert models to generate answer-free hints (subgoals and instructions). These hints should be injected into the rollout process, allowing the policy to explore successful reasoning pathways on difficult problems without relying on inference-time prompts.

  • Capability: The system will overcome local optima in complex reasoning tasks, enabling it to successfully tackle problems that are initially considered too hard or unsolvable by its current policy distribution, leading to superior performance on benchmarks like VisualPuzzles and TIRBench.

)5. Versatile Multimodal Tool Utilization:

Beacon is trained on a highly diverse dataset (Geometry3K, OlympiadBench, HRScene, etc.), exposing it to a wide variety of visual reasoning scenarios requiring different tools (cropping for counting, rotation for orientation, numeric calculation).

  • Improvement: Train models on broad datasets that explicitly cover the spectrum of required code actions (crop, draw line, draw box, numeric calculation, rotation) and their appropriate contexts.

  • Capability: The system will exhibit generalized tool proficiency; it won't just know how to crop an image but will know precisely when and how to use cropping versus calculating a derived property.

Sources

Related papers