FaithEyes: Towards Faithful Tool Use via Multi-Agent Process-Image Self-Verification
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "FaithEyes: Towards Faithful Tool Use via Multi-Agent Process-Image Self-Verification".
Jane: Agentic vision-language models (VLMs) have shown great potential for multimodal reasoning by interleaving textual reasoning with explicit tool calls for image manipulation, yet these models frequently use tools unfaithfully,
Tom: First, who's behind it and why it matters.
Title and authors: Tom: So, looking at the full name, FaithEyes: Towards Faithful Tool Use via Multi-Agent Process-Image Self-Verification tells us exactly what this paper is trying to do—it’s about using multiple agents to self-verify if the images the AI pulls out for its work are actually useful. It’s not just about getting a right answer; it's about making sure the process of getting that answer is sound.
Jane: Exactly, Tom. The authors are aiming to solve a problem where models might call tools like image cropping or manipulation without really looking at what those tools produce, which wastes computing power and can mislead the model if the tool returns junk. They are proposing a framework where the model checks its own work before moving on.
Lu: The way they frame it, using a multi-agent setup where the AI acts as both the main reasoning agent and a subagent judging its own output, seems like a clever way to enforce that self-correction during the process. It moves beyond just training the model to answer questions and into training it to be critically aware of its own actions.
Meng: That sounds complex; I need more detail on how this multi-agent structure actually prevents those unfaithful tool calls we discussed earlier, especially concerning the reward scaling aspect mentioned in their abstract. What’s the practical hurdle there?
Lalam: The idea of injecting an explicit judgment into the reasoning context—giving both a helpful and unhelpful result with a rationale—that feels like it gives us something tangible to work with during training, which is much better than just letting the reward system guess usefulness.
The paper's summary: Tom: To sum up what they’ve done, FaithEyes introduces a system where the model judges its own tool calls by injecting that judgment directly into the reasoning context so it can adjust its next steps immediately. It’s essentially giving the AI an internal editor that says, "Hey, this image you cropped isn't helpful for this question."
Jane: That self-verification is crucial because it tackles those two issues they pointed out: the tool reward doesn't distinguish useful calls from useless ones, and the feedback from the tool itself offers no signal about how good that call actually was. They solve this by scaling the reward based on a helpfulness ratio derived from that judgment.
Lu: The training pipeline is also key here; they use supervised fine-tuning first to get the model comfortable with writing code, then reinforcement learning with GRPO to teach it how to produce those faithful calls through that custom tool reward structure. It’s a very structured way to build this self-awareness from the ground up.
Meng: So, if I understand correctly, the main takeaway is that instead of just hoping the tool call was good, we now explicitly force the model to evaluate its own image processing step and use that evaluation to shape its future actions. That shifts the focus from just generating outputs to validating inputs before proceeding.
Lalam: It means we are moving away from a system where models might be relying on prior knowledge because they can't properly vet the evidence they retrieve through tools, which is a big step toward more reliable multimodal reasoning in general.
The paper's improvements: Tom: The biggest improvement they push is making tool use intentional by scaling that reward using a helpfulness ratio, which directly fights reward hacking where models just call tools without meaning any of it. It’s about rewarding the quality of the evidence retrieved, not just the quantity of tools used.
Jane: And they also improve reasoning itself by allowing the model to act on this explicit verdict—if an image is unhelpful, it can repair the crop or switch to a different approach instead of blindly following an ineffective tool output. This means their subsequent reasoning steps are guided by a real assessment of the visual data.
Lu: It’s interesting that they designed this self-judging framework without needing an external model at inference time; that keeps the evaluation consistent during testing, which is a major practical advantage over systems that might rely on a separate checker every single time.
Meng: From a deployment view, if we can ensure the AI only focuses computation on process images that are actually useful for answering the prompt, it’s going to make our inference pipelines much more efficient and cheaper in the long run. That reduction in wasted compute is a huge practical win for us right now.
Lalam: The overall improvement is that this approach leads to process images that genuinely capture the queried evidence, which means we get much more faithful results because the visual data driving the final decision is actually relevant to what the model was asked about.
Conclusion: Tom: So, wrapping up FaithEyes on "Faithful Tool Use via Multi-Agent Process-Image Self-Verification," we see a framework that injects self-judgment into tool observation and scales the reward based on helpfulness ratios to stop models from just making decorative calls. It really moves us toward more dependable multimodal reasoning systems.
Jane: Indeed, Tom. The implication is that by forcing the model to evaluate its own visual evidence before moving forward, we can build agentic workflows where intermediate steps are actually sound and purposeful, which is essential for building trustworthy AI applications.
Lu: I think the real excitement here is how this self-verification mechanism opens up new avenues for creating more nuanced agents that can handle complex visual queries by dynamically assessing the quality of their retrieved information.
Meng: For us in engineering, the practical implication is a clearer path toward building leaner models that don't waste processing cycles on irrelevant visual data, which translates directly into better performance on real-world tasks.
Lalam: I’m really optimistic about this direction; if we can successfully scale this self-judging concept, it could fundamentally improve how we design agent systems across the board by making them inherently more careful about their tool usage.
Haoqing Wang, Xingrun Xing, Wei Xia, Ziheng Li, Yehui Tang
Samsung Research · Peking University
cs.CV
Submitted: 2026-07-30
Updated: 2026-09-29
Code: https://github.com/Mosi-AI/FaithEyes
Importance score: 90/100
The gist: Agentic vision-language models (VLMs) have shown great potential for multimodal reasoning by interleaving textual reasoning with explicit tool calls for image manipulation, yet these models
Key concepts
- FaithEyes
- A framework that uses multiple agents to self-verify if the images an AI pulls out for its work are actually useful. It focuses on making the process of getting an answer sound rather than just achieving a correct final result.
- Unfaithful Tool Use
- When models call tools, such as image cropping, without properly looking at what those tools produce. This wastes computing power and can mislead the model if the tool returns poor or junk output.
- Self-Verification
- The process where an AI model checks its own work before moving on. FaithEyes achieves this by injecting an explicit judgment into the reasoning context, allowing the AI to act as an internal editor and adjust its next steps based on whether a retrieved image is helpful.
- Reward Scaling via Helpfulness Ratio
- A method used to improve tool use by scaling rewards based on a helpfulness ratio derived from self-judgment. This directly fights reward hacking by rewarding the quality of the evidence retrieved, not just the quantity of tools used.
Terminology
Summary
Agentic vision-language models (VLMs) have shown great potential for multimodal reasoning by interleaving textual reasoning with explicit tool calls for image manipulation, yet these models frequently use tools unfaithfully, often employing decorative or irrelevant process images that waste computation. This paper introduces FaithEyes, a multi-agent self-judging framework designed to ensure tool faithfulness by injecting explicit judgments into the reasoning context and scaling the tool reward based on a helpfulness ratio.
The Problem: Unfaithful Tool Use
Recent studies have revealed that agentic VLMs often use tools unfaithfully, meaning they operate on incorrect evidence. A prominent symptom is that only about half of samples actually contain at least one tool call cropping or revealing the queried target, while the rest return decorative or misaligned process images. This occurs because existing reward designs fail to distinguish useful from useless calls, and tool feedback carries no signal of usefulness, leading to reward hacking
where models invoke tools without meaningfully engaging with their outputs.
FaithEyes Framework: Multi-Agent Self-Judging
FaithEyes is a multi-agent self-judging framework where the model itself acts as a subagent to judge the main agent's tool calls. This design eliminates dependence on an external model at inference, ensuring train-test consistency.
The framework involves two key mechanisms:
-
Injecting Judgement into Context: The judgement is
injected into the reasoning context as part of the tool observation to help subsequent reasoning.
When a process image is judged as helpful, both the image and its rationale are returned; if unhelpful, only the judgment is returned. This allows the main agent toact on an explicit verdict and either repairs the crop or makes a justified fallback.
-
Scaling Tool Reward: The judgement result is used to
scale the tool reward by a helpful-tool ratio to suppress reward hacking.
The tool reward is computed as:
rtool = 1 − nfail + nunhelpful / ntool, if ntool > 0
Training Pipeline: SFT + RL with GRPO
The training utilizes a two-stage pipeline: Supervised Fine-Tuning (SFT) followed by Reinforcement Learning (RL). The SFT stage is crucial for cold-start initialization,
equipping the model with three essential capabilities:
-
Code-based problem solving ability.
-
Judgment—the ability to judge whether a process image is helpful and articulate the rationale behind its assessment.
-
Interactive reasoning—the ability to adapt subsequent reasoning based on the tool observation, including judgment and images.
The RL stage employs the Group Relative Policy Optimization (GRPO) algorithm, combining four reward types: accuracy reward, format reward, consistency reward, and tool reward. The tool reward is specifically designed to target faithfulness by scaling the bonus by the fraction of tools that are both executable and judged helpful.
Training Dynamics and Results
The training dynamics show that the tool reward (helpful-tool ratio) grows steadily and stabilizes at a high level, reflecting that the policy learns to output faithful tool calls rather than decorative ones.
The average number of tool calls gradually increases and converges to roughly one call per trajectory. Furthermore, the ablation studies confirm that both mechanisms are complementary: removing either judgement injection or reward scaling degrades performance, but removing both is weakest on all metrics.
Conclusion
FaithEyes successfully addresses the two coupled causes of unfaithful tool use by injecting an explicit usefulness signal into the tool observation to help reasoning and using that signal to scale the tool reward. Training with this framework attains competitive or superior accuracy across visual perception and reasoning benchmarks, while markedly improving tool faithfulness.
The research demonstrates that this approach produces process images which actually capture the queried evidence,
leading to a genuinely more faithful tool-use policy.
(Self-Correction Check: The summary is structured as requested, uses key phrases from the text, avoids meta commentary, and focuses only on the paper's content. The length is appropriate for a detailed abstract/summary.)
How it works
FaithEyes is a multi-agent self-judging framework where the model itself serves as a subagent to judge each process image it produces. This design eliminates dependence on an external model at inference, ensuring train-test consistency.
The framework involves two key mechanisms:
-
Injecting Judgement into Context: The judgement is
injected into the reasoning context as part of the tool observation to help subsequent reasoning.
When a process image is judged as helpful, both the image and its rationale are returned; if unhelpful, only the judgment is returned. This allows the main agent toact on an explicit verdict and either repairs the crop or makes a justified fallback.
-
Scaling Tool Reward: The judgement result is used to "scale the tool reward by a helpful-tool ratio to suppress reward hacking.
Improvements for AI systems
Here are the specific improvements that can be made to AI systems, based on the FaithEyes framework, and what those improved systems can achieve:
-
Improved Tool Faithfulness through Self-Judgement: The core improvement is moving from an unguided reward system to a multi-agent self-judging framework (FaithEyes).
-
Contextual Reasoning Enhancement: The model's reasoning process is explicitly guided by the judgment of whether a retrieved image is actually relevant. This prevents the model from relying on irrelevant or decorative process images that lead to
reward hacking.
-
Adaptive Tool Observation: Instead of simply appending raw tool outputs, the system dynamically incorporates a structured verdict (helpful/unhelpful + rationale) into the reasoning context. This allows subsequent reasoning steps to explicitly decide whether to correct an off-target tool call or fall back on prior knowledge.
-
Reward Alignment and Suppression of Hacking: The tool reward is scaled by a
helpful-tool ratio
derived from the subagent's judgment, rather than being granted simply for the presence of a tool call when an answer is correct. This directly incentivizes the model to produce tools that genuinely contribute evidence to the final answer. -
Robust Training Pipeline: Implementing a two-stage SFT + RL pipeline ensures cold-start initialization for code writing and judgment capabilities, followed by reinforcement learning (GRPO) that specifically targets faithfulness through the custom tool reward structure.
These improved AI systems can achieve the following:
-
Accurately answer visual questions by ensuring that every tool call is purposeful, targeting the precise region of interest in an image (e.g., accurately cropping to a specific object rather than an arbitrary area).
-
Significantly reduce computational waste and inference cost by discarding irrelevant or decorative process images during reasoning, focusing computation only on evidence-bearing outputs.
-
Develop more robust visual reasoning capabilities because the model learns to
interrogate
the image dynamically—it knows when a tool's output is a reliable clue versus noise—leading to higher accuracy on complex visual perception tasks (especially those involving small objects in high-resolution images). -
Produce more faithful and interpretable agentic workflows where intermediate visual evidence is explicitly verified by the model itself, rather than implicitly assumed.
Sources
- Qwen3-VL Technical Report
- AgentVerse: Facilitating Multi-Agent Collaboration and Exploring Emergent Behaviors
- Thinking with Generated Images
- A Survey on the Optimization of Large Language Model-based Agents
- ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding
- Advancing vision-language models in front-end development via data synthesis
- Large Language Model based Multi-Agents: A Survey of Progress and Challenges
- ControlThinker: Unveiling Latent Semantics for Controllable Image Generation through Visual Reasoning
- GPT-4o System Card
- Mini-o3: Scaling Up Reasoning Patterns and Interaction Turns for Visual Search
- LLaVA-OneVision: Easy Visual Task Transfer
- Encouraging Divergent Thinking in Large Language Models through Multi-Agent Debate
- Decoupled Weight Decay Regularization
- MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
- ChatDev: Communicative Agents for Software Development
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Thinking with Images for Multimodal Reasoning: Foundations, Methods, and Future Frontiers
- Kwai Keye-VL Technical Report
- Multi-Agent Collaboration Mechanisms: A Survey of LLMs
- From Illusion to Intention: Visual Rationale Learning for Vision-Language Reasoning
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models