A Low Grounding Score Is Not an Ungrounded Judge: Identifying the Perceptibility Confound in Multimodal Oversight
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "A Low Grounding Score Is Not an Ungrounded Judge".
Tom: Model judges now supervise multimodal systems at scale, filtering training data, selecting outputs, and supplying the reward that shapes multimodal reasoning models.
Jane: First, who's behind it and why it matters.
Paper summary: Tom: Welcome back everyone to the show! Today we're talking about this fascinating paper from arXiv titled "A Low Grounding Score Is Not an Ungrounded Judge: Identifying the Perceptibility Confound in Multimodal Oversight." Jane, you have a minute to set the stage for us on what this research is actually about.
Jane: Absolutely, Tom. So, essentially, this paper tackles a problem where we rely on these AI judges to oversee multimodal systems at scale—filtering data and shaping reasoning models. The core idea they're pushing is that we need to scrutinize whether a low Verdict Grounding Score means the judge is actually ungrounded or if there's something else going on.
Tom: That’s a big question, Jane. What exactly is the main claim of this paper regarding these grounding scores? I want to know what they are trying to prove about how these scores work in practice.
Jane: The central thesis is that the Verdict Grounding Score, or VGS, can't be read as a simple measure of whether a judge ignores an image or if it never processes an edit. They show it actually suffers from something they call a "perceptibility confound," where the score gets capped by how easily an edit is perceptible to the judge’s decision-relevant reading.
Lu: That's really interesting because if that cap exists, it means the VGS itself isn't giving us a clean signal about grounding. It suggests that we might be misinterpreting what a low score actually signifies in these systems.
Meng: From an engineering standpoint, if we can’t trust the score directly because of this confound, it changes how we design our auditing protocols for these large multimodal models. It means the initial metrics might be misleading us about the system's actual reliability.
Lalam: If we think about what this means for culture and how AI is trusted, it suggests that our current way of checking these systems might be flawed if we only look at one metric like VGS in isolation. We need a richer diagnostic approach to understand the judge’s behavior fully.
Tom: So, to put it simply, the paper argues that a low VGS isn't automatically a red flag for an ungrounded judge; instead, it can just be due to how perceptible an edit is and how that affects the score calculation. It opens up a whole new way of looking at these oversight mechanisms.
Paper summary: Jane: Exactly. The authors introduce the Grounded-Error Detection Rate, or GEDR, as a way to bypass this confusion by measuring read accuracy on an unedited image instead of relying solely on the counterfactual probe for grounding assessment.
Tom: That GEDR seems like the key differentiator here. So what's the main thrust of this research in terms of what it achieves with these new measures?
Jane: The paper formally establishes a relationship between GEDR and VGS, showing how they can be used together to isolate the causal contribution of visual evidence to a judge’s final verdict. This allows them to identify that missing quantity—the causal effect of image content on the verdict—and show it's more robust than it first appeared.
Lu: The way they formalize that causal effect by using interventions that alter the image while holding text fixed is powerful; it moves beyond just observing what a judge does to understanding *why* they do it. It’s like moving from watching a car drive to understanding the physics behind its motion.
Meng: I'm curious about how this theoretical identification of the causal contribution translates into something practical for deployment. If we can precisely measure that causal link, we can build much more targeted safety checks instead of just relying on broad confidence scores.
Lalam: For me, this implies a culture where auditing isn't just about flagging failures, but about deeply diagnosing the mechanism behind those failures so we can improve the underlying model behavior directly. That sounds like a very constructive way to approach AI development.
Tom: So, we’re moving from seeing a low score and guessing the problem to using these new metrics like VGS and GEDR together to get a concrete answer about whether an auditor is truly missing something important or if they're just running into this perceptibility issue.
Jane: That’s precisely what the paper demonstrates with Theorem seven which provides a checkable certificate of a false alarm based on the relationship between VGS and GEDR. It lets auditors know when their low score is just noise from an unedited image probe.
Paper summary: Tom: And the practical suggestion they make is to pair that counterfactual score with a detection probe on the unedited image to upper-bound it, which effectively certifies those false alarms without adding extra cost. That’s a very tangible piece of advice for anyone working in this space.
Jane: It moves the conversation from just reporting a number to adopting a diagnostic protocol that integrates multiple checks, which is really helpful for system reliability assessment. Now, let's move on to what this whole discussion suggests about the future implications of these findings.
Tom: Right, this paper by Rasul Khanbayov and Hasan Kurban really challenges how we interpret performance metrics in multimodal oversight systems. We’ve talked about the technical mechanisms behind VGS and GEDR, but what does it actually mean for the wider world of AI application?
Lu: I see massive potential here for creating more robust AI agents that operate in complex, real-world environments where visual understanding is critical. If we can better calibrate our trust signals, we can build systems that are safer because they have a clearer understanding of their own limitations regarding visual evidence.
Meng: Practically speaking, this means when we deploy these multimodal models in areas like autonomous systems or complex data analysis tools, our safety evaluations won't just be based on pass/fail rates; they’ll involve understanding the specific conditions under which the judge might fail due to perceptual biases. That level of specificity is what engineers need.
Lalam: I think this points toward a future where AI development isn't just about making models bigger, but about building better interpretability tools that help us understand exactly *how* they are reasoning visually, rather than just accepting the output at face value.
Tom: So, we’re talking about a shift in how we validate these systems—moving away from simple scoring toward a more nuanced diagnostic framework that accounts for perceptual factors. Jane, what are your thoughts on the broader impact of this specific finding?
Jane: The implication is that we need to stop treating a single grounding score as the absolute truth about an AI judge’s capability. Instead, we have to adopt a layered approach where scores are contextualized by detection probes on unedited images. This helps prevent unnecessary discarding of reliable overseers due to measurement noise.
Paper summary: Tom: That makes sense from a user perspective, too—if we can reduce false alarms caused by this confound, the systems that get approved for deployment will be more trustworthy because they won't be rejected unfairly by an overly sensitive metric.
Lu: The way they’ve identified the missing causal contribution using those assumptions about exchangeability and misreading is a strong theoretical foundation, and it opens up avenues for other researchers to apply similar causal identification techniques to different multimodal tasks.
Meng: From a practical impact angle, this suggests that future model development might need to incorporate these kinds of diagnostic checks directly into the training loop or fine-tuning process, rather than just as an external audit step after deployment.
Lalam: It really underscores the importance of transparency in AI architecture; we need to be able to see and measure these perceptual confounders so we can build more resilient and trustworthy AI systems that genuinely improve our capabilities.
Tom: So, to wrap up this deep dive into "A Low Grounding Score Is Not an Ungrounded Judge: Identifying the Perceptibility Confound in Multimodal Oversight," we’ve seen how VGS can be misleading due to perceptual factors, and how GEDR provides a necessary countermeasure. Jane, what's your final thought on the practical implications of this work?
Jane: My final thought is that adopting the suggested diagnostic protocol—pairing the counterfactual score with a detection probe—is a concrete, low-cost way to certify those false alarms we discussed. It’s a practical step toward making our oversight processes more accurate and less prone to misinterpretation of the data.
Tom: Indeed, it gives us a clear path forward for auditors and developers who want to get reliable insights into multimodal AI performance without getting stuck in the ambiguity of simple grounding scores. Lu, Meng, Lalam, your input has been incredibly valuable throughout this discussion on this paper.
Lu: I'm excited about the potential for applying these causal identification techniques to other complex reasoning tasks; it feels like a fundamental tool for understanding visual evidence in AI systems.
Meng: I'm looking forward to seeing how this framework translates into more rigorous testing procedures that we can actually implement on our platforms.
Lalam: This research gives us a much sharper lens for thinking about the trust layer in AI, helping us build systems that are not just functional but truly dependable in their visual reasoning.
Conclusion: Tom: So, we've been talking about how those low scores in multimodal oversight systems aren't always bad news, and now we get to look at the actual title and who put this research out there.
Jane: I think it helps to frame it as a direct challenge to how we interpret those numbers; the authors are directly addressing the confusion around what a low Grounding Score actually means.
Lu: The authors are really drilling down into that core issue, showing that the ambiguity they found in previous work wasn't just noise but a specific technical hurdle called the perceptibility confound.
Meng: From my side, I’m interested in seeing if this confusion translates to actual system behavior; is it just a theoretical fix or does it give us something tangible we can build on?
Lalam: I think this paper speaks to the culture of AI evaluation; if we can understand why a score is low, we can stop blindly trusting that score and start asking better questions about the model's visual understanding.
Tom: Exactly, Jane, so to recap, these authors are using a specific metric called the Verdict Grounding Score to show that it’s being tricked by how easily an edit is noticed before it actually changes the judge's mind.
Jane: That’s right; they introduce this concept of a perceptual confound and then prove there's a way to measure read accuracy on unedited images, which we call GEDR, to untangle that confusion.
Lu: The authors make some really solid assumptions about how clean and corrupted traces interact in the text channel, which is crucial for proving their identification of that missing causal contribution.
Meng: Those assumptions are what make it rigorous; they aren't just guessing what’s happening but building a mathematical bridge between the visual evidence and the final verdict.
Lalam: It shows that we need richer diagnostic protocols rather than relying on just one number to judge whether an AI is reliable in complex visual tasks.
Tom: So, if you take the title, "A Low Grounding Score Is Not an Ungrounded Judge," it really captures the essence of their argument that a low score doesn't automatically mean failure.
Jane: And their conclusion boils down to a practical rule: auditors should always pair that counterfactual score with a detection probe on the unedited image to check for those false alarms.
Lu: It’s fascinating how they connect these dots—from the theory about edit-miss rates to the experimental validation across multiple vision-language judges.
Meng: I think if we can implement this diagnostic protocol, it gives us a much better way to identify where our AI systems are actually failing in a real deployment scenario.
Lalam: This work really impacts how we design the trust layer; it shifts our focus from just observing outputs to understanding the underlying mechanisms of visual perception within the AI.
Tom: It definitely sets up a lot of interesting avenues for future research, especially for other multimodal systems where visual evidence is key to decision-making.
Jane: And it gives us a clear direction for how we should be looking at these metrics going forward, moving beyond simple pass or fail judgments.
Rasul Khanbayov, Hasan Kurban
College of Science and Engineering, Hamad Bin Khalifa University
cs.CV
Submitted: 2026-09-09
Updated: 2026-09-09
Code: https://github.com/KurbanIntelligenceLab/groundJudge
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 88/100
The gist: Model judges now supervise multimodal systems at scale, filtering training data, selecting outputs, and supplying the reward that shapes multimodal reasoning models.
Key concepts
- Verdict Grounding Score (VGS)
- This score measures whether a judge's final decision changes when an edit is made, compared to changes in a control. The core finding is that VGS cannot tell if the judge ignored the image or if the edit never reached its decision-relevant reading, creating a 'perceptibility confound'.
- Grounded-Error Detection Rate (GEDR)
- GEDR measures how accurately a judge reads an unedited image. It serves as a diagnostic tool to resolve the ambiguity of VGS. If GEDR is high enough, it helps exclude scenarios where low VGS is merely due to the judge being 'image-blind' or missing edits.
Terminology
Summary
Model judges now supervise multimodal systems at scale, filtering training data, selecting outputs, and supplying the reward that shapes multimodal reasoning models. The core finding is that a low Grounding Score (VGS) does not necessarily indicate an ungrounded judge; instead, it can be caused by a perceptibility confound
where the score is capped by how perceptible an edit is to the judge's decision-relevant reading, leading to false alarms when auditors discard usable overseers.
The Verdict Grounding Score (VGS) and its Confound
The paper introduces the Verdict Grounding Score (VGS) as a measure of grounding, defined as the probability that a judge's verdict changes under a truth-flipping edit, relative to the rate of change under a truth-preserving control. The core confound is identified in Figure 1(a), where VGS cannot distinguish between two scenarios: (i) the judge ignores the image excluded from the edit, or (ii) the edit never reached its decision-relevant reading. This ambiguity is resolved by introducing Grounded-Error Detection Rate (GEDR), which measures a judge's read accuracy on an unedited image. The relationship is formalized as: The edit-miss rate equals (GEDR − VGS) − Δnull.
Identification of the Missing Quantity
The paper proves that the missing quantity—the causal contribution of visual evidence to the verdict—is not merely bounded but exactly identified from measurements already collected by an audit protocol. This identification relies on four key assumptions:
-
(a) The clean and corrupted traces are exchangeable in the text channel, requiring image separation for distinction.
-
(b) Misreading the original does not land exactly on the injected value (Pr[β(X) = av] = 0).
-
The relationship between GEDR and VGS is established via Theorem 5:
GEDR = π1 and Δflip = π2, hence π1 − π2 = (GEDR − VGS) − Δnull.
-
The ordering holds under specific conditions:
VGS ≤ GEDR if and only if δflip ≤ η1 + Δnull.
Certification of False Alarms
The theoretical framework allows for a concrete, checkable certificate of a false alarm. Theorem 7 establishes that when editing does not raise the judge’s read accuracy (pi2 ≤ pi1), the shortfall is entirely accounted for by missed edits and oversensitivity: VGS ≤ GEDR − Δnull ≤ GEDR.
This means that if an auditor uses a low VGS score as grounds to discard a judge, they are certified wrong. The paper proposes a practical rule: pair the counterfactual score with a detection probe on the unedited image, which upper-bounds it and certifies its false alarms at no extra cost.
Experimental Validation and Findings
The researchers audited nine vision-language judges across four datasets (ChartQA, HallusionBench, MathVista, ScienceQA). The results show that the predicted ordering holds strictly across our entire primary pool,
with only a few exceptions. Crucially, the analysis reveals that "GEDR > 1/2 in 27 of 36 cells and in 15 of the primary ones, excluding most judges from being classified as image-blind. The study also quantified the error: on the primary pool,
the typical judge there acts on only about half of the edits whose attribute it can otherwise resolve."
Key Takeaways for Auditing
The paper concludes that an image-side counterfactual grounding score should never be reported alone. Instead, auditors must adopt a diagnostic protocol: "Score the clean trace against one carrying an injected visual error on the unedited image, report that rate beside the counterfactual score, and treat any cell where the counterfactual score is low while the detection rate clears the image-blind threshold as a false alarm rather than a finding." This approach transforms an unobserved confound into an estimand.
The Gist
A verdict can change under an edit only if the judge’s decision-relevant reading of the image changes, so VGS is capped by the rate at which that reading moves (Theorem 3). A judge whose reading the edit never reaches scores low, exactly like a judge that never looks, and VGS cannot separate them (Fig. 1a). GEDR can, because a GEDR above the image-blind rate excludes (i) and certifies false alarms. The rule that follows is that an image-side counterfactual score should never be reported alone: a detection probe on the unedited image upper-bounds it, certifies its false alarms, and costs nothing extra to run.
How it works
Improvements for AI systems
Based on the scientific paper, here are specific, actionable improvements for AI systems, categorized by where they fit in the machine learning pipeline:
The core finding is that a low Grounding Score
(VGS) does not reliably indicate that a model ignores the image; it can be caused by either (1) ignoring the image entirely or (2) failing to perceive an edit because it was never registered or was outweighed by textual priors. The paper identifies this as the Perceptibility Confound.
Here are the specific improvements:
- Release a Mandatory Dual-Probe Auditing Protocol for Multimodal Judges:
The system should not rely on a single counterfactual grounding score (VGS). Instead, an auditor must run two simultaneous probes on any low VGS result:
A. The Verdict Grounding Score (VGS): Measures the judge's response to a truth-flipping edit while holding the text fixed.
B. The Detection Probe (GEDR): Measures the judge's accuracy on the original, unedited image against an injected visual error (detection rate).
A system should only be flagged as image-blind
if it fails both probes simultaneously (i.e., GEDR is below a measured threshold, like 1/2, AND VGS is low).
- Implement a
False Alarm Certification
Rule for Oversight Systems:
Auditors must adopt the rule: A low VGS score should never be reported alone. It must be paired with the detection probe (GEDR) on the unedited image. If GEDR exceeds an image-blind threshold, the low VGS is not evidence of image ignorance; it is instead attributed to missed edits
or verdict instability.
- Develop a Diagnostic Tool for Identifying Judge Failure Modes:
The system should be designed to diagnose whether its failure is due to:
A. Image Blindness (Low GEDR): The judge ignores the visual input entirely.
B. Missed Edits (Low VGS, High GEDR): The judge perceives the image but fails to integrate the new information into its reasoning trace because of text-channel bias or prior knowledge.
- Improve Reward Model Robustness via
Artifact-Matched Control
:
When training or evaluating models that use reward models, include an artifact-matched control condition (using a content-preserving re-encoding, as opposed to a simple blank image) to measure intrinsic verdict instability (Δnull). This helps distinguish between judges that are simply sensitive to the presence of visual input versus those whose reasoning is fundamentally unstable under content change.
- Refine Model Training/Fine-tuning Strategies:
When fine-tuning multimodal models, incorporate read-attribute
training signals. Explicitly train the model to report its internal reading of specific visual attributes (e.g., the bar is blue
) and then test if that reported attribute value changes under counterfactual edits. This helps mitigate the text-channel bias where models favor fluent answers over visually correct ones, as noted in Section 2.1 and Table A15.
The improved AI system can:
-
Accurately distinguish between a judge that is fundamentally incapable of seeing (image-blind) and one that is merely failing to perceive subtle visual changes (missed edits).
-
Provide auditors with a
false alarm certificate
to prevent the premature discarding of usable oversight models when they exhibit low scores. -
Be more robust against text-channel biases by being trained to prioritize visual grounding over linguistic fluency when making decisions.
Abstract
Model judges now supervise multimodal systems at scale, filtering training data, selecting outputs, and supplying the reward that shapes multimodal reasoning models. Trusting one means first checking that it uses its evidence, and that check is itself worth scrutinizing, so we ask whether a counterfactual probe of visual grounding measures what it claims to. The probe edits the image so the ground truth flips, holds the reasoning trace fixed, and asks whether the verdict follows. We formalize it as the Verdict Grounding Score and show it cannot be read the way such scores are read. A verdict responds only to an edit that reaches the judge's decision-relevant reading, so the score is capped by how perceptible the edit is, and unless editing makes the attribute easier to read, the error is one-sided: the score can only make a judge look less grounded than it is. The practical failure is therefore a false alarm, an auditor discarding a usable overseer. Under assumptions we state, we show this missing quantity is not merely bounded but identified from three quantities the same audit protocol already collects, which makes the false-alarm rate directly measurable rather than merely a concern. Auditing nine judges, we find the predicted ordering holds strictly across our entire primary pool, and the typical judge there acts on only about half of the edits whose attribute it can otherwise resolve. Applying a conservative rejection threshold certifies several cells as false alarms outright, the clearest being a judge that detects the injected error essentially every time while still scoring as if it had not used the image at all. The rule that follows is that an image-side counterfactual score should never be reported alone: a detection probe on the unedited image upper-bounds it, certifies its false alarms, and costs nothing extra to run.
Sources
- MLLM-as-a-Judge: Assessing Multimodal LLM-as-a-Judge with Vision-Language Benchmark
- Multimodal RewardBench: Holistic Evaluation of Reward Models for Vision Language Models
- VisualPRM: An Effective Process Reward Model for Multimodal Reasoning
- Reward Under Attack: Analyzing the Robustness and Hackability of Process Reward Models
- VIABLE: A Visually Impaired Assistance Benchmark for VLM-as-a-Judge Evaluation
- Vision Language Models are Biased
- Arbitration Failure, Not Perceptual Blindness: How Vision-Language Models Resolve Visual-Linguistic Conflicts
- Are VLMs Seeing or Just Saying? Uncovering the Illusion of Visual Re-examination
- CounterCount: A Diagnostic Framework for Counting Bias in Vision Language Models
- Mitigating Perceptual Judgment Bias in Multimodal LLM-as-a-Judge via Perceptual Perturbation and Reward Modeling
- Discretizing Reward Models
- FigureQA: An Annotated Figure Dataset for Visual Reasoning
- Reward Auditor: Inference on Reward Modeling Suitability in Real-World Perturbed Scenarios
- VLM Judges Can Rank but Cannot Score: Task-Dependent Uncertainty in Multimodal Evaluation
- MJ1: Multimodal Judgment via Grounded Verification
- Ill-Posed by Design: Probing Evidence Use in VLMs
- Treble Counterfactual VLMs: A Causal Approach to Hallucination
- Do Medical Vision Language Models Actually See? A Counterfactual Grounding Framework and Hard-Negative Contrastive Training for Visually-Reliant Medical VLMs
- VRPRM: Process Reward Modeling via Visual Reasoning
- Grounding the Score: Explicit Visual Premise Verification for Reliable Vision-Language Process Reward Models
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models