VERDICT: Training-Free Step-Wise Verification of Multimodal Reasoning via Disagreement-Aware Consensus

arXiv:2608.10665 · cs.AI, cs.CV, cs.GT · Submitted 2026-08-11 · Read on arXiv

Rohit Sinha, Kunal Tilaganji, Tanuja Ganu, Nagarajan Natarajan, Amit Sharma, Vineeth Balasubramanian

Indian Institute of Technology, Hyderabad · Microsoft Research

cs.AI, cs.CV, cs.GT

Submitted: 2026-08-11

Updated: 2026-08-12

Comments: European Conference on Computer Vision 2026

License: http://creativecommons.org/licenses/by-sa/4.0/

Importance score: 75/100

Terminology

Summary

Summary

This paper introduces VERDICT (VERification via Disagreement-Informed Coupled Thresholding), a training-free, domain-agnostic, step-wise verification approach for multimodal large language models (MLLMs). The core problem addressed is that MLLMs often generate reasoning chains containing subtle errors that lead to incorrect answers. Existing verification methods have notable limitations: domain-specific critics (e.g., process reward models) require expensive labeled supervision and show inconsistent cross-task performance, while training-free aggregation methods (e.g., simple averaging) miss the key insight that disagreement among disparate verifiers itself carries diagnostic information about a reasoning step's validity.

The paper's key insight is that "when a reasoning step is truly valid, disparate judges evaluating it from different perspectives, such as visual grounding, logical consistency, and contextual relevance, should be able to converge. When they cannot agree, even after accounting for each other's views, the step is likely unstable." Rather than treating verification as a classification or regression task, the authors formalize it as a coupled scoring problem among disparate, frozen verifiers, interpretable as a coordination game with a unique closed-form equilibrium.

The VERDICT framework operates as follows: at each reasoning step, the base model generates n=3 candidate continuations via temperature sampling. Three frozen, modality-specialized agents—Visual (V), Logical (L), and Contextual (C)—independently score each candidate with a scalar confidence in [0,1]. These agents are all built on Qwen2.5-VL-7B-Instruct and operate in complete isolation, never seeing other agents' scores. The coupled scoring formulation models each agent's objective as balancing agreement with other agents against fidelity to its own modality-specific judgment, governed by a stubbornness parameter λi. The system of equations admits a closed-form solution, computed via a direct linear solve, which yields consensus scores for each candidate.

The acceptance criterion is a dual threshold: a candidate is accepted if and only if its mean consensus confidence exceeds τ=0.6 and its consensus dispersion falls below ϵ=0.1. Among accepted candidates, the one with the highest mean consensus confidence is selected. When no candidate passes both criteria, a fallback mechanism ranks candidates by s̄∗ − ∆∗, balancing confidence against remaining disagreement. The stubbornness parameters are set to λV=1.5, λL=1.0, λC=0.8, reflecting the intuition that the visual agent should be most resistant to consensus pressure on perception-heavy tasks.

A key theoretical contribution is Proposition 1, which establishes that consensus dispersion is not recoverable from weighted averages or any separable function of individual scores. The paper demonstrates this with a numerical example: score vectors (0.9, 0.2, 0.9) and (0.7, 0.6, 0.7) share the same raw mean (2/3), yet the consensus formulation produces opposite acceptance decisions (REJECT vs. ACCEPT) due to different dispersion values (0.13 vs. 0.02). This coupling, where each agent's adjusted score depends on all others' raw scores, is what distinguishes VERDICT from simpler aggregation methods.

The paper evaluates VERDICT across six benchmarks: 3DSRBench, CV-Bench-3D, CV-Bench-2D, BLINK, MMStar, and AI2D. Results show VERDICT consistently improves over the base model on every benchmark, with gains ranging from +2.90 on 3DSRBench to +5.95 on CV-Bench-3D, and achieves an average accuracy of 70.15% compared to the base model's 66.31%. Crucially, VERDICT never degrades below baseline on any tested benchmark, a property that no domain-specific critic can claim. Domain-specific critics like Sherlock and VisionSR1 show significant performance variability, degrading below baseline on at least two benchmarks each.

Compared to domain-agnostic baselines using the same three agents, VERDICT consistently outperforms Mean scoring by +0.68 to +2.57 points across all benchmarks, confirming that the disagreement structure captured by consensus computation provides a verification signal that score averaging discards. The Variance baseline, which explicitly attempts to incorporate disagreement by rejecting high-variance candidates, still underperforms VERDICT, indicating that treating all variance equally is insufficient.

Ablation studies reveal that the framework operates through two complementary mechanisms: rejection (filtering) and selection (ranking). Removing intelligent ranking (No Selection) costs more (−1.17 to −2.17 points) than removing filtering (No Rejection, −0.26 to −1.08 points), indicating that consensus-adjusted ranking is the primary driver of improvements. The Raw Average baseline, which uses the same dual-criterion architecture but replaces consensus scores with simple means, consistently underperforms VERDICT by 2.59–2.95 points, confirming that the coupled scoring framework produces better-calibrated scores than naive aggregation.

Stubbornness sensitivity analysis shows a clear hierarchy: λV has the widest performance range (2.90 points on 3DSRBench), λL an intermediate range (1.40), and λC the narrowest (1.19). This ordering mirrors the asymmetric design where the parameter with the most impact on spatial reasoning is set highest. Critically, no setting degrades below the base model, and all curves converge to the same peak accuracy at the default configuration. Stubbornness assignment analysis confirms that the ordering λV > λL > λC is not arbitrary: swapping V↔C costs −1.76 points on 3DSRBench, while swapping L↔C costs only −0.41 points.

Threshold sensitivity analysis reveals that the operating point (τ=0.6, ϵ=0.1) lies on a stable plateau, not an edge. The dispersion tolerance ϵ exhibits an optimal plateau spanning a full order of magnitude (ϵ ∈ [0.1, 1.0]), and the confidence threshold τ=0.6 is individually optimal on five of six benchmarks. The joint threshold sweep (49 configurations) confirms that the operating point lies on a plateau, with gradual changes of 1–2 points in any direction.

The paper also demonstrates that the consensus formulation adds value beyond aggregation, especially with weaker judges. When judge model scale decreases from 7B to 2B, the margin between VERDICT and Mean aggregation widens on five of six benchmarks, from +0.68 at 7B to +1.21 at 2B on 3DSRBench. At 2B scale, VERDICT recovers 1.2–4.6 points over naive averaging, confirming that the gains reflect algorithmic structure rather than judge quality alone.

Diagnostic analysis shows that consensus dispersion is a genuine, non-separable diagnostic signal. High dispersion is enriched 1.4–1.6× for incorrect chains, and ROC analysis yields AUC = 0.65–0.66, compared to 0.57–0.58 for raw score variance. The signal is moderate and diagnostic rather than oracular, which explains why VERDICT benefits from soft integration (ranking + dual filtering) rather than hard binary classification.

The framework generalizes across model families at zero adaptation cost. Applying VERDICT to four different base model families (Qwen2.5-VL-7B, InternVL, LLaVA, Kimi) yields consistent gains of +2.45 to +4.00 points on 3DSRBench. The consensus compresses base-model variability from a 1.33-point range to 0.22 points, confirming plug-in generality.

Computational cost analysis shows that VERDICT's total cost is 13–50× lower than domain-specific alternatives, which require 100–1000 GPU-hours of training. VERDICT's inference overhead is 3.80× under sequential execution, but this includes 3.0× for candidate diversity (shared with any best-of-N strategy) and only 0.80× for consensus verification. VERDICT achieves the highest compute efficiency among all training-free methods at 1.01 accuracy points per unit of overhead, compared to 0.63 for Mean scoring and 0.26 for Self-Synthesizer, the strongest response-level method.

The paper concludes that when disparate judges disagree about a reasoning step, the pattern of that disagreement is itself a verification signal that principled aggregation can exploit. The framework has certain limitations: when all agents are confidently wrong, consensus converges on an incorrect assessment, and no filtering recovers from a base model that produces no viable candidates. The authors note that VERDICT generalizes beyond multimodal reasoning to reward model ensembles, multi-agent evaluation, and compositional code verification.

Improvements for AI systems

Improvements to AI Systems:

  1. Disagreement-Aware Verification Module: Integrate VERDICT’s coupled scoring framework into any LLM/MLLM pipeline as a plug-in, post-hoc verifier. The system will compute consensus scores from multiple frozen, modality-specialized judges (e.g., visual, logical, contextual) and use dual thresholds (confidence >0.6, dispersion <0.1) to accept/reject reasoning steps. This enables the AI to detect subtle errors in its own chains-of-thought without retraining, improving answer accuracy by 4% on average across spatial, visual, and multimodal reasoning benchmarks.

  2. Training-Free Self-Correction with Fallback Ranking: Implement the fallback mechanism (rank by s̄* − ∆*) to select the best candidate when no step passes both thresholds. The improved AI can generate multiple candidate continuations (n=3), score them via consensus, and either confidently select the best or gracefully degrade to a balanced confidence-disagreement ranking. This prevents catastrophic failures and ensures the AI never performs worse than its base model on any task.

  3. Modality-Weighted Stubbornness Tuning: Adopt the asymmetric stubbornness parameters (λV=1.5, λL=1.0, λC=0.8) to control how much each judge resists consensus pressure. The improved AI can adapt its verification to be more resistant to visual consensus on perception-heavy tasks, more flexible on logical tasks, and most accommodating on contextual tasks—yielding up to +5.95 accuracy points on spatial reasoning benchmarks without manual per-task tuning.

  4. Non-Separable Disagreement Signal Extraction: Replace naive averaging or variance-based rejection with the closed-form consensus solution. The improved AI will exploit dispersion patterns (e.g., scores like 0.9, 0.2, 0.9 vs. 0.7, 0.6, 0.7) that simple means treat identically, enabling correct rejection of unstable steps and acceptance of stable ones. This yields 1.4–1.6× enrichment for incorrect chains and AUC improvements from 0.57 to 0.66 over raw variance.

  5. Zero-Cost Model-Agnostic Integration: Apply VERDICT’s framework to any base MLLM (Qwen, InternVL, LLaVA, Kimi) without fine-tuning. The improved AI system will show consistent gains (+2.45 to +4.00 points) across model families, compressing performance variability from a 1.33-point range to 0.22 points, making it a universal reliability booster for existing deployed models.

  6. Compute-Efficient Verification: Use VERDICT’s 0.80× inference overhead (beyond candidate generation) to achieve 1.01 accuracy points per unit of overhead—outperforming Mean scoring (0.63) and Self-Synthesizer (0.26). The improved AI can afford real-time verification in production, with total cost 13–50× lower than domain-specific critics that require 100–1000 GPU-hours of training.

  7. Robustness to Weak Judges: Deploy VERDICT with smaller judge models (e.g., 2B instead of 7B) and still recover 1.2–4.6 points over naive averaging. The improved AI can operate in resource-constrained environments (edge devices, mobile) while maintaining verification quality, as the algorithmic structure—not judge scale—drives the gains.

  8. Generalizable Verification Beyond Multimodal Reasoning: Extend the coupled scoring formulation to reward model ensembles, multi-agent evaluation, and compositional code verification. The improved AI can apply the same disagreement-informed thresholding to any multi-judge setting, enabling robust step-wise validation in code generation, mathematical proof checking, and agent-based planning.

Abstract

Multimodal large language models often generate reasoning chains containing subtle errors that lead to incorrect answers. Current verification approaches have notable limitations. Existing approaches either require expensive labelled supervision with inconsistent cross-task performance or aggregate scores from multiple sources by simple aggregations, missing a key insight: when these scores disagree, that disagreement itself carries important information about whether a reasoning step is truly valid or not. We formalise this as a coupled scoring problem among disparate, frozen verifiers, interpretable as a coordination game with a unique closed-form equilibrium where agreement signals valid steps while disagreement reveals instability. Towards this end, we propose a training-free domain-agnostic step-wise verification approach we call VERDICT: VERification via Disagreement-Informed Coupled Thresholding. To our knowledge, VERDICT is the first training-free verifier that makes the structure of cross-modal disagreement explicit and actionable. It computes consensus scores through a closed-form solution, enabling both disagreement-aware filtering and stability-conscious ranking of reasoning steps. Evaluated across six benchmarks, consistently improves over the base model by up to +5.95%, and performs competitively with domain-specific critics that demand extensive supervision, demonstrating that cross-modal agreement provides robust verification signals without task-specific adaptation and Training-Free Verification

Sources

Related papers