LookBack: Where and How to Score LVLM Responses via Visual Reference Usage
Yonsei University
cs.CV, cs.AI, cs.CL
Submitted: 2026-08-12
Updated: 2026-08-31
Comments: 19 pages, 10 figures. Code: https://github.com/bscho333/LookBack
Code: https://github.com/bscho333/LookBack
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 75/100
The gist: L OOK BACK is a training-free LVLM response scoring method that augments token likelihood with a visual lookback score, a lightweight measure of how strongly each response token refers to image
Terminology
Summary
L OOK BACK is a training-free LVLM response scoring method that augments token likelihood with a visual lookback score, a lightweight measure of how strongly each response token refers to image tokens. The paper identifies a gap between image-conditioning and image-sensitivity: the confidence score reflects a response’s textual plausibility rather than its agreement with the image.
Specifically, removing the input image barely changes confidence-based selection, suggesting that output-space confidence primarily captures textual plausibility rather than agreement with the image.
The method combines output-space confidence with visual reference usage, calibrating token confidence with visual lookback scores and aggregating token scores at the response level with visual relevance weights. Across four benchmarks and three models, L OOK BACK consistently improves Best-of-N selection over existing baselines with negligible additional overhead. The paper states: Across all setups, L OOK BACK improves the average score from 65.37% to 68.62%, achieving a 4.97% relative gain over random selection.
The method requires no auxiliary verifier, training, or extra inference passes, as it uses token likelihood and visual lookback scores which can be freely obtained during generation.
The paper concludes: L OOK BACK consistently improves Best-of-N selection over both linguistic and vision side baselines, while requiring no auxiliary verifier, training, or extra inference passes.
Improvements for AI systems
Improvements to AI Systems:
- Image-Grounded Response Scoring for Multimodal Generation
-
Integrate LOOKBACK’s visual lookback score into any LVLM’s decoding or reranking pipeline (e.g., LLaVA, GPT-4V, Gemini) to replace pure likelihood-based selection.
-
The improved system can automatically reject hallucinated or image-agnostic responses during Best-of-N sampling, even when the response is textually fluent but visually inconsistent.
- Training-Free Hallucination Mitigation in Real-Time Chat
-
Use the visual lookback score as a lightweight, per-token gate during autoregressive generation to down-weight tokens that do not attend to relevant image regions.
-
The improved system can reduce object misattribution, attribute errors, and unsupported claims in live multimodal conversations without fine-tuning or extra inference passes.
- Calibrated Confidence for Visual Question Answering (VQA)
-
Replace raw softmax confidence with the LOOKBACK-calibrated score (token likelihood × visual reference weight) for answer selection and abstention.
-
The improved system can decide when to say “I don’t know” based on whether the answer truly references the image, not just whether it is a plausible sentence—improving reliability in medical imaging, autonomous driving, and accessibility tools.
- Best-of-N Self-Improvement for Multimodal Reasoning
-
Apply LOOKBACK as a verifier-free reward signal for self-consistency or majority voting over multiple sampled responses.
-
The improved system can select the most image-faithful reasoning chain in tasks like visual math, diagram interpretation, or scene understanding, boosting accuracy by 5% relative (as shown) without any training.
- Zero-Shot Adaptation to New Domains
-
Since LOOKBACK requires no training or auxiliary models, deploy it directly on new LVLMs or new image distributions (e.g., satellite imagery, medical scans) to improve response selection.
-
The improved system can immediately benefit from better image grounding in specialized domains where labeled data or verifiers are scarce.
- Efficient Ensemble of Linguistic and Visual Signals
-
Combine LOOKBACK with existing linguistic confidence (e.g., perplexity) and visual grounding (e.g., attention maps) into a single, interpretable score for ranking candidate responses.
-
The improved system can outperform both pure language-based and pure vision-based baselines, as demonstrated, while maintaining negligible overhead—useful for high-throughput APIs.
- Bias Reduction in Multimodal Evaluation
-
Use LOOKBACK as an automated evaluation metric to compare LVLMs on image-grounded tasks, replacing human or LLM-judge scores that often favor fluent but unfaithful outputs.
-
The improved system can provide a more objective, reproducible benchmark for model development, highlighting gaps in image sensitivity that current metrics miss.
Abstract
Large Vision-Language Models (LVLMs) integrate visual perception with language generation, enabling responses that span image understanding and complex reasoning. However, LVLMs do not just inherit the text-level hallucinations; they also hallucinate against the image, producing fluent responses ungrounded in what they see. This makes LVLM response scoring inherently harder, and our diagnostics show that existing confidence-based metrics adopted from LLMs are insufficient for LVLMs. Specifically, removing the input image barely changes confidence-based selection, suggesting that output-space confidence primarily captures textual plausibility rather than agreement with the image. To address this gap, we propose LookBack, a training-free LVLM response scoring method that augments token likelihood with visual lookback score, a lightweight measure of how strongly each response token refers to image tokens. Across four benchmarks and three models, LookBack consistently improves Best-of- N selection over existing baselines with negligible additional overhead.
Sources
- GPT-4 Technical Report
- Revisiting the Uniform Information Density Hypothesis in LLM Reasoning
- Qwen2.5-VL Technical Report
- VRPRM: Process Reward Modeling via Visual Reasoning
- The Curious Case of Neural Text Degeneration
- Universal Self-Consistency for Large Language Model Generation
- Language Models (Mostly) Know What They Know
- Training Verifiers to Solve Math Word Problems
- Training-free LLM Verification via Recycling Few-shot Examples
- Solving math word problems with process- and outcome-based feedback
- AMBER: An LLM-free Multi-dimensional Benchmark for MLLMs Hallucination Evaluation
- VisualPRM: An Effective Process Reward Model for Multimodal Reasoning
- VAUQ: Vision-Aware Uncertainty Quantification for LVLM Self-Evaluation
- Gemini: A Family of Highly Capable Multimodal Models
- InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
- LLaMA: Open and Efficient Foundation Language Models
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models