LookBack: Where and How to Score LVLM Responses via Visual Reference Usage

arXiv:2608.11847 · cs.CV, cs.AI, cs.CL · Submitted 2026-08-12 · Read on arXiv

Yonsei University

cs.CV, cs.AI, cs.CL

Submitted: 2026-08-12

Updated: 2026-08-31

Comments: 19 pages, 10 figures. Code: https://github.com/bscho333/LookBack

Code: https://github.com/bscho333/LookBack

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 75/100

The gist: L OOK BACK is a training-free LVLM response scoring method that augments token likelihood with a visual lookback score, a lightweight measure of how strongly each response token refers to image

Terminology

Summary

L OOK BACK is a training-free LVLM response scoring method that augments token likelihood with a visual lookback score, a lightweight measure of how strongly each response token refers to image tokens. The paper identifies a gap between image-conditioning and image-sensitivity: the confidence score reflects a response’s textual plausibility rather than its agreement with the image. Specifically, removing the input image barely changes confidence-based selection, suggesting that output-space confidence primarily captures textual plausibility rather than agreement with the image. The method combines output-space confidence with visual reference usage, calibrating token confidence with visual lookback scores and aggregating token scores at the response level with visual relevance weights. Across four benchmarks and three models, L OOK BACK consistently improves Best-of-N selection over existing baselines with negligible additional overhead. The paper states: Across all setups, L OOK BACK improves the average score from 65.37% to 68.62%, achieving a 4.97% relative gain over random selection. The method requires no auxiliary verifier, training, or extra inference passes, as it uses token likelihood and visual lookback scores which can be freely obtained during generation. The paper concludes: L OOK BACK consistently improves Best-of-N selection over both linguistic and vision side baselines, while requiring no auxiliary verifier, training, or extra inference passes.

Improvements for AI systems

Improvements to AI Systems:

  1. Image-Grounded Response Scoring for Multimodal Generation
  • Integrate LOOKBACK’s visual lookback score into any LVLM’s decoding or reranking pipeline (e.g., LLaVA, GPT-4V, Gemini) to replace pure likelihood-based selection.

  • The improved system can automatically reject hallucinated or image-agnostic responses during Best-of-N sampling, even when the response is textually fluent but visually inconsistent.

  1. Training-Free Hallucination Mitigation in Real-Time Chat
  • Use the visual lookback score as a lightweight, per-token gate during autoregressive generation to down-weight tokens that do not attend to relevant image regions.

  • The improved system can reduce object misattribution, attribute errors, and unsupported claims in live multimodal conversations without fine-tuning or extra inference passes.

  1. Calibrated Confidence for Visual Question Answering (VQA)
  • Replace raw softmax confidence with the LOOKBACK-calibrated score (token likelihood × visual reference weight) for answer selection and abstention.

  • The improved system can decide when to say “I don’t know” based on whether the answer truly references the image, not just whether it is a plausible sentence—improving reliability in medical imaging, autonomous driving, and accessibility tools.

  1. Best-of-N Self-Improvement for Multimodal Reasoning
  • Apply LOOKBACK as a verifier-free reward signal for self-consistency or majority voting over multiple sampled responses.

  • The improved system can select the most image-faithful reasoning chain in tasks like visual math, diagram interpretation, or scene understanding, boosting accuracy by 5% relative (as shown) without any training.

  1. Zero-Shot Adaptation to New Domains
  • Since LOOKBACK requires no training or auxiliary models, deploy it directly on new LVLMs or new image distributions (e.g., satellite imagery, medical scans) to improve response selection.

  • The improved system can immediately benefit from better image grounding in specialized domains where labeled data or verifiers are scarce.

  1. Efficient Ensemble of Linguistic and Visual Signals
  • Combine LOOKBACK with existing linguistic confidence (e.g., perplexity) and visual grounding (e.g., attention maps) into a single, interpretable score for ranking candidate responses.

  • The improved system can outperform both pure language-based and pure vision-based baselines, as demonstrated, while maintaining negligible overhead—useful for high-throughput APIs.

  1. Bias Reduction in Multimodal Evaluation
  • Use LOOKBACK as an automated evaluation metric to compare LVLMs on image-grounded tasks, replacing human or LLM-judge scores that often favor fluent but unfaithful outputs.

  • The improved system can provide a more objective, reproducible benchmark for model development, highlighting gaps in image sensitivity that current metrics miss.

Abstract

Large Vision-Language Models (LVLMs) integrate visual perception with language generation, enabling responses that span image understanding and complex reasoning. However, LVLMs do not just inherit the text-level hallucinations; they also hallucinate against the image, producing fluent responses ungrounded in what they see. This makes LVLM response scoring inherently harder, and our diagnostics show that existing confidence-based metrics adopted from LLMs are insufficient for LVLMs. Specifically, removing the input image barely changes confidence-based selection, suggesting that output-space confidence primarily captures textual plausibility rather than agreement with the image. To address this gap, we propose LookBack, a training-free LVLM response scoring method that augments token likelihood with visual lookback score, a lightweight measure of how strongly each response token refers to image tokens. Across four benchmarks and three models, LookBack consistently improves Best-of- N selection over existing baselines with negligible additional overhead.

Sources

Related papers