How Do VLMs Behave When Blind or Misled? Behavioral Evaluation of VLMs on Scientific Figures

arXiv:2608.13267 · cs.CL, cs.AI, cs.CV, cs.LG · Submitted 2026-08-13 · Read on arXiv

Paul Osemudiame Oamen, Owusu-Banahene Osei, Ananya Mukherjee, Christian Greisinger, Steffen Eger, Pius Onobhayedo, Wei Zhao

University of Aberdeen · International Institute of Information Technology Hyderabad · University of Technology Nuremberg · University of Southern California

cs.CL, cs.AI, cs.CV, cs.LG

Submitted: 2026-08-13

Updated: 2026-08-14

Comments: 25 pages including appendix. Project website: https://scifigbench.nlp4sci.com

Project page: https://scifigbench.nlp4sci.com

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 95/100

The gist: SciFigBench is a diagnostic vision-language model (VLM) benchmark for scientific figure understanding that jointly evaluates perception, reasoning, and behavioral reliability under uncertainty.

Terminology

Summary

SciFigBench is a diagnostic vision-language model (VLM) benchmark for scientific figure understanding that jointly evaluates perception, reasoning, and behavioral reliability under uncertainty. It contains 250 figures from 187 arXiv papers (2023–2025) spanning bar charts, line plots, and pie charts, with high-quality human annotations totaling over 600 hours of annotation effort. The benchmark is extended via image transformations, reasoning questions, resistance probes, caption-bias probes, and confirmed selective-blur targets, producing over 34,000 evaluation setups for stress testing.

The paper proposes the Admittance–Resistance–Inductance (A-R-I) framework to evaluate whether models acknowledge insufficient evidence (Admittance), resist misleading context or false premises (Resistance), and infer cautiously from partial information (Inductance). The framework distinguishes three behavioral regimes based on the presence of underlying evidence and its recoverability from surrounding context.

Results reveal substantial behavioral differences among models. GPT-5.2 achieves the highest description quality (MQM 91.6) with strong reasoning accuracy (78.4%), yet hallucinates unreadable content in 96% of cases. Gemini 3.1 Pro, a comparably capable model (MQM 90.2, reasoning 81.0%), admits uncertainty in 71% of such cases and achieves the strongest resistance score (0.91). These findings show that high perception and reasoning accuracy alone do not guarantee behavioral reliability, a dimension critical for deployment in scientific workflows.

The paper's contributions are: (i) introducing SciFigBench with 250 annotated figures, MQM-based open-ended descriptions, and 1,000 figure-grounded reasoning questions covering counting, computation, comparison, and pattern analysis; (ii) developing a controlled stress-test suite with transformed figures, caption-bias settings, and false-premise probes; and (iii) proposing the A-R-I framework decomposing behavior into uncertainty acknowledgment, resistance to misleading context, and inference from partial evidence through selective-blur and false-premise probes.

Key findings include: rotation is the most damaging perceptual transform (average drop of 19.4 MQM points); Gemini 3.1 Pro leads on nearly every behavioral dimension despite placing second on quality; GPT-5.2 ranks first on description quality but falls to fourth on active admittance; inexist probes grounded in presupposition embedding are the most effective deception technique across all models; caption dependency appears tied to instruction tuning rather than model capability; and models show a must answer bias where direct questions compress responses toward definitive answers while descriptions allow natural hedging.

Improvements for AI systems

Improvements to AI Systems:

  1. Implement an Uncertainty-Aware Answering Module
  • Add a trained classifier that detects when a query lacks sufficient evidence (e.g., unreadable chart regions, ambiguous axes) and forces the model to output I cannot determine this from the given figure instead of hallucinating.

  • Integrate a confidence threshold: if perceptual confidence < 0.5, switch from direct answer to hedged response (e.g., Based on visible data, likely X, but unreadable areas may affect this).

  1. Add a False-Premise and Misleading-Context Resistance Layer
  • Pre-process user prompts to detect embedded false presuppositions (e.g., What is the trend of the blue line? when no blue line exists). If detected, the system responds with a correction (e.g., There is no blue line; the lines are red and green) before answering.

  • Train on adversarial caption-bias pairs to learn to ignore misleading captions when they conflict with visual evidence, using a cross-modal consistency check.

  1. Introduce a Must-Answer Bias Suppressor
  • Modify the decoding strategy to distinguish between open-ended description tasks (allow hedging) and direct question tasks (currently compress toward definitive answers). Add a prompt-type tag that relaxes the answer compression for questions with low evidence, enabling natural hedging like The value appears to be 30, but the label is blurred.
  1. Implement Rotation and Transform Robustness via Pre-Training Augmentation
  • During fine-tuning, include rotated, scaled, and skewed versions of figures (up to ±45°) with explicit labels for orientation. This reduces the 19.4-point MQM drop on rotated inputs by teaching the model to re-orient axes and legends before reasoning.
  1. Add a Selective-Blur Inference Mode
  • For partial occlusion (e.g., blurred data points), the system should explicitly state what is missing and reason from visible portions only, using a two-step process: (a) identify occluded regions via a segmentation model, (b) generate a cautious inference with a caveat (e.g., If the hidden bar follows the visible trend, it would be 40, but this is speculative).
  1. Create a Behavioral Reliability Score for Deployment
  • Output a composite score (0–1) combining admittance rate, resistance score, and inductance quality alongside every answer. This allows downstream scientific workflows to reject low-reliability outputs automatically, even if the answer text appears confident.

What the Improved AI System Can Do:

  • For scientific figure analysis: It will never fabricate data from unreadable regions; it will say I cannot read the value and explain why.

  • Against misleading prompts: It will refuse to answer false-premise questions and instead correct the user's assumption.

  • Under image distortions: It will maintain high accuracy on rotated charts by internally normalizing orientation.

  • With partial data: It will provide cautious, clearly-labeled estimates rather than false certainty.

  • In deployment: It will flag its own low-confidence outputs, enabling human review or automated rejection in high-stakes workflows.

Abstract

Existing vision-language model (VLM) benchmarks emphasize perception and reasoning accuracy (how well VLMs describe and reason about what they see in an image), with limited attention to behavioral reliability under uncertainty (how they behave when visual evidence is missing or misleading). We introduce SciFigBench, a diagnostic VLM benchmark for scientific figure understanding that jointly evaluates perception, reasoning, and behavioral reliability under uncertainty. It contains 250 figures with high-quality human annotations across three evaluation aspects, totaling 600+ hours of annotation effort. We further extend these figures via image transformations, reasoning questions, resistance probes, caption-bias probes, and confirmed selective-blur targets, producing over 34,000 evaluation setups for stress testing. We further propose the Admittance-Resistance-Inductance (A-R-I) framework to evaluate whether models acknowledge insufficient evidence, resist misleading context, and infer cautiously from partial information. Our results reveal substantial behavioral differences among models. GPT-5.2 achieves the highest description quality (MQM 91.6) with strong reasoning accuracy (78.4%), yet hallucinates unreadable content in 96% of cases, whereas Gemini 3.1 Pro, a comparably capable model (MQM 90.2, reasoning 81.0%), admits uncertainty in 71% of such cases and achieves the strongest resistance score (0.91). These findings show that high perception and reasoning accuracy alone do not guarantee behavioral reliability, a dimension critical for deployment in scientific workflows.

Sources

Related papers