Beyond the Verdict: Evidence-Aligned Evaluation of Visual Prompt-Injection Guardrails

arXiv:2609.05535 · cs.CV, cs.AI · Submitted 2026-09-02 · Read on arXiv

cs.CV, cs.AI

Submitted: 2026-09-02

Updated: 2026-09-02

Comments: Accepted at the Second Workshop on Benchmarking Evidence-Aligned Multimodal Reasoning (BEAM2) at ECCV 2026 (Oral Presentation)

License: http://creativecommons.org/licenses/by/4.0/

The gist: Verdict-only evaluation does not reveal whether a vision-language model (VLM) used the visual evidence that should support its decision.

Terminology

Abstract

Verdict-only evaluation does not reveal whether a vision-language model (VLM) used the visual evidence that should support its decision. We study this problem in web-agent guardrails, where a VLM judges whether on-screen text conflicts with a user instruction. We introduce Mind2Web-Injection, a benchmark of 9,954 instruction-screenshot pairs with instruction-relative labels, pixel-exact evidence boxes, and matched image-side counterfactuals. Across six VLMs, two models with nearly identical average precision differ ninefold in Evidence-Aligned Detection (EAD), the fraction of attacks both detected and correctly localized. To test whether a verdict depends on the command cited as evidence, we replace the instruction with one that endorses that command. Qwen3-VL-32B, the strongest open-weight localizer, returns aligned in only 58.7% of cases, whereas GPT-5.6-luna does so in 99.9%. To diagnose these failures, we propose two training-free interventions. ReadGate improves grounding without changing verdicts, while CmdCompare tests whether explicit instruction-command comparison resolves instruction-side inconsistency. These results motivate reporting verdict correctness, evidence localization, and counterfactual responsiveness separately.

Related papers