Do Vision-Language Models Understand Visual Persuasiveness? A Diagnosis via Visual Persuasive Factors
cs.CL, cs.CV
Submitted: 2025-11-21
Updated: 2026-09-10
Comments: EMNLP 2026 Findings (39 pages); Code available at https://github.com/gyuwon12/visual-persuasive-factors
Code: https://github.com/gyuwon12/visual-persuasive-factors
License: http://creativecommons.org/licenses/by/4.0/
The gist: Visual persuasion uses images to shape cognition, emotion, and behavior, with its effects depending on both visual attributes and semantic context.
Terminology
Abstract
Visual persuasion uses images to shape cognition, emotion, and behavior, with its effects depending on both visual attributes and semantic context. Despite recent progress, it remains unclear whether Vision-Language Models (VLMs) understand visual persuasiveness. This motivates us to ask: can VLMs assess whether an image persuasively supports an intended message, which visual factors shape this judgment, and do they align with human judgments? Through empirical analyses on image-message pairs where human raters consistently agree on the persuasiveness judgment, we show that VLMs exhibit a recall-oriented bias: they over-predict images as persuasive while achieving high recall. We introduce Visual Persuasive Factors (VPFs), a taxonomy informed by cognitive psychology for quantifying visual cues that shape persuasive judgments. Our factor-level analysis reveals that VPFs distinguish human persuasiveness judgments, whereas VLMs only partially reproduce these patterns, often generating false positives by treating persuasion-relevant cues as sufficient evidence. Building on this insight, we evaluate VPF-guided interventions and find that properly framed VPF knowledge can improve performance, but merely specifying visual cues or adding step-by-step reasoning is insufficient. By analyzing model rationales at the level of functional reasoning steps, we further identify a central bottleneck in connecting object identification to semantic message alignment.
Sources
- Measuring and Improving Persuasiveness of Large Language Models
- DeepGaze IIE: Calibrated prediction in and out-of-domain for state-of-the-art saliency modeling
- GPT-4o System Card
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Qwen2.5-VL Technical Report
- Gemma 3 Technical Report
- MetaMorph: Multimodal Understanding and Generation via Instruction Tuning
- T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT
- Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering