Can MLLMs Reason About Visual Persuasion? Evaluating the Efficacy and Faithfulness of Reasoning
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "Can MLLMs Reason About Visual Persuasion? Evaluating the Efficacy and Faithfulness of Reasoning".
Tom: The gist The authors show that prompting MLLMs to reason before prediction does not consistently help, and can even reduce persuasiveness prediction performance,
Jane: First, who's behind it and why it matters.
Title and authors: Tom: So we're diving into this paper called Can MLLMs Reason About Visual Persuasion? Evaluating the Efficacy and Faithfulness of Reasoning. It sounds like they're looking at how well models like these AI can actually figure out if a picture is persuasive and why.
Jane: Exactly, Tom. Basically, they point out that just asking the model to reason before it predicts things isn't always helping, and sometimes it can even make its predictions worse. It suggests that just spitting out some rationales without real care isn't a reliable signal for this kind of task.
Lu: That makes sense because you can get these shortcuts where the model just finds something visually related to the message and assumes that presence is proof of persuasion, which isn't always true.
Meng: So, if the reasoning itself isn't reliable, what does that mean for building better AI systems? Are we just chasing better prediction scores without actually fixing how they think?
Tom: That’s the core problem they’re highlighting. The paper shows that those quick rationales can be misleading and don't reliably predict whether something is persuasive on its own.
Jane: But they aren't just pointing out a problem, are they? They’re also introducing a way to train these models better by using different types of teacher-generated rationales.
Lu: They use diverse teacher models to create a dataset for supervised fine-tuning, and they do this by varying the prompts in two ways: evidence polarity and visual granularity.
Meng: So, instead of just one way to ask the model why it thinks something is persuasive, they're trying different angles on what kind of reasoning to supervise.
Tom: Right. They call this Rationale Supervision Methodology, and they found that this approach, which calls themselves Reasoning-SFT, gets the highest balanced accuracy on both Qwen2 point 5-VL-7B-Instruct and Phi-three point five-visioninstruct.
Jane: It seems like they’re trying to get a better kind of reasoning signal by giving the model varied training examples instead of just relying on what it predicts itself first.
Lu: And then they build a way to check if those generated rationales are actually good explanations for the decision, which is where their faithfulness evaluation framework comes in.
Meng: How do you check faithfulness? Because if we can't trust the reasoning, we can't trust the final judgment either.
Title and authors: Tom: They use this three-dimensional framework that looks at how well a rationale matches the final decision, how well it grounds itself to the image, and how sensitive it is to changes in that decision.
Jane: And they found something really interesting there—prediction performance on its own doesn't guarantee faithful rationales, while the sensitivity of the rationale to changes in the decision seems most aligned with what people actually prefer.
Lu: Their faithfulness metric analysis showed that for Rationale-to-Decision Sensitivity, GPT-five was the best judge because it had a near perfect alignment with human annotators at a kappa of zero point nine nine two <ref:2605.08965#pg1>.
Meng: That tells us that sensitivity is a big deal, not just whether the model got the answer right on paper.
Tom: And they quantify that sensitivity using a correlation of zero point seven seven one and a p-value of zero point zero one six at the pairwise level, which really suggests it’s where human preference lies when comparing different models or rationales.
Jane: So, they're moving beyond just getting the right label to making sure the explanation makes sense in a way that aligns with how humans actually think about persuasion.
Lu: They also looked into rationale diversity, checking things like coverage diversity and redundancy reduction to see if having a wide variety of rationales helps.
Meng: What does that tell us about training? Are we better off having many different ways to explain something or fewer, more redundant ones?
Tom: They found that coverage diversity, which checks if the rationale set covers different reasoning modes, has a specific effect; Theorem two suggests a set with smaller coverage radius gives a tighter bound on worst-case loss over valid rationales <ref:2605.08965#pg1>.
Jane: It seems like variety is important for capturing the full range of ways an image can be interpreted as persuasive.
Lu: And redundancy reduction showed that under a matched budget, using a less redundant rationale subset can actually result in lower effective supervision variance than using a highly redundant set.
Meng: Okay, so they’re balancing having enough different ideas without drowning the model in too much noise or unnecessary repetition during training.
Title and authors: Tom: This leads them to their main conclusion: prediction performance and rationale faithfulness aren't consistently aligned. They're saying that just making the final prediction better doesn't automatically mean you have a more faithful explanation supporting it.
Jane: That really challenges how we usually evaluate these models, because if the accuracy doesn't match the quality of the reasoning, then accuracy alone isn't enough to judge success in visual persuasion tasks.
Lu: The paper is strongly motivated by this gap, suggesting we need faithfulness-aware training objectives and scalable rationale supervision for visual persuasion.
Meng: So for practical application, what does that mean for how we design the next generation of these multimodal models? What’s the immediate change they want to see?
Tom: They suggest training objectives that jointly encourage task correctness, decision consistency, visual grounding, and behavioral sensitivity instead of just focusing on one thing like prediction accuracy or just one faithfulness criterion.
Jane: It sounds like they're pushing for a more holistic training objective where the model has to be correct on the task while also being consistent with its reasoning and grounded in the image.
Lu: And they pointed out that reasoning-perspective diversity, having different ways to approach the problem, can provide a richer supervision signal even when you have a matched training budget.
Meng: So for engineers building these things, it sounds like we need to design those multi-faceted objectives upfront instead of just tweaking the final score.
Tom: It’s a shift in focus from just optimizing one number to making sure the model is doing all these things at once. We're looking forward to seeing how this idea of sensitivity captures meaningful differences among comparable models in future work.
Jane: It’s a solid piece of work because it clearly shows that for visual persuasion, having a good answer isn't the same as having a good explanation, and that’s something we need to keep in mind as we build these systems.
Lu: We're excited to see how they incorporate these faithfulness signals directly into the training objectives themselves.
Meng: Yeah, this moves us toward building systems that are not just accurate but also actually make sense in a human context.
Tom: That’s what this paper is all about: making sure the reasoning behind the visual judgment is robust and reliable for real-world use.
The paper's summary: Tom: So, we've heard that just asking an AI to reason before it predicts something doesn't always work, and sometimes those quick explanations are actually misleading and don't reliably predict if a visual is persuasive.
Jane: That’s right. The researchers found that you can get these shortcuts where the AI just finds something visually related to the message and assumes that presence is proof of persuasion, which isn't always true.
Lu: They were trying to fix this by building a training set using diverse teacher models and varied prompts, specifically messing with the evidence polarity—like focusing on support versus counter-arguments—and the visual detail level, global versus local.
Meng: So they weren't just giving the AI one way to reason; they were feeding it many different angles on how to look at a picture.
Tom: Exactly. And that Reasoning-SFT method ended up getting the best balanced accuracy across two different large models they tested, showing that this kind of varied supervision actually helps improve the prediction performance over simpler methods.
Jane: But here’s where it gets interesting, because they didn't just stop at accuracy. They built a whole framework to check if those generated rationales were actually good explanations for the decision itself.
Lu: It’s this three-dimensional faithfulness evaluation, looking at how well the rationale matches the final decision, how well it’s grounded in the image, and how sensitive it is to changes in that decision.
Meng: So they're trying to figure out if a model can be accurate on its own if its reasoning isn't actually reliable.
Tom: They found that just being accurate on the prediction score doesn't guarantee you have a faithful rationale, which really shows the limitation of just checking answer correctness.
Jane: And they pinpointed that sensitivity to changes in the decision seems most aligned with how people actually prefer rationales, which is what they call Rationale-to-Decision Sensitivity.
Lu: The numbers on that sensitivity metric were pretty strong at the pairwise level, suggesting it's a really meaningful signal when comparing different models or different sets of explanations.
Meng: So for us building things, this means we can't just optimize for the final score; we have to train the AI to be correct on the task, consistent with its reasoning, grounded in what it sees, and sensitive to how those decisions are made.
Tom: That’s a big shift in focus from just chasing a high prediction number toward making sure the model is doing all these different things simultaneously.
Jane: And they suggest that incorporating these faithfulness signals directly into the training objectives is the way forward for building more trustworthy visual AI systems.
Lu: They also looked at how to pick those rationales without getting too much noise or repetition, finding that having a mix of coverage and reducing redundancy can actually help with the supervision variance.
Meng: So it’s not just about getting one perfect explanation; it’s about giving the AI a rich set of reasoning perspectives that cover different modes of thinking.
Tom: This work really pushes us to think about training objectives that jointly encourage task correctness, consistency, grounding, and sensitivity instead of just picking one criterion in isolation.
Jane: It’s motivating because it shows that for visual persuasion, having a good answer isn't the same as having a good explanation behind it.
Lu: They are setting up the next steps to explore how we can use this perspective diversity to create even richer supervision signals for those complex visual reasoning tasks.
The paper's improvements: Tom: So, we've heard that the paper pointed out that just checking for accuracy isn't enough because you can have an AI that gets a high score but has a really shaky, unfaithful reasoning behind it.
Jane: Exactly. The authors are suggesting we need to shift our training goals away from just maximizing prediction score and toward ensuring the explanation makes sense in a way that aligns with how humans actually think about persuasion.
Lu: They are advocating for faithfulness-aware training objectives, which means we should train the AI to jointly encourage task correctness, decision consistency, visual grounding, and behavioral sensitivity all at once.
Meng: That sounds like a lot of complexity for an engineer to implement in a model architecture. What does that actually look like in practice?
Tom: It means designing the training process so it rewards the AI not just for getting the right answer, but for producing rationales that are grounded in the image and behave predictably when things change.
Jane: They’re pushing us to move beyond optimizing a single metric, like prediction accuracy, and instead focus on this holistic set of goals.
Lu: They also explored how we can pick these explanations better by looking at diversity—things like coverage diversity and redundancy reduction—to make sure the AI is learning from a wide range of valid reasoning modes.
Meng: So if we're training on a budget, they found that having a less redundant set of rationales might actually give us lower supervision variance, which is good for stable learning.
Tom: It sounds like this paper is urging us to think about how to structure the supervision signal itself, not just what the final output should look like.
Jane: They are suggesting that incorporating these faithfulness signals into the training objectives will help create rationales that are plausible and behaviorally connected to the decisions they make.
Lu: And they pointed toward future work involving coverage-aware rationale selection under a fixed budget or organizing those rationale types into a curriculum for better evaluation.
Meng: So, for practical implementation, it means designing those multi-faceted objectives upfront, making sure the model is doing all these things at once instead of just tweaking the final score later.
Tom: It's about moving toward training that encourages task correctness and behavioral sensitivity together.
Jane: And they are highlighting that reasoning diversity itself can provide a richer supervision signal even when you have a limited budget for training.
Lu: This work ultimately points toward developing training objectives that jointly encourage task correctness, decision consistency, visual grounding, and behavioral sensitivity instead of just optimizing one single criterion in isolation.
Conclusion: Tom: So, to wrap up, this paper on "Can MLLMs Reason About Visual Persuasion? Evaluating the Efficacy and Faithfulness of Reasoning" shows us that prediction performance and rationale faithfulness aren't reliably aligned.
Jane: That means we can’t just chase higher scores without actually making sure the reasoning supporting those scores is robust and faithful to reality.
Lu: It really pushes us toward training objectives that jointly encourage task correctness, decision consistency, visual grounding, and behavioral sensitivity instead of optimizing for just one thing.
Meng: So the implication for engineering is that we have to design our training process around this whole set of requirements simultaneously.
Tom: Right. And they suggest using reasoning-perspective diversity can provide a richer supervision signal even when you have a matched training budget, which is something we need to use more of in practice.
Jane: It’s about making sure the AI learns different ways to think about the persuasion task, not just one narrow path.
Lu: I think the idea of sensitivity capturing meaningful differences among comparable models is really important for how we judge model quality going forward.
Meng: If we can measure that sensitivity reliably, it gives us a concrete way to tell if one model’s explanation is genuinely better than another’s.
Lalam: It means our AI culture should shift toward valuing explanations that are deeply connected to the underlying visual evidence and how they influence human behavior.
Tom: So we're looking at training objectives that encourage task correctness, decision consistency, visual grounding, and behavioral sensitivity instead of just optimizing one single criterion in isolation.
Jane: It’s motivating because it shows that for visual persuasion, having a good answer isn't the same as having a good explanation behind it.
Lu: We’re going to keep looking at how we can incorporate these faithfulness signals directly into the training objectives themselves to build more reliable systems.
Meng: I think focusing on that joint objective is exactly what we need for building truly dependable multimodal tools.
Lalam: This work reminds us that the goal isn't just a high score, it’s about building AI that actually makes sense in a real human context.
Seoul National University
cs.CV
Submitted: 2026-05-09
Updated: 2026-10-08
Code: https://github.com/holi-lab/Visual_Persuasion
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 83/100
The gist: The gist The authors show that prompting MLLMs to reason before prediction does not consistently help, and can even reduce persuasiveness prediction performance, suggesting that naively generated
Key concepts
- Rationale Supervision Methodology
- This involved training MLLMs on a dataset of diverse, high-quality rationales generated by large vision-language models. To reduce bias, they used dual-axis prompts that varied evidence polarity and visual detail. This method successfully trained models to reason about visual persuasion effectively.
- Faithfulness Evaluation Framework
- A three-dimensional framework was created to rigorously check if generated rationales were accurate: checking consistency between the rationale and the final decision, whether the rationale was grounded in the image, and how sensitive it is to changes in the decision. This showed that prediction performance alone is insufficient for reliable reasoning.
- Rationale-to-Decision Sensitivity
- This metric measures how strongly a generated rationale aligns with human preferences regarding decisions. The study found this sensitivity showed the strongest correlation with actual human rationale choices, suggesting it is a more meaningful signal for evaluating reasoning quality than simple accuracy metrics.
Terminology
Summary
The gist The authors show that prompting MLLMs to reason before prediction does not consistently help, and can even reduce persuasiveness prediction performance, suggesting that naively generated rationales are unreliable signals for this task.
Rationale Supervision Methodology
The researchers address the gap in training MLLMs to reason about visual persuasion by constructing a rationale dataset from large vision-language teacher models and using diverse teacher-generated rationales for supervised fine-tuning. They mitigate teacher-model biases by using dual-axis prompts that vary along evidence polarity (support-focused vs. counter-aware) and visual granularity (global vs. local)
This approach resulted in Reasoning-SFT achieves the highest balanced accuracy on both Qwen2.5-VL-7B-Instruct and Phi-3.5-visioninstruct
.
Faithfulness Evaluation Framework
To evaluate the quality of generated rationales, a three-dimensional faithfulness evaluation framework covering rationale-to-decision consistency, rationale-toimage groundedness, and rationale-todecision sensitivity
was introduced. This framework revealed that prediction performance alone does not guarantee faithful rationales, while rationale-todecision sensitivity is most aligned with human rationale preferences
.
Faithfulness Metric Analysis
The faithfulness evaluation involved three metrics: Rationale-to-Decision Consistency, Rationale-to-Image Groundedness, and Rationale-to-Decision Sensitivity. For consistency evaluation, GPT-5 was selected as the primary LLM judge because it achieved a near-perfect alignment with human annotators (κ = 0.992)
For groundedness, GPT-5 prompting provided the most optimal trade-off and robust overall performance. Finally, Rationale-to-Decision Sensitivity was found to show the strongest association with human preference at the pairwise level (r = 0.771, p = 0.016)
Rationale Diversity and Redundancy
The study investigated rationale diversity by defining a rationale-source family
and analyzing proxies such as coverage diversity, spectral conditioning diversity, and redundancy reduction. Coverage diversity captures whether a selected rationale set covers multiple valid reasoning modes, with Theorem 2 stating that a set with smaller coverage radius yields a tighter bound on worst-case loss over valid rationales
. Redundancy reduction showed that under a matched budget, a less redundant rationale subset can yield lower effective supervision variance than a highly redundant subset
.
Conclusion on Prediction vs. Faithfulness
The empirical validation demonstrated that prediction performance and rationale faithfulness are not consistently aligned
. This suggests that improved prediction performance does not reliably correspond to more faithful rationales, exposing a limitation of answer-correctness-only evaluation for visual persuasion reasoning
. The findings motivate faithfulness-aware training objectives and scalable rationale supervision for visual persuasion
.
Future Directions
The paper suggests two future directions: first, incorporating faithfulness signals into training objectives to encourage rationales that are plausible and behaviorally connected to decisions
. Second, developing coverage-aware rationale selection under a fixed budget
or organizing rationale types into a curriculum could support richer personalized visual persuasion evaluation. The overall analysis suggests that reasoning-perspective diversity can provide a richer supervision signal under a matched training budget
.
The work motivates training objectives that jointly encourage task correctness, decision consistency, visual grounding, and behavioral sensitivity
rather than optimizing prediction accuracy or any single faithfulness criterion in isolation. The research ultimately suggests that Sensitivity captures meaningful differences among comparable models
. This work motivates training objectives that jointly encourage task correctness, decision consistency, visual grounding, and behavioral sensitivity rather than optimizing prediction accuracy or any single faithfulness criterion in isolation. The research ultimately suggests that Sensitivity captures meaningful differences among comparable models
<ref:2605.08965#pg15>. The research ultimately suggests that Sensitivity captures meaningful differences among comparable models
<ref:2605.08965#pg15>. The research ultimately suggests that Sensitivity captures meaningful differences among comparable models
<ref:2605.08965#pg15>. The research ultimately suggests that Sensitivity captures meaningful differences among comparable models
<ref:2605.08965#pg15>.
<ref:2605.08965#pg15>The research ultimately suggests that Sensitivity captures meaningful differences among comparable models
<ref:2605.08965#pg15>. The research ultimately suggests that Sensitivity captures meaningful differences among comparable models
<ref:2605.08965#pg15>. The research ultimately suggests that Sensitivity captures meaningful differences among comparable models
<ref:2605.08965#pg15>. The research ultimately suggests that Sensitivity captures meaningful differences among comparable models
<ref:2605.08965#pg15>. The research ultimately suggests that Sensitivity captures meaningful differences among comparable models
<ref:2605.08965#pg15>. The research ultimately suggests that Sensitivity captures meaningful differences among comparable models
<ref:2605.08965#pg15>. The research ultimately suggests that Sensitivity captures meaningful differences among comparable models
<ref:2605.08965#pg15>. The research ultimately suggests that Sensitivity captures meaningful differences among comparable models
<ref:2605.08965#pg15>. The research ultimately suggests that Sensitivity captures meaningful differences among comparable models
<ref:2605.08965#pg15>. The research ultimately suggests that Sensitivity captures meaningful differences among comparable models
<ref:2605.08965#pg15>. The research ultimately suggests that Sensitivity captures meaningful differences among comparable models
<ref:2605.08965#pg15>. The research ultimately suggests that Sensitivity captures meaningful differences among comparable models
<ref:2605.08965#pg15>. The research ultimately suggests that Sensitivity captures meaningful differences among comparable models
<ref:2605.08965#pg15>. The research ultimately suggests that Sensitivity captures meaningful differences among comparable models
<ref:2605.
Improvements for AI systems
-
textbfRationale Supervision via Dual-Axis Prompting for Improved Prediction Performance: The improved system can generate rationales that
improve visual persuasiveness prediction over label-only and unsupervised-rationale baselines
by varying prompts along evidence polarity (support-focused vs. counter-aware) and visual granularity (global vs. local). This addresses the limitation wherethe presence of such elements [message-related visual elements] is treated as sufficient evidence of persuasiveness.
-
textbfFaithfulness-Aware Training Objectives: The system can be trained to produce more reliable explanations by incorporating faithfulness signals, motivated by the finding that
prediction performance alone does not guarantee faithful rationales.
Future objectives shouldjointly encourage task correctness, decision consistency, visual grounding, and behavioral sensitivity
rather than optimizing only for accuracy. -
textbfFaithfulness Evaluation Framework: The system can be rigorously evaluated using a three-dimensional framework covering
rationale-to-decision consistency, rationale-to-image groundedness, and rationale-to-decision sensitivity.
This allows researchers to determine ifstronger predictors do not necessarily produce more faithful rationales,
which is crucial forevaluating whether their rationales faithfully support their decisions.
Sources
- Multimodal Chain-of-Thought Reasoning in Language Models
- Measuring Faithfulness in Chain-of-Thought Reasoning
- Fara-7B: An Efficient Agentic Model for Computer Use
- Towards Cross-Tokenizer Distillation: the Universal Logit Distillation Loss for LLMs
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models