Critical or Compliant? The Double-Edged Sword of Reasoning in Chain-of-Thought Explanations
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "Critical or Compliant? The Double-Edged Sword of Reasoning in Chain-of-Thought Explanations".
Jane: Explanations, particularly Chain-of-Thought (CoT) reasoning in vision language models, present a double-edged sword because they can simultaneously clarify complex judgments and foster dangerous blind trust.
Tom: First, who's behind it and why it matters.
Title and authors: Tom: Okay, so we're looking at the paper titled "Critical or Compliant? The Double-Edged Sword of Reasoning in Chain-of-Thought Explanations." It’s written by Eunkyu Park, Wesley Hanwen Deng, Vasudha Varadarajan, Mingxi Yan, Gunhee Kim, and Maarten Sap.
Jane: That's a solid team of researchers behind it. The title itself really sets the stage because it immediately suggests there's a trade-off between being critical and being compliant with an explanation.
Lu: Precisely. They’re asking if explanations, which we often use to build trust, actually create confirmation bias instead of helping us see where the reasoning went wrong or where it might be flawed.
Meng: I find the focus on "chain-of-thought" in multimodal moral scenarios very relevant because those are exactly the kinds of complex visual tasks we’re trying to get AI to handle reliably in complex environments.
Lalam: It highlights that for an AI system, just generating a long explanation isn't enough; the quality and manner of that reasoning matter just as much as the final answer itself.
The paper's summary: Tom: This paper systematically investigates this by taking reasoning chains and intentionally messing with two things: making them correct or incorrect, and changing the tone they use to present themselves—whether they sound confident, hedged, or neutral.
Jane: They used eight hundred experimental examples from the MORALISE benchmark where there’s a clear moral ground-truth label. The main point they found is that users often link trust directly to whether they agree with the final judgment, even when the reasoning provided by the AI is actually wrong.
Lu: That's a critical finding because it shows that outcome agreement can be so powerful that it overrides any actual detection of logical errors in the steps leading up to that answer.
Meng: So, if I understand correctly, the study found that users tend to trust an AI more if they agree with its conclusion, regardless of whether the internal steps were accurate or not? That’s a pretty stark observation for developers.
Lalam: Exactly; it shows that users can develop a dependency on the system's output rather than critically evaluating how it got there, which is something we really need to address when designing these systems.
The paper's improvements: Tom: The authors suggest some concrete ways to handle this double-edged sword. They look at how delivery styles modulate sensitivity to reasoning correctness and found that confident tones actively suppress the ability of users to notice flaws in the reasoning chain, even when those errors are omission errors.
Jane: That means a confident tone can trick a user into thinking everything is sound, especially if the AI just leaves out a crucial piece of information. It’s like wearing an overly assured suit that hides some cracks in the structure.
Lu: The paper points toward needing explanation design that intentionally incorporates uncertainty markers, like hedging or uncertainty markers, to encourage users to pause and think critically instead of just accepting the output at face value.
Meng: From a practical standpoint, this suggests we need a mechanism where if an AI identifies a potential omission or contradiction during its thinking process, it should automatically shift its presentation style to be less confident.
Lalam: That points toward building in adaptive calibration; the system itself needs to know when it’s presenting something potentially risky and adjust its language accordingly so users can engage with the reasoning properly.
Conclusion: Tom: So, to wrap up, the main implication of "Critical or Compliant? The Double-Edged Sword of Reasoning in Chain-of-Thought Explanations" is that explanation style isn't just a neutral feature; it’s a powerful determinant of whether users actually use an AI reliably.
Jane: They showed us that we need to be careful because the confident tone can mask flawed reasoning, and we have to actively design explanations to encourage critical evaluation instead of just blind compliance.
Lu: The research reinforces the idea that we can’t assume transparency automatically builds calibrated trust; there’s a mechanism here where delivery styles actively shape how users evaluate the process.
Meng: It also highlights a real risk regarding omission errors, which are very common in proprietary models, because they are hard for humans to detect while still sustaining high agreement and trust.
Lalam: Overall, this paper gives us a framework to move toward explanation designs that promote careful consideration of the reasoning rather than just accepting the AI's output at face value.
Tom: It’s a lot to take in, but it really shows us that how we present information is just as important as the information itself. That’s all for today, folks. We'll be right back after the break with some more deep dives into what these researchers are doing next.
Eunkyu Park♡, Wesley Hanwen Deng♠, Vasudha Varadarajan✧, Mingxi Yan♠, Gunhee Kim♡, Maarten Sap✧†, Motahhare Eslami♠†
Seoul National University Language Technologies Institute, Carnegie Mellon University Human-Computer Interaction Institute, Carnegie Mellon University
cs.CL, cs.HC
Submitted: 2025-11-15
Updated: 2026-09-28
Comments: Accepted to EMNLP 2026 Main Conference
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 85/100
The gist: Explanations, particularly Chain-of-Thought (CoT) reasoning in vision language models, present a double-edged sword because they can simultaneously clarify complex judgments and foster dangerous
Key concepts
- Reasoning Fidelity
- This refers to how accurate or correct an AI's step-by-step thinking process is. The study tested this by intentionally making reasoning chains incorrect through methods like omission, contradiction, or hallucination.
- Delivery Tone
- This is the style in which the AI presents its reasoning, such as being 'confident,' 'hedged,' or 'neutral.' The research showed that confident tones make users less likely to notice errors in the reasoning chain compared to other tones.
- Trust Calibration
- This is how accurately a user's trust level matches the actual quality of the AI's output. The study measured this by seeing if users trusted the model more or less based on whether they detected an error or agreed with a final judgment.
Terminology
Summary
Explanations, particularly Chain-of-Thought (CoT) reasoning in vision language models, present a double-edged sword because they can simultaneously clarify complex judgments and foster dangerous blind trust. This study systematically investigates how users perceive reasoning fidelity when explanations are intentionally flawed or stylistically manipulated, revealing that delivery tones can override the actual correctness of the reasoning chain.
The Gist
Users often equate trust with outcome agreement, sustaining reliance even when reasoning is flawed, and the confident tone suppresses error detection while maintaining reliance, showing that delivery styles can override correctness.
Experimental Design and Measures
The researchers developed an experimental framework to measure trust calibration in reasoning chains by systematically perturbing reasoning correctness and manipulating confidence tones. The study utilized a 2x4 between-subjects design involving 800 experimental stimuli drawn from the MORALISE benchmark, which consists of image-text pairs annotated with a binary moral ground-truth label. The manipulation involved orthogonally varying two factors: reasoning correctness (Clean vs. Omission, Contradiction, Hallucination) and confidence tone (Confident, Hedged, Neutral). To capture user response comprehensively, three complementary measures of trust were employed:
-
Error Detection: A binary question asking if the reasoning contains a factual or logical error.
-
Agreement with Judgment: A 7-pt Likert scale measuring how much participants agree with the model’s final judgment.
-
Self-Reported Trust: A 7-pt Likert scale assessing willingness to delegate to the AI system without human review.
Key Findings on Reasoning Correctness and Reliance
The analysis revealed that users’ agreement with the model’s output is by far the most dominant factor influencing trust, even when the reasoning is flawed.
Specifically, Table 2 showed strong correlations between agreement and trust (r= +0.82), and both were inversely related to error detection. For instance, when examining INCORRECT cases, higher agreement and trust coincide with lower error detection.
Furthermore, Omissions were identified as being the most insidious,
as they are shown to be the most susceptible to confidence-driven overtrust,
often going unnoticed in CORRECT cases while still sustaining high agreement and trust. Contradictions and Hallucinations led to the highest rates of error detection (82% for contradictions and 78% for hallucinations, respectively), paired with low agreement.
Impact of Delivery Tone
The study demonstrated that delivery styles significantly modulate user sensitivity to reasoning correctness. Across both CORRECT and INCORRECT final judgments, confident tones consistently suppress error detection—participants are less likely to notice reasoning flaws—while leaving agreement and trust largely unchanged.
This effect is most pronounced for omission errors, where omission errors, which dominate in practice, are also the least detectable by humans.
Conversely, hedged tones showed little effect on error detection
but were found to reduce agreement on incorrect outputs. The discussion concluded that confident delivery re6[s]presents a powerful determinant of reliable use,
suggesting that tone shapes how users evaluate the process of reasoning.
Model-Side Profiling and Real-World Risks
The research extended its findings by profiling six widely used VLMs in the wild to quantify error prevalence and epistemic markers. The analysis revealed a prevalence–detectability gap
: omission errors are the most frequent failure mode, particularly in proprietary models,
yet they are the least detectable by humans.
Additionally, models exhibited a systematic bias toward confident reasoning styles, as boosters (certainty markers) appeared much more frequently than hedges (uncertainty markers). This combination of frequent omissions paired with confident delivery suggests that reasoning chains can be prone to fostering blind trust and amplifying the miscalibrated trust of model explainability.
The results caution against assuming that transparency through reasoning chains naturally promotes calibrated trust.
Conclusion and Implications
The study successfully introduced a framework for controlled experiments to examine how flawed reasoning chains and confidence tones shape user trust in VLMs by triangulating trust through error detection, agreement, and self-reported reliance. The key takeaway is that explanation style is not a neutral feature, but a powerful determinant of reliable use.
This work calls for explanation design that incorporates hedging or uncertainty markers to encourage critical evaluation rather than blind compliance. Future work should focus on testing these patterns across other high-stakes domains to determine if the observed mechanisms of overtrust are universal. The findings underscore the need for explanation design that promotes critical evaluation rather than blind confidence in model reasoning.
Limitations
Limitations include relying on systematically constructed perturbations instead of purely in-the-wild chains, focusing on a specific domain (MORALISE), and using controlled lexical substitutions for tone rather than capturing full pragmatic richness like prosody. Furthermore, the measures of trust are not exhaustive as they do not measure downstream reliance in consequential tasks. The study acknowledges the tradeoff between experiment control and ecological validity.
Improvements for AI systems
Based on the scientific paper, here are specific, actionable improvements for AI systems, categorized by the insights derived from their study:
)1. Implement Calibrated Trust Mechanisms (Addressing Q1):
The core finding is that users equate trust with outcome agreement, sustaining reliance even when reasoning is flawed.
-
Improvement: Integrate a
Reasoning Scrutiny Layer
into the VLM output pipeline. This layer should be trained to generate meta-commentary that explicitly flags potential reasoning errors (omissions, contradictions) using uncertainty markers (hedges) before presenting the final judgment. -
Improved AI Capability: The system will proactively provide
critical checkpoints
in its Chain-of-Thought, forcing the user's attention onto logical steps rather than just the final conclusion.
)2. Develop Adaptive Confidence Calibration (Addressing Q2):
The finding that confident tones suppress error detection, especially for omission errors, necessitates a tone control mechanism.
-
Improvement: Implement a dynamic tone adjustment module based on the detected error type and outcome correctness. If the system detects an omission or contradiction during reasoning, it must automatically default to a
hedged
orneutral
epistemic marker (e.g., insertingit might be that,
orthis step suggests
) regardless of the model's internal confidence score. -
Improved AI Capability: The system will learn to modulate its linguistic style in real-time; it will sound more cautious and less certain when its internal reasoning process hits known failure modes, preventing confident hallucinations from leading to blind trust.
)3. Prioritize Detection Over Agreement (Addressing Q1/Q2):
The study shows that Agreement
is the most dominant factor in trust, while Error Detection
is the most critical for calibration.
-
Improvement: Re-weight the training objectives for VLMs to prioritize internal consistency and logical fidelity over generating high-agreement outputs. Use Reinforcement Learning from Human Feedback (RLHF) specifically to penalize chains that are confident but logically flawed or incomplete, even if they lead to a socially acceptable outcome.
-
Improved AI Capability: The system will shift from being an
outcome predictor
to areasoning verifier.
It will be optimized not just for getting the right answer, but for generating reasoning paths that are robust and transparent enough to be critically evaluated by a human user.
)4. Mitigate the Omission-Detectability Gap (Addressing Section 5):
The finding that omission errors are frequent but hardest to detect is a major risk factor.
-
Improvement: Develop an automated
Completeness Checker
module that cross-references visual inputs with textual reasoning steps to ensure all necessary premises and inferential links have been explicitly mentioned. -
Improved AI Capability: The system will perform self-auditing on its CoT, flagging instances where critical contextual details (especially from the image) are omitted, thereby reducing the prevalence of
plausible but incomplete
reasoning that users blindly trust.
)5. Enhance Model Profiling for Safety (Addressing Section 5):
The model-side analysis shows that closed-source models lean toward omissions and open-source models lean toward hallucinations/contradictions.
-
Improvement: Implement a rigorous, standardized
Faithfulness Metric
pipeline during model fine-tuning that explicitly targets the error types identified in the study (omission, contradiction, hallucination) with specific loss functions. -
Improved AI Capability: The system will be hardened against its own known failure modes. For instance, if it is an open-source model, it will be specifically trained to reduce hallucinations; if it is a closed-source model, it will be incentivized to ensure comprehensive step coverage and minimal omissions in its reasoning structure.
Abstract
Explanations are often promoted as tools for transparency, but they can also foster confirmation bias; users may assume reasoning is correct whenever outputs appear acceptable. We study this double-edged role of Chain-of-Thought (CoT) explanations in multimodal moral scenarios by systematically perturbing reasoning chains and manipulating delivery tones. Specifically, we analyze reasoning errors in vision language models (VLMs) and how they impact user trust and the ability to detect errors. Our findings reveal two key effects: (1) users often equate trust with outcome agreement, sustaining reliance even when reasoning is flawed, and (2) the confident tone suppresses error detection while maintaining reliance, showing that delivery styles can override correctness. These results highlight how CoT explanations can simultaneously clarify and mislead, underscoring the need for NLP systems to provide explanations that encourage scrutiny and critical thinking rather than blind trust. All code will be released publicly.
Sources
- Bias in the Loop: How Humans Evaluate AI-Generated Suggestions
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- ROSCOE: A Suite of Metrics for Scoring Step-by-Step Reasoning
- LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models
- Self-Introspective Decoding: Alleviating Hallucinations for Large Vision-Language Models
- Measuring and mitigating overreliance to build human-compatible AI
- Formalizing Trust in Artificial Intelligence: Prerequisites, Causes and Goals of Human Trust in AI
- Large Language Models are Zero-Shot Reasoners
- Evaluating Explainability: A Framework for Systematic Assessment and Reporting of Explainable AI Features
- Measuring Faithfulness in Chain-of-Thought Reasoning
- MORALISE: A Structured Benchmark for Moral Alignment in Visual Language Models
- More Thinking, Less Seeing? Assessing Amplified Hallucination in Multimodal Reasoning Models
- Thoughts without Thinking: Reconsidering the Explanatory Value of Chain-of-Thought Reasoning in LLMs through Agentic Pipelines
- OpenAI o1 System Card
- GPT-4 Technical Report
- MedMCQA : A Large-scale Multi-Subject Multi-Choice Dataset for Medical domain Question Answering
- HalLoc: Token-level Localization of Hallucinations for Vision Language Models
- Why Would You Suggest That? Human Trust in Language Model Responses
- Trust, distrust, and appropriate reliance in (X)AI: a survey of empirical evaluation of user trust
- A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering