Label Over Logic? How Source Cues Bias Human Fallacy Judgments More Than LLMs
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Label Over Logic? How Source Cues Bias Human Fallacy Judgments More Than LLMs".
Jane: The paper was written by Mahjabin Nahar, Nafis Irtiza Tripto, Aiping Xiong, Ting-Hao ‘Kenneth’ Huang and Dongwon Lee from The Pennsylvania State University Department of Computer and Information Science, The Pennsylvania State University Institute for Human-Centered AI, The Pennsylvania State University School of Social Sciences and Arts & Sciences, The Pennsylvania State University College of Engineering and Computing and The Pennsylvania State University.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Summary of Key Findings: Tom: Moving past that initial bias, we need to look at what happened when judging the logical accuracy itself. The core summary shows a huge difference between human judgment and how LLMs handle this task.
Jane: The data clearly demonstrates that while humans are highly susceptible to source cues, LLMs remain remarkably stable across those same different source labels. This is a massive point of contrast to our own behavior, showing us where the gap in performance lies.
Lu: I think that stability reflects the programmed nature of LLMs—they execute logic based on patterns and rules rather than social heuristics. They are performing the task without assigning any social weight to the author's identity, which is a huge difference from our messy human cognitive processes.
Meng: My concern is how we can use this stable behavior in practical systems; since LLMs aren't swayed by source labels, they could be a reliable tool for first-pass evaluation, helping us filter out the inconsistencies introduced by human bias.
Lalam: We can leverage that consistency to create more culturally objective evaluations of information quality, moving away from favoring certain types of authorship toward a consistent standard of logical rigor.
Tom: The researchers also looked at how we felt about our own judgment, and while humans were consistently confident in their scores across all source conditions, the LLMs had higher confidence scores overall. This is interesting because it suggests that LLMs might be more certain about their conclusions than we are when they make a judgment.
Jane: It's a bit of a problem if the AI is highly confident but wrong, isn't it? The fact that both groups maintained high levels of certainty means we have to be careful about relying on confidence as an indicator of quality, especially since the LLMs might be overconfident.
Practical Improvements and Solutions: Tom: This paper suggests several paths forward for us, especially given how common AI-assisted content is now; since humans are so vulnerable to source cues, we need systems that actively encourage content-focused evaluation over misleading cues about who wrote it.
Jane: We also saw interesting insights into the various forms of collaboration; for example, people interpreted "human with AI assistance" as human ideas polished by AI. This distinction is vital because our understanding of hybrid authorship isn't uniform and we need to address that cultural gap.
Lu: The idea that source labels can unintentionally encourage us to take shortcuts—relying on who said it rather than what was actually said—is a major area for future research, particularly when we look at more complex, real-world contexts where these cues are misleading.
Meng: I think the practical implication is designing interfaces that aren't just showing a label but actively guiding users toward logical scrutiny. We need to present the fallacy alongside clear explanations of why it is flawed so that the source doesn't distract from the logical error.
Lalam: We should also look at how our cultural expectations of "human agency" play into this; we value human effort and experience, and that bias needs to be addressed in how we structure information to align with objective goals.
Tom: The researchers are proposing a collaborative model where human-LLM partnerships can mitigate the vulnerabilities unique to either humans or LLMs. We can utilize the strengths of each other's approach to create a safer environment for everyone.
Jane: It's a sophisticated idea—using AI not as a replacement for human judgment, but as a complementary check against source-driven biases that we are all susceptible to, even if we feel very confident in our own judgments.
Deeper Look at Complementarity: Tom: This paper really shows us that human and LLM biases aren't just different; they are complementary, which is the next big step in how we interpret this research. The data suggests a specific synergy between the two systems.
Jane: It’s fascinating that when humans were susceptible to certain errors like hasty generalization, the LLMs were more influenced by things like appeal to tradition. This difference means we aren't seeing a simple substitution of one thing for another; we are seeing distinct strengths.
Lu: I think this complementary nature allows us to build systems that address specific types of flawed reasoning—one model might catch the emotional bias, while the other catches the formal logical error. It’s a powerful division of labor.
Meng: From an operational standpoint, this suggests we can run parallel checks in a pipeline; because humans are susceptible to certain errors and LLMs have their own set of patterns, we don't need one replacement for another.
Lalam: We can use this complementarity to build systems that understand both cultural biases and logical rigor, which is essential for creating content that resonates with our human experience while maintaining objective truth.
Tom: It’s a really important insight: the fact that human and LLM biases are complementary means we have a roadmap for how to support each other in this complex process.
Final Wrap-up and Conclusion: Tom: So, looking at the whole picture, the paper "Label Over Logic? How Source Cues Bias Human Fallacy Judgments More Than LLMs" shows us that source perception is a real bias for humans when evaluating logical reasoning. We have seen how we trust human-labeled content more easily, even with errors.
Jane: And we also learned that while LLMs are largely stable and less biased by the author's identity, their high confidence level doesn't necessarily mean they are correct, which is a crucial distinction for us to understand in a world full of AI-assisted content.
Lu: The paper highlights that human and LLM biases are complementary; one might be susceptible to hasty generalizations while the other is more influenced by tradition—these differences can actually be our strength when we work together.
Meng: This means we don't need AI to replace human judgment, but rather, we need robust pipelines that support both as a check against source-driven bias and systematic errors in our digital spaces.
Lalam: By understanding these vulnerabilities, the paper suggests that we are building systems capable of supporting critical thinking and understanding both cultural biases and logical rigor for everyone.
Tom: It’s a really important discussion for our listeners, reminding us that how we see who is speaking often matters as much as what they are saying. We'll be wrapping up this segment by acknowledging the authors, Mahjabin Nahar and his team.
Jane: Thank you to the authors for "Label Over Logic? How Source Cues Bias Human Fallacy Judgments More Than LLMs" research; it gives us a lot to think about as we move toward more AI-mediated online environments.
Lu: It’s truly amazing work, showing that our own cognitive biases are mirrored in the way different models process information.
Meng: I just hope that this research allows us to build better tools for content moderation and public understanding.
Lalam: The future needs to be a place where human experience is valued alongside logical consistency, guided by what these findings teach us about source perception.
The Pennsylvania State University Department of Computer and Information Science, The Pennsylvania State University Institute for Human-Centered AI, The Pennsylvania State University School of Social Sciences and Arts & Sciences, The Pennsylvania State University College of Engineering and Computing · The Pennsylvania State University
cs.HC, cs.AI
Submitted: 2026-05-28
Updated: 2026-09-04
Comments: to appear in EMNLP 2026
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 85/100
The gist: This paper investigates how perceived authorship—the source label attached to content—influences human judgment regarding logical reasoning, specifically examining whether source cues bias human
Key concepts
- Source Cues Bias
- This refers to the tendency for humans to judge content based on who is the author or source, rather than judging the actual logical content. The paper shows humans rely on these cues, which can lead them to take shortcuts and accept flawed reasoning.
- LLM Stability
- Large Language Models (LLMs) are remarkably stable when presented with different source labels. They process information based on patterns and rules rather than social heuristics or author identity, making them less susceptible to the biases that affect human judgment.
- Complementary Biases
- The research suggests that human and LLM biases are not just different but complementary. One system might be prone to hasty generalization, while the other is influenced by tradition. This synergy allows each other to catch errors in a collaborative system.
Terminology
Summary
This paper investigates how perceived authorship—the source label attached to content—influences human judgment regarding logical reasoning, specifically examining whether source cues bias human evaluation more than Large Language Models (LLMs) when assessing content containing logical fallacies. As AI-generated and AI-assisted content becomes increasingly prevalent,
understanding this source perception is critical for designing robust human-AI collaboration pipelines and mitigating potential biases in online discourse.
Methodology of the Study
The study employed a mixed-design experiment involving 505 participants, who were randomly assigned to one of five source conditions:
-
Human (H)
-
AI (A)
-
Human with AI assistance (H+AI)
-
AI with human assistance (A+H)
-
Control/No disclosure
Participants evaluated 32 news title–comment pairs from the CoCoLoFa dataset, which contained eight common fallacy types. The judgments were compared against three LLMs: GPT-5.2, Gemini 2.5 Flash, and Claude Sonnet 4.5, who were evaluated under the exact same source conditions using identical stimuli and rating scales (perceived logical accuracy and confidence).
Human Susceptibility to Source Bias
The findings revealed that human evaluators were significantly susceptible to source-label bias. Specifically, perceived logical accuracy was strongly influenced by both condition and fallacy presence, resulting in a significant Condition times Fallacy interaction (p <.001). Participants rated comments containing logical fallacies as significantly less accurate than non-fallacious ones across all conditions. However, the penalty for fallacies was substantially smaller in the Human (M = 3.35) and Human+AI (M = 3.37) conditions compared to the Control (M = 2.40), AI (M = 2.53), and AI+Human (M = 2.59) conditions, indicating a smaller gap between fallacies
when human involvement was perceived as primary.
LLMs: Stability and Comparative Strictness
In contrast, LLMs demonstrated remarkable stability across source labels, with no significant pairwise differences in their judgments for either fallacious or non-fallacious comments. While LLMs were resistant to source-label bias, they differed significantly from humans in evaluative strictness. Humans assigned higher ratings to certain fallacy types (e.g., hasty generalization), whereas LLMs tended to assign higher ratings to appeal to nature and appeal to tradition. LLMs were also generally stricter than humans,
especially when evaluating flawed reasoning, suggesting that their source-agnostic evaluation patterns are robust across prompting strategies.
Implications for Human-AI Collaboration
The study suggests that human judgment is highly sensitive to perceived authorship, particularly favoring content labeled as written by a human or human with AI assistance, which receives higher trust and evaluation ratings.
These results indicate that source perception functions as a heuristic cue whose influence on human judgment is highly context-dependent and often contradictory across settings.
Furthermore, the differing error patterns between the two groups—humans being susceptible to specific fallacies while LLMs are influenced by others—suggest that humans and LLMs may have complementary weaknesses,
highlighting the value of human–LLM collaboration for decision-making in AI-mediated environments.
Improvements for AI systems
Based on the findings of Nahar et al., the primary vulnerability lies in human heuristic reliance on source cues, while LLMs exhibit model-specific biases. To improve AI systems and support effective human-AI collaboration, I propose implementing the following specific architectural and operational improvements:
Improvement: Integrate a standardized, structured evaluation interface that forces users into a high-cognitive load state when assessing content.
- Mechanism: When presented with an article/comment, the system automatically prompts the user to complete
Pre-Judgment Scaffolding
before rating accuracy or trust. This scaffolding requires users to explicitly answer:
-
What is the core claim of this statement?
(Identifying the central assertion.) -
What evidence is provided for this claim?
(Separating claims from support.) -
Does the relationship between the evidence and the claim rely on emotional appeal, established custom, or a general assumption about nature?
(Explicitly identifying potential fallacies.)
- What the Improved System Can Do: This mode forces users away from rapid, heuristic judgments (which were observed in Human/Human+AI conditions) and increases cognitive engagement. By forcing explicit identification of logical flaws before rating accuracy, it reduces the tendency to assign high trust and high accuracy merely because a human or human-assisted source is perceived as credible.
Sources
- Penalizing Transparency? How AI Disclosure and Author Demographics Shape Human and AI Judgments About Writing
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- ExpertPrompting: Instructing Large Language Models to be Distinguished Experts
- OpenAI GPT-5 System Card
- Large Language Models are overconfident and amplify human bias
- Label Effects: Shared Heuristic Reliance in Trust Assessment by Humans and LLM-as-a-Judge
Related papers
- EduGage: A Multimodal Dataset and Benchmark for Sensor-Based Momentary Assessment of Engagement in Self-Guided Video Learning
- EvoDesign: Agentic Editable Diagram Creation via Design Expertise Evolution
- HAGI++: Head-Assisted Gaze Imputation and Generation
- Linking Behaviour and Perception to Evaluate Meaningful Human Control over Partially Automated Driving
- Review of Explainable Decision Support and Adaptive Human-Machine Interfaces for Automation Transparency in Maritime Autonomous Surface Ships
- Towards Cognitive Process-Aware Proactive Writing Support