When to Call an Apple Red: Humans Follow Introspective Rules, VLMs Don't
Jonathan Nemitz, Carsten Eickhoff, Junyi Jessy Li, Kyle Mahowald, Michal Golovanevsky, William Rudman
cs.CL, cs.AI, cs.CV
Submitted: 2026-08-19
Updated: 2026-08-20
Code: https://github.com/wrudman/when
Terminology
Sources
- Faithfulness vs. Plausibility: On the (Un)Reliability of Explanations from Large Language Models
- A Closer Look at Bias and Chain-of-Thought Faithfulness of Large (Vision) Language Models
- Reasoning Models Don't Always Say What They Think
- Measuring and Improving Chain-of-Thought Reasoning in Vision-Language Models
- Explanation-Driven Counterfactual Testing for Faithfulness in Vision-Language Model Explanations
- Pixels Versus Priors: Controlling Knowledge Priors in Vision-Language Models through Visual Counterfacts
- Alignment faking in large language models
- Counterfactual Simulation Training for Chain-of-Thought Faithfulness
- Measuring Faithfulness in Chain-of-Thought Reasoning
- More Thinking, Less Seeing? Assessing Amplified Hallucination in Multimodal Reasoning Models
- Seeing but Not Believing: Probing the Disconnect Between Visual Attention and Answer Correctness in VLMs
- On the Faithfulness of Visual Thinking: Measurement and Enhancement
- SPD-Faith Bench: Diagnosing and Improving Faithfulness in Chain-of-Thought for Multimodal Large Language Models
- Do Vision & Language Decoders use Images and Text equally? How Self-consistent are their Explanations?
- FaithCoT-Bench: Benchmarking Instance-Level Faithfulness of Chain-of-Thought Reasoning
- Balancing Faithfulness and Performance in Reasoning via Multi-Listener Soft Execution
- Privileged Self-Access Matters for Introspection in AI
- Multi-Agent Simulator Drives Language Models for Legal Intensive Interaction
- On Robustness and Chain-of-Thought Consistency of RL-Finetuned VLMs
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering