Does a model's stated reason for rejecting a candidate do any work?
cs.CL, cs.AI, cs.LG
Submitted: 2026-09-24
Updated: 2026-09-24
Terminology
Sources
- Explanation-Driven Counterfactual Testing for Faithfulness in Vision-Language Model Explanations
- Measuring Faithfulness in Chain-of-Thought Reasoning
- A Positive Case for Faithfulness: LLM Self-Explanations Help Predict Model Behavior
- Walk the Talk? Measuring the Faithfulness of Large Language Model Explanations
- The Disagreement Problem in Explainable Machine Learning: A Practitioner's Perspective
- "My Answer is C": First-Token Probabilities Do Not Match Text Answers in Instruction-Tuned Language Models
- Look at the Text: Instruction-Tuned Language Models are More Robust Multiple Choice Selectors than You Think
- The Probabilities Also Matter: A More Faithful Metric for Faithfulness of Free-Text Explanations in Large Language Models
- A Causal Lens for Evaluating Faithfulness Metrics
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- The Out-of-Distribution Problem in Explainability and Search Methods for Feature Importance Explanations
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering