An Empirical Study of Counterfactual Self-Explanations in LLMs

arXiv:2609.17119 · cs.CL · Submitted 2026-09-15 · Read on arXiv

cs.CL

Submitted: 2026-09-15

Updated: 2026-09-15

Code: https://github.com/gianniskalyvas/cfself-explanations

License: http://creativecommons.org/licenses/by/4.0/

The gist: Large language models can easily generate explanations for their own outputs, but such self-explanations are not necessarily faithful to the model's behavior.

Terminology

Abstract

Large language models can easily generate explanations for their own outputs, but such self-explanations are not necessarily faithful to the model's behavior. We study this issue through counterfactual self-explanations, where a model minimally edits an input so that its own prediction changes. Across sentiment analysis and natural language inference, we evaluate ten instruction-tuned models from the LLaMA-3 and Qwen-2.5 families, measuring faithfulness, minimality, and alignment with human-annotated rationales. Our results show that model scale is the strongest determinant of explanation quality: larger models are substantially more likely to generate counterfactuals that flip their own predictions and target decision-relevant evidence. In contrast, the rationale-guided condition produces edit-minimal counterfactuals that are also more human-aligned. However, it does not consistently improve faithfulness. Overall, counterfactual self-explanations can provide useful behavioral evidence about model decisions, but their reliability depends strongly on model capacity and should be empirically validated rather than assumed.

Sources

Related papers