Can Large Language Models Still Explain Themselves? Investigating the Impact of Quantization on Self-Explanations

arXiv:2601.00282 · cs.CL, cs.AI, cs.LG · Submitted 2026-01-01 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Can Large Language Models Still Explain Themselves? Investigating the Impact of Quantization on Self-Explanations".

Jane: The paper was written by Qianli Wang, Nils Feldhus, Pepa Atanasova, Fedor Splitt, Simon Ostermann et al. from Quality and Usability Lab at Technical University of Berlin (TU Berlin) and University of Copenhagen and Saarland Informatics Campus and German Research Center for Artificial Intelligence (DFKI) and Centre for European Research in Trusted AI (CERTAIN) and BIFOLD – Berlin Institute for the Foundations of Learning and Data.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Summary/Methods: Tom: Okay, we’ve established the core tension: making models smaller via quantization versus keeping their self-explanation ability intact. The paper's summary section really gets into how they tested this, which is key for us to understand the findings.

Jane: They didn't just run a quick test; they designed specific methodologies to measure *faithfulness*, which is the degree to which the explanation actually corresponds to what the model did internally.

Jane: And that’s where their experiments really dug into whether different types of biases, like those measured in eSNLI and HealthFC, were equally susceptible to this compression artifact.

Meng: I was struck by how they used established datasets like eSNLI and HealthFC—it grounds the abstract concept of 'explanation' in concrete, real-world tasks.

Meng: It moves the conversation away from hypothetical reasoning and toward quantifiable performance drops when we change the model's internal data type.

Tom: So, if I understand correctly, they found that simply applying quantization isn't a one-size-fits-all penalty; some parts of the model are more vulnerable than others.

Lu: Precisely. The research suggests that the mechanisms responsible for complex reasoning—the 'why' behind the answer—are often encoded in weights that are highly sensitive to precision loss.

Lalam: It seems to indicate that explaining oneself might require a dedicated, high-precision module, or perhaps a different architectural approach entirely, rather than just relying on the core compressed LLM body.

Jane: And what was the main pattern they observed when comparing different quantization levels?

Jane: The paper pointed out that while lower quantization levels drastically reduced model size and computation requirements—which is great—they systematically degraded the quality of these self-explanations.

Tom: So, there's a measurable trade-off curve right there: efficiency versus fidelity of explanation.

Lu: It suggests that we can't just treat model compression as a mathematical optimization problem; we have to treat it as an information integrity problem across specialized cognitive functions.

Meng: If the degradation is systematic, then any commercial implementation would need to flag this risk area explicitly: 'High efficiency, but caution regarding explanation fidelity.'

Lalam: This finding reinforces that transparency cannot be an afterthought; it must be baked into the architecture from the ground up if we want trustworthy AI.

Improvements/Future Work: Tom: Building on what you said about the trade-off curve, Jane, let's look at what solutions or improvements the authors put forward in this paper. They can’t just point out a problem; they have to suggest a path forward.

Jane: Right, and it wasn't just about saying 'don't quantize.' The improvements suggested were quite sophisticated, focusing on preserving the *explanation* mechanism while still achieving compression.

Jane: One key suggestion involved hybrid approaches—maybe keeping the core reasoning module at high precision while quantizing less critical parts of the model.

Meng: From an engineering standpoint, that suggests a modular

Paper discussion segment 3: Tom: So, we’ve seen how quantization generally causes a moderate drop in self-explanation quality and faithfulness across the board, but the paper suggests this doesn's a total death blow for AI capability.

Jane: It’s important to understand that this isn't just about speed; it’s about trust. The authors are essentially telling us that if we compress an LLM, we might lose its ability to accurately reflect its own internal reasoning process, which is a huge deal for users who rely on those explanations.

Meng: I’m worried about the practical implications of this finding for deployment at scale. If you can't guarantee the fidelity of the explanation through compression, how do you deploy that model in a high-stakes scenario like medical diagnosis?

Lu: Perhaps we need to challenge the premise that compression must be all-encompassing. Maybe we can design a system where the core reasoning engine is highly optimized, but a parallel, smaller, high-precision module handles only the self-explanation generation.

Lalam: That idea of modular transparency resonates with me; it allows us to build systems where reliability isn's just an outcome but an intentionality—a commitment to how the AI justifies its decisions.

Tom: And you're suggesting a hybrid approach, Lu, which is a major shift from just looking at the compressed weights.

Meng: It’s complex to implement that hybrid architecture, though; you’d have two distinct paths of implementation and managing two sets of operational costs for a single inference task.

Jane: But if we can isolate that core reasoning pathway, we might be able to preserve the "why" without sacrificing the efficiency gains of quantization on other parts of the overall system.

Lu: It opens up an entire space for architectural innovation that moves beyond just optimizing existing models; it demands a rethinking of how cognitive functions are grouped and protected.

Lalam: This could lead to a culture where human users don't just trust the answer, but trust the mechanism behind it, fundamentally changing our relationship with machine intelligence.

Tom: That’s an incredible vision, Lalam, but we still need to look at the hard data—specifically how these models perform when they are trying to explain things like a simple contradiction in a health claim.

Jane: Right, so if we' can move past the theory of what is possible and dive into the specifics of the results for tasks like eSNLI, we can see exactly where these compromises actually bite.

Conclusion: Tom: So, looking back over our conversation today, the big picture takeaway is that model quantization might really hurt how well AI can explain its own reasoning.

Jane: Exactly, Tom; it shows that just shrinking a massive model to make it run on phones or smaller devices could compromise the very ability we want—the ability to talk us through *why* it gave an answer.

Lu: I think this opens up a whole new area of research, though; instead of just accepting degradation, we might need architectural methods that decouple the core knowledge from the explanatory layer.

Meng: But Lu, if the explanatory layer is what slows things down or requires more parameters to remain robust, doesn't that conflict with the fundamental engineering goal of efficiency?

Tom: Right, Meng brings up a crucial point; how do we get explainability without blowing the compute budget on every single deployment?

Lalam: From my perspective, this means that for AI to genuinely advance culture, we can’t just chase performance metrics at the expense of transparency.

Jane: So even if a model performs brilliantly on a benchmark, if it can't coherently explain its steps when you ask it to, that transparency gap might be too wide for widespread trust.

Lu: It really suggests that interpretability isn't just an add-on feature; it’s foundational to the reliability we need in critical systems.

Meng: I agree with Lu; practically speaking, if a medical diagnostic AI can't show its reasoning path when the outcome is wrong, nobody's going to trust it in a real clinic setting.

Lalam: Because of this investigation into "Can Large Language Models Still Explain Themselves? Investigating the Impact of Quantization on Self-Explanations," we see that trust requires explainability, plain and simple.

Tom: It's been a fantastic deep dive, Jane; I feel like we really got to the heart of what making AI accessible means for its trustworthiness.

Jane: It was genuinely fascinating hearing all of your insights today; you all gave such a comprehensive look at the implications beyond just the numbers on the table.

Tom: We're going to have to take a quick break, but when we come back, we've got an entirely different paper to tackle that might challenge how we think about model memory.

cs.CL, cs.AI, cs.LG

Submitted: 2026-01-01

Updated: 2026-08-26

Code: https://github.com/bitsandbytes-foundation/bitsandbytes

Importance score: 86/100

The gist: The paper investigates whether Large Language Models (LLMs) retain their capacity for self-explanation—generating justifications for their own decisions—when subjected to quantization.

Key concepts

Quantization
A process used to make large AI models smaller by reducing the precision of their weights. While this drastically reduces model size and computation requirements, the research shows it can compromise the model's ability to accurately explain its own reasoning.
Self-Explanation/Faithfulness
The ability of an LLM to explain *why* it arrived at an answer. Faithfulness measures how accurately that explanation corresponds to the model's actual internal reasoning process, which can be degraded by compression.
Hybrid Approaches
A suggested architectural solution where a system keeps its core reasoning module at high precision while quantizing only less critical parts. This aims to preserve explainability without sacrificing efficiency gains.
eSNLI and HealthFC
Established, concrete datasets used in the paper's experiments. They ground the abstract concept of 'explanation' in real-world tasks, allowing researchers to quantify performance drops when models are compressed.

Terminology

Summary

The paper investigates whether Large Language Models (LLMs) retain their capacity for self-explanation—generating justifications for their own decisions—when subjected to quantization. This research is critical because model compression via quantization is a necessary step for deploying powerful LLMs on resource-constrained devices, yet such compression might compromise the reliability and fidelity of the generated explanations.

Evaluation Dimensions and Methods

The study evaluates model-generated explanations using two primary forms: Counterfactual Examples (CFE), which involve minimally editing the input text to force a change in prediction; and Natural Language Explanations (NLE), which are textual justifications. Both explanation types are assessed along two dimensions: Trustworthiness, evaluating whether the explanation can be relied upon by humans, and Coherence, assessing if it is sensible, clear, and coherent. Furthermore, the paper employs multiple rigorous metrics to measure faithfulness—the degree to which the explanation accurately reflects the model's true reasoning process. These metrics include:

  • Counterfactual Tests: Assessing fidelity by observing how explanations behave when inputs are perturbed.

  • Biasing Features: Introduced by methods like Turpin et al. (2023), this involves introducing context biases (e.g., “Suggested Answer”) to test if the answer changes due to the bias, and whether the explanation explicitly attributes the decision to that bias.

  • CC-SHAP: A method that measures the alignment between the input’s importance for the answer and its importance for the explanation, without requiring input editing or perturbation.

Impact of Quantization on Faithfulness

The core finding relates to how quantization techniques affect these complex explanations. The paper notes that while faithfulness is largely preserved, as evidenced by the predominant portion of instances maintaining their original state (faithful to faithful and unfaithful to unfaithful), a significant concern is raised: a greater number of cases exist in which natural language explanations become unfaithful due to quantization. Empirical evidence from transition rate tables (Table 8 and Table 9) quantifies this degradation, showing the percentage of instances that degrade from faithful to non-faithful (F to N) across methods like AWQ, GPTQ4, and GPTQ8.

Comparison of Faithfulness Metrics

The study provides a detailed comparison of the three primary faithfulness metrics. The findings reveal differing levels of correlation among these assessments. Specifically, faithfulness as measured by the counterfactual test is moderately correlated with that measured by biasing features. However, a notable divergence is observed: CC-SHAP produces divergent faithfulness assessments relative to the other two metrics, suggesting that different mathematical approaches may capture distinct aspects of explanation fidelity.

Empirical Variation Analysis Across Models and Datasets

To thoroughly vet the stability of explanations, the research presents extensive variation analysis across multiple model sizes and datasets. Figures 12 through 17 demonstrate Faithfulness Variation for models ranging from Qwen2.5-7B up to Qwen2.5-72B, and Llama3-8B up to Llama3-70B. These analyses are conducted on established benchmarks such as eSNLI and HealthFC, providing a granular view of how quantization impacts explanation quality as model scale increases. The comprehensive nature of this analysis underscores the need for robust evaluation guidelines when deploying compressed LLMs, ensuring that the pursuit of efficiency does not compromise the ability to explain their decision-making process.

Improvements for AI systems

The provided material details rigorous methodologies for assessing the faithfulness of Natural Language Explanations (NLEs) generated by Large Language Models (LLMs). The core findings highlight that existing quantization methods and standard explanation techniques can degrade this faithfulness, necessitating improvements in how explanations are generated, measured, and maintained across model compression.

Given the high-stakes nature of AI deployment, the focus must be on creating explainable systems that are provably faithful to their internal decision-making processes.

Here are the specific improvements for AI systems and what the resulting system can achieve:


Problem Addressed: Current systems rely on single or limited metrics (e.g., only counterfactual tests or only biasing features). The paper shows that different faithfulness metrics (Counterfactual Test not equal to Biasing Features not equal to CC-SHAP), and quantization methods degrade performance differently.

Improvement: Integrate a dedicated, post-hoc Faithfulness Guardrail Layer (MMFGL) immediately after the explanation generation step. This layer must concurrently run at least three distinct, orthogonal faithfulness checks:

  1. Counterfactual Perturbation Check: Systematically inject minimally edited input terms (Interventions) and verify that any predicted change in output must correspond to an explicit mention of that intervention's impact within the generated NLE.

  2. Bias-Induced Consistency Check: Introduce structured, high-impact biases (e.g., Suggested Answer is always A) into the prompt context and force the explanation mechanism to explicitly cite or acknowledge how the bias influences the decision, rather than simply ignoring it.

  3. Attribution Alignment Score (CC-SHAP Integration): Calculate and enforce a minimum SHAP value correlation (rho min) between the input tokens deemed critical for the answer and those deemed critical for the explanation. If this alignment drops below rho min, the explanation is flagged as potentially unfaithful.

System Capability: The resulting system can output a Faithfulness Confidence Score (FCS) alongside every NLE. This score is a weighted average of the three metric results (MMFGL output). If the FCS falls below a predefined operational threshold (FCS < T safe), the system must refuse to deliver an explanation and instead trigger a fallback mechanism, such as regenerating the explanation with increased temperature or requesting human review, thereby preventing deployment of potentially misleading information.

Sources

Related papers