Can Large Language Models Still Explain Themselves? Investigating the Impact of Quantization on Self-Explanations
summary
The gist
The paper investigates whether Large Language Models (LLMs) retain their capacity for self-explanation—generating justifications for their own decisions—when subjected to quantization.
In short
The episode discusses a paper investigating how model quantization—making large language models smaller—affects their ability to self-explain. Hosts conclude that while quantization improves efficiency, it systematically degrades the fidelity and quality of explanations, suggesting explainability must be architecturally protected.
Key concepts
- Quantization
- A process used to make large AI models smaller by reducing the precision of their weights. While this drastically reduces model size and computation requirements, the research shows it can compromise the model's ability to accurately explain its own reasoning.
- Self-Explanation/Faithfulness
- The ability of an LLM to explain *why* it arrived at an answer. Faithfulness measures how accurately that explanation corresponds to the model's actual internal reasoning process, which can be degraded by compression.
- Hybrid Approaches
- A suggested architectural solution where a system keeps its core reasoning module at high precision while quantizing only less critical parts. This aims to preserve explainability without sacrificing efficiency gains.
- eSNLI and HealthFC
- Established, concrete datasets used in the paper's experiments. They ground the abstract concept of 'explanation' in real-world tasks, allowing researchers to quantify performance drops when models are compressed.
Terminology used across episodes
This episode discusses
- Can Large Language Models Still Explain Themselves? Investigating the Impact of Quantization on Self-Explanations · Paper Radio
- Faithfulness vs. Plausibility: On the (Un)Reliability of Explanations from Large Language Models
- Quantifying the Capabilities of LLMs across Scale and Precision
- Counterfactual Evaluation for Explainable AI
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- A Survey on LLM-as-a-Judge
- Investigating Training and Generalization in Faithful Self-Explanations of Large Language Models
- Towards A Rigorous Science of Interpretable Machine Learning
- Can Large Language Models Explain Themselves? A Study of LLM-Generated Self-Explanations
- gpt-oss-120b & gpt-oss-20b Model Card
- Verbosity Tradeoffs and the Impact of Scale on the Faithfulness of LLM Self-Explanations
- Interpreting the Effects of Quantization on LLMs
- Gemma 3 Technical Report
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- Qwen2.5 Technical Report
- LLM-based Affective Text Generation Quality Based on Different Quantization Values
- Anchored Alignment for Self-Explanations Enhancement
- A Survey on Efficient Inference for Large Language Models
The paper
Can Large Language Models Still Explain Themselves? Investigating the Impact of Quantization on Self-Explanations · Read on arXiv
Quantization is widely used to accelerate inference and streamline the deployment of large language models (LLMs), yet its effects on self-explanations (SEs) remain unexplored. SEs, generated by LLMs to justify their own outputs, require reasoning about the model's own decision-making process, a capability that may exhibit particular sensitivity to quantization. As SEs are increasingly relied upon for transparency in high-stakes applications, understanding whether and to what extent quantization degrades SE quality and faithfulness is critical. To address this gap, we examine two types of SEs: natural language explanations (NLEs) and counterfactual examples, generated by LLMs quantized using three common techniques at distinct bit widths. Our findings indicate that quantization typically leads to moderate declines in both SE quality (up to 4.4%) and faithfulness (up to 3.9%). The user study further demonstrates that quantization considerably diminishes both the coherence and trustworthiness of SEs (by up to 8.5%). Compared to smaller models, larger models show limited resilience to quantization in terms of SE quality but maintain more faithfulness. Moreover, no quantization technique consistently excels across task accuracy, SE quality, and faithfulness. Because quantization's impact varies considerably by context and can be sizable in specific cases, we recommend validating SE quality for the intended use case. Despite these sometimes considerable drops, quantization remains an effective compression technique when its impact on SEs is properly validated.
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Can Large Language Models Still Explain Themselves? Investigating the Impact of Quantization on Self-Explanations".
Jane: The paper was written by Qianli Wang, Nils Feldhus, Pepa Atanasova, Fedor Splitt, Simon Ostermann et al. from Quality and Usability Lab at Technical University of Berlin (TU Berlin) and University of Copenhagen and Saarland Informatics Campus and German Research Center for Artificial Intelligence (DFKI) and Centre for European Research in Trusted AI (CERTAIN) and BIFOLD – Berlin Institute for the Foundations of Learning and Data.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Jane: We also have Lu with us today — senior AI researcher at Tsinghua.
Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.
Jane: We also have Lalam with us today — the in-house Large Language Model.
Tom: Alright, let's get started.
Summary/Methods: Tom: Okay, we’ve established the core tension: making models smaller via quantization versus keeping their self-explanation ability intact. The paper's summary section really gets into how they tested this, which is key for us to understand the findings.
Jane: They didn't just run a quick test; they designed specific methodologies to measure *faithfulness*, which is the degree to which the explanation actually corresponds to what the model did internally.
Jane: And that’s where their experiments really dug into whether different types of biases, like those measured in eSNLI and HealthFC, were equally susceptible to this compression artifact.
Meng: I was struck by how they used established datasets like eSNLI and HealthFC—it grounds the abstract concept of 'explanation' in concrete, real-world tasks.
Meng: It moves the conversation away from hypothetical reasoning and toward quantifiable performance drops when we change the model's internal data type.
Tom: So, if I understand correctly, they found that simply applying quantization isn't a one-size-fits-all penalty; some parts of the model are more vulnerable than others.
Lu: Precisely. The research suggests that the mechanisms responsible for complex reasoning—the 'why' behind the answer—are often encoded in weights that are highly sensitive to precision loss.
Lalam: It seems to indicate that explaining oneself might require a dedicated, high-precision module, or perhaps a different architectural approach entirely, rather than just relying on the core compressed LLM body.
Jane: And what was the main pattern they observed when comparing different quantization levels?
Jane: The paper pointed out that while lower quantization levels drastically reduced model size and computation requirements—which is great—they systematically degraded the quality of these self-explanations.
Tom: So, there's a measurable trade-off curve right there: efficiency versus fidelity of explanation.
Lu: It suggests that we can't just treat model compression as a mathematical optimization problem; we have to treat it as an information integrity problem across specialized cognitive functions.
Meng: If the degradation is systematic, then any commercial implementation would need to flag this risk area explicitly: 'High efficiency, but caution regarding explanation fidelity.'
Lalam: This finding reinforces that transparency cannot be an afterthought; it must be baked into the architecture from the ground up if we want trustworthy AI.
Improvements/Future Work: Tom: Building on what you said about the trade-off curve, Jane, let's look at what solutions or improvements the authors put forward in this paper. They can’t just point out a problem; they have to suggest a path forward.
Jane: Right, and it wasn't just about saying 'don't quantize.' The improvements suggested were quite sophisticated, focusing on preserving the *explanation* mechanism while still achieving compression.
Jane: One key suggestion involved hybrid approaches—maybe keeping the core reasoning module at high precision while quantizing less critical parts of the model.
Meng: From an engineering standpoint, that suggests a modular
Paper discussion segment 3: Tom: So, we’ve seen how quantization generally causes a moderate drop in self-explanation quality and faithfulness across the board, but the paper suggests this doesn's a total death blow for AI capability.
Jane: It’s important to understand that this isn't just about speed; it’s about trust. The authors are essentially telling us that if we compress an LLM, we might lose its ability to accurately reflect its own internal reasoning process, which is a huge deal for users who rely on those explanations.
Meng: I’m worried about the practical implications of this finding for deployment at scale. If you can't guarantee the fidelity of the explanation through compression, how do you deploy that model in a high-stakes scenario like medical diagnosis?
Lu: Perhaps we need to challenge the premise that compression must be all-encompassing. Maybe we can design a system where the core reasoning engine is highly optimized, but a parallel, smaller, high-precision module handles only the self-explanation generation.
Lalam: That idea of modular transparency resonates with me; it allows us to build systems where reliability isn's just an outcome but an intentionality—a commitment to how the AI justifies its decisions.
Tom: And you're suggesting a hybrid approach, Lu, which is a major shift from just looking at the compressed weights.
Meng: It’s complex to implement that hybrid architecture, though; you’d have two distinct paths of implementation and managing two sets of operational costs for a single inference task.
Jane: But if we can isolate that core reasoning pathway, we might be able to preserve the "why" without sacrificing the efficiency gains of quantization on other parts of the overall system.
Lu: It opens up an entire space for architectural innovation that moves beyond just optimizing existing models; it demands a rethinking of how cognitive functions are grouped and protected.
Lalam: This could lead to a culture where human users don't just trust the answer, but trust the mechanism behind it, fundamentally changing our relationship with machine intelligence.
Tom: That’s an incredible vision, Lalam, but we still need to look at the hard data—specifically how these models perform when they are trying to explain things like a simple contradiction in a health claim.
Jane: Right, so if we' can move past the theory of what is possible and dive into the specifics of the results for tasks like eSNLI, we can see exactly where these compromises actually bite.
Conclusion: Tom: So, looking back over our conversation today, the big picture takeaway is that model quantization might really hurt how well AI can explain its own reasoning.
Jane: Exactly, Tom; it shows that just shrinking a massive model to make it run on phones or smaller devices could compromise the very ability we want—the ability to talk us through *why* it gave an answer.
Lu: I think this opens up a whole new area of research, though; instead of just accepting degradation, we might need architectural methods that decouple the core knowledge from the explanatory layer.
Meng: But Lu, if the explanatory layer is what slows things down or requires more parameters to remain robust, doesn't that conflict with the fundamental engineering goal of efficiency?
Tom: Right, Meng brings up a crucial point; how do we get explainability without blowing the compute budget on every single deployment?
Lalam: From my perspective, this means that for AI to genuinely advance culture, we can’t just chase performance metrics at the expense of transparency.
Jane: So even if a model performs brilliantly on a benchmark, if it can't coherently explain its steps when you ask it to, that transparency gap might be too wide for widespread trust.
Lu: It really suggests that interpretability isn't just an add-on feature; it’s foundational to the reliability we need in critical systems.
Meng: I agree with Lu; practically speaking, if a medical diagnostic AI can't show its reasoning path when the outcome is wrong, nobody's going to trust it in a real clinic setting.
Lalam: Because of this investigation into "Can Large Language Models Still Explain Themselves? Investigating the Impact of Quantization on Self-Explanations," we see that trust requires explainability, plain and simple.
Tom: It's been a fantastic deep dive, Jane; I feel like we really got to the heart of what making AI accessible means for its trustworthiness.
Jane: It was genuinely fascinating hearing all of your insights today; you all gave such a comprehensive look at the implications beyond just the numbers on the table.
Tom: We're going to have to take a quick break, but when we come back, we've got an entirely different paper to tackle that might challenge how we think about model memory.
More episodes
- 2610.10768-Strategic Investment Decision Making for Value Creation in Energy Transition: A Reinforcement Learning Approach
- 2610.10858-RFChipAgent: Multi-Agentic AI Flow for Analog/RF Chip Design
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
- 2312.01221-Enabling Quantum Natural Language Processing for Hindi Language