Explainability in Practice: A Survey of Explainable NLP Across Various Domains

summary

Video file (mp4)

In short

The episode surveys 'Explainability in Practice,' discussing how explainable NLP must be tailored to specific domains rather than using a one-size-fits-all approach. Hosts discuss the paper's seven domains, emphasizing that explanations must meet the unique needs of stakeholders (e.g., doctors vs. chatbot users).

Key concepts

Domain-First Approach
Instead of developing explanation methods and then finding problems for them, this approach starts with a specific problem or domain (like medicine or finance) and determines what kind of explanation is actually needed by the stakeholders.
Two-Tier Evaluation Protocol
The paper proposes a structured method for evaluating explainability. Tier one requires reporting common technical metrics (fidelity, faithfulness), while tier two demands domain-specific validation, such as expert clinician review or regulatory audit.
Fidelity vs. Faithfulness
These are key concepts in explainability. Fidelity measures how accurately the explanation reflects the model's output, while faithfulness assesses whether the explanation truly matches the model's internal reasoning process.

Terminology used across episodes

This episode discusses

The paper

Explainability in Practice: A Survey of Explainable NLP Across Various Domains · Read on arXiv

Hadi Mohammadi, Robert A. Bagheri, Anastasia Giachanou, Daniel L. Oberski

Utrecht University

Natural Language Processing (NLP) is now embedded in critical sectors including healthcare, finance, and customer relationship management, where models such as GPT-4o, Gemini, and BERT increasingly inform decisions. The black-box nature of these models has created an urgent need for transparency. This review examines explainable NLP (XNLP) as it is actually deployed, working through seven application domains: medicine, finance, systematic reviews, customer relationship management, chatbots, social and behavioral science, and human resources. For each domain, we ask what kind of explanation the setting needs, which methods are used there, and how they are evaluated. A structured cross-domain synthesis then contrasts how those requirements diverge. We compare the main explanation method families on scope, evidence of faithfulness, and computational cost. We also propose a two-tier evaluation protocol that separates a shared technical core of metrics from the domain-specific validation layer through which those metrics have to be read. The review also addresses areas that remain underrepresented in the XNLP literature, including real-world applicability, the gap between fidelity and faithfulness, and the role of human judgment in assessing explanations. It closes with research directions, among them personalized explanations, human-in-the-loop evaluation, and mechanistic interpretability for large language models.

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Explainability in Practice: A Survey of Explainable NLP Across Various Domains".

Jane: The paper was written by Hadi Mohammadi, Robert A. Bagheri, Anastasia Giachanou and Daniel L. Oberski from Utrecht University.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Title: Tom: Welcome back to the show, everyone. Today we're digging into a paper that's got a pretty straightforward title but a massive scope: "Explainability in Practice: A Survey of Explainable NLP Across Various Domains." Jane, I have to say, when I first read that title, I thought, "Okay, another survey." But this one feels different.

Jane: Tom, I had the exact same reaction. But the word that jumps out at me is "Practice." This isn't just a list of algorithms. It's a deep look at how explainability actually gets used in the real world, in places like hospitals and banks. The authors, Mohammadi, Bagheri, Giachanou, and Oberski from Utrecht University, they've organized the whole paper around seven application domains, not around the math.

Tom: Seven domains, right. Medicine, finance, systematic reviews, customer relationship management, chatbots, social and behavioral science, and human resources. That's a lot of ground to cover. And what I love is that they're asking a really practical question for each one: what kind of explanation does this setting actually need?

Jane: Exactly. Because a doctor needs something completely different from a chatbot user. A doctor needs to know which symptoms or medical codes drove a diagnosis. A chatbot user just needs to know why the bot suggested a particular product. The paper makes this point really clearly with the RETAIN model for heart failure prediction, which uses attention to show which hospital visits mattered most.

Tom: And that's the thing, Jane. It's not just about having an explanation. It's about having the *right* explanation for the right audience. The paper even has this great comparison table that shows how the requirements diverge. Finance needs regulatory compliance, so they need audit trails. Social science needs cultural sensitivity because what counts as hate speech in one country might not in another.

Jane: Right, and that's where the title really earns its keep. "Across Various Domains" isn't just a subtitle. It's the whole point. The authors are arguing that a one-size-fits-all approach to explainability is doomed to fail. You can't just bolt on a SHAP analysis and call it a day.

Tom: So, Lu, you've been quiet. What's your take on this domain-first framing? Is this the right way to think about the field?

Lu: I think it's the most honest way. For years, we've been developing explanation methods and then looking for problems to apply them to. This paper flips that. It starts with the problem, the stakeholder, the regulatory pressure, and then asks what method fits. That's a much more mature way to do research.

Tom: Mature, I like that. And it sets the stage for the rest of the paper, where they actually walk through each domain and show what's being done. We're going to get into the meat of that next. Stick around.

Summary: Jane: So we've set the stage with the title and the domain-first approach. Now let's talk about what the paper actually found when it looked inside those seven domains. Tom, the thing that struck me most was the sheer variety of methods being deployed.

Tom: Absolutely. In medicine, you've got attention mechanisms on recurrent neural networks, like RETAIN, which I mentioned. But you also have rule-based systems like EliIE for extracting clinical trial criteria. And then in finance, you've got SHAP and LIME everywhere, but also these ensemble methods for fraud detection that hit ninety-nine percent accuracy on the IEEE-CIS dataset.

Jane: And that's the fascinating part. The paper shows that the methods aren't just different, they're suited to different kinds of questions. In systematic reviews, for example, they're using active learning models to prioritize which papers to screen. The paper cites a simulation study with over twenty-nine thousand runs showing active learning beats random screening. That's not about explaining a single prediction, it's about explaining the whole screening process.

Tom: Right, and that's a different kind of explainability. It's about trust in the workflow, not just trust in one output. Meng, you work on systems day in and day out. Does this match what you see in practice?

Meng: It does, and it's refreshing to see it written down. When we deploy models, the question is never "what's the SHAP value for this token?" It's "can the fraud analyst see why this transaction was flagged in the next five seconds?" The paper's discussion of chatbots makes this concrete. They cite research showing that usability explains fifty-nine percent of the variance in user trust. That's a huge number.

Jane: fifty-nine percent is massive. And it ties into what the paper says about the evaluation gap. They argue that technical metrics like fidelity and faithfulness are often disconnected from practical utility. A model can have a perfect fidelity score, but if the explanation doesn't help a clinician make a better decision, what's the point?

Tom: Exactly. And they have this great example from the social science domain. The HateXplain dataset, which has human rationales for hate speech detection, shows that models trained with those rationales get better accuracy, zero point six nine eight, but also reduce unintended bias. So the explanation isn't just a nice-to-have, it's actively improving the model.

Lu: That's the key insight for me. The paper isn't just cataloging methods. It's showing that explanations, when done right, change the model's behavior. It's not a post-hoc add-on. It's part of the learning signal.

Meng: But it also shows the cost. The paper is honest about the computational overhead. SHAP, for example, is expensive. In real-time fraud detection, you can't wait for exact Shapley values. So there's this tension between fidelity and speed that the paper lays out really well.

Jane: And that tension is exactly what leads us to the paper's proposed solutions. They don't just describe the problem, they suggest a two-tier evaluation protocol. We should get into that, because that's where the paper gets really actionable.

Improvements: Tom: So we've talked about what the paper found. Now let's get into what it suggests we do about it. Jane, the two-tier evaluation protocol is a big one.

Jane: It really is. The idea is simple but powerful. Tier one is a shared technical core that every XNLP study should report: fidelity, faithfulness, completeness. These are the numbers that let you compare methods across domains. But tier two is the domain-specific validation layer. For medicine, that means expert clinician review. For finance, it's regulatory audit.

Tom: And that's the improvement. Instead of everyone inventing their own evaluation, you have a common baseline, and then you add the layer that matters for your stakeholders. The paper even has a table, Table thirteen that maps each domain to its primary metrics and validation methods.

Meng: I like that because it gives engineers a checklist. When I'm building a system, I can look at that table and know what I need to report. But I also want to push back a little. The paper talks about the gap between fidelity and faithfulness. Can you actually measure faithfulness in practice, or is it still a research problem?

Lu: That's the million-dollar question. The paper is very honest about this. They discuss the chain-of-thought faithfulness problem, where LLMs generate plausible reasoning that doesn't match their actual computation. Turpin et al. showed that models often give correct answers while ignoring the biasing information in the prompt. So the explanation is a story, not a mechanism.

Tom: And that's why the paper pushes mechanistic interpretability as a future direction. Sparse autoencoders, for example, are trying to find actual features in the model's activations, not just verbal rationales. Anthropic's work on Claude three Sonnet found features for concrete things like the Golden Gate Bridge and abstract things like sycophancy.

Jane: But the paper also cautions that SAEs have their own problems. They can fail to reliably detect concepts, and steering behavior through them underperforms simpler baselines. So it's not a silver bullet.

Meng: So what's the practical path forward? If I'm building a system today, do I use chain-of-thought or do I use a surrogate model?

Lu: The paper suggests a hybrid. For high-stakes decisions, you don't rely on the model's self-explanation alone. You use intervention-based validation. You perturb the reasoning steps and see if the output changes. If it doesn't, the reasoning isn't faithful. That's a concrete test you can run.

Tom: And that connects to the broader point about human-in-the-loop evaluation. The paper argues that we need to fold human judgment into model selection, not just as a final check, but as part of the training objective. They cite work on optimizing for explanations that people can actually use.

Jane: Right, and that's the bridge to the future directions. Personalized explanations, adaptive explainability, systems that learn what format works for each user. That's where this is heading. And it's a much richer vision than just "make the model transparent."

Conclusion: Tom: Alright, we've covered a lot of ground on "Explainability in Practice: A Survey of Explainable NLP Across Various Domains." Let's pull it together. Jane, what's the one thing you want listeners to remember?

Jane: I think it's that explainability is not a single thing. It's a set of practices that have to be tailored to the domain, the stakeholder, and the risk. The paper shows that a doctor, a fraud analyst, and a chatbot user all need different kinds of explanations. And the authors give us a framework for figuring out what those are.

Tom: And the two-tier evaluation protocol is the practical takeaway. Report the technical core, then add the domain validation. That's something every researcher and engineer can adopt tomorrow.

Lu: I'd add that the paper is honest about what we don't know. The fidelity-faithfulness gap is real, and chain-of-thought reasoning is not a reliable window into model internals. That's a warning and an opportunity.

Meng: From my side, the paper gives engineers a roadmap. We know what metrics to report, we know what validation to run, and we know the computational costs. That's actionable.

Tom: And the future directions are exciting. Personalized explanations, mechanistic interpretability, hybrid neuro-symbolic systems. This is a field that's still finding its footing, but the map is getting clearer.

Jane: So we say goodbye to this paper, but we're not done with the conversation. The next one is waiting, and I have a feeling it's going to push us even further. Thanks for listening, everyone.

Tom: See you on the next episode.

More episodes

← Home