Explainability in Practice: A Survey of Explainable NLP Across Various Domains

arXiv:2502.00837 · cs.CL, cs.AI · Submitted 2026-08-12 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Explainability in Practice: A Survey of Explainable NLP Across Various Domains".

Jane: The paper was written by Hadi Mohammadi, Robert A. Bagheri, Anastasia Giachanou and Daniel L. Oberski from Utrecht University.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Title: Tom: Welcome back to the show, everyone. Today we're digging into a paper that's got a pretty straightforward title but a massive scope: "Explainability in Practice: A Survey of Explainable NLP Across Various Domains." Jane, I have to say, when I first read that title, I thought, "Okay, another survey." But this one feels different.

Jane: Tom, I had the exact same reaction. But the word that jumps out at me is "Practice." This isn't just a list of algorithms. It's a deep look at how explainability actually gets used in the real world, in places like hospitals and banks. The authors, Mohammadi, Bagheri, Giachanou, and Oberski from Utrecht University, they've organized the whole paper around seven application domains, not around the math.

Tom: Seven domains, right. Medicine, finance, systematic reviews, customer relationship management, chatbots, social and behavioral science, and human resources. That's a lot of ground to cover. And what I love is that they're asking a really practical question for each one: what kind of explanation does this setting actually need?

Jane: Exactly. Because a doctor needs something completely different from a chatbot user. A doctor needs to know which symptoms or medical codes drove a diagnosis. A chatbot user just needs to know why the bot suggested a particular product. The paper makes this point really clearly with the RETAIN model for heart failure prediction, which uses attention to show which hospital visits mattered most.

Tom: And that's the thing, Jane. It's not just about having an explanation. It's about having the *right* explanation for the right audience. The paper even has this great comparison table that shows how the requirements diverge. Finance needs regulatory compliance, so they need audit trails. Social science needs cultural sensitivity because what counts as hate speech in one country might not in another.

Jane: Right, and that's where the title really earns its keep. "Across Various Domains" isn't just a subtitle. It's the whole point. The authors are arguing that a one-size-fits-all approach to explainability is doomed to fail. You can't just bolt on a SHAP analysis and call it a day.

Tom: So, Lu, you've been quiet. What's your take on this domain-first framing? Is this the right way to think about the field?

Lu: I think it's the most honest way. For years, we've been developing explanation methods and then looking for problems to apply them to. This paper flips that. It starts with the problem, the stakeholder, the regulatory pressure, and then asks what method fits. That's a much more mature way to do research.

Tom: Mature, I like that. And it sets the stage for the rest of the paper, where they actually walk through each domain and show what's being done. We're going to get into the meat of that next. Stick around.

Summary: Jane: So we've set the stage with the title and the domain-first approach. Now let's talk about what the paper actually found when it looked inside those seven domains. Tom, the thing that struck me most was the sheer variety of methods being deployed.

Tom: Absolutely. In medicine, you've got attention mechanisms on recurrent neural networks, like RETAIN, which I mentioned. But you also have rule-based systems like EliIE for extracting clinical trial criteria. And then in finance, you've got SHAP and LIME everywhere, but also these ensemble methods for fraud detection that hit ninety-nine percent accuracy on the IEEE-CIS dataset.

Jane: And that's the fascinating part. The paper shows that the methods aren't just different, they're suited to different kinds of questions. In systematic reviews, for example, they're using active learning models to prioritize which papers to screen. The paper cites a simulation study with over twenty-nine thousand runs showing active learning beats random screening. That's not about explaining a single prediction, it's about explaining the whole screening process.

Tom: Right, and that's a different kind of explainability. It's about trust in the workflow, not just trust in one output. Meng, you work on systems day in and day out. Does this match what you see in practice?

Meng: It does, and it's refreshing to see it written down. When we deploy models, the question is never "what's the SHAP value for this token?" It's "can the fraud analyst see why this transaction was flagged in the next five seconds?" The paper's discussion of chatbots makes this concrete. They cite research showing that usability explains fifty-nine percent of the variance in user trust. That's a huge number.

Jane: fifty-nine percent is massive. And it ties into what the paper says about the evaluation gap. They argue that technical metrics like fidelity and faithfulness are often disconnected from practical utility. A model can have a perfect fidelity score, but if the explanation doesn't help a clinician make a better decision, what's the point?

Tom: Exactly. And they have this great example from the social science domain. The HateXplain dataset, which has human rationales for hate speech detection, shows that models trained with those rationales get better accuracy, zero point six nine eight, but also reduce unintended bias. So the explanation isn't just a nice-to-have, it's actively improving the model.

Lu: That's the key insight for me. The paper isn't just cataloging methods. It's showing that explanations, when done right, change the model's behavior. It's not a post-hoc add-on. It's part of the learning signal.

Meng: But it also shows the cost. The paper is honest about the computational overhead. SHAP, for example, is expensive. In real-time fraud detection, you can't wait for exact Shapley values. So there's this tension between fidelity and speed that the paper lays out really well.

Jane: And that tension is exactly what leads us to the paper's proposed solutions. They don't just describe the problem, they suggest a two-tier evaluation protocol. We should get into that, because that's where the paper gets really actionable.

Improvements: Tom: So we've talked about what the paper found. Now let's get into what it suggests we do about it. Jane, the two-tier evaluation protocol is a big one.

Jane: It really is. The idea is simple but powerful. Tier one is a shared technical core that every XNLP study should report: fidelity, faithfulness, completeness. These are the numbers that let you compare methods across domains. But tier two is the domain-specific validation layer. For medicine, that means expert clinician review. For finance, it's regulatory audit.

Tom: And that's the improvement. Instead of everyone inventing their own evaluation, you have a common baseline, and then you add the layer that matters for your stakeholders. The paper even has a table, Table thirteen that maps each domain to its primary metrics and validation methods.

Meng: I like that because it gives engineers a checklist. When I'm building a system, I can look at that table and know what I need to report. But I also want to push back a little. The paper talks about the gap between fidelity and faithfulness. Can you actually measure faithfulness in practice, or is it still a research problem?

Lu: That's the million-dollar question. The paper is very honest about this. They discuss the chain-of-thought faithfulness problem, where LLMs generate plausible reasoning that doesn't match their actual computation. Turpin et al. showed that models often give correct answers while ignoring the biasing information in the prompt. So the explanation is a story, not a mechanism.

Tom: And that's why the paper pushes mechanistic interpretability as a future direction. Sparse autoencoders, for example, are trying to find actual features in the model's activations, not just verbal rationales. Anthropic's work on Claude three Sonnet found features for concrete things like the Golden Gate Bridge and abstract things like sycophancy.

Jane: But the paper also cautions that SAEs have their own problems. They can fail to reliably detect concepts, and steering behavior through them underperforms simpler baselines. So it's not a silver bullet.

Meng: So what's the practical path forward? If I'm building a system today, do I use chain-of-thought or do I use a surrogate model?

Lu: The paper suggests a hybrid. For high-stakes decisions, you don't rely on the model's self-explanation alone. You use intervention-based validation. You perturb the reasoning steps and see if the output changes. If it doesn't, the reasoning isn't faithful. That's a concrete test you can run.

Tom: And that connects to the broader point about human-in-the-loop evaluation. The paper argues that we need to fold human judgment into model selection, not just as a final check, but as part of the training objective. They cite work on optimizing for explanations that people can actually use.

Jane: Right, and that's the bridge to the future directions. Personalized explanations, adaptive explainability, systems that learn what format works for each user. That's where this is heading. And it's a much richer vision than just "make the model transparent."

Conclusion: Tom: Alright, we've covered a lot of ground on "Explainability in Practice: A Survey of Explainable NLP Across Various Domains." Let's pull it together. Jane, what's the one thing you want listeners to remember?

Jane: I think it's that explainability is not a single thing. It's a set of practices that have to be tailored to the domain, the stakeholder, and the risk. The paper shows that a doctor, a fraud analyst, and a chatbot user all need different kinds of explanations. And the authors give us a framework for figuring out what those are.

Tom: And the two-tier evaluation protocol is the practical takeaway. Report the technical core, then add the domain validation. That's something every researcher and engineer can adopt tomorrow.

Lu: I'd add that the paper is honest about what we don't know. The fidelity-faithfulness gap is real, and chain-of-thought reasoning is not a reliable window into model internals. That's a warning and an opportunity.

Meng: From my side, the paper gives engineers a roadmap. We know what metrics to report, we know what validation to run, and we know the computational costs. That's actionable.

Tom: And the future directions are exciting. Personalized explanations, mechanistic interpretability, hybrid neuro-symbolic systems. This is a field that's still finding its footing, but the map is getting clearer.

Jane: So we say goodbye to this paper, but we're not done with the conversation. The next one is waiting, and I have a feeling it's going to push us even further. Thanks for listening, everyone.

Tom: See you on the next episode.

Hadi Mohammadi, Robert A. Bagheri, Anastasia Giachanou, Daniel L. Oberski

Utrecht University

cs.CL, cs.AI

Submitted: 2026-08-12

Updated: 2026-08-13

Comments: 32 pages, 5 figures, 15 tables, 257 references. Under review at the Journal of Information Science. Supplementary materials and structured data: https://github.com/mohammadi-hadi/xnlp-survey

Code: https://github.com/mohammadi-hadi/xnlp-survey

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 73/100

Key concepts

Domain-First Approach
Instead of developing explanation methods and then finding problems for them, this approach starts with a specific problem or domain (like medicine or finance) and determines what kind of explanation is actually needed by the stakeholders.
Two-Tier Evaluation Protocol
The paper proposes a structured method for evaluating explainability. Tier one requires reporting common technical metrics (fidelity, faithfulness), while tier two demands domain-specific validation, such as expert clinician review or regulatory audit.
Fidelity vs. Faithfulness
These are key concepts in explainability. Fidelity measures how accurately the explanation reflects the model's output, while faithfulness assesses whether the explanation truly matches the model's internal reasoning process.

Terminology

Summary

Summary

This paper is a survey of Explainable Natural Language Processing (XNLP) as it is deployed in practice, examining seven application domains: medicine, finance, systematic reviews, customer relationship management (CRM), chatbots, social and behavioral science, and human resources. For each domain, the authors ask what kind of explanation the setting needs, which methods are used, and how they are evaluated. The paper makes four main contributions: (1) it compiles and examines research across the seven domains, including a structured cross-domain comparison in Table 10; (2) it examines the full range of evaluation approaches, giving mathematical definitions of key metrics like fidelity and discussing human-centered measures like user trust, and compares the main explanation method families in Table 12; (3) it addresses open questions regarding bias, privacy, data availability, and the balance between performance and interpretability; and (4) it proposes research avenues such as personalized explanations, human-in-the-loop evaluation, and mechanistic interpretability for LLMs.

The paper begins by reviewing modeling techniques for XNLP, starting with traditional models like Bag of Words (BoW) and TF-IDF, which are simple and interpretable but struggle with contextual relationships. It then covers embedding models like Word2Vec and GloVe, which capture semantic details but add complexity, and discusses explainability strategies such as visualization of vector spaces and gradient-based saliency maps. The paper then examines transformer-based models like BERT, RoBERTa, and ALBERT, discussing attention-weight visualization, probing tasks, attribution methods like Integrated Gradients, and model-agnostic methods like LIME and SHAP. A key caveat is that attention-weight visualization can mislead, as weights could highlight neutral tokens while downplaying salient words.

The paper then discusses LLMs and mechanistic interpretability, focusing on sparse autoencoders (SAEs) which decompose neural activations into sparse, interpretable features. Anthropic's work on Claude 3 Sonnet demonstrated that SAEs can extract millions of interpretable features, spanning both concrete entities and abstract notions. However, the paper notes serious challenges: SAEs may fail to reliably detect concepts they purportedly encode, and steering model behavior through SAE features underperforms simpler baselines. The paper also addresses the Chain-of-Thought (CoT) faithfulness problem, where LLMs frequently produce plausible-sounding reasoning that does not correspond to their true decision-making mechanisms. The paper notes that larger models exhibited less faithful reasoning than smaller ones, and that CoT outputs may not accurately reflect the model's actual computational process.

For each application domain, the paper provides detailed analysis. In medicine, XNLP is used for EHR classification, disease risk assessment, and mental health analysis. The RETAIN model for heart failure prediction uses a two-level attention mechanism to show which clinical encounters and diagnoses mattered most. Domain-specific challenges include HIPAA/GDPR privacy constraints, concept drift in longitudinal EHR data, and the need for explanations to fit into time-pressured clinical workflows. In finance, XNLP supports risk assessment, fraud detection, and firm valuation. A stacking ensemble of XGBoost, LightGBM, and CatBoost on the IEEE-CIS Fraud Detection dataset reaches 99% accuracy, with SHAP used for feature selection and explanation. Challenges include adversarial gaming of explanations and real-time scalability. In systematic reviews, XNLP tools like RobotReviewer address bias assessment, and active learning models substantially shorten the time to discover relevant records. In CRM, XNLP enhances sentiment analysis and customer support automation, with a fine-tuned BERT classifier reaching 92% accuracy on aspect detection. In chatbots, explainability significantly influences user acceptance and satisfaction, with chatbot usability explaining 59% of variance in user trust. In social and behavioral science, XNLP is used for hate speech detection, sexism detection, and fake news detection, with benchmark datasets like HateXplain providing human rationales. In HR, XNLP is used for resume screening, employee sentiment analysis, and bias auditing, with documented biases in AI resume screening making explainability essential.

The paper then discusses critical aspects of XNLP, including evaluation metrics. Quantitative metrics include fidelity, which measures how accurately an explanation captures the true behavior of a model, with a simplified formula given as Fidelity(E, M, D) = (1/D) * sum of indicator functions; coherence, which assesses logical consistency using metrics like BLEU; and completeness, which determines whether an explanation includes all salient factors, with a simplified Shapley-based formula provided. Qualitative metrics include user trust, satisfaction, and transparency. The paper emphasizes the distinction between fidelity and faithfulness: fidelity measures whether an explanation accurately predicts the model's outputs, while faithfulness asks whether the explanation reflects the actual causal mechanisms the model uses. The paper notes that an explanation can have high fidelity while having low faithfulness, citing sanity checks that show saliency methods produce convincing-looking maps that barely change when model weights are randomized.

The paper proposes a two-tier evaluation protocol: Tier 1 comprises core technical metrics that every XNLP study should report (fidelity, faithfulness, completeness), and Tier 2 is a domain validation layer chosen according to domain-specific requirements, such as expert clinical review in medicine, regulatory audit in finance, fairness audits in HR, and user studies in CRM and chatbot settings. The paper also discusses rationalization techniques, including extractive rationalization (LIME, Grad-CAM, SHAP) and abstractive rationalization (free-form natural language explanations), noting that post-hoc rationalizations can mislead if they do not reflect true causal pathways.

The paper concludes with future research directions, including reinforcement learning and chain-of-thought reasoning, hybrid neuro-symbolic systems, explainable dialogue and social media analytics, personalized and adaptive explainability, and rational AI (RAI). The paper emphasizes that XNLP's future hinges on bridging performance and transparency, using advanced techniques while upholding rigorous evaluation and open-science practices. The cross-domain synthesis shows that trust is the convergent goal across all seven domains, whereas risk tolerance, temporal dynamics, stakeholder diversity, and explanation granularity diverge substantially, concluding that one-size-fits-all XNLP is insufficient and what is needed is a shared technical core of explanation methods and metrics combined with a domain-specific validation layer.

Improvements for AI systems

Based on the paper, here are the specific improvements I can make to AI systems, along with what the improved system can then do:

Improvement: I will build an evaluation framework that separates a shared technical core (Tier 1) from domain-specific validation (Tier 2). Tier 1 will always report fidelity (Eq. 1), faithfulness via perturbation/intervention tests, and completeness. Tier 2 will add domain-specific checks: expert clinical review for medicine, regulatory audit for finance, fairness audits for HR, and user studies for CRM/chatbots.

What the improved system can do: It can automatically select the appropriate validation layer based on deployment context, ensuring that explanations are not just technically sound but also practically useful and legally compliant in the specific domain where the AI is deployed.


Summary of what the improved AI system can do: It can produce faithful, domain-appropriate, bias-audited, real-time, and user-trusted explanations. It will refuse to show misleading reasoning, will explain at the right granularity for the stakeholder, will be robust to adversarial gaming, and will be validated against both technical metrics and real-world usability. This makes it safe for deployment in medicine, finance, HR, and social media moderation.

Abstract

Natural Language Processing (NLP) is now embedded in critical sectors including healthcare, finance, and customer relationship management, where models such as GPT-4o, Gemini, and BERT increasingly inform decisions. The black-box nature of these models has created an urgent need for transparency. This review examines explainable NLP (XNLP) as it is actually deployed, working through seven application domains: medicine, finance, systematic reviews, customer relationship management, chatbots, social and behavioral science, and human resources. For each domain, we ask what kind of explanation the setting needs, which methods are used there, and how they are evaluated. A structured cross-domain synthesis then contrasts how those requirements diverge. We compare the main explanation method families on scope, evidence of faithfulness, and computational cost. We also propose a two-tier evaluation protocol that separates a shared technical core of metrics from the domain-specific validation layer through which those metrics have to be read. The review also addresses areas that remain underrepresented in the XNLP literature, including real-world applicability, the gap between fidelity and faithfulness, and the role of human judgment in assessing explanations. It closes with research directions, among them personalized explanations, human-in-the-loop evaluation, and mechanistic interpretability for large language models.

Sources

Related papers