Explainable Machine Learning in Healthcare: Methods, Interpretation, and Applications for Clinical Research

arXiv:2608.07522 · cs.CY, cs.LG · Submitted 2026-07-13 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Explainable Machine Learning in Healthcare: Methods, Interpretation, and Applications for Clinical Research".

Jane: The paper was written by Krishna Padmanabhan, Minxin Lu, Dai Feng, Natalia Kan-Dobrosky, Sai Konduri et al. from Madrigal Pharmaceuticals Inc and Boston University School of Medicine and AbbVie Inc and St. Elizabeth Healthcare and Thermo Fisher Scientific.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Title: Tom: Alright, welcome back to the show, everyone. Today we're cracking open a paper that's been making the rounds, and it's called "Explainable Machine Learning in Healthcare: Methods, Interpretation, and Applications for Clinical Research." Jane, I gotta say, the title alone is a mouthful, but it's hitting on something that's been nagging at me for a while.

Jane: Oh, absolutely, Tom. And I think the key word there is "Explainable." We've all heard the horror stories about these black-box models that make predictions but nobody can figure out why. This paper is basically saying, "Hey, if we're going to use machine learning to decide who gets what treatment, we need to be able to explain our reasoning."

Tom: Right, and it's not just about being curious. I mean, think about a doctor looking at a risk score for a patient. If the machine says "high risk" but can't say why, is the doctor supposed to just trust it? That's a tough sell in a hospital room.

Jane: Exactly. And that's why this paper is so timely. They're not just complaining about the problem; they're offering a toolkit. They walk through methods like SHAP and LIME, which are basically ways to peek inside the black box and see which features are pulling the prediction up or down.

Tom: So it's like giving the doctor a translator for the machine's language. I love that. But Jane, who are the folks behind this? It's a big author list.

Jane: It is a big team, and that's actually a strength. You've got people from biostatistics at Madrigal Pharmaceuticals, folks from Boston University School of Medicine, statisticians from AbbVie, and even a pulmonary-critical care physician from St. Elizabeth Healthcare. That mix of industry, academia, and clinical practice tells me they're trying to bridge a real gap.

Tom: A clinician on the team is huge. That means they've actually had to explain a model's output to a patient or a colleague, not just in theory but in practice. That grounds the whole paper in reality.

Jane: For sure. And the corresponding author is Minxin Lu from Boston University. They've structured this as a primer, which I appreciate. It's not a dense mathematical treatise; it's a guide for biomedical scientists who need to use these tools.

Tom: A primer for the people on the front lines of clinical research. That's a smart move. So the title is promising a practical guide, and it sounds like they're delivering. What's the first thing they dive into?

Jane: Well, they start by making the case for why this matters, and they use some pretty stark examples of what happens when models go unexplained. I think that's where we should look next.

Tom: Yeah, let's get into the summary and see how they set the stage. Stick around, folks.

Summary: Tom: So Jane, we're back with "Explainable Machine Learning in Healthcare." We just talked about the team behind it, but now let's get into what the paper actually claims. What's the core message?

Jane: The core message is that explainability isn't a nice-to-have; it's a requirement for safe and effective clinical AI. They frame it around three concepts: transparency, interpretability, and explainability. Transparency is about how the model works internally, interpretability is about the overall logic, and explainability is about justifying a single prediction.

Tom: That's a helpful breakdown. So transparency is the machinery, interpretability is the manual, and explainability is the specific answer to "why did you say this about my patient?"

Jane: Exactly. And they back this up with real-world failures. There's the famous pneumonia model that learned hospital-specific artifacts instead of actual pathology, and another one that wrongly concluded asthma patients had lower risk because it didn't account for the aggressive care they receive.

Tom: Oh, those are the cautionary tales. The model wasn't learning medicine; it was learning the data's quirks. That's terrifying when you think about deploying that in a hospital.

Jane: It is. And that's why they argue we need methods to catch these issues. They also mention the regulatory angle, which is a big deal. The FDA, Health Canada, and the UK's MHRA have all published guidance on good machine learning practices, and the EU has its AI Act with transparency requirements.

Tom: So regulators are starting to demand explainability. That's a huge shift. It's not just about best practices anymore; it's about compliance.

Jane: Right. And the paper's summary sets up their approach: they're going to show us a bunch of methods, explain how they work, and demonstrate them on a real dataset. They picked the Heart Disease Dataset, which has about one thousand one hundred eighty-nine samples with clinical features like age, cholesterol, and ST slope.

Tom: A familiar dataset for anyone in the field. It's a good choice because the variables are clinically meaningful. You can actually reason about whether the model's logic makes sense.

Jane: Exactly. And they're using a Random Forest model for the demonstration. They're upfront that they didn't do a train-test split because the goal is to show the explanation methods, not to build a production-ready model.

Tom: That's an important caveat. We're not supposed to draw clinical conclusions from their examples; we're supposed to learn the tools. Got it.

Jane: And they organize the methods into categories: model-specific, model-agnostic, and intrinsically interpretable models. That's the roadmap for the rest of the paper.

Tom: Model-specific means it only works for one type of model, model-agnostic means it works for any model, and intrinsically interpretable means the model itself is simple enough to understand. That's a clean way to think about it.

Jane: Clean and practical. So now that we know the plan, let's see how they actually execute it. The next part should show us the methods in action.

Tom: Let's do it. We're just getting to the good stuff.

Improvements: Tom: Welcome back. We're digging into "Explainable Machine Learning in Healthcare," and Jane, we've covered the who and the what. Now I want to know: what are they actually proposing we do differently?

Jane: Good question. The paper's real contribution is a structured, practical guide. They're not inventing new algorithms; they're organizing existing ones into a coherent framework that a working scientist can actually use. They even include a table that maps clinical questions to the right method.

Tom: A cheat sheet. I love that. So if I'm a researcher asking "which variables matter most," what do they point me to?

Jane: They point you to feature importance and SHAP beeswarm plots. If you're asking "why did this specific patient get this prediction," you'd use a SHAP waterfall plot or LIME. And if you want to see how changing a variable affects risk across the population, you'd use Partial Dependence Plots and ICE plots.

Tom: So it's about matching the question to the tool. That's the improvement over just saying "use explainable AI." They're giving you a decision tree for choosing the explanation method itself.

Jane: Precisely. And they don't just list the methods; they show them on the Heart Disease dataset. For example, they have a SHAP waterfall plot for a specific patient, Patient twenty-two showing that chest pain type decreased the predicted risk while ST slope increased it.

Tom: That's concrete. You can see the contributions adding up to the final prediction. It makes the abstract concept tangible.

Jane: And they're careful to discuss limitations. For SHAP, they note it assumes features contribute independently, which may not hold in biology. For LIME, they point out that results can be unstable near decision boundaries. They're not overselling any single method.

Tom: That honesty is refreshing. A lot of papers just hype their favorite tool. This one is saying, "Here's what each method is good at, and here's where it falls short."

Jane: Exactly. And they also emphasize that explanations describe associations, not causation. That's a crucial caveat for clinical researchers who might be tempted to interpret a high SHAP value as proof that a variable causes the outcome.

Tom: That's a trap I could see people falling into. The model says "high cholesterol increases risk," but that's a correlation learned from data, not a biological mechanism.

Jane: Right. So the improvement they're suggesting is a mindset shift: use these tools to validate and understand your models, but always keep their limitations in mind. It's about building trust through scrutiny, not blind faith.

Tom: I like that framing. So they're giving us the methods and the guardrails. What happens when they actually put this into practice? Let's look at the first page and see how they set the stage.

First Page: Tom: So we're back with "Explainable Machine Learning in Healthcare," and Jane, we've talked about the framework and the methods. Let's zoom in on the very beginning of the paper. What's the hook?

Jane: The introduction makes a bold claim: the clinical value of machine learning depends on our ability to interpret the outputs. They cite a survey where eighty-four percent of physicians say they need proper training on AI tools, but only twelve percent actually use them for assistive diagnostics.

Tom: That gap is striking. Doctors want to use these tools, but they don't trust them because they can't understand them. And the paper says that's a training problem, not just a technical problem.

Jane: Exactly. They're arguing that proficiency in explainability is critical for biomedical scientists. It's not enough to build a model with good accuracy; you have to be able to communicate what it's doing.

Tom: And they bring in those real-world examples again, like the pneumonia model that learned hospital artifacts. That example is so powerful because it shows that even a model with high accuracy can be completely wrong in its reasoning.

Jane: Right. And they tie it to regulatory pressure. The FDA, Health Canada, and MHRA have all published guiding principles for good machine learning practice, and the EU's AI Act mandates transparency. So this isn't just academic curiosity; it's becoming a legal requirement.

Tom: That's a big deal. If you're building a clinical decision support tool, you might be legally obligated to explain its predictions. That changes the engineering calculus entirely.

Jane: It does. And the paper positions itself as a bridge between advanced machine learning methodology and clinical applicability. They want to help people move from "I have a model" to "I can explain my model to a skeptical clinician."

Tom: And they're doing it with a hands-on demonstration, not just theory. They're using the Heart Disease Dataset to show each method in action, which makes it feel approachable.

Jane: Exactly. The first page sets up the problem, the stakes, and the audience. It's saying, "If you're a biomedical scientist who needs to use machine learning responsibly, this paper is for you."

Tom: So they're targeting the people who are actually going to deploy these models in clinical research. That's a smart audience to write for.

Jane: Absolutely. And I think that's what makes this paper stand out. It's not just a review; it's a practical guide with worked examples and honest discussions of limitations.

Tom: Well, we've covered the intro, the methods, and the framework. Let's wrap this up and see what the big takeaway is for our listeners.

Conclusion: Tom: Alright, we're wrapping up our discussion of "Explainable Machine Learning in Healthcare." Jane, give us the final word. What should our listeners remember?

Jane: I think the biggest takeaway is that explainability is a toolbox, not a single solution. The paper shows that you need different tools for different questions: SHAP for patient-level explanations, PDPs for population-level effects, and rule-based models when you want the logic to be transparent by design.

Tom: And they're not saying one method is the best. They're saying, "Know your question, pick the right tool, and understand its limitations."

Jane: Exactly. And that's a more mature approach than just saying "use explainable AI." It's about being thoughtful and rigorous about how you build trust in these systems.

Tom: The paper also makes a strong case that this is a translational challenge, not just a technical one. It's about bridging the gap between data science and medical practice.

Jane: Right. And they mention that the field is evolving toward causal explanations and uncertainty quantification. So this primer is a foundation, not the final word.

Tom: I appreciate that they're honest about that. They're not claiming to have solved everything; they're giving us the tools we have today and pointing to where things are heading.

Jane: And for anyone working in clinical research, that's incredibly valuable. It gives you a starting point for making your models more transparent and accountable.

Tom: Well said. We've covered the authors, the methods, the demonstrations, and the implications. I think our listeners have a solid understanding of what this paper offers.

Jane: Absolutely. And I hope it inspires more people to think about explainability as a core requirement, not an afterthought.

Tom: That's a great note to end on. Thanks for joining us, everyone. We'll be back next time with another paper to dissect.

Jane: Until then, keep asking why. It's the most important question in machine learning.

Krishna Padmanabhan, Minxin Lu, Dai Feng, Natalia Kan-Dobrosky, Sai Konduri, Heather J. Litman, Achilleas Livieratos

Madrigal Pharmaceuticals Inc · Boston University School of Medicine · AbbVie Inc · St. Elizabeth Healthcare · Thermo Fisher Scientific

cs.CY, cs.LG

Submitted: 2026-07-13

Comments: 29 pages, 6 figures

Journal ref: J. Am. Med. Inform. Assoc. 33 (2026) ocag077

DOI: 10.1093/jamia/ocag077

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 62/100

The gist: This paper provides a practical and methodologically grounded overview of explainable machine learning (XML) approaches in healthcare, with emphasis on their interpretation and application in

Terminology

Summary

This paper provides a practical and methodologically grounded overview of explainable machine learning (XML) approaches in healthcare, with emphasis on their interpretation and application in clinical research and decision support. By moving beyond traditional predictive models, this primer aims to foster trust, transparency, and informed clinical decision-making, ultimately bridging the gap between data science and medical practice.

The paper presents a structured review of commonly used XML methodologies, including global and local interpretability tools such as SHapley Additive exPlanations (SHAP), Local Interpretable Model-Agnostic Explanations (LIME), Partial Dependence Plots (PDP), and Individual Conditional Expectation (ICE) plots. For each method, the paper explains the underlying mechanism at a high level, visualizes representative outputs, and provides structured guidance on interpretation, appropriate use, and limitations, illustrated using the publicly available Heart Disease dataset.

The paper notes that XML techniques provided intuitive visual and quantitative insights into how predictors influence model predictions. Global methods characterized population-level feature effects, whereas local methods revealed patient-level contributions useful for individualized interpretation. The worked examples demonstrate how XML outputs can identify nonlinear relationships, detect interaction effects, and reveal heterogeneity in predicted risk across patients, addressing key challenges in translating ML predictions into interpretable outputs for clinical research.

The paper concludes that XML tools offer valuable interpretability for ML models and support more transparent and accountable ML applications in clinical research. By providing a methodologically grounded overview alongside practical implementation examples and structured guidance on each method's strengths and limitations, this primer helps bridge the gap between advanced ML methodology and clinical applicability. Thoughtful adoption of XML approaches may facilitate better understanding, communication, and critical evaluation of ML predictions in healthcare research, ultimately supporting evidence-based clinical decision-making.

The paper emphasizes that the integration of Machine learning (ML) into healthcare is transforming how biomedical informaticians, clinical investigators, clinicians, and healthcare data scientists process data, diagnose conditions, and personalize treatments. ML offers a powerful tool for synthesizing multidimensional datasets, identifying patterns, and predicting outcomes at a large scale. However, the clinical value of ML tools depends on practitioners’ ability to interpret outputs within biological and pathophysiological contexts. For example, sepsis risk forecast must be reconciled with bedside observations and patient history to avoid alarm fatigue. Despite strong interest, 84% of physicians find it important to receive proper training/education on AI tools being used for the adoption of AI into their practice, only 12% of clinicians report using them for assistive diagnostics. Proficiency in explainability techniques is therefore critical for biomedical scientists to translate predictions into interpretable outputs and communicable model behaviors that support clinical judgements.

XML focuses on making the ML decision-making more understandable, analogous to how domain experts translate complex medical knowledge into comprehensible explanations for patients. XML covers three related concepts: transparency, interpretability, and explainability. Transparency focuses on the model’s internal processes, including model structure, training data, population coverage, and performance characteristics. Interpretability addresses how a model works and refers to the overall model’s behavior and logic. Explainability provides the rationale behind an individual prediction and answers “what is the model saying?”.

XML was introduced to address “black-box” models, such as deep learning systems with millions of parameters, that are difficult to interpret in clinical contexts. The risk of deploying non-explainable ML systems has been demonstrated by several real-world examples: Zech et al. (2018) proved that a deep learning model for pneumonia detection relied on hospital-specific artifacts on the chest X-rays rather than actual pathology; Caruana et al. (2015) demonstrated that a black-box risk model incorrectly learned that asthma patients had lower complication risk, because it failed to account for the aggressive care these patients typically receive; and Adamson & Smith. (2018) showed that an algorithm trained predominantly on Western populations can perform poorly in underrepresented groups and exacerbate health disparities. These examples underscore the importance of interpretability tools to make ML systems more understandable.

The integration of AI/ML into healthcare has promoted global regulatory agencies to establish frameworks to ensure the use of ML is safe, transparent, and accountable. In 2021, the guiding principles for Good ML Practice was jointly released by the U.S. Food and Drug Administration (FDA), Health Canada, and the United Kingdom’s Medicines and Healthcare products Regulatory Agency (MHRA). In the same year, FDA also released an Action Plan for AI/ML-Based Software as a Medical Device (SaMD). The EU Artificial Intelligence Act introduced requirement for AI transparency, mandating AI providers to disclose their use, purpose, and limitations to support informed adoption and accountability. In January 2025, the FDA published a draft guidance outlining a seven-step framework for evaluating the credibility of AI models.

Biomedical scientists must be able to evaluate models confidently, explain the model’s output, identify bias in training data, ensure fairness of model process, and confirm that a model behaves as expected, all of which are essential to developing safer and more accountable applications. While prior reviews have addressed XML methods broadly or within specific clinical domains, this work offers a method-oriented primer aimed at a broader informatics audience with structured guidance on interpretation and limitations of each method.

The XML methods detailed in this primer were chosen to maximize clinical research utility based on three criteria. First, the paper aimed for broad applicability by prioritizing model-agnostic approaches (SHAP, LIME, PDP, ICE) that provide consistent explanations across algorithms, while also including established model-specific measures (Gini importance, regression coefficients) that remain foundational in clinical research. Second, the paper focused on tabular clinical data, the most common format for risk prediction in clinical informatics; accordingly, modality-specific techniques such as class activation maps (CAM) and layer-wise relevance propagation (LRP), essential for deep learning in medical imaging, were excluded to maintain a focused scope. Finally, the paper selected methods that balance global explanations of population-level predictors with local explanations of individual patient risk.

The utility of XML is best demonstrated by looking at a dataset that reflects clinical complexity. The paper selects the Heart Disease Dataset for its diversity and real-world relevance. It contains 1,189 samples drawn from five hospital systems with 11 clinically interpretable features, including demographic, physiological, and diagnostic variables, that resemble structured clinical data commonly used in predictive modeling. This structure allows explanation methods to be illustrated in a transparent and easily interpretable manner for biomedical scientists in cardiology risk assessment models. While this example focuses on tabular clinical variables, the XML techniques are model-agnostic and broadly applicable to other data modalities used in healthcare machine learning, including imaging, genomic data, and high-dimensional electronic health record (EHR) features. Because the models used here are purely predictive, the explanations describe how the model uses features to generate predictions. These explanations should be interpreted as describing associations learned by the model rather than causal relationships between variables and outcomes.

The paper demonstrates the XML methods using a Random Forest prediction model on the Heart Disease dataset. Since the focus is on explainability, no test-train split or optimization of the ML algorithm itself was performed. For each method, the paper explains the mechanism at a high level, visualizes outputs, and provides interpretation, limitations, and guidance on appropriate use.

Model-specific explanation methods take advantage of the structure and parameters of a particular model type. They “open up the black box” and use model internals, such as coefficients, tree splits, or neural network weights, to provide an explanation. A common first question about any predictive model results is which variables matter most. In linear models, standardized coefficients offer a straightforward measure of feature importance. Decision trees and tree-based ensembles, such as random forest, gradient boosting and XGBoost, capture complex non-linear relationships. Feature importance in tree models is generally defined by the reduction in prediction error or impurity when a variable is used for splitting. In random forest, the impurity-based feature importance is calculated by summing Gini impurity, or other split criteria, reduction across all trees, whenever that feature is used to split a node. Features that yield larger reductions in impurity (increase in model accuracy) receive higher importance scores. An alternative model-agnostic measure is permutation importance, which measures how randomizing a feature’s values drops the model’s accuracy.

In the coronary artery disease prediction example, Random Forest and Logistic Regression rank features by predictive importance. Both models assign high importance to features fundamental to cardiac evaluation, such as ST slope and chest pain characteristics. This convergence across the two methodological distinct methods suggests internal consistency of highlighted important variables within the dataset. The variation in rankings reflects the differences in ways these models encode feature relationships. The Gini importance of Random Forest captures non-linear relationships such as threshold effects and interaction effects. In contrast, the coefficients of Logistic Regression describe linear, monotonic, independent contributions to the log-odds. It should be noted that convergence in feature importance does not indicate predictive performance, generalizability to new patients requires evaluation with performance metrics on a held-out test set.

Model-specific methods should not be used when explanations need to be compared across different models, as they only apply to a single model type. They can also be misleading if the model is biased or misspecified, since the explanations will reflect those issues. Feature importance indicates which variables contribute most to the model’s predictions, but it does not indicate the direction of association (whether higher values increase or decrease predicted risk), threshold effects (at what values features become influential), or interactions between variables. These insights require complementary tools such as partial dependence plots or SHAP values.

Model-agnostic methods work for any ML model and do not require knowing the internal structure or weights. They usually work by perturbing inputs, observing outputs, and computing relationships. Important model-agnostic techniques include SHAP and LIME for local (patient level) explanation and Partial Dependence and ICE plots for global (overall) explanation for feature effect pattern visualization.

SHAP is a widely-used method and often considered the “gold standard” for explaining individual predictions by quantifying each feature's contribution to the total deviation from the population average. SHAP feature contributions are additive and sum precisely to the prediction difference. For example, if the average risk is 10% and a patient’s risk is 15%, SHAP might show: Age (+2%), Blood pressure (+4%), Current medications (-1%) = +5%. The SHAP waterfall plot for a randomly selected Patient 22 shows how the Random Forest model’s prediction was constituted by features that decreased (blue) and increased (red) the predicted risk score. The largest negative SHAP contribution was chest pain type (-0.15), partially offset by ST slope (+0.14) and other factors. The sum of the SHAP values is-0.1915, which indicates a 36.37% lower predicted risk compared to the “average” patient. The SHAP values represent the contribution of each feature to the model's prediction on the probability scale.

SHAP provides theoretically grounded patient-level insights by explicitly quantifying each feature’s contribution to the model prediction. This is particularly useful for cases with mixed risk profiles, where SHAP clarifies which variables most influenced the model’s decision. SHAP should not be used to compare individual patients across subgroups, as attributions are not calibrated across different patient populations. SHAP should not be used to show the general relationship between a variable and predicted risk. This is better accomplished by Partial Dependence Plots (PDP). Moreover, the exponential computational complexity of standard SHAP makes it impractical for large high-dimensional datasets without optimization. LIME is often a more efficient choice for these scenarios. SHAP assumes that features contribute independently to predictions, which may not reflect true biological interactions. SHAP values capture the model associations, not necessarily causation; a high SHAP value could signal an unmeasured proxy, not a direct cause. Furthermore, the average patient in the data may not represent typical clinical cases. Additionally, SHAP explanations can be unstable for patients near decision boundaries.

SHAP Beeswarm plot aggregates SHAP values to reveal population-level patterns of feature impact distributions, directions, and interactions. Each dot represents the SHAP value of one patient, with a color gradient from purple to yellow indicating normalized feature value from low to high. The horizontal position shows the impact: negative values (left) lower disease risk and positive values (right) increase the risk. ST slope shows a clear discriminatory pattern, with low values associated with strongly negative contributions and high values with strongly positive contributions to the predicted risk. Sex shows a bimodal clustering pattern, which shows that for categorical predictors with strong and consistent effects, their SHAP values tend to cluster by categories. Higher age values are associated with higher model-predicted risk, but its variable effect suggests interaction with other factors. Resting ECG and fasting blood sugar cluster near zero, which illustrates that SHAP can reveal variables that contribute minimally to risk prediction in the presence of other strong predictors. Oldpeak and maximum heart rate show dispersed, patient-specific effects, underscoring the value of patient-specific explanations.

SHAP Beewarm plot is a good tool to visualize the distribution of variable impact across variable values, rather than relying on a single number for average impact. This plot can show the magnitude, direction, and spread of effects for each variable across different patients. SHAP Beewarm can be more complex to interpret than SHAP waterfall plot. For specific patients’ risk explanations, SHAP waterfall plot is a better option since it explicitly decomposes the predictions into additive contributions. Moreover, when there are certain clinically important subgroups with very few patients, SHAP Beewarm plot may be unsuitable, since the data points for rare but important categories may be overshadowed by the majority class in the summary plot, masking their clinical significance. Summary plots can mask patient-level details and feature interactions, particularly when the impact is bimodal or skewed. A feature with low average impact can be critical for specific subgroups, risking underestimation of its clinical importance. The color coding reflects only the raw feature values without clinical context: a high cholesterol might be normal or represent good control on statin therapy rather than risk.

LIME addresses the same patient-level question as SHAP - why did this patient receive this prediction - but uses a different approach. It fits a surrogate model in the neighborhood around a specific case. It generates perturbed samples around the patients’ feature values, obtains the black-box model predictions for them, weights these samples by proximity, and then trains a sparse linear model (or some other interpretable models) whose coefficients approximate local feature importance. LIME explained Patient 22's low predicted heart disease risk by decomposing it into feature contributions: the absence of exercise-induced angina (-0.13) and chest pain type 2 (-0.02) contribute negatively to the predicted risk, while a flat ST slope (+0.19) contributes positively. Oldpeak contributes a small positive effect while maximum heart rate provides a small negative effect. For simplicity, the plot displays only the top 5 features obtained from sparse local linear approximation. This illustration shows how a complex model's local behavior around a single observation can be approximated, and individual predictions can be interpreted as a weighted sum of feature contributions.

LIME offers quick, straightforward, model-agnostic explanations by listing the most locally important features for a single prediction. It helps answer “why this prediction?” and is useful when Shapley values are too computationally expensive. LIME should not be used when we need consistent explanations for the overall predictions instead of specific cases, as LIME results can vary depending on how samples are generated. It can also be misleading when the simple local model does not accurately represent the true behavior of the underlying model. As a surrogate, LIME may misrepresent models with highly non-linear decision boundaries. Synthetic data generation introduces instability for patients in data-sparse regions or near decision boundaries. Small changes in local sampling or perturbation can produce different explanations for the same patient. The sparse feature selection might exclude clinically relevant factors. LIME’s result depends on the neighborhood size and perturbation strategy; poor parameter selection may capture noise instead of meaningful model behavior. Unlike SHAP, LIME explanations are not additive and do not sum to the total prediction difference from baseline, making them difficult to verify for completeness.

Beyond individual patients, biomedical scientists often want to know how a single variable affects predicted risk across the full population. Partial dependence plots (PDP) answer this by showing the marginal effect of features on predictions, averaging over the population. PDP answer: “If feature X was set to a specific value for everyone, what would the model predict on average?” Repeating this for a range of X values traces how the predictions change, offering a visualization for interactions and non-linearities effects. This allows biomedical scientists to verify expected patterns or identify counterintuitive trends. However, averaging may mask subgroup differences or interactions, which are better revealed by ICE plots.

PDP plots visualize features’ average effects on predicted probabilities across the top 3 predictors. ST slope shows a threshold effect: value 1 (upsloping) show low risk, flat and downsloping at values 2-3 show higher risk, with a slight decline at value 3. Chest pain type reveals a clearly distinct risk profile, with type 4 demonstrating substantially higher risk than types 1-3. Oldpeak shows a non-linear association with model-predicted risk: risk is lowest at-1, increases gradually with mild ST depression (-1-2mm), then steeply increases beyond 2mm, with a plateau at 0.7 for severe depression (>3mm). These illustrates that PDP can reveal non-linear, step-wise, or diminishing marginal effects which are important to understand individual variable’s contribution to the risk predictions. These patterns in the model prediction are broadly consistent with commonly described diagnostic patterns in cardiology. However, this consistency does not constitute clinical validity. Whether these relationships generalize requires formal evaluation on a held-out test set.

PDP should not be used when features are highly correlated, as it assumes feature independence. This may lead to visualization of impossible variable combinations and unrealistic relationships. PDP can also be misleading when strong interactions exist, since averaging can hide important patterns in the model; when datasets are small and not having enough points to average over reliably; when you need causal answers since PDP shows association, not cause-and-effect. PDPs can obscure heterogeneity and interaction effects by averaging across all instances. PDP also assumes feature independence, which is frequently invalid in clinical practice. For example, the impact of cholesterol might depend on age, blood pressure, and other factors, but PDPs cannot reveal these complex interdependencies. It is a good practice to complement PDPs with ICE plots or to plot PDPs stratified by a second feature (or use 2D PDP) to check for interaction effects.

Population averages can mask important patient-level variation. ICE plots address this by showing individual prediction curves, revealing how changes in one feature affect specific cases. ICE plots reveal heterogeneity masked in population averages. They answer: for this patient, how does the model’s predicted risk change as feature X is varied while other features are held fixed. Parallel ICE lines indicate an additive, globally consistent effect (no major interactions) of the features; diverging or crossing ICE lines indicate interactions, where feature’s effect depends on the patient’s characteristics.

ICE plots visualize individual prediction trajectories across four key predictors. Each gray line corresponds to one patient, with Patient 22 highlighted in red, and the yellow line showing the population average PDP. ST slope shows heterogeneous trajectories between values 1 and 2, suggesting that the change from uploading to flat affects observations differently, and mostly parallel trajectories between values 2 and 3, suggesting a more globally consistent effect. Patient 22 follows the same directional trend as the population, while maintaining a consistently lower predicted risk than the population average across all ST slope values. Chest pain type exhibits heterogeneity and some patients have sharp risk increase between types 3 and 4, including Patient 22. Oldpeak displays even greater patient-specific variation. ICE curves diverge around 0: some are U-shaped, others monotonic. This crossing pattern is an illustration of how ICE plots can detect subgroup-level interaction effects that PDPs would smooth over. Higher maximum heart rate values are associated with lower model-predicted risk, with varying slopes across observations. The variation in trajectory gradients illustrates that ICE plots can quantify the degree to which a feature’s effect is consistent versus individual-specific.

ICE plots are useful for precision medicine and model validation, confirming whether population-level effects apply to individuals. Divergence patterns in ICE highlight feature interactions that PDP alone will miss. ICE plots should not be used when you have too many instances, since the overlapping lines become unreadable clutter. They also should not be used when a clean global summary is needed. High-dimensional interactions are hard to display in ICE plots. Apparent heterogeneity can stem from noise, overfitting, or data limitations, rather than true clinical heterogeneity.

In summary, PDPs and ICE plots focus on marginal feature effects on model predictions. They show: “All else being equal, increasing X in the model input is associated with higher or lower model predictions on average”. This is analogous to biomedical scientists examining a nomogram or risk calculator that shows the incremental risk associated with different levels of a factor. Interpretation should remain within the bounds of the observed data to avoid unreliable extrapolation.

In high-stakes clinical applications, inherently interpretable models are often preferable to post-hoc explanations of black-box models, which may not reliably reflect model logic. Interpretable models enable detection of biases in the training data and uncover discriminatory behavior in ML models. Traditional models, such as logistic regression, rule-based models, and simple decision trees, are often considered more interpretable than complex black-box models because their prediction logic can be inspected directly. These should be the primary choice in high-stakes clinical settings where model transparency is a prerequisite for safety and regulatory compliance. They are ideal when the relationship between features is expected to be relatively linear or when a simple, deployable scoring system is required. Classical interpretable models should not be used when relationships are highly nonlinear or involve complex interactions, as they may oversimplify the data. They can also be misleading if their assumptions are violated, leading to poor model fit and incorrect conclusions. Classical models often struggle to capture complex, non-linear interactions common in high-dimensional healthcare data.

Rule-based models output human-readable decision rules or formulas instead of any computed quantities. These include decision trees, decision lists, and sparse rule sets. The advantage is that the model is self-explanatory – the logic by which it makes predictions can be directly inspected. Many clinical scoring systems and risk calculators fall into this category, either as simple scoring rules or flowcharts. RuleFit algorithm is an example of such model. RuleFit was developed by Friedman and Popescu (2008), which combines the predictive power of ensembles with interpretability. It first generates many candidate rules by training small decision trees, each path is a rule (e.g., “age > 70 and BP > 140 -> outcome”). It then performs a regularized linear regression to fit and assign weights to those binary rules (true/false flags). The result is a set of if-then rules with coefficients. For instance, in the heart disease example, RuleFit might generate a rule such as: "If Age > 60 AND Chest Pain Type = 4 AND ST Slope = 2, then Risk +1.2". This allows a clinician to directly access the model's logic.

RuleFit is particularly effective when biomedical scientists seek if-then logic that mirrors diagnostic flowcharts while maintaining the predictive power. Rule-based ensembles should not be used when the feature space is high-dimensional, since the number of rules explodes and interpretability is lost. They can also be misleading if the rules capture spurious patterns or do not generalize well to small changes in the data. The primary drawback is rule overlap, where a single patient may trigger multiple, sometimes conflicting rules, making the ultimate prediction harder to trace than a simple decision tree. RuleFit produces an interpretable model that expresses the outcome as a weighted sum of human-readable rules. It tells the biomedical scientists which combinations of features are informative and how they are combined. This can sometimes mirror the way biomedical scientists reason, but with data-driven thresholds and selections. Bayesian Rule Lists (BRL) is another method that extends from this concept, to learn a sparse decision list – essentially a series of ordered if-else rules – using Bayesian inference to impose simplicity.

The paper illustrates explainable machine learning methods through a single demonstration task based on a random forest model applied to structured clinical data. This choice was intentional to emphasize conceptual understanding of explanation techniques in a setting that closely mirrors many clinical risk prediction models. Readers can inspect and directly interpret XML outputs in terms of named, recognizable variables (e.g., ST slope, chest pain type). However, since the primary aim is XML methods demonstration, no test-train split was implemented, and readers should not draw clinical conclusions from these examples.

Modern healthcare machine learning increasingly involves high-dimensional modalities such as medical imaging, genomics, and multimodal EHR data. The explainability techniques investigated within this work are also widely used in these contexts, although their visualization and interpretation may differ. For example, the model-agnostic methods (SHAP, LIME, PDP, ICE) are broadly applicable to deep learning models as well. Future work could extend these demonstrations to multimodal or imaging-based models to further illustrate these techniques in contemporary clinical AI applications. Other XML methods specific to deep learning (eg: Saliency Maps, Layer-wise Relevance Propagation) that complement the model-agnostic approaches discussed here are reviewed in [45].

The complexity of ML models in clinical decision-making strengthens the need for explainable systems. XML offers biomedical scientists a toolkit to reason with models, connect them with clinical knowledge, and detect biases or unintended shortcuts in model logic. Explainability is key in model auditing and fairness diagnostics. However, no single explanation method suffices for all clinical contexts. Instead, biomedical scientists must adopt a toolbox approach: global methods (e.g. feature importance, PDPs) allow population-level validation for model predictions, while local methods (e.g. SHAP, LIME) support patient-level interpretability. Agreement between methods increases confidence that a feature is consistently influential in model predictions, while divergences highlight factors that need closer examinations.

Limitations specific to each XML method are discussed within their respective sections. Beyond these limitations, a common challenge across XML approaches is their potential to produce misleading explanations when input features are correlated or when the underlying model is misspecified, yielding outputs that may appear interpretable but lack reliability. Furthermore, the explanations generated by XML methods in this study should be interpreted as reflecting associations learned by the model, rather than causal relationships between variables and outcomes. Nonetheless, these methods remain valuable when applied with appropriate caveats, and mitigation strategies exist to address their limitations.

The field continues to evolve with several promising directions, including causal explanations (What if this patient’s blood pressure were lower?), uncertainty quantification, and interactive explanation systems. The ultimate goal of XML in healthcare is to foster a collaborative environment where biomedical scientists can actively shape model development, curate features, and contextualize predictions. The biostatistics community has a unique opportunity and responsibility to develop explanation methods that are both rigorous and usable, bridging the algorithmic sophistication of modern ML with the interpretive needs of clinical practice.

In summary, the integration of explainable ML into healthcare is not merely a technical challenge; it is a translational one. By grounding explanation methods in statistical intuition and presenting them in clinically resonant terms, we can foster a new generation of biomedical scientists capable of navigating and shaping the ML-enabled future of healthcare.

Improvements for AI systems

Based on the paper, here are the specific improvements I can implement in AI systems, along with what the improved system can do:


Improvement: Implement SHAP values as the default local explanation engine, ensuring that feature contributions sum exactly to the prediction difference from the baseline (e.g., Age +2%, Blood pressure +4%, Medications −1% = +5% total risk).

What the improved system can do: For any individual patient, it can output a precise, auditable breakdown of why a risk score was assigned, allowing clinicians to verify that no hidden factors are driving the prediction.

Improvement: Automatically cross-validate global explanations (feature importance, PDPs) against local explanations (SHAP, LIME) for each prediction. Flag cases where they disagree (e.g., a feature ranked low globally but high locally).

What the improved system can do: It can detect when a model is relying on a feature that is only important for a specific subgroup, preventing clinicians from over-trusting population-level summaries for individual patients.

Improvement: Integrate Individual Conditional Expectation (ICE) plots with automatic detection of crossing or diverging lines, which indicate feature interactions that PDPs would mask.

What the improved system can do: It can automatically alert the user when a variable's effect is not uniform across patients (e.g., Oldpeak increases risk for some patients but decreases it for others), prompting a deeper clinical review rather than a single average effect.

Improvement: Add a RuleFit or Bayesian Rule List module that generates human-readable if-then rules (e.g., "If Age > 60 AND Chest Pain Type = 4 AND ST Slope = 2, then risk increases by +1.2") as a secondary explanation layer.

What the improved system can do: It can provide a transparent, inspectable logic path for regulatory submission or for clinicians who prefer flowchart-style reasoning, without sacrificing the predictive power of the black-box model.

Improvement: Implement a stability metric that quantifies how much SHAP or LIME explanations change when the input is slightly perturbed (e.g., ±5% in continuous features). Flag patients near decision boundaries where explanations are unreliable.

What the improved system can do: It can warn clinicians when a prediction is borderline and the explanation is unstable, reducing the risk of overconfidence in ambiguous cases.

Improvement: Before generating PDPs or SHAP values, compute feature correlation matrices. If high correlation exists (e.g., >0.7), either warn the user or switch to conditional explanations (e.g., PDP stratified by the correlated feature).

What the improved system can do: It can avoid generating misleading explanations that assume feature independence, which is a common failure mode in clinical data (e.g., age and cholesterol are often correlated).

Improvement: Enhance SHAP beeswarm plots with automatic subgroup stratification (e.g., by sex, age group, or comorbidity). Highlight rare but clinically important subgroups that may be overshadowed by the majority class.

What the improved system can do: It can surface hidden disparities in model behavior (e.g., The model relies on chest pain type for men but on ST slope for women), supporting fairness audits and reducing health disparities.

Improvement: Attach confidence intervals to SHAP values and PDP curves using bootstrap resampling or Monte Carlo dropout.

What the improved system can do: It can tell a clinician not just this feature increased risk by 3% but this feature increased risk by 3% (95% CI: 1–5%), preventing overinterpretation of noisy estimates.

Improvement: Generate a structured, human-readable report for each prediction that includes: (a) the top contributing features with direction and magnitude, (b) a comparison to the population average, (c) a stability score, and (d) any relevant interaction warnings.

What the improved system can do: It can produce a document that is directly usable in clinical notes, patient discussions, or regulatory filings, without requiring the clinician to interpret raw plots.

Improvement: Add a persistent disclaimer and a causal mode toggle. In default mode, all explanations are labeled as model associations, not causal. In causal mode, the system only uses methods that explicitly model counterfactuals (e.g., What if this patient's blood pressure were lower?) and requires additional validation data.

What the improved system can do: It can prevent clinicians from mistakenly treating model explanations as causal evidence, which is a critical error in treatment decisions.

The improved system can:

  • Explain why a specific patient received a prediction, with verifiable, additive contributions.

  • Detect and warn when explanations are unstable, unreliable, or based on correlated features.

  • Reveal hidden interactions and subgroup-specific effects that average summaries miss.

  • Provide human-readable rules for high-stakes decisions without losing predictive power.

  • Generate audit-ready reports with uncertainty bounds and fairness checks.

  • Prevent causal misinterpretation by clearly distinguishing association from causation.

These improvements directly address the paper's core message: XML is not a single tool but a toolbox, and the system must guide the user to the right tool for the right clinical question while flagging when explanations may be misleading.

Related papers