A Comparative Study of Feature Selection Methods for EHR Diagnosis Codes in Opioid Use Disorder Prediction

arXiv:2608.04180 · cs.LG · Submitted 2026-08-18 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Title Not Provided in Snippet".

Jane: The paper was written by Leslie Marino, Cari F. Besserman and Xia Zhao from Columbia University and Suffolk County Department of Health and Stony Brook Medicine and Center for Behavioral Health Statistics and Quality and Substance Abuse and Mental Health Services Administration and Centers for Disease Control and Prevention (CDC) and U.S. Agency for Healthcare Research and Quality (AHRQ).

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Paper summary: Tom: So, this study really sets up a head-to-head comparison between five different ways to pick features from the EHR data. We’re looking at Recurrence Enrichment, NTK Sensitivity, LightGBM-SHAP, Elastic Net, and LLM Semantic.

Jane: It’s like testing five different kinds of filters to see which one gives you the cleanest water for your prediction model. They are all trying to find the most informative diagnosis codes that predict OUD risk.

Lu: I'm fascinated by how they are using a unified framework, making sure that because we are comparing these methods, the differences in performance come solely from *how* we selected the features, not from any technical inconsistencies in data processing.

Meng: From an engineering standpoint, this is crucial; it isolates the feature selection strategy as the variable to optimize. We are essentially asking which feature selection method creates a high-quality input dataset for our final model.

Lalam: My take is that the paper suggests we need a balance between predictive power and reliability, and stability—the consistency of how those features are chosen—is just as important as raw accuracy.

Page 1 of the paper: Tom: The introduction sets the stage by highlighting the opioid crisis in the United States, which is a massive public health problem. It establishes that EHR data is our primary tool to address it.

Jane: It’s not just that we use EHR data; they are pointing out how difficult raw diagnosis-code representations are because they are sparse and high-dimensional at the patient level.

Lu: The authors stress that these codes, while useful, come with issues like imbalanced frequencies and noise, making the full set unstable for predictive modeling.

Meng: This points to the fundamental problem we face in any big data project—we have too much information that is actually redundant or just confusing the signal.

Lalam: The paper clearly defines feature selection not as a convenience, but as a central methodological problem that must be solved before any classifier can even start training.

Page 2 of the paper: Tom: Moving into the methodology, we see they have a lot of prior work mentioned in this field, like Deep Patient and KESER, showing how much research has gone into EHR modeling.

Jane: It's important to note that they are moving beyond just these existing methods and conducting a systematic comparison across three methodological categories: statistical ranking, machine-learning embedded selection, and LLM-assisted identification.

Lu: I’m particularly interested in how the paper frames this gap—since most previous work focused on representation learning, they are directly tackling the specific challenge of subset selection for diagnosis codes.

Meng: The practical implication here is that we’re moving away from "dump everything into the model" and instead actively curating a set of candidates to improve our workflow.

Lalam: By comparing these different approaches, they want to build a blueprint for robust EHR-based prediction models, considering factors like stability alongside clinical relevance.

Page 3 of the paper: Tom: The first method, Recurrence Enrichment, is a filter-based approach that looks at how often a diagnosis appears across different encounters for OUD-positive patients.

Jane: It’s not just looking at if the code exists; it’s counting the recurrence, which makes sense because repeated diagnoses might be more indicative of an ongoing issue than a single occurrence.

Lu: The formulas show they are calculating a population-normalized mean recurrence, which is a clever way to quantify that elevation in OUD-positive patients relative to others.

Meng: From an engineering perspective, this gives us a clear score—the d score—which is just the difference between positive and negative groups, making it highly quantifiable for ranking.

Lalam: This suggests that even basic statistical filters can capture meaningful patterns if they are designed to account for repeated events in the patient’s history.

Page 4 of the paper: Tom: Next, we have NTK-Motivated Early Gradient Sensitivity, which is a very clever way to look at what happens *before* training even starts.

Jane: It ranks features by their gradient magnitude at initialization, essentially looking for initial "sensitivity" in the data to predict OUD.

Lu: This is theoretically powerful because, based on NTK theory, high sensitivity early on suggests that feature will have significant expected influence during the entire subsequent training process.

Meng: The challenge here is that using a pre-training gradient can be noisy, but it seems like a highly efficient way to get a predictive ranking without having to train and iterate through all of those full model cycles.

Lalam: This approach shows how we can use theoretical insights from machine learning dynamics to guide our feature selection process, moving beyond just simple frequency counts.

Page 5 of the paper: Tom: The third method is LightGBM-SHAP, which uses a gradient boosting model to find nonlinear relationships and then measures the average contribution of each diagnosis code.

Jane: It’s not just about correlation; it’s about how much that specific diagnosis contributes to the final prediction in a complex, non-linear way.

Lu: SHAP values are really powerful because they capture conditional dependencies—how a single code' helps predict OUD given other codes are present—which standard linear models might miss entirely.

Meng: This is where LightGBM shines; it handles complexity and sparsity extremely well, providing us with a measure of the average contribution d across the validation set.

Lalam: The use of SHAP allows us to see not just *if* a feature is important, but *how* it provides that predictive power within the complex system, offering a deep level of interpretability.

Page 6 of the paper: Tom: We’ve seen statistical and model-based methods, now let’s look at Elastic Net Regularized Logistic Regression. This method uses regularization to manage complexity.

Jane: It's using two types of penalties—L1 and L2—to ensure the model is sparse while stabilizing estimates, which is essential when dealing with highly correlated diagnosis codes.

Lu: The goal here is controlled sparsity; we want the model to only keep the most robust features, and Elastic Net allows us to control that trade-off using the alpha parameter.

Meng: For an engineer, this provides a very stable way to handle collinearity—when two or more diagnosis codes are essentially saying the same thing—by preventing one from dominating and making decisions unstable.

Lalam: This is about finding a reliable middle ground, ensuring that we get enough predictive power while keeping the model lean and robust against feature correlation.

Page 7 of the paper: Tom: The next page introduces the LLM Semantic approach, using Claude Sonnet four point five to evaluate codes based on clinical plausibility, which is a very novel idea for feature selection.

Jane: It's not just looking at statistics; it’s using AI to assess if a code makes clinical sense as an OUD precursor or if it represents a consequence of OUD itself.

Lu: The two-stage process is genius—first evaluating per-code, and then the LLM selects the final top one hundred based on criteria like relevance and non-redundancy.

Meng: This addresses the qualitative gap in other methods; we are leveraging clinical knowledge embedded in the AI model to find features that data alone might overlook.

Lalam: The LLM approach helps us build a vision of what *should* be predictive, based on medical expertise, which is a powerful way to complement data-driven findings.

Page 8 of the paper: Tom: So now we are looking at the results and how they compare across various feature budgets, from K=fifty up to K=five hundred. This is where the real impact shows.

Jane: We see that performance improves rapidly as more features are allowed, but then it starts leveling off, which is what the diminishing returns look like.

Lu: The most striking finding here is that around three hundred features, most of them—the diagnostically informative ones—are already captured, showing a robust saturation point in the complexity.

Meng: That suggests we can deploy a much more compact feature set for OUD prediction without losing significant predictive power compared to using every single code.

Lalam: This is great news for implementation because it shows that operational efficiency and high performance are achievable simultaneously with a moderate vocabulary size.

Conclusion: Tom: We’ve covered the methodology and the results, showing how these five different feature selection methods perform in predicting OUD using EHR data.

Jane: Overall, "A Comparative Study of Feature Selection Methods for EHR Diagnosis Codes in Opioid Use Disorder Prediction" provides a clear roadmap for selecting robust features.

Lu: The key findings are that NTK Sensitivity provided the best overall balance of accuracy and stability, which is a huge win for reliable AI applications.

Meng: And the fact that performance plateaus around three hundred features means we have a practical limit on how much more data we need to gather to make the biggest impact.

Lalam: My final thought is that this work guides us toward building clinical tools that are not only accurate but also consistent, allowing for a truly reliable advancement in public health interventions.

Leslie Marino, Cari F. Besserman, Xia Zhao

Columbia University · Suffolk County Department of Health · Stony Brook Medicine · Center for Behavioral Health Statistics and Quality · Substance Abuse and Mental Health Services Administration · Centers for Disease Control and Prevention (CDC) · U.S. Agency for Healthcare Research and Quality (AHRQ)

cs.LG

Submitted: 2026-08-18

Updated: 2026-08-19

Comments: Accepted at the AMIA 2026 Annual Symposium. Author list corrected to match the accepted version

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 81/100

The gist: A Comparative Study of Feature Selection Methods for EHR Diagnosis Codes in Opioid Use Disorder Prediction Problem Statement and Context The study addresses the challenge of using electronic health

Key concepts

Feature Selection
The process of choosing informative diagnosis codes from large, sparse EHR data to use in a prediction model. This is necessary because raw codes can be redundant or confusing, requiring a curated set of candidates.
EHR Data
Electronic Health Record data, which serves as the primary tool for addressing the opioid crisis. The episode discusses the challenges associated with using raw diagnosis-code representations from this data.
NTK Sensitivity
A feature selection method that ranks features by their gradient magnitude at initialization, before training begins. This theoretical approach aims to find initial 'sensitivity' that suggests significant influence during subsequent model training.
LLM Semantic Approach
A novel feature selection method where a Large Language Model (Claude Sonnet four point five) is used to evaluate diagnosis codes based on clinical plausibility, assessing if a code makes medical sense as an OUD precursor.

Terminology

Summary

A Comparative Study of Feature Selection Methods for EHR Diagnosis Codes in Opioid Use Disorder Prediction

Problem Statement and Context

The study addresses the challenge of using electronic health record (EHR) data for predicting opioid use disorder (OUD), noting that input variables are often high-dimensional, sparse, noisy, and redundant. In this setting, diagnosis-code features—which are routinely collected, clinically interpretable, and standardized—are widely used to summarize patient morbidity burden. However, the raw representation of these codes is problematic because they are extremely sparse and high-dimensional at the patient level and are influenced by coding conventions, terminology mappings, and site- or time-specific documentation practices. Furthermore, diagnosis-code frequencies are strongly imbalanced, making the full space noisy, redundant, difficult to interpret, and potentially unstable for predictive modeling.

Methodology

The researchers conducted a systematic comparison of five distinct feature selection paradigms applied to diagnosis-code features derived from the Cerner Health Facts database. The study cohort included 11,791,858 patients; 137,214 (1.16%) were identified as having OUD (operationalized using the ICD-10-CM category F1x). To ensure a unified evaluation framework, all ICD-9-CM codes were harmonized to ICD-10-CM and further truncated to the first three characters.

The five methods evaluated are detailed below:

  1. Recurrence Enrichment: A filter-based approach that prioritizes diagnosis features based on their recurrence across encounters for OUD-positive patients, calculating the population-normalized mean recurrence (d+ and d-). Diagnoses are ranked by the enrichment score d = d+ -. Only diagnoses meeting a minimum patient support threshold (n d 50 were retained.

  2. NTK-Motivated Early Gradient Sensitivity: This method ranks diagnosis features by their gradient contribution at model initialization, without iterative training. The features are ranked by the aggregated absolute gradient S d = sum g d(e), which is theorized to indicate greater expected influence during subsequent training.

  3. LightGBM–SHAP: A tree-based approach where a LightGBM model is trained on encounter-level binary diagnosis features. Feature importance is quantified using SHAP (SHapley Additive exPlanations) values (phi d(e)), and diagnoses are ranked by their mean absolute SHAP value over the validation set, d.

  4. Elastic Net Regularized Logistic Regression: A logistic regression model fitted with Elastic Net regularization. Features are ranked by the absolute value of the fitted coefficient d, utilizing the L1 component for sparsity and the L2 component for stabilizing estimates under feature correlation.

  5. Large Language Model–Guided Semantic Prior (LLM Semantic): A two-stage LLM strategy using Claude Sonnet 4.5. Stage 1 evaluates each ICD-10 code based on its description and cohort prevalence, excluding codes that represent consequences or treatments of established OUD to prevent label leakage. Stage 2 selects the final top 100 features based on clinical relevance, specificity, and non-redundancy.

For downstream modeling, the researchers constructed feature subsets under multiple budgets (K in 50, 100, 150, 200, 300, 500). The input vocabulary was restricted to these top-K codes in a BERT-based transformer encoder model trained from scratch for binary OUD prediction.

Evaluation Metrics

Model performance was assessed using three metrics: the Area Under the Precision-Recall Curve (AUPRC) as the primary metric due to class imbalance, the Area Under the Receiver Operating Characteristic Curve (AUROC), and the Optimal F1 score.

Key Results and Findings

  • Overall Predictive Performance: The results demonstrate that performance improves with larger feature budgets with diminishing returns beyond a moderate size. Specifically, NTK Sensitivity achieved the highest overall BERT performance, reaching an AUPRC of 0.291, AUROC of 0.901, and Optimal F1 of 0.319 at K = 500. LightGBM-SHAP produced competitive results (AUPRC = 0.266). Recurrence Enrichment and LLM Semantic yielded the lowest AUPRC values among all methods (AUPRC about 0.191 at K=500).

  • Feature Budgeting: Across all five methods, performance gains became marginal after approximately 300 features, suggesting that a moderate diagnosis vocabulary of 300 codes may be sufficient to capture the most diagnostically informative signal for OUD prediction. The full-vocabulary upper bound (1,908 features; AUPRC = 0.458) remained substantially higher than all feature-selected configurations.

  • Stability Analysis: NTK Sensitivity exhibited near-deterministic behavior, achieving mean Jaccard similarity of 0.983 plus or minus 0.011 at K=100. Recurrence Enrichment also showed high stability (0.950 plus or minus 0.019 at K=100). Elastic Net demonstrated substantially lower and more variable stability (e.689 plus or minus 0.685 at K=100).

  • Cross-Method Overlap: Substantial intersections were observed among NTK Sensitivity, LightGBM-SHAP, and Recurrence Enrichment, with 61 diagnosis codes selected by at least three methods. LLM Semantic introduced a complementary set of features not consistently identified by purely data-driven methods, exhibiting low pairwise overlap (Jaccard 0.156).

Conclusion

The study concludes that NTK-motivated early gradient sensitivity achieved the best overall downstream performance while also demonstrating near-deterministic stability. The findings provide practical guidance for selecting compact, informative, and stable diagnosis feature sets for EHR modeling under realistic computational and deployment constraints.

Improvements for AI systems

As a diligent AI researcher focused on maximizing predictive accuracy and model robustness in high-stakes clinical environments, I have analyzed this paper to identify specific architectural and methodological improvements for any downstream AI system designed for OUD prediction.

The following improvements are highly specific, actionable, and address the identified limitations in the current state-of-the-art models presented in this study.


Improvement: Instead of selecting features based on a single data-driven metric (e.g., frequency or tree importance), the system must implement a hybrid pipeline that fuses the statistical rigor of NTK Sensitivity with the clinical context provided by LLM-guided selection.

  • Process:
  1. Initial Ranking: Run both NTK Sensitivity (to capture early gradient influence) and LLM Semantic Prior (to capture clinically plausible, non-obvious links).

  2. Weighted Fusion: Assign a weight to the NTK score (Score NTK) and a weight to the LLM confidence/relevance score (Score LLM). The final ranking for each code d is determined by: Final Rank proportional to alpha((d)) + beta(Confidence LLM(d)).

  3. Selection: Select the top-K features, where K should be dynamically adjusted based on observed predictive gain (see Improvement #2).

  • System Capability: The improved system can achieve superior stability and clinical relevance simultaneously. It will select features that are both mathematically sensitive to the model's initialization (NTK) and clinically justifiable, allowing it to identify critical, low-frequency long-tail diagnoses that purely data-driven methods often overlook.

Improvement: The system must move away from a fixed feature budget (K=300 or K=500). Instead, it should implement a dynamic stopping criterion based on the marginal gain in predictive performance (AUPRC).

  • Process:

1.Start with a small initial set of features (K start).

2.Iteratively add the next highest-ranked feature from the combined NTK/Semantic list.

3.Monitor the change in AUPRC (AUPRC) relative to the complexity increase (K).

4.Stop adding features when AUPRC falls below a predefined threshold (e.g., 0.01) for subsequent additions, or when the feature-to-prediction gain curve plateaus significantly (as observed near K=300).

  • System Capability: The improved system will achieve optimal computational efficiency and maximal predictive power. It ensures that computational resources are not wasted on redundant or low-signal features beyond the point of diminishing returns, guaranteeing a compact model while maintaining peak performance.

Improvement: The current reliance solely on diagnosis codes is insufficient for robust risk prediction. The system must evolve to incorporate other critical EHR domains as input features, specifically targeting the missing data points identified in the limitations section of the paper.

  • Process:
  1. Temporal Feature Mapping: For each patient encounter, extract time-since-last-event metrics (e.g., time since last prescription, time since last lab test).

  2. Medication/Procedure Encoding: Encode medication classes (via DrugBank mapping) and procedure codes into categorical embeddings or counts.

  3. Dynamic Feature Weighting: Assign a weighted feature set F Total = F Diagnosis, F Medication, F Lab where weights are determined by the clinical relevance (LLM score) and statistical sensitivity (NTK score).

  • System Capability: The improved system can achieve holistic, multi-domain risk assessment. By incorporating medication and lab data, it moves beyond merely capturing what the patient has (diagnosis) to understanding how they are being treated and how their physiology is responding, drastically increasing predictive power.

Improvement: The system must explicitly penalize feature selection processes that exhibit high variance in their rankings across different subsets of training data, addressing the instability observed in Elastic Net and LightGBM-SHAP.

  • Process:

1.During the feature ranking phase (especially for NTK/Shapley), calculate a measure of rank consistency across N bootstrap resamples (e.g., average Jaccard similarity of the top 50 features).

2.Apply a regularization term to the ranking function that penalizes high variance in stability measures, ensuring that features selected are robust and not merely artifacts of specific data splits.

  • System Capability: The improved system will provide reproducible and trustworthy feature sets. This is critical for clinical deployment; the system guarantees that if it were trained on a slightly different cohort, the core predictive features would remain highly consistent, reducing model drift and increasing regulatory confidence.

Related papers