Learning Under Treatment-Induced Label Indeterminacy with Expert Annotations of Counterfactual Outcomes: A Case Study in Neurological Prognostication
Xiaobin Shen, Chloe Y. H. Huang, Jonathan Elmer, George H. Chen
Carnegie Mellon University · University of Pittsburgh
cs.LG
Submitted: 2026-08-12
Updated: 2026-08-14
Comments: Machine Learning for Healthcare (MLHC) 2026
Code: https://github.com/xiaobin-xs/learning-under-label-indeterminacy
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 75/100
The gist: This paper addresses the problem of treatment-induced label indeterminacy in clinical prediction models, using post-cardiac-arrest neurological prognostication as a case study.
Terminology
Summary
This paper addresses the problem of treatment-induced label indeterminacy in clinical prediction models, using post-cardiac-arrest neurological prognostication as a case study. The authors note that "Clinical prediction models are often developed as if the outcome of interest were cleanly observed for every patient. This assumption fails when treatment decisions make the clinically relevant outcome permanently unobservable. In their cohort of 2,497 patients, 1,429 patients had outcomes rendered indeterminate by treatment decisions such as
withdrawing or limiting life-sustaining therapies, which immediately led to death, so we do not know what would have happened otherwise. These patients were reviewed by independent clinical experts who
provided their guesses of counterfactual outcomes about what would have happened to the patients," and are referred to as uncertain cases. The remaining 1,068 patients, for whom clinically relevant outcomes (e.g., regaining consciousness) are observed, are referred to as certain cases.
The paper makes three key contributions. First, it formalize[s] predictive modeling under treatment-induced label indeterminacy when we have target labels that differ between certain and uncertain cases.
Second, it "introduce[s] a split evaluation framework that evaluates certain cases using standard observed-label metrics (e.g., AUROC, Brier score) and uncertain cases using a notion of agreement or 'alignment' with clinical expert guesses of counterfactual outcomes. Third, it
empirically characterize[s] the resulting tradeoff across various tabular models and neural networks, and introduce[s] a simple two-parameter neural objective that makes movement along this tradeoff explicit: the weight placed on observed poor-outcome cases and the degree of alignment to expert guesses on counterfactual outcomes of uncertain cases."
The proposed neural model uses a weighted binary cross-entropy loss on certain cases, where ωbad ≥ 1 controls the relative emphasis placed on observed poor-outcome cases,
and an auxiliary alignment loss on uncertain cases: Lalign(θ) = (1/UE) Σ (fθ(xj) − p̃j)2,
where p̃j is the patient-level expert-derived reference probability. The full objective is L(θ) = Lcertain(θ) + λalign Lalign(θ),
where λalign ≥ 0 controls the strength of uncertain-case alignment.
The main empirical finding is a clear certain–uncertain tradeoff. The authors state: "Across model configurations, certain-case AUROC changes relatively little, whereas uncertain-case MAE spans a wide range. At the same time, the right panel shows that lower uncertain-case MAE is generally associated with worse certain-case Brier score. In other words, models with similar discrimination on the observed-label cohort can place uncertain patients very differently on the scale for probability of recovery, and those changes are accompanied by a clear tradeoff in certain-case probability accuracy. This reliability gap is largely invisible if evaluation is restricted to AUROC or other ranking-based metrics alone."
For example, Table 1 shows that "the standard observed-only neural model and the strongly aligned neural model have nearly identical certain-case AUROC (0.882 vs. 0.884), yet their certain-case Brier scores and uncertain-case MAE differ dramatically (0.116 vs. 0.407 and 0.410 vs. 0.053, respectively). The authors conclude:
Thus, ranking performance (like AUROC) alone would suggest little difference between these models, even though they behave very differently on both the certain-case probability scale and the uncertain cohort."
The paper also examines prediction-scale behavior, finding that In the observed-only models, uncertain cases receive relatively high recovery probabilities and overlap substantially with certain good cases.
Incorporating expert information shifts uncertain-case predictions downward. The authors note: "The moderately aligned neural model yields an intermediate regime in which uncertain cases are more clearly separated from certain good cases without collapsing all predictions toward zero. By contrast, the strongly aligned neural model compresses uncertain cases and certain poor cases into a narrow low-probability region and also lowers predictions for certain good cases, consistent with the deterioration in certain-case Brier score."
The authors emphasize that the main failure mode in this setting is miscalibration of classifier probabilities
and that "The key object is therefore not a single best model, but the certain–uncertain tradeoff frontier. Better agreement with expert-informed uncertain-case assessments generally came at the cost of worse probability accuracy on certain cases, making model selection an explicit design choice rather than a hidden consequence of standard supervised learning."
The paper also includes extensive appendices with cohort construction details, feature definitions, expert assessment protocols, full tabular baseline results (including random forest, XGBoost, and TabPFN variants), sensitivity analyses for the expert-score mapping, and full neural sweep heatmaps. The authors note limitations: "The analysis is retrospective and based on a single-center cohort, so the certain/uncertain partition and the resulting tradeoff may partly reflect local practice patterns. The uncertain-case target labels come from expert guesses of counterfactual outcomes, which might not actually be accurate, so better alignment with these target labels should not be interpreted as recovery of ground-truth counterfactual probabilities."
Improvements for AI systems
Improvements to AI systems:
-
Add a dual-objective training framework with explicit uncertainty-aware loss terms. The system can be trained using a weighted binary cross-entropy loss on cases with observed outcomes (with a tunable weight ωbad for poor outcomes) plus an auxiliary alignment loss (mean squared error) on cases with treatment-induced indeterminate outcomes, where the target is an expert-derived counterfactual probability. This allows the model to explicitly balance discrimination on certain cases against calibration to expert judgment on uncertain cases, rather than treating all labels as equally reliable.
-
Implement a split evaluation protocol that reports both ranking metrics (AUROC) and probability-calibration metrics (Brier score) separately for certain and uncertain cases. The improved system will surface the certain–uncertain tradeoff frontier during model selection, preventing users from choosing a model based solely on AUROC when it may be severely miscalibrated on uncertain cases (e.g., AUROC 0.882 vs. 0.884 with Brier scores 0.116 vs. 0.407).
-
Introduce a two-parameter neural objective (ωbad, λalign) that makes the tradeoff explicit and tunable at deployment time. The system can generate a family of models along the tradeoff frontier, allowing clinicians to select a model that matches their risk tolerance—e.g., a moderately aligned model that separates uncertain cases from certain good cases without collapsing all predictions toward zero, versus a strongly aligned model that prioritizes expert agreement at the cost of certain-case probability accuracy.
-
Add a calibration-aware prediction layer that detects and flags miscalibration on indeterminate cohorts. The system will monitor the distribution of predicted probabilities for uncertain cases and warn users if predictions overlap heavily with certain good cases (as seen in observed-only models), indicating that the model is not appropriately incorporating counterfactual uncertainty.
-
Enable counterfactual-aware data augmentation using expert guesses as soft labels. Instead of discarding patients with indeterminate outcomes, the system will use their expert-derived probabilities as auxiliary training targets, improving the model’s ability to generalize to real-world scenarios where treatment decisions (e.g., withdrawal of life support) make outcomes unobservable.
-
Provide a model-selection dashboard that visualizes the certain–uncertain tradeoff frontier (e.g., certain-case Brier score vs. uncertain-case MAE) across all hyperparameter configurations, so users can explicitly choose where to operate on the frontier rather than relying on a single default model.
What the improved AI system can do:
-
Predict neurological recovery probabilities after cardiac arrest while explicitly accounting for patients whose outcomes are rendered unobservable by treatment decisions.
-
Offer clinicians a tunable knob (λalign) to control how much the model aligns with expert counterfactual judgments, with real-time feedback on the resulting certain-case calibration cost.
-
Avoid the hidden failure mode of miscalibrated probabilities on indeterminate patients, which standard AUROC-based evaluation would miss.
-
Generate a spectrum of models—from observed-only to strongly expert-aligned—so that users can select the operating point that best matches their clinical context and tolerance for false confidence on uncertain cases.
Abstract
Clinical prediction models are often developed as if the outcome of interest were cleanly observed for every patient. This assumption fails when treatment decisions make the clinically relevant outcome permanently unobservable. As a case study of this problem, we consider post-cardiac-arrest neurological prognostication using a cohort of 2,497 patients, including 1,429 patients whose outcomes were rendered indeterminate by treatment decisions. These patients with indeterminate outcomes were reviewed by independent clinical experts, who provided their guesses of counterfactual outcomes about what would have happened to the patients. We refer to these patients as uncertain cases. We also have patients for whom we observe their clinically relevant outcomes; we refer to these patients as certain cases. We propose a framework for evaluating prediction models that explicitly splits the evaluation between certain and uncertain cases. Here, we cannot easily evaluate both types of cases in a uniform manner as the available target labels differ. We then propose a simple prediction model that uses target labels from both certain and uncertain cases in a manner that allows us to trade off between them. Across the proposed neural model and a collection of tabular baselines, models with similar certain-case AUROC can nevertheless differ substantially in both certain-case Brier score and their probability estimates for uncertain cases. Improving alignment with target labels of uncertain cases for our proposed model generally comes at the cost of worse accuracy on certain cases, highlighting an explicit tradeoff that standard evaluation conceals. These results show that when treatment decisions determine whether clinically meaningful outcomes remain observable, conventional evaluation metrics can miss important failure modes in the very patients for whom prognostic support matters most.
Sources
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks