Detecting and explaining clinical-omics inconsistencies to improve patient cohort stratification: an application to Parkinson's disease
cs.LG
Submitted: 2025-07-04
Updated: 2026-09-08
Comments: 21 pages, 2 figures
Code: https://github.com/JoseAdrian3/MLASDO
License: http://creativecommons.org/licenses/by/4.0/
The gist: Discrepancies between clinical diagnoses and omics profiles within a characterized cohort may reflect misdiagnosis, hidden subgroups or prodromal disease states.
Terminology
Abstract
Discrepancies between clinical diagnoses and omics profiles within a characterized cohort may reflect misdiagnosis, hidden subgroups or prodromal disease states. We propose MLASDO, a tool to detect and characterize such discrepancies before downstream analyses. MLASDO (1) detects outliers using two unsupervised methods, thereby flagging potential poor-quality samples; and (2) identifies and characterizes anomalous samples (ASs), i.e., individuals whose molecular profile resembles the opposite clinical class. We applied MLASDO to the Parkinson's Progression Markers Initiative (PPMI) and Parkinson's Disease Biomarkers Program (PDBP) Parkinson's disease (PD) cohorts. In PPMI, it detected 26 outliers and 12 ASs: 5 anomalous healthy controls (AHCs) and 7 anomalous PD cases (APDs). AHCs exhibited higher cerebrospinal fluid (CSF) A 1-42 levels than controls (P < 0.0396), suggesting resistance to cognitive decline. AHC-specific genes were enriched for the MAPK pathway (P<0.0145), implicated in PD pathogenesis. One AHC later received a different neurological disorder, three months after enrollment. In PDBP, it identified 10 outliers and 8 ASs: 2 AHCs and 6 APDs. One AHC exhibited 24 clinical features consistent with a PD-like phenotype, including severe motor impairment. Results are compiled in an interactive report for inspection and querying, highlighting clinically meaningful individuals otherwise overlooked in conventional analyses.
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks