Conditional Independence Tests for Constraint-Based Causal Discovery: A Survey
Pavel Averin, Theodoros Moysiadis, Ioannis Katakis
University of Nicosia
stat.ML, cs.LG
Submitted: 2026-08-11
Updated: 2026-08-12
Comments: 33 pages. Published in Transactions on Machine Learning Research (07/2026). https://openreview.net/forum?id=3jzafJK8Tz
Journal ref: Transactions on Machine Learning Research (07/2026)
Project page: https://christophm.github.io/interpretable-ml-book
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 85/100
The gist: This survey reviews conditional independence (CI) testing with emphasis on assumptions, robustness, and scalability in high-dimensional and mixed-type settings common in biomedical domains.
Terminology
Summary
This survey reviews conditional independence (CI) testing with emphasis on assumptions, robustness, and scalability in high-dimensional and mixed-type settings common in biomedical domains. The survey organizes widely used CI methods into six families: partial-correlation, contingency-table, regression, nearest-neighbor, kernel, and machine-learning–based. Special emphasis is provided on the robustness layers that address the limitations of these families. For each family, the survey examines when CI decisions reflect the data-generating distribution and when they fail. By this, we link test-level properties, including power decay with conditioning set size and asymmetric type I/II error consequences, to graph-level errors in skeleton recovery and v-structure orientation. The survey also compares adoption across major R and Python libraries and summarizes open challenges, including mixed-type CI testing without discretization, small-sample error control, and strategies for improving scalability of CI-testing.
Improvements for AI systems
Improvements to AI Systems:
- Robustness-Aware CI Testing Module
-
Integrate a meta-layer that detects when a chosen CI test’s assumptions (e.g., linearity, normality, no hidden confounders) are violated, and automatically switches to a more robust family (e.g., from partial-correlation to kernel-based or ML-based tests).
-
Improved capability: AI can now perform causal discovery or feature selection in messy biomedical data (e.g., gene expression with non-linear interactions) without silently producing false edges or wrong v-structure orientations.
- Adaptive Power-Scaling for High-Dimensional Conditioning
-
Implement a dynamic test selection mechanism that estimates the decay in statistical power as the conditioning set size grows, and recommends a fallback (e.g., using nearest-neighbor tests for sparse dependencies, or ML-based tests for dense, non-linear dependencies).
-
Improved capability: AI can reliably test conditional independence in datasets with hundreds of covariates (e.g., EHR data) while controlling false positives, even when the true dependency is weak and high-order.
- Mixed-Type Data Handling Without Discretization
-
Build a unified CI test that uses a copula-based or rank-based transformation for continuous variables and a separate likelihood-ratio test for categorical variables, then combines them via a meta-test that preserves type I error.
-
Improved capability: AI can directly analyze mixed clinical data (e.g., continuous lab values + categorical diagnoses) without losing information from binning, leading to more accurate graph recovery in personalized medicine.
- Small-Sample Error Control via Bootstrap-Calibrated Tests
-
Add a calibration layer that uses residual bootstrap or permutation-based null distributions to adjust p-values for small n (e.g., n < 50), specifically for kernel and ML-based CI tests that otherwise over-reject.
-
Improved capability: AI can perform causal inference on rare-disease cohorts or small pilot studies, where traditional asymptotic tests are unreliable, and still maintain valid type I error rates.
- Scalability-Aware Test Selection for Large Graphs
-
Implement a cost-benefit scheduler that estimates computational cost (e.g., O(n 2) for kernel tests vs. O(n) for partial-correlation) and statistical accuracy trade-off, then chooses the fastest test that meets a user-defined error budget for each edge in a graph.
-
Improved capability: AI can run CI testing on entire patient networks (millions of edges) in reasonable time, while prioritizing high-risk edges for more robust (but slower) tests, enabling real-time clinical decision support.
- Asymmetric Error Consequence Mitigation
-
Add a decision-theoretic layer that assigns asymmetric costs to false positives (e.g., adding a wrong causal edge) vs. false negatives (missing a true edge) based on downstream task (e.g., drug target identification vs. risk factor screening), and adjusts the CI test threshold accordingly.
-
Improved capability: AI can tailor its causal discovery output to the specific application—e.g., being conservative for treatment effect estimation but more sensitive for hypothesis generation—reducing costly misdirections in biomedical research.
- Library-Agnostic CI Test Interface
-
Develop a standardized API that wraps R (e.g.,
pcalg,bnlearn) and Python (e.g.,scipy,sklearn,causal-learn) implementations, with automatic checks for input type compatibility and assumption pre-screening. -
Improved capability: AI systems can seamlessly switch between different CI test implementations without code rewrite, enabling rapid benchmarking and deployment across diverse biomedical pipelines.
What the improved AI system can do:
-
Perform reliable causal discovery on high-dimensional, mixed-type, small-sample biomedical data (e.g., genomics, EHR, imaging) with controlled error rates.
-
Automatically adapt its testing strategy to data complexity and computational constraints, avoiding both false causal claims and missed true dependencies.
-
Provide interpretable confidence in graph edges, with explicit warnings when assumptions are violated and when power is low.
-
Scale to large-scale patient networks while maintaining statistical rigor, and tailor its output to the clinical or research decision context.
Abstract
Conditional Independence (CI) tests are the statistical engine of constraint-based causal discovery: in algorithms such as PC (Peter-Clark) and FCI (Fast Causal Inference), skeleton pruning and key orientations follow directly from CI decisions. This survey reviews CI testing with emphasis on assumptions, robustness, and scalability in high-dimensional and mixed-type settings common in biomedical domains. The survey organizes widely used CI methods into six families: partial-correlation, contingency-table, regression, nearest-neighbor, kernel, and machine-learning-based. Special emphasis is provided on the robustness layers that address the limitations of these families. For each family, the survey examines when CI decisions reflect the data-generating distribution and when they fail. By this, we link test-level properties, including power decay with conditioning set size and asymmetric type I/II error consequences, to graph-level errors in skeleton recovery and v-structure orientation. The survey also compares adoption across major R and Python libraries and summarizes open challenges, including mixed-type CI testing without discretization, small-sample error control, and strategies for improving scalability of CI-testing.
Sources
- Permutation-Based Rank Test in the Presence of Discretization and Application in Causal Discovery with Mixed Data
- Efficient Ensemble Conditional Independence Test Framework for Causal Discovery
- Multi-Agent Causal Discovery Using Large Language Models
- Mixed Graphical Models for Causal Analysis of Multi-modal Variables
- gCastle: A Python Toolbox for Causal Discovery
Related papers
- Behavior of prediction performance metrics with rare events
- Optimal Estimation of Generic Dynamics by Path-Dependent Neural Jump ODEs
- A Posterior-Dynamics Framework for Imaging Inverse Problems with Pretrained Diffusion Priors
- One Permutation Is All You Need: Fast, Deterministic Feature Importance and Model Stress-Testing
- Online Conformal Prediction for Non-Exchangeable Panel Data
- Deep Time-Series Forecasting in 10 Years: A Survey