Conditional Independence Tests for Constraint-Based Causal Discovery: A Survey

arXiv:2608.11156 · stat.ML, cs.LG · Submitted 2026-08-11 · Read on arXiv

Pavel Averin, Theodoros Moysiadis, Ioannis Katakis

University of Nicosia

stat.ML, cs.LG

Submitted: 2026-08-11

Updated: 2026-08-12

Comments: 33 pages. Published in Transactions on Machine Learning Research (07/2026). https://openreview.net/forum?id=3jzafJK8Tz

Journal ref: Transactions on Machine Learning Research (07/2026)

Project page: https://christophm.github.io/interpretable-ml-book

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 85/100

The gist: This survey reviews conditional independence (CI) testing with emphasis on assumptions, robustness, and scalability in high-dimensional and mixed-type settings common in biomedical domains.

Terminology

Summary

This survey reviews conditional independence (CI) testing with emphasis on assumptions, robustness, and scalability in high-dimensional and mixed-type settings common in biomedical domains. The survey organizes widely used CI methods into six families: partial-correlation, contingency-table, regression, nearest-neighbor, kernel, and machine-learning–based. Special emphasis is provided on the robustness layers that address the limitations of these families. For each family, the survey examines when CI decisions reflect the data-generating distribution and when they fail. By this, we link test-level properties, including power decay with conditioning set size and asymmetric type I/II error consequences, to graph-level errors in skeleton recovery and v-structure orientation. The survey also compares adoption across major R and Python libraries and summarizes open challenges, including mixed-type CI testing without discretization, small-sample error control, and strategies for improving scalability of CI-testing.

Improvements for AI systems

Improvements to AI Systems:

  1. Robustness-Aware CI Testing Module
  • Integrate a meta-layer that detects when a chosen CI test’s assumptions (e.g., linearity, normality, no hidden confounders) are violated, and automatically switches to a more robust family (e.g., from partial-correlation to kernel-based or ML-based tests).

  • Improved capability: AI can now perform causal discovery or feature selection in messy biomedical data (e.g., gene expression with non-linear interactions) without silently producing false edges or wrong v-structure orientations.

  1. Adaptive Power-Scaling for High-Dimensional Conditioning
  • Implement a dynamic test selection mechanism that estimates the decay in statistical power as the conditioning set size grows, and recommends a fallback (e.g., using nearest-neighbor tests for sparse dependencies, or ML-based tests for dense, non-linear dependencies).

  • Improved capability: AI can reliably test conditional independence in datasets with hundreds of covariates (e.g., EHR data) while controlling false positives, even when the true dependency is weak and high-order.

  1. Mixed-Type Data Handling Without Discretization
  • Build a unified CI test that uses a copula-based or rank-based transformation for continuous variables and a separate likelihood-ratio test for categorical variables, then combines them via a meta-test that preserves type I error.

  • Improved capability: AI can directly analyze mixed clinical data (e.g., continuous lab values + categorical diagnoses) without losing information from binning, leading to more accurate graph recovery in personalized medicine.

  1. Small-Sample Error Control via Bootstrap-Calibrated Tests
  • Add a calibration layer that uses residual bootstrap or permutation-based null distributions to adjust p-values for small n (e.g., n < 50), specifically for kernel and ML-based CI tests that otherwise over-reject.

  • Improved capability: AI can perform causal inference on rare-disease cohorts or small pilot studies, where traditional asymptotic tests are unreliable, and still maintain valid type I error rates.

  1. Scalability-Aware Test Selection for Large Graphs
  • Implement a cost-benefit scheduler that estimates computational cost (e.g., O(n 2) for kernel tests vs. O(n) for partial-correlation) and statistical accuracy trade-off, then chooses the fastest test that meets a user-defined error budget for each edge in a graph.

  • Improved capability: AI can run CI testing on entire patient networks (millions of edges) in reasonable time, while prioritizing high-risk edges for more robust (but slower) tests, enabling real-time clinical decision support.

  1. Asymmetric Error Consequence Mitigation
  • Add a decision-theoretic layer that assigns asymmetric costs to false positives (e.g., adding a wrong causal edge) vs. false negatives (missing a true edge) based on downstream task (e.g., drug target identification vs. risk factor screening), and adjusts the CI test threshold accordingly.

  • Improved capability: AI can tailor its causal discovery output to the specific application—e.g., being conservative for treatment effect estimation but more sensitive for hypothesis generation—reducing costly misdirections in biomedical research.

  1. Library-Agnostic CI Test Interface
  • Develop a standardized API that wraps R (e.g., pcalg, bnlearn) and Python (e.g., scipy, sklearn, causal-learn) implementations, with automatic checks for input type compatibility and assumption pre-screening.

  • Improved capability: AI systems can seamlessly switch between different CI test implementations without code rewrite, enabling rapid benchmarking and deployment across diverse biomedical pipelines.

What the improved AI system can do:

  • Perform reliable causal discovery on high-dimensional, mixed-type, small-sample biomedical data (e.g., genomics, EHR, imaging) with controlled error rates.

  • Automatically adapt its testing strategy to data complexity and computational constraints, avoiding both false causal claims and missed true dependencies.

  • Provide interpretable confidence in graph edges, with explicit warnings when assumptions are violated and when power is low.

  • Scale to large-scale patient networks while maintaining statistical rigor, and tailor its output to the clinical or research decision context.

Abstract

Conditional Independence (CI) tests are the statistical engine of constraint-based causal discovery: in algorithms such as PC (Peter-Clark) and FCI (Fast Causal Inference), skeleton pruning and key orientations follow directly from CI decisions. This survey reviews CI testing with emphasis on assumptions, robustness, and scalability in high-dimensional and mixed-type settings common in biomedical domains. The survey organizes widely used CI methods into six families: partial-correlation, contingency-table, regression, nearest-neighbor, kernel, and machine-learning-based. Special emphasis is provided on the robustness layers that address the limitations of these families. For each family, the survey examines when CI decisions reflect the data-generating distribution and when they fail. By this, we link test-level properties, including power decay with conditioning set size and asymmetric type I/II error consequences, to graph-level errors in skeleton recovery and v-structure orientation. The survey also compares adoption across major R and Python libraries and summarizes open challenges, including mixed-type CI testing without discretization, small-sample error control, and strategies for improving scalability of CI-testing.

Sources

Related papers