Interpretable Causal Discovery via Causal-Effect Constraints
Yale University · University of California, Los Angeles
cs.LG, cs.AI, stat.ML, stat.OT
Submitted: 2026-08-12
Updated: 2026-08-12
Comments: Accepted by UAI 2026
Code: https://github.com/ZCX031116/MLS-Framework
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 100/100
Terminology
Summary
Affiliation: Computer Science Dept., Yale University; Computer Science Dept., University of California, Los Angeles
The paper addresses the challenge of conditional causal discovery — the task of uncovering causal structures that are consistent with user-specified constraints, particularly extreme causal-effect events. The authors note that causal discovery is often motivated not only by prediction, but by the need to obtain interpretable and actionable explanations of how a system works.
Rather than simply asking which graph is most probable given data, practitioners may ask which causal mechanisms could explain a particular domain-relevant phenomenon,
such as an unusually large causal effect from one variable to another.
The authors formulate this as a Bayesian inference problem targeting the posterior over causal graphs and parameters conditional on an event such as a causal-effect constraint.
This is computationally challenging because the event of interest may be rare under the unconstrained posterior, especially when the causal effect threshold t is large, or when we impose a combination of constraints.
The paper's contributions are:
-
Formulation of conditional causal discovery as posterior inference under extreme-effect constraints, targeting both constrained posterior samples and posterior tail probabilities.
-
A practical rare-event inference procedure using adaptive multilevel splitting (AMS) that can estimate small posterior tail probabilities while producing representative graph–parameter samples from the corresponding constrained posterior.
-
Validation on linear-Gaussian benchmarks with dimensions d ∈ 4, 8, 16, 32 and a Sachs dataset case study.
The paper considers Bayesian networks (G, θ) over d variables, where G is a DAG and θ parameterizes conditional distributions. The joint distribution factorizes as:
p(X G, θ) = ∏i p(Xi pa G(Xi), θi)
For the linear-Gaussian structural equation model, variables satisfy X = XB + ε, where B is the weighted adjacency matrix. The signed total causal effect from Xi to Xⱼ has the closed form:
CEiⱼ(G, B) = ((I − B)−1)iⱼ
Since G is acyclic, (I − B)−1 = I + B + B2 +... + B(d−1), meaning the causal effect aggregates the products of edge weights along all directed paths from Xi to Xⱼ.
The authors define a graph–weight state as Z = (G, B). A conditional query is specified by a score function h: Z → R and a final level λ*. The score-level event, its posterior probability, and the constrained posterior are:
A λ* = Z: h(Z) ≥ λ*, p* = π(A λ*), π*(Z) = π(Z A λ*)
For a single ordered pair (i, j) and threshold t > 0, the right-tail event E+iⱼ(t) is represented by h(Z) = CEiⱼ(Z) and λ* = t, while the left-tail event E−iⱼ(t) uses h(Z) = −CEiⱼ(Z). This unified representation avoids requiring separate algorithms for positive and negative extreme effects.
The framework also supports conjunctions of multiple constraints. For a collection of C constraints, the normalized margin is:
m c(Z) = δ c(CEi cⱼ c(Z) − t c) / κ c
and the aggregate score is h C(Z) = min 1≤c≤C m c(Z).
The core algorithmic contribution is adapting AMS to the joint graph–parameter space. The rare-event probability decomposes as:
p* = π(A K) = ∏ k=0 K−1 π(A k+1 A k)
Each factor is much larger than p* itself, making it estimable with moderate particle counts. The algorithm:
-
Initializes particles from the structural prior p(G) and conjugate posterior p(B G, D).
-
Mutates the population using MCMC targeting the level-truncated posterior π λ(Z) ∝ π(Z)1 h(Z) ≥ λ.
-
Selects the next level as an empirical quantile: λ k+1 = min Q 1−ρ(scores), λ*.
-
Estimates the conditional factor β̂ k+1 = S/N where S is the set of surviving particles.
-
Resamples survivors with replacement.
The final probability estimate is p̂* = ∏ β̂ k, and after the final mutation step, particles approximate samples from the constrained posterior.
The mutation kernel alternates between:
-
Structure moves (with probability p struct): proposing new graphs via Structure-MCMC or PARNI-DAG, then refreshing only coefficient blocks whose parent sets changed.
-
Weight moves (with probability 1 − p struct): keeping the graph fixed and resampling one incoming coefficient block from its conditional posterior.
The acceptance probability for level λ k is:
a λ(Z, Z′) = 1 h(Z′) ≥ λ k × min(1, π(Z′)q(Z Z′) / (π(Z)q(Z′ Z)))
Proposals violating the current score-level constraint are rejected immediately.
The authors validated tail-probability estimation against exhaustive enumeration (all 543 DAGs for d = 4). Results showed that all methods except DiBS are broadly consistent with the enumeration baseline across the d = 4 cases.
DiBS overestimates tail probabilities by orders of magnitude on these tail events
and was excluded from subsequent experiments.
The authors constructed constraints from known ground-truth causal effects and tested whether conditioning concentrates posterior mass toward the true mechanism. Results showed:
-
For d = 4: Mean SHD decreased from 0.325 ± 0.164 (weak constraints) to 0.115 ± 0.053 (strong constraints), with Pr(SHD = 0) increasing from 0.790 to 0.885.
-
For d = 8: Mean SHD decreased from 5.342 ± 0.442 to 2.290 ± 0.199, with Pr(SHD = 0) increasing from 0.040 to 0.108.
The authors conclude: stronger, more informative multi-effect constraints consistently concentrate the conditional posterior toward the ground-truth mechanism.
Both MLS variants produce smooth, monotone-decaying tail curves as the target level λ* becomes more stringent,
while unconditional baselines quickly fail to generate samples satisfying the more extreme targets.
At d = 32, PARNI-DAG retains relatively stable run-to-run behavior, whereas Structure-MCMC exhibits noticeably larger variability in the extreme tail,
suggesting that reliable deep-tail estimation in larger graphs benefits from more informed structure proposals.
The authors applied the framework to the Sachs protein-signaling dataset (n = 7466 measurements, d = 11 proteins). They studied two target pairs:
-
PIP3 → PIP2 (indices 6 → 5), which has a clean decomposition into direct and Plcg-mediated routes.
-
Erk → Akt (indices 1 → 10), a cross-module influence.
Key findings:
-
Cond-PIP: "The conditional posterior yields a highly concentrated explanation for PIP3 → PIP2: the direct route and the Plcg-mediated route appear in essentially all conditional samples, and the conditional mean effect is approximately 0.776."
-
Cond-ERK:
Erk → Akt remains direct-dominated, while PIP3 → PIP2 becomes almost always reachable but is supported by many alternative, partly cancelling paths.
-
Non-symmetry: "Conditioning on PIP3 → PIP2 does not force Erk → Akt to become extreme; for example, Pr(CE Erk,Akt(Z) > 0 D, E PIP(t PIP)) ≈ 0.498."
-
Cond-Joint:
Joint conditioning increases the frequency of the mediated Erk → Plcg → Akt route relative to Cond-ERK, highlighting Plcg as a shared mediator under co-extreme coupling.
The estimated posterior masses of the tail events were approximately 2.2 × 10−7 (Cond-PIP), 1.3 × 10−7 (Cond-ERK), and 5.2 × 10−7 (Cond-Joint).
The overall cost of the algorithm is O(KNmC MH + KN log N), where K is the number of levels, N is particle count, m is MCMC steps per level, and C MH is the average proposal cost. The key advantage is that the leading cost grows approximately logarithmically in 1/p* rather than linearly in 1/p*,
making multilevel splitting useful in the deep-tail regimes where unconditional sampling degenerates.
The authors note: Our current implementation assumes linear-Gaussian SEMs under causal sufficiency. Extending the framework to nonlinear mechanisms, latent confounding, and interventional data remains an important direction.
They also suggest that incorporating richer constraint families, such as path-specific effects or qualitative monotonicity constraints, could enable deeper mechanistic analysis of complex scientific systems.
Improvements for AI systems
Based on this paper, I can improve AI systems in the following ways:
1. Rare-Event Bayesian Inference Engine
-
Implement adaptive multilevel splitting (AMS) as a general-purpose module for any probabilistic model where target events have posterior probability < 10−6
-
The system can estimate tail probabilities and generate representative samples from constrained posteriors without exhaustive enumeration or importance sampling degeneracy
-
This enables AI systems to answer
what causal structures could explain this extreme outcome?
rather than onlywhat is the most likely structure?
2. Constraint-Aware Causal Discovery
-
Add a conditioning layer on top of existing causal discovery algorithms (e.g., NOTEARS, DiBS, PC) that accepts user-specified causal-effect thresholds (e.g., "CE(X→Y) > 5") and returns only graphs consistent with those constraints
-
The system can handle conjunctions of multiple constraints (e.g., "CE(A→B) > 3 AND CE(C→D) < −2") by using the normalized margin aggregation
-
This makes causal discovery interpretable for domain experts who care about specific mechanisms, not just global graph fit
3. Mechanistic Explanation Generator
-
Given observational data and a target variable pair, the system outputs: (a) posterior probability of extreme effect, (b) a ranked list of causal paths (direct, indirect, mediated) that explain the extreme effect, (c) confidence intervals for each path's contribution
-
For example, on protein signaling data, it can identify that
PIP3→PIP2 extreme effect is explained by both direct and Plcg-mediated routes appearing in 100% of conditional samples
— providing actionable mechanistic insight
4. Scalable Deep-Tail Estimator for High-Dimensional Models
-
Replace naive Monte Carlo with AMS for any Bayesian model where the event probability is extremely small (e.g., < 10−7) and dimension is moderate (d ≤ 32)
-
The system automatically selects between Structure-MCMC (for d ≤ 16) and PARNI-DAG (for d > 16) based on empirical variability in tail estimates
-
This enables AI systems to reason about rare but catastrophic events (e.g., extreme drug side effects, financial crashes) in high-dimensional causal systems
5. Non-Symmetric Conditional Reasoning
-
The system can answer asymmetric queries like "If I condition on CE(X→Y) > t, what is the implied distribution of CE(Y→X)?" without assuming symmetry
-
This is useful for AI systems that need to reason about feedback loops or bidirectional influences where extreme effects in one direction don't imply the reverse
-
The system quantifies how conditioning on one causal effect changes the posterior over other, potentially unrelated effects
6. Automated Constraint Feasibility Checker
-
Before running expensive inference, the system estimates whether a user-specified constraint is even plausible given the data (e.g., returns "this constraint has posterior probability < 10−10, likely infeasible")
-
This prevents wasted computation and helps users refine their queries
-
The system can suggest alternative thresholds that are more data-supported while still being scientifically meaningful
7. Causal Effect Decomposition Visualizer
-
For any conditioned causal effect, the system decomposes it into contributions from all directed paths (direct, 2-step, 3-step, etc.) and shows which paths dominate under the constraint
-
This helps AI systems explain why an extreme effect occurs (e.g.,
the effect is large because of a strong 3-step path through variables X, Y, Z, not because of a direct edge
) -
The system can compare path decompositions between unconstrained and constrained posteriors to highlight mechanistic shifts
Abstract
Causal discovery aims to uncover the underlying causal relationships given data generated from a system. The goal, however, is not merely to predict causal edges given data, but also to be able to interpret and explain either observed or hypothesized phenomena, such as a particularly large causal effect. We consider this task of conditional causal discovery and cast it as a Bayesian inference problem, in which we target the posterior over causal graphs and parameters conditional on an event such as a causal-effect constraint. Unfortunately, this poses a computational challenge: existing approaches to Bayesian causal discovery struggle when the event has small posterior mass. To address this, we adapt rare-event estimation techniques to perform inference the joint graph-parameter space. Our method gradually drives a particle population toward the constrained region while maintaining samples that approximate the conditional posterior. Empirical evaluation on synthetic graphs validates the accuracy of our approach at small and large scales, and we show in a case study on the Sachs protein dataset how our method can be used to aid scientific exploration by providing pathway-level summaries.
Sources
- Variational Causal Networks: Approximate Bayesian Inference over Causal Structures
- Benchmarking Bayesian Causal Discovery Methods for Downstream Treatment Effect Estimation
- Effective Bayesian Causal Inference via Structural Marginalisation and Autoregressive Orders
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks