Operationalising Relative Causal Knowledge: Backbone Identifiability from Private Reports on a Shared Outcome

arXiv:2608.10664 · cs.AI · Submitted 2026-08-11 · Read on arXiv

Fabrizio Russo, Mark Somers

Imperial College London · Fifty One Degrees Ltd

cs.AI

Submitted: 2026-08-11

Updated: 2026-08-12

Comments: Accepted at Causal Decision Making Workshop, UAI2026. 4 pages + Appendix (13 total)

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 95/100

Terminology

Summary

arXiv: 2608.10664v1 [cs.AI], 11 Aug 2026


The paper addresses a foundational gap in the Relativity of Causal Knowledge (RCK) framework. RCK, introduced by D'Acunto and Battiloro (2025), explains how a network of agents with different structural causal models can exchange causal knowledge through a shared interventionally consistent abstraction called a backbone. The authors state: RCK is mathematically defined given a backbone, but the choice of backbone becomes an identification problem.

The central question is: when is the backbone determined by the agents' private causal knowledge? The paper studies the prior identification question that this transport mechanism presupposes. In the basic two-agent common-effect case, two private causes influence one shared outcome, and each agent identifies only the single-cause causal marginal relevant to its own perspective.

The authors frame this through a concrete education example: Agent ρ studies teacher quality (Zρ), Agent σ studies neighbourhood environment (Zσ), and both affect adult earnings (X). The teacher value-added literature estimates roughly 1.3% higher adult earnings per standard-deviation improvement in teacher quality, while the neighbourhood-mobility literature estimates roughly 0.5% higher earnings per standard-deviation better county. Both local findings can be correct, but they do not determine the edge value in the backbone needed for RCK communication.

The backbone in this two-agent setting is represented concretely by the interventionally consistent kernel:

Q(H zρ, zσ) = P(X ∈ H do(Zρ = zρ, Zσ = zσ))

This joint-intervention kernel represents the edge value χτ. With χτ fixed, restriction and extension maps have a common object through which to transport abstracted causal knowledge.

The private reports are obtained by marginalising over the other agent's private cause:

Pρ(· zρ):= ∫ Q(· zρ, zσ) PZσ(dzσ)

Pσ(· zσ):= ∫ Q(· zρ, zσ) PZρ(dzρ)

The identification problem fixes the reported kernels (Pρ, Pσ) and asks whether they determine a unique, interventionally consistent Q. A candidate backbone Q̄ is compatible if it satisfies both projection constraints.

The paper proves a kernel-level non-identifiability theorem under three assumptions:

  1. Existence: There exists at least one compatible backbone kernel Q0.

  2. Non-degeneracy: There are measurable sets Aρ and Aσ with positive private-cause probability, and bounded nonzero mean-zero functions u(Zρ) and v(Zσ) supported on Aρ and Aσ respectively.

  3. Local overlap (minorisation): For some probability measure ν with 0 0, Q0(H zρ, zσ) ≥ ε ν(H) for every measurable event H and all (zρ, zσ) ∈ Aρ × Aσ.

Conclusion: There exists a radius δ > 0 and a one-parameter family of distinct kernels Qt: t ∈ (−δ, δ) such that every Qt induces the same private reports Pρ and Pσ, but for t ≠ s the kernels Qt and Qs disagree on the joint interventional distribution for some intervention pair (zρ, zσ).

Construction: The perturbation family is defined as:

Qt(· zρ, zσ) = Q0(· zρ, zσ) + t·u(zρ)·v(zσ)·S(·)

where S is a signed measure on the outcome space with total mass zero. The product u(zρ)v(zσ) is the hidden interaction direction: it is invisible to both one-cause reports because each report integrates out one centred factor, but it is visible before marginalisation.

The proof proceeds in three steps:

  • Proposition 4: Constructs a nontrivial signed perturbation S on the outcome space with total mass zero (using a set B with 0 < ν(B) < 1).

  • Proposition 5: Shows the perturbed family Qt consists of valid probability kernels that reproduce the same private reports, provided t < ε/(∥u∥∞∥v∥∞).

  • Proposition 6: Shows distinct members of the family disagree on some joint intervention, making the ambiguity causally meaningful.

If the shared outcome satisfies X = fρ(Zρ) + fσ(Zσ) + ε with E[ε] < ∞, then the joint interventional mean response is:

m(zρ, zσ) = fρ(zρ) + fσ(zσ) + E[ε]

The response contains no mixed term in (zρ, zσ). The authors emphasise this should not be read as a free modelling convenience but is motivated by the institutional separation between classroom instruction and neighbourhood/family environment, and is empirically revisable.

Even under additive separability, arbitrary observational summaries do not identify the additive backbone. The paper shows this via a counterexample: two different additive models with the same residual variances but different response surfaces. Specifically, with Zρ, Zσ N(0,1), ε N(0,η2), and fσ(z) = z, comparing:

  • fρ(1)(z) = z

  • fρ(2)(z) = (z2 − 1)/√2

Both have Var(fρ(Zρ)) = 1, and both yield identical residual variances Var(Rρ) = Var(Rσ) = 1 + η2, but the additive backbone response surfaces differ: m1(zρ, zσ) = zρ + zσ versus m2(zρ, zσ) = (zρ2 − 1)/√2 + zσ.

"What resolves the ambiguity is causal communication. Once Agent ρ communicates fρ and Agent σ communicates fσ, the backbone mean response m(zρ, zσ) = fρ(zρ) + fσ(zσ) is fixed for every joint intervention pair, up to the shared constant E[ε] when absolute outcome levels rather than contrasts are required."

The paper provides three concrete witnesses of the non-identifiability:

E.1 Binary 2×2 Witness: With pij:= P(X = 1 Zρ = i, Zσ = j), the reported one-cause kernels impose linear constraints with rank three (not four). The null space is spanned by (1, −1, −1, 1), leaving the interaction term unresolved.

E.2 Gaussian Covariance Witness: With vector-valued outcome X = (X1, X2) and binary causes, cell-wise cross-covariances σ12(t)(zρ, zσ) = t(−1)(zρ+zσ) alternate in sign. When averaged over the other cause, these cancel to zero for every t, so all values of t induce the same communicated cross-covariance summaries even though t = 0 implies X1 ⊥ X2 while t ≠ 0 implies dependence.

E.3 Continuous Non-Gaussian Witness: With Laplace base kernel and perturbation seed s(x) = 10,1 − 1−1,0, the construction works on continuous outcomes with non-Gaussian base kernels, showing the obstruction is not a finite-dimensional or Gaussian artefact.

The paper works through the policy implications. With fρ(zρ) = 1.3zρ and fσ(zσ) = 0.5zσ (in percentage points per standard-deviation-year):

  • A full-childhood neighbourhood disadvantage of roughly one standard deviation corresponds to about Δσ:= 18 × 0.5% ≈ 9% lower adult earnings.

  • The compensating teacher-quality exposure solves 1.3(zρ* − z̄ρ) = Δσ, requiring approximately 9/1.3 ≈ 7 teacher-quality standard-deviation years.

  • Spread over a 13-year school career, this is about 7/13 ≈ 0.54 standard deviations above baseline per year.

The authors note: This calculation is not licensed by the local marginals alone. It becomes meaningful only after separability rules out the hidden interaction regime and the agents communicate causally identified response functions.

The paper distinguishes its contribution from several related areas:

  • RCK and abstraction: Unlike D'Acunto and Battiloro (2025), Rubenstein et al. (2017), Beckers and Halpern (2019), and others, the paper does not ask whether transport is well-behaved once an abstraction is available, but whether the target of causal abstractions is identifiable.

  • Expert judgement aggregation: Unlike Bradley et al. (2014), Alrajeh et al. (2018), and Friedenberg and Halpern (2018), the paper does not assume access to full expert models or try to select a collective graph or merged SCM.

  • Transportability and data fusion: Unlike Pearl and Bareinboim (2014) and related work, the unidentified object is not a pooled treatment effect, a merged graph, or a full joint SCM. It is the edge in the network sheaf and cosheaf through which RCK would allow sharing abstracted causal knowledge.

  • ICA and non-Gaussian identification: The positive result does not use non-Gaussianity as the identifying lever; identification comes from additive separability at the shared outcome together with communicated causally identified responses.

The paper's main message is negative generically and positive conditionally:

  1. Negative: Private local causal marginals need not determine a unique edge-level backbone, and the ambiguity is causally meaningful because compatible backbones can disagree on joint interventions.

  2. Positive: Additive separability removes the interaction ambiguity, but agents must still communicate causal response objects rather than observational residual summaries.

The authors conclude: "This gives a narrower but actionable reading of RCK. The framework's transport maps are useful once the shared edge value is available. Our results clarify when private causal knowledge can supply that edge value, and when additional structure or communication is required."

Natural next steps identified: testing separability in pooled studies, extending results to larger agent networks, and identifying weaker equivalence classes of backbones that preserve only the intervention queries needed for a given decision.

Improvements for AI systems

Based on this paper, here are specific improvements for AI systems:

Improvement: Implement a backbone-identification layer that explicitly tracks non-identifiability when aggregating causal models from multiple agents or data sources. Instead of assuming a unique joint causal model, the system maintains a family of compatible backbones (as in Theorem 1) and reports the range of possible intervention effects.

What the improved system can do: When asked What is the effect of teacher quality on earnings given neighbourhood data?, it returns not a point estimate but an interval spanning all compatible backbones, flagging when the answer is underdetermined. It can also identify which additional experiments or communications would resolve the ambiguity.

Improvement: Build a pre-processing module that tests whether a shared outcome is additively separable in the private causes (Proposition 2). The system checks for interaction terms in the joint response surface and quantifies the degree of separability using the perturbation family construction.

Improvement: Design an agent communication protocol where agents exchange full causal response functions (fρ, fσ) rather than summary statistics. The protocol includes a compatibility check: after exchange, agents verify that the communicated functions, when combined, reproduce their observed local marginals.

Improvement: Implement a diagnostic that detects when observational residual summaries (variances, correlations) are being used to infer causal backbone structure. The system flags cases analogous to the counterexample where different additive models yield identical residual summaries but different intervention responses.

Improvement: Add an identifiability-aware confidence layer to policy recommendation systems. For decisions requiring joint interventions (e.g., compensating teacher quality for neighbourhood disadvantage), the system computes the policy recommendation under all compatible backbones and reports the worst-case and best-case outcomes.

Improvement: Implement a clustering algorithm that groups compatible backbones into equivalence classes based on the intervention queries they support (as suggested in the paper's future work). The system identifies which backbones give identical answers for the specific decisions at hand.

Improvement: Build an audit tool that checks whether a multi-agent causal reasoning system is correctly handling the backbone identification problem. It tests whether the system's outputs are invariant under the perturbation family Qt from Theorem 1.

Abstract

The Relativity of Causal Knowledge (RCK) explains how a network of agents with different structural causal models can exchange causal knowledge through a shared interventionally consistent abstraction, or backbone. We ask the prior identification question that this transport mechanism presupposes: when is that backbone determined by the agents' private causal knowledge? In the basic two-agent common-effect case, two private causes influence one shared outcome and each agent identifies only the single-cause causal marginal relevant to its own perspective. We show that, under standard compatibility, non-degeneracy, and local overlap assumptions, those local causal marginals do not identify a unique backbone. Infinitely many joint intervention kernels can induce exactly the same private reports while disagreeing on joint interventions. We then give a conditional recovery result. Additive separability removes the hidden interaction degree of freedom, but observational residual summaries remain insufficient. Identification becomes possible when agents communicate causally identified response functions. An education value-added example illustrates why this is first a communication problem, and only then a policy-composition problem.

Related papers