The dynamic generalized covariance measure for conditional independence testing with nonstationary time series
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "The dynamic generalized covariance measure for conditional independence testing with nonstationary time series".
Jane: The paper was written by Michael Wieck-Sosa, Michel F. C. Haddad and Aaditya Ramdas from Carnegie Mellon University and Queen Mary University of London.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title and Authors: Tom: Welcome back, everyone. We are digging into a brand new paper today, and the title is a mouthful: “The Dynamic Generalized Covariance Measure for Conditional Independence Testing with Nonstationary Time Series.” Jane, I’m going to need you to break that down for me before my brain melts.
Jane: Happy to, Tom. So, imagine you’re watching two stock markets move over time, and you want to know if one really causes the other, or if they just look connected because they both react to the same news. That’s what “conditional independence” means — are X and Y still related once you account for Z?
Tom: And the “nonstationary” part is the kicker, right? That means the rules of the game are changing over time. The market isn’t behaving the same way in January as it does in June.
Jane: Exactly. Most statistical tests assume the world is stable, but economies, weather systems, and brain signals are constantly shifting. This paper tackles that messy reality, and it’s from a team at Carnegie Mellon and Queen Mary University of London.
Tom: So they’re not just adding a tweak to an old test. They’re building something for a world where everything is in flux. That feels huge for anyone trying to make sense of real-world data.
Jane: It is. And the clever part is the “Dynamic” in the title. Instead of pretending the relationship is fixed, they let it evolve. They’re saying, “Let’s check if these two things are linked at each moment in time, given everything else we know at that moment.”
Tom: So, for a listener trying to forecast the economy, this could mean actually trusting the signals they’re using, because the test isn’t being fooled by a changing landscape. I’m already excited to see how they pull this off.
Jane: Me too. And the authors, Michael Wieck-Sosa, Michel Haddad, and Aaditya Ramdas, they’ve put together a framework that feels like it could be the new standard for time series analysis.
Tom: Alright, we’ve got the title and the big idea. Next up, we’re going to get into the actual meat of the paper — how this test works under the hood. Stay with us.
Summary: Tom: Welcome back. We’re still on “The Dynamic Generalized Covariance Measure for Conditional Independence Testing with Nonstationary Time Series.” Jane, we’ve established it’s a big deal, but how does it actually work?
Jane: Okay, so the core trick is to run two regressions. You try to predict X using Z, and you try to predict Y using Z. Then you look at the errors — the parts of X and Y that Z couldn’t explain.
Tom: So if those leftover errors are still correlated, that means X and Y have a connection that Z isn’t responsible for.
Jane: You’ve got it. That’s the “Generalized Covariance Measure” part. But here’s where it gets clever for time series. Because the world is changing, the regression itself has to change over time. You can’t just fit one line through the whole dataset.
Tom: Right, because the relationship between interest rates and housing prices isn’t the same in two thousand eight as it is in two thousand twenty-four. So they let the regression be time-varying.
Jane: Exactly. And then they take those time-varying errors and they look at the cumulative sum of their products over time. If that sum gets too big, it’s evidence that the errors are actually connected.
Tom: And this is where the “Dynamic” part comes in. They’re not just looking at the total sum; they’re watching it grow over time. That way, they can catch a relationship that only appears for a few months and then disappears.
Jane: Precisely. A standard test might average that blip away and miss it entirely. This test is designed to catch those fleeting, time-sensitive connections.
Tom: So it’s like having a motion detector instead of a single photograph. It sees the movement, not just the final position. That’s a powerful upgrade.
Jane: It is. And the theory they’ve built around it is really solid. They show that even when the data is messy and dependent, the test doesn’t cry wolf. It only raises the alarm when there’s a real connection.
Tom: So we’ve got the mechanics. But what does this mean for people actually trying to use this? Let’s bring in Meng to talk about the practical side of things.
Improvements: Tom: We’re back with “The Dynamic Generalized Covariance Measure for Conditional Independence Testing with Nonstationary Time Series.” We’ve covered the what and the how. Now, Meng, I want to know — what does this actually improve for someone like you, building real systems?
Meng: For me, the biggest win is that it works with a single realization of the process. In the real world, you rarely get to run the economy twice. You have one set of data, and you have to make decisions from it.
Jane: That’s such a good point. A lot of statistical methods assume you can repeat the experiment, but you can’t with financial markets or climate data.
Meng: Exactly. And this test is built for that constraint. It also handles the fact that the errors aren’t clean. In my world, the noise is never just random; it’s correlated with the past, and it changes over time. This paper explicitly allows for that.
Tom: So it’s not just a theoretical toy. It’s designed for the grimy, messy data that engineers actually deal with.
Meng: Right. And there’s another improvement I appreciate. The test is “doubly robust.” That means if your regression for X is a bit off, but your regression for Y is really good, you can still get a valid test. You don’t need both to be perfect.
Jane: That’s a huge practical advantage. It gives you a safety net when your models aren’t perfect, which is always.
Meng: And they’ve shown it works with a sieve estimator, which is a flexible way to approximate those time-varying functions. That’s not just theory; it’s a concrete recipe you can implement.
Tom: So instead of a black box, we get a tool with clear instructions. That’s the kind of improvement that moves a paper from “interesting” to “essential reading.”
Meng: Absolutely. It lowers the barrier to entry for using these powerful techniques in production systems.
Jane: And that’s what we want to see — research that can actually be deployed. Let’s hear what Lu thinks about the bigger picture before we wrap up.
Conclusion: Tom: We’ve reached the end of our time with “The Dynamic Generalized Covariance Measure for Conditional Independence Testing with Nonstationary Time Series.” Jane, can you give us the final summary?
Jane: Sure, Tom. This paper gives us a reliable way to ask “is X connected to Y, even after accounting for Z?” when the world is constantly changing. It does this by using time-varying regressions and watching how the errors move together over time.
Meng: And it does it without requiring clean, stationary data or multiple runs of the same experiment. That makes it a practical tool for real-world forecasting and causal discovery.
Tom: Lu, what’s the big takeaway for the field?
Lu: This is a foundation for a new generation of analysis. It opens the door to asking deeper questions about how systems evolve, not just whether they’re connected. It’s a step towards truly understanding dynamic systems.
Lalam: And from a cultural standpoint, this helps us build more responsive and resilient systems — from early warning systems for financial crises to better models of how information spreads. It helps us understand the world as it is, in motion.
Tom: Well said. We’ve covered the title, the mechanics, and the real-world impact. It’s a dense paper, but the ideas are powerful. Thanks for joining us, and we’ll see you next time with another paper.
Michael Wieck-Sosa, Michel F. C. Haddad, Aaditya Ramdas
Carnegie Mellon University · Queen Mary University of London
stat.ME, math.ST, stat.ML, stat.TH
Submitted: 2026-08-10
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 70/100
The gist: The paper addresses the problem of conditional independence testing for nonstationary nonlinear time series.
Terminology
Summary
The paper addresses the problem of conditional independence testing for nonstationary nonlinear time series. The authors state: "Identifying relationships among stochastic processes is a core objective in many fields, such as economics. While the standard toolkit for multivariate time series analysis has many advantages, it can be difficult to capture nonlinear dynamics using linear vector autoregressive models. This difficulty has motivated the development of methods for causal discovery and variable selection for nonlinear time series, which routinely employ tests for conditional independence."
The authors note that "most conditional independence tests lack Type I error control guarantees outside the iid setting. At most, some methods provide guarantees for stationary mixing processes. To the best of our knowledge, only one conditional independence test has been proposed for the setting with one realization of a nonstationary process. They reference Malinsky and Spirtes [MS19], who
introduce a conditional independence test for nonstationary linear vector autoregressions with iid Gaussian errors, focusing on processes with stochastic trends. In contrast, the authors state:
we allow linear and nonlinear processes with general nonstationarity, time-varying regression functions, and non-iid non-Gaussian errors."
The paper's contributions are summarized as follows:
-
First conditional independence test for a single realization of a nonstationary nonlinear process: "We propose the first conditional independence test that can be used with a single realization of a nonstationary nonlinear process. In Theorem 3.1, we show that our test has asymptotic Type I error control, uniformly over a large family of null distributions."
-
Two types of test statistics:
We provide l∞-type and l2-type test statistics to achieve high power against sparse and dense alternatives, respectively.
-
Double robustness property: "Our test statistics are based on the products of the residuals from nonlinear regressions of future or past values of the processes X and Y on the covariates Z at time t. Hence, our test is doubly robust, in the sense that slower convergence rates of one regression estimator can be compensated for by faster convergence rates of the other estimator."
-
Flexible error assumptions: "We allow the errors to be nonstationary and temporally dependent, in contrast with other residual-based conditional independence tests that require iid errors. Moreover, we allow the errors to depend on the covariates, whereas many other residual-based tests require covariate-independent errors, thereby ruling out heteroskedasticity."
-
Compatibility with any regression estimator: "Our test can be used with any regression estimator for nonstationary time series. However, we provide a guarantee for an instantiation of our test based on the sieve regression estimator from Ding and Zhou [DZ21] in Section A of the supplement."
-
Distribution-uniform strong Gaussian approximation:
We introduce a distribution-uniform strong Gaussian approximation in Section C of the supplement, which we use to prove uniformly asymptotic Type I error control.
-
Unconditional independence test: "We introduce an unconditional independence test in Section D of the supplement. Thus, our work is the first to supply the two components used in constraint-based causal discovery algorithms for nonstationary nonlinear time series."
The paper works in a triangular array framework for high-dimensional nonstationary nonlinear time series.
Let (Xt,n, Yt,n, Zt,n)t∈[n] be the observed sequence of length n ∈ N, where [n] = 1,..., n. The dimensions dX = dX,n, dY = dY,n, dZ = dZ,n can grow with n.
Time-offsets are introduced: Negative time-offsets are called lags, and positive time-offsets are called leads.
Sets Ai, Bj, Ck contain the time-offsets of Xt,n,i, Yt,n,j, Zt,n,k under consideration, with the time-offsets Ck to be non-positive so that the covariates are known at time t.
The null hypothesis tested is: Xt,n ⊥⊥ Yt,n Zt,n for all times t ∈ Tn,
which includes all leads and lags of interest.
For a fixed sample size n, distribution P, time t, and dimension/time-offset tuple m = (i, j, a, b) ∈ Dn, the paper decomposes:
Xt,n,i,a = fP,t,n,i,a(Zt,n) + εP,t,n,i,a, Yt,n,j,b = gP,t,n,j,b(Zt,n) + ξP,t,n,j,b,
where fP,t,n,i,a(z) = EP(Xt,n,i,aZt,n = z) and gP,t,n,j,b(z) = EP(Yt,n,j,bZt,n = z) are the time-varying regression functions, and RP,t,n,m = εP,t,n,i,aξP,t,n,j,b denotes the product of these errors.
The test statistic is defined as:
Sn,p(R̂n) = max s∈Tn,L (1/Tn,L) Σ t≤s R̂t,n p,
which is the max lp-norm, p ∈ [2, ∞], achieved by the scaled partial sum process of residual products.
The main idea: "under the null hypothesis (3), the products of the residuals will be small, as long as the regression functions are estimated well. Our test statistic (4) is designed to have power against alternatives in which the covariances of the errors from (1) are non-zero for some times."
The paper emphasizes: "As with the GCM test from Shah and Peters [SP20], our test only has power against alternatives with non-zero expected conditional covariance; see Section D of the supplement for the details. Hence, our test is not, in general, an omnibus test of conditional independence."
Algorithm 1 (the dGCM test) proceeds as follows:
-
Select hyperparameters for regression method via cross-validation
-
For each time t and dimension/time-offset tuple m, obtain estimates f̂t,n,i,a and ĝt,n,j,b, calculate residuals, and calculate the product of residuals R̂t,n,m
-
Select the lag-window size Ln using the minimum volatility method
-
Calculate covariances Σ̂R t,n using a local window
-
Simulate independent Gaussian random vectors R̆(r) t,n N(0, Σ̂R t,n)
-
Calculate test statistics for simulated data
-
Calculate the 1−α empirical quantile of simulated test statistics
-
Reject if the test statistic exceeds the quantile
The paper states: Suppose Assumptions 3.1, 3.2, 3.3, 3.4, 3.5, 3.6 related to the temporal dependence and nonstationarity hold for the sequence of collections of distributions (P∗0,n)n∈N, where P∗0,n ⊂ PCI0,n for each n ∈ N.
The regression estimation errors must satisfy specific convergence rates, including:
-
sup max EP(ŵfP,t,n,i,a2)(1/2) · max EP(ŵgP,t,n,j,b2)(1/2) = o(Tn(−1/2)τn(7)Dn(−3))
-
sup max EP(ŵfP,t,n,i,a2)(1/2) = o(τn 7Dn(−5))
-
sup max EP(ŵgP,t,n,j,b2)(1/2) = o(τn 7Dn(−5))
Then: "lim sup n→∞ sup P∈P∗0,n PP(Sn,p(R̂n) > q̂1−α+νn + τn) ≤ α."
The paper notes: "The dGCM test possesses a property known as rate double robustness, which means that we place stronger convergence rate requirements on the products of the L2(P) norms of the estimation errors than on each one individually."
Assumption 3.1 (Causal representations of observed processes): Each dimension of the observed sequence can be represented as a measurable function of iid inputs: Xt,n,i = GXt,n,i(HtX), Yt,n,j = GYt,n,j(HtY), Zt,n,k = GZt,n,k(HtZ).
Assumption 3.4 (Causal representations of error processes): εP,t,n,i,a = GεP,t,n,i,a(Hεt,a), ξP,t,n,j,b = GξP,t,n,j,b(Hξt,b),
with EP(εP,t,n,i,aHĝt) = 0 and EP(ξP,t,n,j,bHf̂t) = 0.
Assumption 3.5 (Distribution-uniform decay of temporal dependence): Uses the functional dependence measure of Wu [Wu05], requiring polynomial decay: θe,∞P,t,n,l,d(h) ≤ Θ̄∞ · (h ∨ 1)(−β̄∞)
with β̄∞ > 1, and for error products θRP,t,n,m(h, q̄R) ≤ Θ̄R · (h ∨ 1)(−β̄R)
with β̄R > 3, q̄R > 4.
Assumption 3.6 (Distribution-uniform total variation condition for nonstationarity): Requires a bound on the total variation of the causal mechanism of error products over time.
The paper provides a guarantee for an instantiation of our test based on the sieve regression estimator from Ding and Zhou [DZ21] in the setting of locally stationary time series.
The sieve estimator uses basis functions (Legendre polynomials in the simulations) for both time and covariate values, with the numbers of basis functions for time and for the covariate values, denoted by c̃n and d̃n, respectively, are chosen to increase with the sample size n at some rate.
Theorem A.1 states: "Suppose that Assumptions A.1, A.2, A.3, A.4, A.5, A.6, A.7 all hold... Further, suppose that we use the sieve time-varying regression estimator from Section A.3 with the basis functions ϕl1(u), φl2(z) chosen to be mapped Legendre polynomials, where the numbers of basis functions are chosen to satisfy c̃n = O(log(Tn)), d̃n = O(log(Tn)). Then Assumptions 3.1, 3.2, 3.3, 3.4, 3.5, 3.6 hold for (P∗0,n)n∈N, and the sieve estimators will achieve the convergence rates required by Theorem 3.1."
The paper studies the setting with dX = 1, dY = 1, dZ = 1 and no time-offsets, so A = 0, B = 0, C = 0, and Tn = [n]. We test the null hypothesis Xt,n ⊥⊥ Yt,n Zt,n for all times t ∈ Tn.
The data generating process uses:
-
A tvAR(1) covariate process:
Zt,n = θZ(t/n)Zt−1,n + ηtZ
withθZ(u) = 0.35 + 0.2 cos(2πu)
-
Regression functions:
fK(z, u) = (0.5 + 0.25 cos(2πu)) exp(−z2) sin(Kz)
andgK(z, u) = (0.3 + 0.15 sin(πu)) exp(−z2) cos(Kz)
with K ∈ 1, 2, 3, 4 -
Correlated shocks with correlation ρ ∈ 0, 0.3, 0.6, 0.9
The paper compares Sieve-dGCM with the GCM test [SP20] using a generalized additive model, and the residual prediction test (RPT) [SB18; HPM18]. Results show: Our test holds the level even with fairly small sample sizes, and gains power as we increase the correlation and sample size. The other tests fail to hold the level.
The paper investigates how stock markets in the United States, United Kingdom, Hong Kong, and Japan are linked
using "daily log returns based on the adjusted closing prices of the S&P 500, FTSE 100, Hang Seng, and Nikkei 225 from January 2022 to March 2025," yielding n = 845 observations.
Key findings: "We retain the null hypotheses that S&P(t) is independent of Nikkei(t) (BH-adjusted p-value 0.074) and that FTSE(t) is independent of Nikkei(t) (BH-adjusted p-value 0.055). All other null hypotheses of independence are rejected at the significance level α = 0.05."
For conditional independence: "We retain the null hypothesis that S&P(t) is independent of HangSeng(t) given FTSE(t) (BH-adjusted p-value 0.797). However, we reject the null that S&P(t) is independent of FTSE(t) given HangSeng(t) (BH-adjusted p-value 0.005)."
The paper concludes: "The dGCM test shows promise for detecting conditional dependencies among nonstationary nonlinear time series while controlling the Type I error in finite samples. Specifically, we find that the Sieve-dGCM test can hold the level if the sample size is large enough to reliably estimate the time-varying regression functions."
Improvements for AI systems
Based on the paper, I can improve AI systems in the following specific ways:
Improvement: Implement the Dynamic Generalized Covariance Measure (dGCM) test as a new statistical tool for AI systems that analyze time series data.
What the improved AI system can do:
-
Test whether two time series variables are conditionally independent given a third variable, even when the underlying processes are nonstationary (e.g., financial markets, climate data, sensor readings)
-
Detect nonlinear relationships that traditional linear VAR models miss
-
Handle a single realization of a nonstationary process (no need for multiple trials)
-
Control Type I error rates uniformly across a large family of null distributions
-
Work with temporally dependent and heteroskedastic errors, unlike many existing tests
These improvements enable AI systems to handle the complex, nonstationary, nonlinear time series data that are common in economics, finance, climate science, and engineering, while providing statistically rigorous guarantees that are essential for high-stakes applications.
Sources
- Simultaneous Sieve Inference for Time-Inhomogeneous Nonlinear Time Series Regression
- A Note on Physical Dependence and Mixing Conditions for Triangular Arrays
- A Scalable Conditional Independence Test for Nonlinear, Non-Gaussian Data
- The robusTest package: two-sample tests revisited
- Time-varying correlation network analysis of non-stationary multivariate time series with complex trends
- Granger Causality in Extremes
- Doubly robust and computationally efficient high-dimensional variable selection
- Cross-Fitting and Fast Remainder Rates for Semiparametric Estimation
- Causal inference for temporal patterns
- Nonparametric Tests of Conditional Independence for Time Series
- Distribution-uniform anytime-valid sequential inference and the Robbins-Siegmund distributions
- Bootstrapping High Dimensional Time Series
Related papers
- Doubly robust inference via calibration
- Bayesian Empirical Bayes: Simultaneous Inference from Probabilistic Symmetries
- Flexible Nonparametric Inference for Causal Effects under the Front-Door Model
- Deployment of AI-Assisted Interventions: Capacity Constraints and Noisy Compliance
- A Survey on Archetypal Analysis
- Dynamic Spatial Bayesian Machine Learning Model: Applications to Intergenerational Economic Mobility and Geographic Income Inequality in the United States