Concomitant DAG Learning: On the Roles of Noise Adaptivity, Sparsity, and Non-negativity

arXiv:2605.23537 · stat.ML, eess.SP · Submitted 2026-05-22 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "Concomitant DAG Learning".

Tom: The gist The paper introduces CoLiDE, a novel convex score function for sparsity-aware DAG inference that jointly estimates the DAG adjacency matrix and exogenous noise levels,

Jane: First, who's behind it and why it matters.

Paper summary: Tom: To recap, we just touched on how this paper introduces CoLiDE, focusing on joint estimation of DAG adjacency matrices and exogenous noise levels for linear Structural Equation Models, or SEMs. The authors are looking at the challenge where you need a sparse graph but the noise level is unknown and potentially variable.

Jane: Exactly. They propose this score function as an alternative to previous methods where you had to guess the regularization weight based on some prior knowledge about how much noise there was in the system, which we know isn't always true. This new approach makes it adaptive to those noise levels.

Lu: The core claim is that this adaptivity removes the coupling between your sparsity regularization parameter and those unknown exogenous noise levels, which supposedly leads to minimum recalibration effort across different problem instances or distribution shifts. That’s a pretty strong statement about generalization, if it holds up.

Meng: From an engineering standpoint, that sounds promising because it simplifies the pipeline. Instead of needing a separate tuning step for sparsity and noise estimation, you just run this one unified process to get the DAG and the noise parameters simultaneously. That cuts down on complexity in deployment.

Lalam: It’s about making the learning process more resilient to real-world data imperfections, which is where most AI models struggle—when the assumptions we build them on aren't perfectly met by messy input data. This paper shows how to bake that resilience directly into the estimation objective.

Conclusion: Tom: So, wrapping up the discussion on "Concomitant DAG Learning: On the Roles of Noise Adaptivity, Sparsity, and Non-negativity," we’ve seen how CoLiDE attempts to solve that noise adaptation problem by jointly estimating structure and variance. The paper highlights how this unified approach offers robustness against heteroscedastic noise profiles in linear SEMs.

Jane: And the authors also explore the impact of adding non-negativity constraints on edge weights, leading to an algorithm called NOMAD when applied under those conditions. It shows that even adding structure like non-negativity can lead to better characterization of acyclicity.

Lu: The implication for the broader field is that we are seeing research exploring how sparsity regularization, noise modeling, and structural constraints interact in a unified optimization framework for causal discovery. It opens up new avenues for multi-objective optimization techniques in this area.

Meng: For me, it means we need to look at practical scalability. The paper mentions continuous relaxation approaches inspired by things like the Cayley-Hamilton theorem and DAGMA’s performance metrics, which suggests there's still work needed on making these solvers efficient enough for massive datasets.

Lalam: Looking ahead, this pushes us toward developing online adaptive algorithms that can track time-varying connectivity in the network while processing signals on the fly, which is a huge goal for real-time causal inference systems. That’s where the next big leap might be.

Tom: We covered how CoLiDE addresses noise and sparsity, and how non-negativity plays a role in finding acyclicity. It really shows that combining these elements in one score function gives you better performance across different metrics than using them separately.

University of Rochester · Rey Juan Carlos University · Elastic

stat.ML, eess.SP

Submitted: 2026-05-22

Updated: 2026-10-08

Comments: Submitted to the IEEE Signal Processing Magazine Special Issue: From Signals to Causes: Methodological Advances in Causal Inference

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 88/100

The gist: The gist The paper introduces CoLiDE, a novel convex score function for sparsity-aware DAG inference that jointly estimates the DAG adjacency matrix and exogenous noise levels, offering robustness

Key concepts

Convex Score Function
A mathematical function used in optimization to guide the learning process of a DAG. CoLiDE proposes a specific convex score function that helps simultaneously estimate both the graph's connections (adjacency matrix) and the unknown noise levels in the data, ensuring a stable and well-behaved optimization.
Noise Adaptivity
This is a key feature where the score function adjusts its estimation based on whether the noise variance is constant (homoscedastic) or varies across different data points (heteroscedastic). By making it noise-adaptive, the method avoids needing to recalibrate parameters when the noise profile changes.
Acyclic Constraint H(W) = 0
This constraint mathematically enforces that the learned structure must be a Directed Acyclic Graph (DAG), meaning there are no cycles in the causal relationships. The paper uses continuous relaxations of this constraint to solve the complex problem, aiming for a smooth optimization path.
Non-negative Edge Weights
This assumption restricts the DAG structure such that all edge weights must be greater than or equal to zero. When applicable, this constraint simplifies the optimization landscape and allows for an exact characterization of acyclicity using techniques like NOMAD.

Terminology

Summary

The gist The paper introduces CoLiDE, a novel convex score function for sparsity-aware DAG inference that jointly estimates the DAG adjacency matrix and exogenous noise levels, offering robustness against heteroscedastic noise profiles.<ref:2605.23537#pg11>

Problem Formulation and Context

Directed acyclic graphs (DAGs) are used to encode causal relationships in signal processing and machine learning applications, but inferring DAGs from observational data is an NP-hard problem due to the combinatorial acyclicity constraint<ref:2605.23537#pg5> The goal is to learn the latent DAG G ∈ D as the solution to a constrained optimization problem min G(W) S(G(W); X) subject to G(W) ∈ D<ref:2605.23537#pg5> For a linear SEM, the data matrix X can be written in compact matrix form as X = W⊤X + Z, where Z is a vector of mutually independent, exogenous noises<ref:2605.23537#pg5> A proper score function typically encompasses a loss or data-fidelity term ensuring alignment with the SEM as well as regularizers to promote desired properties of G<ref:2605.23537#pg5>

Score Function and Noise Adaptivity

The paper introduces a novel convex score function S(W, σ; X) to jointly estimate the DAG adjacency matrix and exogenous noise standard deviation(s)<ref:2605.23537#pg11> This score function is inspired by concomitant sparse linear regression [21], [23] and aims to address the parameter finetuning predicament caused by unknown exogenous noise levels<ref:2605.23537#pg5> A key message is that by making the score function noise adaptive, one effectively removes the coupling between the sparsity regularization parameter and the unknown exogenous noise levels, which leads to minimum (or no) recalibration effort across diverse problem instances or distribution shifts<ref:2605.23537#pg5> In the homoscedastic setting where all exogenous noises z1,..., zd in the linear SEM (1) have identical variance σ2, CoLiDE-EV is formulated as min W,σ≥σ0 1/2nσ∥X − W⊤X∥2F + dσ2 + λ∥W∥1 z⟩:=S(W,σ;X) subject to H(W) = 0<ref:2605.23537#pg5>

Optimization and Continuous Relaxation

The continuous optimization approach tackles the DAG learning problem (2) by enforcing a smooth acyclicity constraint H(W) = 0<ref:2605.23537#pg5> The methodology is inspired by the literature of concomitant sparse linear regression [21], [23] dating back to seminal work by Huber in the context of robust location and scale estimation<ref:2605.23537#pg5> The objective function in (8) is still a nonconvex optimization problem due to the acyclicity constraint H(W) = 0<ref:2605.23537#pg5> Inspired by the Cayley-Hamilton theorem, the general family H(W) = Pd k=1 ck tr((W◦W)k), ck > 0, was studied in [39] and DAGMA outperformed prior continuous relaxation methods in terms of (nonlinear) DAG recovery and computational efficiency due to several factors<ref:2605.23537#pg5>

Non-Negative Edge Weights

The paper discusses the impact of additional DAG structure on top of sparsity, namely, non-negativity of edge weights<ref:2605.23537#pg5> In applications where this assumption is tenable, non-negativity constraints have favorable optimization landscape and statistical implications<ref:2605.23537#pg5> Inspired by DAGMA’s log-determinant penalty in (7), non-negativity enables an exact characterization of acyclicity over the domain Ws+ = 1/2n∥X − W⊤X∥2F + λX i,j Wij z⟩:=S+(W;X) subject to W ≥ 0, ρ(W) The resulting iterative algorithm is termed Non-negative Optimization via Multipliers for Acyclic Digraphs (NOMAD), which we encountered with the Sachs dataset<ref:2605.23537#pg5>

Performance and Conclusions

CoLiDE-NV, although overparametrized for homoscedastic problems, performs remarkably well, either being on par with CoLiDE-EV or the secondbest alternative in Table I<ref:2605.23537#pg5> The heteroscedastic scenario presents further challenges where the Gaussian case is known to be non-identifiable<ref:2605.23537#pg5> CoLiDE-NV is the clear winner, outperforming the alternatives in virtually all variations<ref:2605.23537#pg5> The results illustrate that CoLiDE’s advantage over its competitors is not restricted to SHD alone, but equally extends to all other relevant metrics introduced in “Performance evaluation metrics”<ref:2605.23537#pg5> NOMAD achieves a SHD of 12 for the Sachs dataset, which is the lowest achieved SHD among continuous optimization-based techniques applied to the Sachs problem<ref:2605.23537#pg5> The work concludes with an outline of CoLiDE extensions and other frontier topics that offer exciting research opportunities at the confluence of (graph) SP, ML, optimization, and causal inference<ref:2605.23537#pg5>

Research Outlook

A wide variety of potential research avenues naturally follows from the developments presented here<ref:2605.23537#pg5> In terms of computational complexity, there is room for improving the scalability of some of the solvers described via parallelization and decentralized implementations<ref:2605.23537#pg5> Online adaptive algorithms that can track the (possibly) time-varying connectivity structure of the acyclic network and achieve both memory and computational savings by processing the signals on-the-fly are naturally desirable, but so far largely unexplored<ref:2605.23537#pg5> There is hope as progress on benign optimization landscape analyses is being made<ref:2605.23537#pg5> The envisioned problem is naturally a multi-objective optimization; hence, amenable to scalarization or multiple-gradient descent algorithms for which descent directions can be analytically derived and used to update graph and task parameters in tandem<ref:2605.23537#pg5>

Acknowledgment

This work was supported by the NSF under Award ECCS 2231036, by the Spanish AEI Grants PID2022-136887NB-I00 and PID2023-149457OB-I00, and by the Community of Madrid (via grants CAM-URJC F1180 (CP2301), TEC-2024/COM-89, and Madrid ELLIS Unit)<ref:2605.23537#pg5>

References

[1] K. Bello, B. Aragam, and P. Ravikumar, “DAGMA: Learning DAGs via M-matrices and a log-determinant acyclicity characterization,” in Proc. Adv. Neural. Inf. Process. Syst., vol. 35, 2022, pp. 8226–8239<ref:2605.23537#pg5>

[11] A Ghassami, A Yang, N Kiyavash, and K Zhang, “Characterizing distribution equivalence and structure learning for cyclic and acyclic directed graphs,” in Proc. Int. Conf. Mach. Learn., 2020, pp. 3494–3504<ref:2605.23537#pg5>

[19] M Massias, O Fercoq, A Gramfort, and J Salmon, “Generalized concomitant multi-task lasso for sparse multimodal regression,” in Proc. Int. Conf. Artif. Intell. Statist., 2018, pp. 998–1007<ref:2605.23537#pg5>

[21] E Ndiaye, O Fercoq, A Gramfort, V Leclere, and J Salmon, “Efficient smoothed concomitant lasso estimation for high dimensional regression,” in Journal of Physics: Conference Series, vol. 904, 2017<ref:2605.23537#pg5>

[26] J Peters, D Janzing, and B Scholkopf, Elements of Causal Inference: Foundations and Learning Algorithms. The MIT Press, 2017<ref:2605.23537#pg5>

[30] S S Saboksayr, G Mateos, and M Tepper, “CoLiDE: Concomitant linear DAG estimation,” Proc. Int.

Improvements for AI systems

  1. Bold score function adaptation for noise robustness: Implement CoLiDE to jointly estimate DAG adjacency matrices and exogenous noise standard deviations, which is described as a noise-adaptive procedure that is robust (both in terms of DAG learning performance and parameter fine-tuning) to possibly heteroscedastic exogenous noise profiles.

  2. Non-negative structure exploitation: Utilize the non-negativity constraint by casting the problem into a formulation where "non-negativity induces a benign landscape with three useful properties: (i) the true DAG W0 is the unique global minimizer of the augmented Lagrangian; (ii) there are no spurious interior stationary points; and (iii) every acyclic KKT point is W0."

  3. Scalable optimization for large graphs: Employ Non-negative Optimization via Multipliers for Acyclic Digraphs (NOMAD) which uses an iterative sequence of unconstrained problems with a penalty term to tackle the combinatorial acyclicity constraint, achieving per-iteration cost of O(d3), on par with state-of-the-art DAG learning methods.

  4. Time-series causal inference: Apply the NOMAD framework to time series data modeled by SVARMs, allowing for the inference of instantaneous dependencies and lagged matrices, with near-perfect F1-score in recovery of W (and improved estimation of Aτ) when exploiting non-negativity.

  5. Online/Streaming inference: Develop an online adaptive algorithms that can track the (possibly) time-varying connectivity structure by utilizing mini-batch or online updates, allowing for real-time causal structure discovery in streaming data while maintaining computational efficiency.

Abstract

Directed acyclic graphs (DAGs) constitute a central modeling tool to enable principled reasoning about cause-effect interactions in complex systems. However, since the causal structure underlying a group of variables is often unknown and interventions may be infeasible or ethically challenging to implement, there is a need to address the task of inferring DAGs from observational data. However, most classical structure identification approaches face two key obstacles: the combinatorial challenge of enforcing acyclicity, which severely limits scalability, and identifiability challenges arising from latent confounding or heterogeneous noise. This tutorial offers an overview of recent signal processing and optimization advances that address these issues by recasting DAG structure learning as a continuous, score-based estimation problem over adjacency matrices. We begin with a didactic introduction to structural equation models and the formulation of causal graph recovery, followed by a historical survey of score-based methods ranging from early combinatorial search schemes and greedy heuristics to modern continuous frameworks that leverage smooth characterizations of acyclicity. Building on this foundation, we describe concomitant DAG estimation methods that jointly infer sparse causal structure and exogenous noise levels, improving robustness under heteroscedasticity and distribution shifts by rendering the estimator noise adaptive. All in all, the tutorial introduces readers to challenges and opportunities for signal processing research at the crossroads of causal inference, high-dimensional statistics, and scalable graph learning, while outlining emerging directions including online, nonlinear, and neural causal discovery.

Sources

Related papers