High-dimensional networks and mean squared error for possibly misspecified models
Lourens Waldorp
University of Amsterdam
stat.ML, cs.LG
Submitted: 2026-08-13
Updated: 2026-08-14
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
Importance score: 75/100
The gist: The paper "High-dimensional networks and mean squared error for possibly misspecified models" by Lourens Waldorp addresses the challenge of estimating networks (Gaussian graphical models) when the
Terminology
Summary
The paper High-dimensional networks and mean squared error for possibly misspecified models
by Lourens Waldorp addresses the challenge of estimating networks (Gaussian graphical models) when the number of parameters (edges) exceeds the number of observations (high-dimensional setting). The goal is to perform reliable neighbourhood selection for each node in a network, specifically to avoid false positive edges (spurious connections).
The paper shows that in high-dimensional settings, it is possible to obtain a conservative estimate of a node's neighbourhood (i.e., low false positive rate). The approach uses nodewise regression, where each node is regressed on all others, and the ridge estimator is used to handle the case where p > n (more parameters than observations). The key findings are:
-
Mean squared error (MSE) and bias-variance trade-off: The paper decomposes the test MSE into squared bias and variance. It shows that the ridge parameter α directly controls the variance:
by increasing α we see that we are lowering the variance V (J), and hence the test MSE.
The squared bias behaves differently:the squared test bias could increase by increasing α.
This analysis explains thedouble descent
phenomenon, where the MSE decreases again after the interpolation point (p = n). The peak at the interpolation point is described asan artefact of choosing the wrong ridge parameter α = 0.0001.
-
Model misspecification: When the true model is nonlinear (e.g., a sigmoid function) but a linear model is used for estimation, the MSE includes an additional misspecification term. The paper shows that this term can be bounded:
Mf (J) ≤ σO(K)
(Proposition B.5). In the misspecified case, the global minimum of the test MSE can shift to a largely overparameterised model, which can lead to selecting too many edges. -
Minimum description length (MDL) for model selection: The paper proves that MDL, which explicitly accounts for the volume of the parameter space, leads to correct neighbourhood selection or smaller (no false positives) in high-dimensional settings. The key result is Proposition 5.1: "If Eσ̂J2 > γ, for some γ > 0, then with probability at least 1 − dp exp(−(p − d)) it holds that Ŝ ⊆ S; equivalently, the probability of false positives goes to 0.
This is because the penalty term involving the p-dimensional ball Bp (β2)
will become large, and hence apply a stronger penalty with increasing dimension." -
Simulation results: Simulations confirm the theoretical results. For correctly specified linear models, MDL has a false positive rate of 0 regardless of dimension. For misspecified models, MDL and Lasso maintain low false positive rates, while other methods (AIC, Lasso-ds) show increased false positive rates, especially at high dimensions. The paper concludes:
we recommend using MDL for nodewise selection in the high-dimensional setting for Gaussian graphical modelling.
The paper's contribution is showing that while standard model selection methods (AIC, BIC, Lasso) lead to spurious edges in high-dimensional networks, MDL's explicit penalty for the volume of the parameter space counteracts the low MSE at high dimensions, achieving the goal of low false positive rates in both correctly specified and misspecified linear models.
Improvements for AI systems
Improvements to AI Systems:
- Adaptive Regularization for High-Dimensional Models
-
Implement a ridge parameter α that is automatically tuned based on the bias-variance decomposition of test MSE, rather than fixed at a small value (e.g., α = 0.0001). This avoids the
double descent
peak at the interpolation point (p = n) by selecting α that minimizes expected test MSE, not just training error. -
Resulting capability: AI systems can train models with p >> n (e.g., genomics, medical imaging) without spurious overfitting, achieving stable generalization.
- Misspecification-Aware Model Selection
-
Add a misspecification bound term (e.g., Mf(J) ≤ σO(K)) to the loss function during model selection. This term penalizes models where the assumed linear structure deviates from true nonlinear relationships (e.g., sigmoid, exponential).
-
Resulting capability: AI systems can detect when their underlying assumptions are wrong and automatically shift to more complex architectures (e.g., neural networks) or add a regularization term to prevent overconfident edge selection in causal inference or graph learning.
- Minimum Description Length (MDL) for Sparse Graph Learning
-
Replace AIC/BIC/Lasso penalties with MDL-based penalties that explicitly account for the volume of the parameter space (the p-dimensional ball Bp). This ensures that as dimension grows, the penalty grows proportionally, driving false positive rates to zero.
-
Resulting capability: AI systems performing feature selection, causal discovery, or network inference (e.g., in social networks, brain connectivity) will output only true edges, even when the number of candidate edges vastly exceeds samples. This is critical for interpretable AI in scientific discovery.
- Conservative Neighbourhood Selection for High-Dimensional Inference
-
Use the theoretical guarantee (Proposition 5.1) to set a threshold γ for the estimated error variance σ̂J2. If σ̂J2 > γ, the system can declare a node's neighbourhood with zero false positives, even if it misses some true edges.
-
Resulting capability: AI systems can provide reliable, conservative predictions in safety-critical domains (e.g., medical diagnosis, fraud detection) where false positives are more costly than false negatives.
- Bias-Variance Decomposition for Model Complexity Control
-
Integrate the paper’s decomposition into automated machine learning (AutoML) pipelines to decide when to stop adding parameters. The system can compute squared bias and variance separately and stop when variance reduction no longer offsets bias increase, avoiding the interpolation peak.
-
Resulting capability: AI systems can automatically determine the optimal model size for any dataset, reducing overfitting in deep learning and improving generalization on small-sample, high-dimensional data.
- Robustness to Nonlinearity in Linear Models
-
When a linear model is used but the true relationship is nonlinear, the system can bound the additional misspecification error and adjust its confidence intervals or prediction intervals accordingly.
-
Resulting capability: AI systems can flag when their linear assumptions are violated and either warn the user or switch to a nonlinear kernel method, improving reliability in real-world applications like economics or ecology.
Abstract
To avoid missing important variables and their connections in networks, more and more variables are included in network analysis. Here we show that in a setting with many more parameters than observations (high-dimensional) it is possible to get a conservative (i.e., low false positive rate) estimate of the neighbourhood for each node (which connections are in the network). A neighbourhood is often estimated with a linear model, and this leads to two interesting cases: (i) If the true model is linear, then neighbourhood selection work reasonably well, and (ii) if the true model is nonlinear, then neighbourhood selection requires a penalty for the high dimensions. Here we show the impact of the ridge parameter on the mean squared error, and how this leads to low test variance and hence to neighbourhoods with large numbers of edges. We connect these insights with results from machine learning, where the so-called double descent (when more parameters are included than observations, the mean squared error goes down a second time) has put the traditional view on model selection upside down. Essentially, for adequate neighbourhood selection in models with a large number of parameters, the volume of the model space needs to be included in the penalty. Most neighbourhood selection methods (e.g., Lasso, AIC, BIC) lead to spurious edges (high false positive rate), but we prove that in the high-dimensional setting, minimum description length leads to correct neighbourhood selection or smaller (low false positive rates) in both cases when either the model is correctly or incorrectly assumed linear
Sources
- Deep learning: a statistical viewpoint
- Double Descent Risk and Volume Saturation Effects: A Geometric Perspective
- A Farewell to the Bias-Variance Tradeoff? An Overview of the Theory of Overparameterized Machine Learning
- Revisiting minimum description length complexity in overparameterized models
- Surprises in High-Dimensional Ridgeless Least Squares Interpolation
- Linear Regression in a Nonlinear World
Related papers
- Behavior of prediction performance metrics with rare events
- Optimal Estimation of Generic Dynamics by Path-Dependent Neural Jump ODEs
- A Posterior-Dynamics Framework for Imaging Inverse Problems with Pretrained Diffusion Priors
- One Permutation Is All You Need: Fast, Deterministic Feature Importance and Model Stress-Testing
- Online Conformal Prediction for Non-Exchangeable Panel Data
- Deep Time-Series Forecasting in 10 Years: A Survey