Eigenvalue-Decomposition Cost Denoising as an Alternative to Predict-then-Optimize for Shortest-Path Problems
econ.EM, cs.LG, math.OC, math.SP
Submitted: 2026-09-14
Updated: 2026-09-14
Comments: 7 pages
License: http://creativecommons.org/licenses/by/4.0/
The gist: Predict-then-optimize methods such as Smart "Predict, then Optimize" (SPO+) of Elmachtoub and Grigas (2022) learn a mapping from contextual features to unknown edge costs and then solve the induced
Terminology
Abstract
Predict-then-optimize methods such as Smart "Predict, then Optimize" (SPO+) of Elmachtoub and Grigas (2022) learn a mapping from contextual features to unknown edge costs and then solve the induced combinatorial problem on the predicted costs. This approach is powerful but relies on the predictive model being well specified: when the true cost-generating process is nonlinear in the features and the predictor is linear, SPO+'s performance degrades as the misspecification grows. We propose and evaluate a structurally different remedy for a specific but common setting: when the decision-maker observes many noisy realizations of the same underlying cost process, the realized cost vectors themselves can be treated as a noisy signal and denoised directly, via eigenvalue decomposition (equivalently, Principal Component Analysis) of their covariance matrix, before ever invoking a predictive model. We instantiate this idea on the 5 times5 grid shortest-path benchmark introduced by Elmachtoub and Grigas (2022), retaining only the top- k eigenvectors of the training cost covariance matrix and projecting new noisy cost observations onto that subspace prior to solving with Dijkstra's (1959) algorithm. We find that the choice of k is decisive: keeping only k = 2 eigenvectors discards real signal and underperforms even the naive noisy-cost baseline, while setting k = 5 to match the true latent feature dimension makes eigenvalue-denoised Dijkstra the best-performing method at every misspecification level tested, outperforming SPO+ by a wide margin under high misspecification.
Sources
Related papers
- SLIM: Stochastic Learning and Inference in Overidentified Models
- High-dimensional censored MIDAS logistic regression for corporate survival forecasting
- Cross-Fitting-Free Debiased Machine Learning with Multiway Dependence
- Can large language models assist choice modelling? Insights into prompting strategies and current models' capabilities
- Causal Inference in Possibly Nonlinear Factor Models
- Mining Causality: AI-Assisted Search for Instrumental Variables