Projection-based multifidelity linear regression for data-scarce applications
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Projection-based multifidelity linear regression for data-scarce applications".
Jane: The paper was written by Vignesh Sella, Julie Pham, Karen Willcox and Anirban Chaudhuri from University of Texas at Austin and Santa Fe Institute.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Jane: We also have Lu with us today — senior AI researcher at Tsinghua.
Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.
Jane: We also have Lalam with us today — the in-house Large Language Model.
Tom: Alright, let's get started.
Title: Tom: Welcome back to the show, everyone. Today we're looking at a paper that's got a real mouthful of a title: "Projection-based multifidelity linear regression for data-scarce applications." Jane, what do you make of that title?
Jane: Tom, I love it because every single word there is doing work. "Data-scarce" tells you the problem, "multifidelity" tells you the strategy, and "projection-based" tells you the trick they use to make it all fit together.
Tom: And the authors are from UT Austin and the Santa Fe Institute — Vignesh Sella, Julie Pham, Karen Willcox, and Anirban Chaudhuri. Karen Willcox has been a big name in reduced-order modeling for years, so this is a group that knows what they're doing.
Jane: Right, and the problem they're tackling is one that keeps coming up in engineering. You've got a high-fidelity simulation that's accurate but expensive, and you've got a low-fidelity simulation that's cheap but less accurate. You want to combine both to build a surrogate model that predicts the expensive one.
Tom: So it's like learning from a really smart but expensive tutor and a cheaper, less reliable one at the same time.
Jane: Exactly. And the "data-scarce" part is crucial. We're talking about having maybe three to ten high-fidelity samples. That's nothing. You can't train a neural network on that. You can barely fit a simple polynomial.
Tom: And that's where the "projection-based" part comes in. The outputs here are huge — we're talking about pressure fields with tens of thousands of values. So they use principal component analysis to compress those outputs into a much smaller set of coordinates.
Lu: If I can jump in here, Tom — the clever bit is that they do the regression in that compressed space, not in the original high-dimensional space. That's what makes the whole thing tractable when you have so few samples.
Jane: Lu's right. And the title also hints at something important: they're not just combining the two data sources naively. They're being careful about how much to trust the low-fidelity data, which is what we'll dig into next.
Tom: So before we get into the weeds, let's set the stage. This is a paper about making the most of very little data by being smart about what you trust and how you compress. And the application they test it on is hypersonic vehicle design, which is about as exciting as aerospace gets.
Jane: And that's the hook for our next segment — how exactly they combine those two fidelity levels and what choices they make about weighting the low-fidelity data. Stay with us.
Summary: Tom: So we're back with "Projection-based multifidelity linear regression for data-scarce applications." Jane, last segment we set the stage. Now let's talk about what the paper actually does.
Jane: The core idea is data augmentation. You take your limited high-fidelity samples, you take your plentiful low-fidelity samples, and you combine them into one training set. Then you do weighted least squares regression on that combined set.
Tom: But it's not just throwing them together, right? They have two different ways of generating the synthetic data from the low-fidelity model.
Jane: Right. The first is direct augmentation — you just use the low-fidelity data as-is. The second is more interesting. They train a linear correction map that learns how to transform low-fidelity outputs into something that approximates high-fidelity outputs.
Lu: And that second approach is really the clever one. You're essentially learning the systematic bias between the two models and correcting for it before you even train your surrogate.
Meng: But from an engineering standpoint, I want to know about the weighting. If you're mixing data from two different sources, you can't treat them equally. The low-fidelity data is less trustworthy.
Jane: Exactly, Meng. And that's where the weighted least squares comes in. They assign weight one to the high-fidelity samples and a smaller weight to the synthetic low-fidelity samples. But here's the question — how small?
Tom: And that's where it gets interesting. They test fixed weights, like zero point zero one or zero point nine, and they also test a proximity-based scheme. The idea there is that low-fidelity samples that are far away from any high-fidelity sample are more valuable because they fill gaps in the input space.
Lu: So a low-fidelity sample sitting right next to a high-fidelity sample is redundant — it's not telling you anything new. But one that's far away is giving you information about a region you haven't explored.
Jane: And they use leave-one-out cross-validation to pick the optimal weight for each training dataset. That's important because the right weight depends on how well the low-fidelity model matches the high-fidelity model in that particular dataset.
Meng: So the weight isn't just a hyperparameter you set once. It's being tuned automatically based on the data you have.
Tom: And the results? On their hypersonic test case, the data augmentation methods get about three to twelve percent improvement in median accuracy over single-fidelity regression. And that's with only three to ten high-fidelity samples.
Jane: The additive method they compare against — which is a more traditional approach based on the Kennedy-O'Hagan framework — doesn't do nearly as well. It's basically no better than single-fidelity.
Lu: That's a striking result. The additive correction approach has been popular for years, but in this ultra-low-data regime, the data augmentation approach is clearly superior.
Tom: So the summary is: use data augmentation, learn a correction map if you can, and let cross-validation pick your weights. That's the recipe. And next segment, we'll talk about what this means for real engineering problems.
Improvements: Tom: We're back with "Projection-based multifidelity linear regression for data-scarce applications." We've covered what the paper does. Now let's talk about why it matters and what it improves on.
Jane: The biggest improvement is robustness. When they trained the single-fidelity model on just high-fidelity data, the accuracy varied wildly depending on which training samples you happened to draw. The multifidelity methods were much more consistent.
Meng: That's huge for engineering practice. If your surrogate model's accuracy depends heavily on luck, you can't trust it for design decisions. You need to know what you're getting.
Tom: And there's another improvement I want to highlight. Because the data augmentation gives you more training samples, you can afford to use a higher-order polynomial basis. The single-fidelity model was stuck with linear regression because with only ten samples, you can't fit a quadratic.
Lu: Right, Tom. With the augmented dataset, they could train a second-order polynomial. That's a real practical advantage — you're not just adding more data, you're enabling a more expressive model.
Jane: And the proximity-based weighting is a genuine improvement over just using a fixed weight. The paper shows that the direct augmentation method is sensitive to the choice of weight — you can lose ten percent accuracy by picking the wrong one. But with the cross-validation approach, you automatically land near the best option.
Meng: So the improvement isn't just in accuracy, it's in making the method usable without a lot of hand-tuning. That's what gets a method adopted in industry.
Tom: And let's not forget the computational cost angle. The low-fidelity model in their test case costs about one one-hundred-twenty-seventh of the high-fidelity model. So eighty low-fidelity samples cost less than one extra high-fidelity sample.
Jane: That's the real win. You're getting a twelve percent accuracy improvement at essentially no additional computational cost. That's the kind of trade-off engineers love.
Lu: And I think the methodological improvement extends beyond this specific application. The idea of learning a linear correction map between fidelity levels in a reduced space — that's portable to other problems where you have a cheap model and an expensive model.
Meng: Though I'd want to see how it handles cases where the low-fidelity model is less correlated with the high-fidelity one. In their test case, the low-fidelity model is pretty good. What if it's really bad?
Jane: That's a fair concern, Meng. The cross-validation weighting should help — if the low-fidelity data is useless, the optimal weight would be near zero. But that's something future work would need to explore.
Tom: And that's exactly the kind of question we'll wrap up with in our final segment — what this means going forward and what questions remain.
Conclusion: Tom: So we've spent some time with "Projection-based multifidelity linear regression for data-scarce applications." Let's pull it all together.
Jane: The paper gives us a practical recipe for building surrogate models when high-fidelity data is extremely scarce. You project your high-dimensional outputs onto a low-dimensional subspace, you augment your limited high-fidelity data with transformed low-fidelity data, and you use weighted least squares with automatically tuned weights.
Lu: And the key result is that this beats both single-fidelity regression and the more traditional additive multifidelity approach. In the three-to-ten high-fidelity sample regime, you get meaningful accuracy gains at almost no extra computational cost.
Meng: From my perspective, the most valuable contribution is the robustness. The multifidelity methods are more consistent across different training data draws, and the automatic weight selection removes a lot of guesswork.
Tom: And the application is genuinely exciting — predicting pressure fields on a hypersonic vehicle at different Mach numbers, angles of attack, and sideslip angles. That's directly relevant to designing vehicles that need to survive extreme flight conditions.
Jane: The future work they mention is also interesting — extending these methods to neural networks and other regression techniques, and exploring different coordinate transformations for the correction map.
Lu: I'd add that the fundamental idea — learning how to correct a cheap model using a few expensive samples, in a reduced space — could apply well beyond aerospace. Any field with expensive simulations and cheap approximations could benefit.
Tom: Well said, Lu. So that's "Projection-based multifidelity linear regression for data-scarce applications" — a solid contribution to the scientific machine learning toolkit. Jane, what's next on the docket?
Jane: We've got a paper on physics-informed neural networks for inverse problems that I think will get us all talking. Let's take a quick break and then jump into that.
Tom: Sounds good. Thanks for listening, everyone. We'll be right back.
Vignesh Sella, Julie Pham, Karen Willcox, Anirban Chaudhuri
University of Texas at Austin · Santa Fe Institute
stat.ML, cs.CE, cs.LG
Submitted: 2026-08-14
Updated: 2026-08-18
Comments: 23 page, 7 figures, submitted to Machine Learning for Computational Science and Engineering special issue Accelerating Numerical Methods With Scientific Machine Learning
Journal ref: Mach. Learn. Comput. Sci. Eng. 1, 47 (2025)
DOI: 10.1007/s44379-025-00049-5
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 68/100
Key concepts
- Data Scarce Applications
- This refers to engineering problems where high-accuracy simulations are extremely expensive and rare. The paper addresses scenarios where only a handful of high-fidelity samples—typically three to ten—are available, making traditional machine learning difficult.
- Multifidelity/Data Augmentation
- This is the core strategy of combining limited, accurate (high-fidelity) data with plentiful, cheaper but less accurate (low-fidelity) data. The method creates a single augmented training set to improve model performance while managing resource constraints.
- Projection-based
- The outputs of these simulations are often very high-dimensional. This technique uses Principal Component Analysis (PCA) to compress those huge output fields into a much smaller, manageable set of coordinates for the regression model.
- Weighted Least Squares Regression
- This is the training method used on the combined data. Instead of treating all samples equally, weights are assigned—high-fidelity samples receive more weight than low-fidelity ones—to account for the lower trustworthiness of the cheaper data.
Terminology
Summary
Summary
This paper develops multifidelity (MF) methods for multiple-input multiple-output linear regression, targeting data-limited applications with high-dimensional outputs. The motivation is that surrogate modeling for systems with high-dimensional quantities of interest remains challenging, particularly when training data are costly to acquire. The authors note that an important challenge in scientific machine learning is to develop methods that can exploit and maximize the amount of learning possible from scarce data,
especially in computational fluid dynamics (CFD), where expensive-to-evaluate high-fidelity (HF) models make many-query problems such as uncertainty quantification, risk analysis, optimization, and optimization under uncertainty computationally prohibitive.
The proposed methods integrate many inexpensive low-fidelity model evaluations with limited, costly high-fidelity evaluations.
The core methodological contribution is the formulation of a data augmentation approach that leverages weighted least squares (WLS) to explicitly incorporate LF data into the regression.
Specifically, the authors introduce and analyze two methods for MF data augmentation: (i) direct data augmentation by combining fidelity sources, and (ii) data augmentation employing an explicit mapping between low- and HF outputs.
They also present two weighting schemes for WLS and perform a sensitivity analysis on the choice of weight.
As a point of comparison, they extend the work in Ref. [16] to create a projection-enabled variation of the additive MF structure following the Kennedy-O'Hagan formulation [18].
The methods operate within a reduced-dimensional subspace obtained through principal component analysis (PCA) to effectively handle both training data scarcity and the high dimensionality (on the order of tens of thousands of quantities of interest) inherent in our problem setting.
The regression problem is formulated as a linear-regression-based surrogate model in the reduced-dimensional space f: Rd → Rk, parameterized by regression coefficients β.
The surrogate model is linear in the regression coefficients and can be trained using either ordinary or weighted least squares, depending on the MF regression methodology.
For the data augmentation approaches, synthetic data are generated in two ways: (1) direct augmentation, where the LF data are used directly as synthetic training data,
and (2) explicit mapping, where a learned linear correction map is applied to the LF data to approximate the HF behavior at X LF.
The explicit mapping approach trains a linear model g between reduced-order representations of the LF and HF outputs in a shared low-dimensional space.
When LF samples are not co-located with HF samples, a LF surrogate model fLF is used to obtain LF predictions at HF input locations. The linear mapping model g is then trained via OLS on the co-located dataset.
To account for fidelity-dependent variance,
the authors apply WLS with distinct weights assigned to HF and synthetic training samples.
They introduce a proximity-based weighting scheme that down-weights LF samples located near HF samples,
since LF samples located near HF samples in the input space can be considered redundant or uninformative since continuity ensures that proximity in the input space yields proximity of the outputs.
The weighting function uses a Heaviside step function with a threshold τ, and the hyperparameter wsyn is selected via leave-one-out cross-validation (LOOCV) using the BFGS algorithm.
The methods are demonstrated on the Initial Concept 3.X (IC3X) hypersonic vehicle
problem, where the quantity of interest is the distributed aerodynamic pressure load over the surface of the vehicle at various flight condition parameters.
The input space consists of Mach number M ∈ [5, 7], angle of attack α ∈ [0, 8], and sideslip angle β ∈ [0, 8]. The surface pressure field is computed using the flow solver Cart3D, with the HF model using finer mesh refinements and the LF model using coarser mesh. The output dimension is m = 55966 nodes. The computational cost ratio of HF to LF is 1/127.
The numerical experiments use a very limited number of HF samples NHF ∈ [3, 10], a LF training sample size of NLF = 80, and a HF testing sample size of NHF test = 50.
A tolerance of ϵ = 0.995 for cumulative energy leads to k = 7 for most LF training datasets and k = 4 for most HF training datasets with NHF > 4. Results are presented over 50 repetitions of the training dataset.
Key findings include: the direct data augmentation method is sensitive to the choice of wsyn for the fixed weighting scheme, with a variation of up to ∼ 10% in median accuracy,
while the explicit map data augmentation method is less sensitive to changes in the sample weight, with a variation of up to 2% in median accuracy.
The LOOCV method for determining the optimal weight performs close to the best fixed weighting scheme option for both data augmentation methods.
The distribution of optimized LF sample weights is generally bimodal in our setting with NHF ≤ 10,
and as the HF sample size increases, the resulting weight distribution shifts toward smaller magnitudes, suggesting that the added HF data reduces the reliance on LF information.
The comparison of methods shows that the additive MF method performs similar to the SF linear regression and does not offer significant increase in accuracy for this application,
while "both the data augmentation techniques (using the optimal wsyn after LOOCV and proximity-based weights) perform better than the additive approach and show significant improvement in accuracy over the SF linear regression for equivalent computational cost. The MF method with explicit map performs best with few samples, while direct augmentation had the highest accuracy with the largest amount of training data. Specifically,
the data augmentation technique using explicit map leads to an improvement of approximately 9.5% compared to the SF model for NHF = 3 HF samples and 3.2% compared to the SF model for NHF = 10 HF samples. Interpolating at NHF = 3.63 (3 HF, 80 LF samples) yields
a 12.4% improvement in accuracy compared to the SF method. Overall,
multifidelity linear regression achieves approximately 3% − 12% improvement in median accuracy compared to single-fidelity methods with comparable computational cost" in the low-data regime of no more than ten high-fidelity samples.
Future work directions include expand these MF regression methods to different underlying regression techniques, such as neural networks and regression trees
and explore different coordinate transformation techniques for the explicit mapping method.
Improvements for AI systems
Based on the paper, here are the specific improvements I can make to AI systems and what the improved system can do:
-
Improvement: Add a data augmentation layer that integrates low-fidelity (LF) and high-fidelity (HF) data using two methods: (a) direct augmentation and (b) explicit linear mapping between LF and HF outputs.
-
What it does: The system can train accurate regression models with as few as 3–10 HF samples by leveraging 80+ inexpensive LF samples, achieving 3–12% higher median accuracy than single-fidelity models at comparable computational cost.
-
Improvement: Embed PCA-based dimensionality reduction (via thin SVD) before regression, projecting outputs (e.g., 55,966 pressure nodes) onto a low-dimensional subspace (k=4–7) with cumulative energy tolerance ϵ=0.995.
-
What it does: The system can handle outputs with tens of thousands of dimensions without ill-posedness, making it feasible for CFD surface pressure fields, structural stress fields, or any high-dimensional physical quantity.
-
Improvement: Implement WLS with fidelity-specific weights: HF samples get weight 1; LF samples get weight w syn (optimized via LOOCV) and are further down-weighted if they are close to HF samples (using a Heaviside step function with Euclidean distance threshold).
-
What it does: The system automatically discounts redundant LF samples near HF points and emphasizes LF samples that fill gaps in input space, reducing heteroscedastic noise and improving robustness across different training data draws.
-
Improvement: Add a LOOCV-based optimizer (using BFGS) to select the optimal LF sample weight w syn* for each training dataset, rather than using a fixed weight.
-
What it does: The system adapts to the specific data distribution, avoiding sensitivity to arbitrary weight choices (which caused up to 10% accuracy variation in fixed-weight schemes) and consistently performs near the best fixed-weight option.
-
Improvement: Train a linear map g: R k → R k that transforms reduced LF states to reduced HF states in the HF PCA basis, using co-located predictions from an LF surrogate model.
-
What it does: The system can generate synthetic HF-like data from LF evaluations, improving accuracy by 9.5% with 3 HF samples and 3.2% with 10 HF samples, while being less sensitive to weight selection than direct augmentation.
-
Improvement: Implement a projection-enabled additive MF method that trains an LF surrogate and a discrepancy model (δ = HF − LF) separately, then combines predictions.
-
What it does: Provides a baseline comparison method that, while not outperforming data augmentation, offers an alternative for cases where LF and HF models have strong linear correlation and co-located data is available.
-
Predict high-dimensional physical fields (e.g., pressure, temperature, stress) with minimal HF data: It can train on 3–10 HF simulations plus 80 LF simulations and achieve median normalized L2 accuracy of 0.85–0.93, versus 0.77–0.90 for single-fidelity models.
-
Operate in ultra-low-data regimes: It remains robust and accurate when HF samples are extremely scarce (NHF ≤ 10), which is common in aerospace design, clinical trials, or rare-event modeling.
-
Automatically balance data fidelity and informativeness: It down-weights redundant LF samples near HF points and up-weights LF samples that fill unexplored input regions, improving generalization.
-
Scale to outputs with 50,000+ dimensions: The PCA projection reduces computational and memory overhead, enabling regression on full CFD mesh outputs without manual feature engineering.
-
Provide uncertainty-aware predictions: The LOOCV-based weight selection and proximity weighting inherently account for heteroscedasticity, giving more reliable error estimates than ordinary least squares.
-
Support multiple fidelity levels: The bifidelity framework can be extended to more than two levels, allowing hierarchical data integration (e.g., coarse, medium, fine mesh simulations).
These improvements are directly implementable in existing surrogate modeling pipelines, physics-informed machine learning frameworks, or any regression system dealing with scarce high-fidelity data and high-dimensional outputs.
Related papers
- Behavior of prediction performance metrics with rare events
- Optimal Estimation of Generic Dynamics by Path-Dependent Neural Jump ODEs
- A Posterior-Dynamics Framework for Imaging Inverse Problems with Pretrained Diffusion Priors
- One Permutation Is All You Need: Fast, Deterministic Feature Importance and Model Stress-Testing
- Online Conformal Prediction for Non-Exchangeable Panel Data
- Deep Time-Series Forecasting in 10 Years: A Survey