Bifidelity Karhunen-Lo`eve Expansion Surrogate with Active Learning for Random Fields

arXiv:2511.03756 · stat.ML, cs.LG, physics.flu-dyn, stat.AP · Submitted 2025-11-05 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Bifidelity Karhunen-Lo`eve Expansion Surrogate with Active Learning for Random Fields".

Jane: Bifidelity KLEs with Active Learning for Random Fields presents a novel surrogate modeling framework that combines Karhunen–Loève expansions and polynomial chaos expansions with an active learning strategy to efficiently construct…

Tom: First, who's behind it and why it matters.

Title and authors: Tom: Let’s talk about the title and who wrote this, which is "Bifidelity Karhunen–Loève Expansion Surrogate with Active Learning for Random Fields." It tells us immediately that they are using a specific mathematical tool, the KLE, combined with a strategy called active learning to handle fields where inputs are uncertain.

Jane: The authors are Aniket Jivani, Cosmin Safta, and Beckett Y. Zhou from institutions like the University of Michigan and Sandia National Laboratories <ref:2511.03756#pg0>. They come from some really strong research backgrounds in applied mathematics and computational science.

Lu: Their background in spectral methods is key here because they leverage the Karhunen–Loève expansion, which is mathematically optimal for representing a function based on its covariance structure <ref:2511.03756#pg2>.

Meng: I'm curious about the active learning part; how do they decide which high-fidelity simulations are worth running when there are so many options available? That decision-making process is where the practical efficiency really gets tested.

Lalam: The authors clearly wanted a method that isn't just brute force simulation, and by combining these concepts, they’re aiming for a model that is both highly efficient and very trustworthy in its predictions.

The paper's summary: Tom: So, the core of the paper explains this bifidelity surrogate modeling framework by breaking down the high-fidelity output into a low-fidelity component and a discrepancy term <ref:2511.03756#pg1>. Essentially, they are trying to capture the main pattern cheaply and then use those few expensive simulations just to fix the systematic errors that the cheap model misses.

Jane: That makes sense when you think about modeling complex things; you can get a rough idea of how something behaves using simple assumptions, and then use targeted, detailed runs only where things get messy or when the rough idea isn't quite right.

Lu: They define the zero-mean component of this bifidelity surrogate by substituting the truncated Karhunen–Loève expansion for the low-fidelity term and a polynomial chaos expansion for the discrepancy part <ref:2511.03756#pg0>. This coupling allows them to build a global polynomial chaos expansion of the entire field with coefficients that change based on where you are in the input space.

Meng: So, they are essentially creating a structure where the low-fidelity data gives you the general shape, and then they use PCE to handle the uncertainty that’s left over from those cheap runs. That’s a sophisticated way to manage complexity.

Lalam: It’s amazing how they create this single surrogate that is informed by both cheap and expensive data simultaneously, which should make our overall modeling pipeline much more scalable for complex problems.

The paper's improvements: Tom: The real innovation here is the active learning strategy they bake in to adaptively pick those high-fidelity simulations, using a Gaussian process regression model guided by an expected improvement criterion <ref:2511.03756#pg1>. They aren't just running simulations randomly; they are intelligently targeting the areas where their current model is weakest.

Jane: That active learning loop means they spend their computational budget wisely, focusing on the regions of the input space that actually cause the most trouble for their model predictions. It’s like having a smart researcher pointing you toward the most important data points to collect next.

Lu: The process involves estimating local surrogate error through cross-validation, modeling those errors with Gaussian process regression, and then using an expected improvement function to select new samples <ref:2511.03756#pg1>. This whole iterative cycle refines the model structure continuously.

Meng: From a practical standpoint, this adaptive sampling is what saves time; instead of running hundreds of expensive simulations blindly, they only run the ones that give us the biggest predictive boost according to that expected improvement function. That’s a massive win for cost control.

Lalam: This iterative refinement capability means the system doesn't just produce one result; it continuously gets smarter by learning from every expensive sample, which should lead to incredibly robust predictions over time.

Conclusion: Tom: So, to wrap up, the paper on "Bifidelity Karhunen–Loève Expansion Surrogate with Active Learning for Random Fields" shows a way to construct computationally affordable models by smartly mixing low-fidelity trends with targeted high-fidelity corrections <ref:2511.03756#pg0>.

Jane: It’s about taking the spectral efficiency of the KLE and coupling it with polynomial chaos expansions, all while using active learning to ensure we only spend our expensive simulation time where it truly matters for accuracy.

Lu: The implication is that we can achieve accurate predictions for field quantities under uncertainty without needing an impossible number of high-fidelity runs, which opens up avenues for simulating more complex physical systems efficiently.

Meng: For the engineering team, this means we can build prototypes faster because the surrogate model itself becomes a highly efficient tool that guides our expensive simulation efforts toward the most critical parameters.

Lalam: This work has implications for how we approach complex modeling; it suggests a future where AI-driven surrogates are not just fast predictors but intelligent construction engines that optimize computational resources for accuracy.

Aniket Jivani, Cosmin Safta, Beckett Y. Zhou, Xun Huan

University of Michigan · Sandia National Laboratories · Georgia Institute of Technology

stat.ML, cs.LG, physics.flu-dyn, stat.AP

Submitted: 2025-11-05

Updated: 2026-10-01

Code: https://github.com/aniketjivani/KLE_UQ_Final

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 92/100

The gist: Bifidelity KLEs with Active Learning for Random Fields presents a novel surrogate modeling framework that combines Karhunen–Loève expansions and polynomial chaos expansions with an active learning

Key concepts

Bifidelity KLE Formulation
This core technique approximates the true high-fidelity output by combining results from low-fidelity (cheap) and discrepancy terms. It defines a bifidelity surrogate that balances these two sources, creating a model that is computationally affordable while still capturing important field trends.
Polynomial Chaos Expansion (PCE)
PCE is used to represent input uncertainties in terms of physical sources. It allows the framework to explicitly link the uncertain input parameters ($\theta$) to the output field values. This enables conditional inference across different fidelity levels, making it easier to understand how uncertainty propagates.
Active Learning Strategy
This adaptive process intelligently selects which high-fidelity simulations need to be run next. It uses a Gaussian process model of the surrogate's error and an expected improvement function to target regions where the current model is least accurate, ensuring computational effort is spent where it yields the biggest predictive gain.

Terminology

Summary

Bifidelity KLEs with Active Learning for Random Fields presents a novel surrogate modeling framework that combines Karhunen–Loève expansions and polynomial chaos expansions with an active learning strategy to efficiently construct accurate, computationally affordable models for field-valued quantities of interest under uncertain inputs. This method leverages inexpensive low-fidelity simulations to capture bulk trends and uses a limited number of high-fidelity simulations, adaptively selected by a Gaussian process regression model guided by an expected improvement criterion, to correct systematic bias and maximize predictive accuracy.

Bifidelity KLE Formulation

The core of the approach is the bifidelity surrogate model, which approximates the high-fidelity (HF) output as a combination of low-fidelity (LF) and discrepancy terms:

yHF(x, θ) = yLF(x, θ) + [yHF(x, θ) − yLF(x, θ)].

This is expanded into mean and zero-mean components to define the bifidelity surrogate:

y˜BF(x, θ) = ˜µBF(x) + ˜y0,BF(x, θ), where µ˜BF(x) = [˜µLF(x) + ˜µ∆(x)] and y˜0,BF(x, θ) = [˜y0,LF(x, θ) + ˜y0,∆(x, θ)]. The zero-mean component is constructed by substituting the truncated Karhunen–Loève expansion (KLE) for the LF term and a polynomial chaos expansion (PCE) for the discrepancy term:

y˜0,BF(x, θ) = ˜y0,LF(x, θ) + ˜y0,∆(x, θ). This results in a global PCE of the bifidelity field with spatially varying coefficients constructed from two independently built surrogates.

Polynomial Chaos Expansion Integration

To explicitly represent inputs in terms of their physical sources of uncertainty and enable conditional inference across fidelity levels, the KLE is coupled with PCEs. Each modal coefficient ζk is treated as a deterministic function of the input parameters θ, expressed as:

ζk(θ) = X∑β∈I bk,βΨβ(ξ1(θ),..., ξns(θ)), where Ψβ are multivariate orthonormal basis polynomials. The coefficients bk are determined by solving a non-intrusive regression problem using the available simulation data, often employing Tikhonov regularization to prevent overfitting and promote sparsity in the coefficient vector bk. This allows for fast evaluation of ζk(θ) for new inputs θ, which is crucial for reconstructing y0(x, θ) via the truncated KLE.

Active Learning Strategy

The framework incorporates an active learning strategy to adaptively select new HF simulations based on the surrogate’s generalization error. The process involves:

  1. Estimating local surrogate error through a k-fold cross-validation procedure, yielding a scalar error metric ε(i) = yHF(x, θi) − y˜BF,s(i)BF (x, θi) / yHF(x, θi).

  2. Modeling these errors using Gaussian process (GP) regression to infer a smooth approximation of the generalization error across the input space: ε ∼ GP(m0(·), k(·, ·)).

  3. Maximizing an expected improvement (EI) acquisition function to select new HF samples, targeting regions of high surrogate error: EI(θ) = E[max(ε − ε∗, 0) ε ∼ N (mε(θ), σ2ε(θ))].

Overall Algorithm and Performance

The complete procedure is summarized in Algorithm 1, which iteratively refines the surrogate model. New HF samples are acquired by maximizing the EI function or employing the Kriging Believer heuristic for batch selection, allowing up to five HF evaluations simultaneously. The bifidelity surrogate is rebuilt at each stage after incorporating new data, and cross-validation errors are recomputed to fit a new GP prior. This iterative refinement continues until a stopping criterion is met, such as reaching a target error tolerance or exhausting the computational budget B. Numerical experiments on 1D analytical benchmarks, 2D convection-diffusion problems, and 3D turbulent jet flows demonstrate consistent improvements in predictive accuracy and sample efficiency relative to single-fidelity and random-sampling approaches.

Forward Uncertainty Quantification (UQ)

The framework is evaluated for forward UQ by comparing the mean and ±1 standard deviation bounds propagated from the true HF model (yHF) against those from the bifidelity surrogate (y˜BF) and LF-KLE surrogates.

Improvements for AI systems

Here are specific improvements to AI systems based on the BF-KLE-AL framework, along with what those improved systems can achieve:


The following improvements focus on integrating the core concepts of Bifidelity KLEs, Polynomial Chaos Expansions (PCEs), and Active Learning into existing AI/ML pipelines for uncertainty quantification (UQ) and surrogate modeling.

  1. A novel, scalable surrogate model capable of handling high-dimensional, correlated input spaces with computational efficiency.

  2. An adaptive sampling strategy that minimizes the required expensive high-fidelity (HF) evaluations necessary to achieve a target level of predictive accuracy.

Specific capabilities of the improved AI systems:

  1. A system that can perform fast, accurate predictions for complex physical phenomena (e.g., CFD simulations, climate modeling, material science) with significantly fewer total computational hours compared to relying solely on expensive HF models or traditional random sampling methods.

  2. The ability to provide reliable uncertainty quantification (UQ) estimates for these complex systems by accurately capturing both the bulk behavior learned from cheap low-fidelity (LF) data and the necessary corrections derived from a minimal, intelligently selected set of high-fidelity samples.

  3. The system can automatically identify hot spots or regions in the input parameter space where the current model is most uncertain or inaccurate, allowing engineers to focus expensive computational resources precisely where they yield the greatest reduction in prediction error (e.g., optimizing a design parameter that is sensitive to turbulent flow conditions).

  4. A robust workflow for iterative refinement: The system can continuously learn from new HF evaluations, adapt its internal model structure (by retraining the KLE and PCE components), and dynamically adjust its sampling strategy to achieve convergence toward the true high-fidelity solution at an optimal cost.

In summary, the improved AI system shifts from being a black-box predictor (relying on expensive HF runs) to being an intelligent, cost-aware surrogate construction engine that rapidly converges on accurate solutions by strategically balancing cheap approximations with targeted, high-value physical insights.

Abstract

We present a bifidelity Karhunen--Loève expansion (KLE) surrogate model for field-valued quantities of interest (QoIs) under uncertain inputs. The QoIs considered here are scalar fields. The approach combines the spectral efficiency of the KLE with polynomial chaos expansions (PCEs) to preserve an explicit mapping between input uncertainties and output fields. By coupling inexpensive low-fidelity (LF) simulations that capture dominant response trends with a limited number of high-fidelity (HF) simulations that correct for systematic bias, the proposed method can enable accurate and computationally affordable surrogate construction. To further improve surrogate accuracy, we develop an active learning strategy that adaptively selects new HF evaluations based on the surrogate's generalization error, estimated via cross-validation and modeled using Gaussian process regression. New HF samples are then acquired by maximizing an expected improvement criterion, targeting regions of high surrogate error. The resulting BF-KLE-AL framework is demonstrated on three examples of increasing complexity: a one-dimensional analytical benchmark, a two-dimensional convection-diffusion system, and a three-dimensional turbulent round jet simulation based on Reynolds-averaged Navier--Stokes (RANS) and enhanced delayed detached-eddy simulations (EDDES). The experiments show that bifidelity gains depend on LF accuracy, discrepancy approximation, and the allocation of simulation cost. Active learning improves prediction over random sampling in several settings, while the cost-matched comparisons identify both favorable regimes and cases where an HF-only surrogate is more accurate.

Sources

Related papers