Kernel Methods for Learning Operators with Multiple Inputs and Outputs
Adrien Weihs, Chunyang Liao, Jingmin Sun, Hayden Schaeffer
University of California Los Angeles · Johns Hopkins University · University of Arkansas
cs.LG, math.ST, stat.ML, stat.TH
Submitted: 2026-08-12
Updated: 2026-08-13
Code: https://github.com/liaochunyang/kernelMO
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 75/100
The gist: This paper introduces a general kernel-based encoder-decoder framework for learning operators with multiple inputs and outputs.
Terminology
Summary
This paper introduces a general kernel-based encoder-decoder framework for learning operators with multiple inputs and outputs. The authors develop a mathematically rigorous approach that separates observation, representation, learning, and reconstruction in operator learning problems. The framework is specialized to multiple operator learning through the introduction of KernelMO, a family of kernel methods with complementary operator-valued and product-space formulations.
The paper makes five principal contributions:
1. A general kernel framework for encoder–decoder learning. The authors develop a kernel-based framework for learning maps between Hilbert spaces within an encoder–decoder architecture. In Theorem 3.3, they show that every kernel defined on the latent surrogate space induces a corresponding kernel on the original input and output spaces.
In Corollary 3.4, they establish an equivalence between learning in the latent reproducing kernel Hilbert space and learning with the induced kernel on the original spaces.
2. An encoder–decoder error decomposition. Theorem 3.7 provides an error decomposition of the form: Approximation error ≲ Representation error + Learning error + Data consistency error,
which separates the contributions of the encoder-decoder architecture, the learned surrogate, and the measurement process.
3. Approximation theory for multi-input, multi-output operator learning. For maps of the form G: X1 × ··· × Xm → Y1 × ··· × Yn, Theorem 3.15 establishes: Approximation error ≲ max Input reconstruction errori + max Output reconstruction errorⱼ + Learning error + Data consistency error.
Importantly, these guarantees scale favorably with the numbers of input and output tasks: the approximation rate is governed by the most challenging individual task rather than deteriorating with the overall problem dimension.
4. Kernel formulations for multiple operator learning. The framework is instantiated for multiple operator learning through two formulations: the operator-valued formulation KernelMO-OV, which directly approximates G: W → G[α]: U → V α∈W, or the product-space formulation KernelMO-PS, which instead approximates the equivalent map G: W × U → V.
5. Comprehensive empirical evaluation. The methods are validated on a diverse collection of multiple operator learning benchmarks arising from parametric PDEs,
demonstrating competitive predictive accuracy, substantial reductions in training and inference time compared with state-of-the-art neural operator architectures.
The paper builds on two mathematical foundations:
Minimum-norm recovery from linear measurements (Theorem 2.1): Given a bounded linear measurement operator L: X → Z, the minimum-norm recovery problem x(S) = argmin ∥x∥X Lx = S has the explicit solution x(S) = L*(LL*)−1S when L is surjective. The regularized version xγ(S) = argmin ∥x∥2X + γ−1∥Lx − S∥2Z has solution xγ(S) = L*(LL* + γIZ)−1S.
Operator-valued reproducing kernel Hilbert spaces (RKHS): The paper uses function-valued RKHSs, where K: X × X → L(Y) is an operator-valued kernel that is Hermitian and non-negative. Theorem 2.4 establishes the bijection between non-negative kernels and RKHSs.
The learning framework works as follows: inputs and outputs are encoded into latent spaces via encoding maps EX: X → ZX and EY: Y → ZY, a surrogate is learned between these representations, and predictions are decoded back to the original output space via decoding maps DX: ZX → X and DY: ZY → Y. The full approximation is defined as G = DY ◦ Ĝ ◦ EX.
The paper shows that the choice of learning method depends on the structure of the measurement spaces ZX and ZY, and involves a trade-off between expressivity, numerical complexity, and theoretical tractability.
Kernel methods are chosen because they provide minimum-norm interpolation, regularized regression, deterministic error bounds, and kernel-based uncertainty quantification can all be written in closed form.
Theorem 3.7 provides the fundamental error decomposition. For any surrogate Ĝ: ZX → ZY, the error satisfies:
∥G(x) − G(x)∥Y ≤ ∥G(x) − DY EY G(DX EX x)∥Y + ∥DY∥op ∥(Genc − Ĝenc,λ)(EX x)∥ZY + ∥DY∥op ∥RU,λη(EX x)∥ZY
The first term is the encoder-decoder reconstruction error, the second is the learning error, and the third is the data-consistency error measuring the mismatch between the measured data and the ideal encoded target.
Under Sobolev regularity assumptions (Assumptions S.1, O.1, O.2, L.1-L.3), Theorem 3.15 provides explicit convergence rates. For each input component j and output component k, the reconstruction errors decay as:
-
Input reconstruction: hX,j(sX,j − tX,j − nX,j(1/pX,j − 1/qX,j)+)
-
Output reconstruction: hY,k(sY,k − tY,k − nY,k(1/pY,k − 1/qY,k)+)
where h denotes the fill distance of the sampling points. The learning error decays as htr(τ − dX/2) + √λ, where τ is the Sobolev regularity of the RKHS and dX is the dimension of the encoded input space.
The paper introduces two formulations for multiple operator learning:
-
Learns the map G: W → G[α]: U → V α∈W
-
Predicts entire operators
-
Generalizes across the parameter space W
-
Kernel matrix size: nα × nα (number of operators)
-
Particularly attractive when one wishes to repeatedly query the same operator for many different input functions
-
Learns the map G′: W × U → V defined by G′(α, u) = Gα
-
Predicts single operator evaluations
-
Generalizes across both parameter space W and input space U
-
Kernel matrix size: (nα·nu) × (nα·nu) (parameter-input pairs)
-
Natural choice when only isolated evaluations are required or when different parameters are associated with different input functions
The paper notes that the operator-valued formulation requires one training sample per operator, whereas the product-space formulation requires one sample for every parameter–input pair.
The methods are evaluated on five parametric PDE benchmarks:
-
Conservation laws: ut + (α1u + α2u2 + α3u3)x = α4uxx
-
Diffusion-reaction-advection: ut = α1uxx + α2ux + α3u α4(1 − u α5)
-
Nonlinear Klein-Gordon: utt = α12uxx − α22α14u − α3u3
-
Parametric diffusion-reaction: ut = (α(x)ux)x + u(1 − u)
-
Parametric wave equation: utt = α2(t)uxx
The learning task is to map (α, u0) → u, where α is a parametric function/parameters and u0 is an initial condition, to the complete solution trajectory u: (0, 2] × [0, 2] → R.
Operator-Valued Learning:
-
The proposed kernel-based methods achieve the lowest prediction errors on both the in-distribution and out-of-distribution datasets
across most benchmarks -
On the conservation law: error reduced
from 1.23% for MIONet to 0.01%
(two orders of magnitude improvement) -
On the nonlinear Klein-Gordon equation: error decreased
from 4.63% for MNO to 0.21%
-
KernelMO-OV methods
generally outperforming the KernelMO-PS variants
Product-Space Learning:
-
KernelMO-PS
consistently outperforms the corresponding single-kernel baseline KernelO
(e.g., 2.13% versus 59.28% for the parametric wave equation) -
Remains competitive with, or superior to, the neural operator baselines across most benchmarks
Out-of-Distribution Performance:
-
The proposed methods also frequently attain the lowest out-of-distribution errors, although the relative improvement depends more strongly on the governing PDE and the type of distribution shift
-
PCA-based variants
exhibit improved robustness under distribution shift, suggesting that the reduced latent representation can provide a useful regularization effect
Kernel Choice:
-
Matérn kernels generally provide the strongest and most stable performance
-
RBF kernels can nevertheless attain extremely small in-distribution errors
The kernel methods demonstrate substantial computational advantages:
The paper concludes that kernel methods constitute a viable alternative to deep neural surrogates for scientific machine learning
with three key advantages: mathematical guarantees, computational efficiency, and leverage limited data sampling.
Future directions include:
-
Developing an approximation theory for learning maps whose inputs and outputs are themselves spaces of operators
-
Extending the framework to settings in which the observation spaces are themselves infinite-dimensional
-
Replacing the prescribed encoders and decoders by adaptive learned representations
-
Investigating the transferability of learned surrogates across different observation operators
-
Leveraging the close connection between kernel methods and Gaussian processes for principled uncertainty quantification
Improvements for AI systems
Based on this paper, I can make the following specific improvements to AI systems:
Improvement: Implement the encoder-decoder error decomposition (Theorem 3.7) as a built-in validation mechanism in AI systems that learn mappings between function spaces (e.g., PDE solvers, climate models). The system can automatically report separate error contributions from representation, learning, and data consistency, allowing users to identify which component needs improvement.
Capability: An AI system that not only predicts solutions to parametric PDEs but also tells you why its prediction might be wrong—whether due to poor input encoding, insufficient training data, or measurement noise—enabling targeted fixes.
Improvement: Replace fixed neural network architectures with the KernelMO framework that automatically selects between operator-valued (OV) and product-space (PS) formulations based on task structure. Use the theoretical guarantee that approximation rates are governed by the most challenging individual task, not the total problem dimension, to dynamically allocate computational resources.
Improvement: Leverage the closed-form solutions and Matérn kernel stability to build a system that quantifies prediction uncertainty via the kernel's RKHS norm. Use the observed out-of-distribution robustness of PCA-based variants to implement automatic distribution-shift detection.
Improvement: Implement the two-order-of-magnitude training speedup (0.5 seconds vs. 158–250 seconds) and 40–80× prediction speedup (0.004 ms/sample) into a system that can be retrained on-the-fly with new experimental data.
Improvement: Use the theoretical result that RBF kernels achieve extremely small in-distribution errors while Matérn kernels provide better out-of-distribution stability to build an adaptive kernel selection mechanism. The system can monitor validation performance and switch kernels or blend them based on the current data regime.
Improvement: Exploit the minimum-norm recovery theory (Theorem 2.1) to build a system that can learn from very few operator samples (the OV formulation requires only one training sample per operator). Use the theoretical convergence rates to determine the minimum number of sensors needed for a target accuracy.
Improvement: Build a meta-learning system that compares kernel-based and neural operator performance on a given task (using the paper's benchmarks as reference) and selects the best approach. Use the theoretical guarantees of kernels for small-data regimes and neural networks for large-data, high-complexity regimes.
Abstract
Learning mappings between infinite-dimensional objects is a central challenge in scientific machine learning. We introduce a general kernel-based encoder-decoder framework for operator learning that separates observation, representation, learning, and reconstruction. We develop this framework for multi-input, multi-output operator learning, where operators map between products of potentially distinct function spaces. Our approximation theory shows that, although the number of inputs and outputs can increase, the convergence rate is governed by the most challenging constituent approximation problem rather than the overall problem dimension. The framework leads to practical kernel methods with closed-form training and inference, combining mathematical tractability with computational efficiency. We further specialize the approach to multiple operator learning by introducing KernelMO, a family of kernel methods with complementary operator-valued and product-space formulations. Across five families of parametric partial differential equations, the proposed methods achieve competitive or state-of-the-art predictive accuracy while reducing training and inference costs relative to neural operator architectures and deep learning based models, offering an efficient and lightweight alternative.
Sources
- VICON: Vision In-Context Operator Networks for Multi-Physics Fluid Dynamics Prediction
- Fourier Neural Operator for Parametric Partial Differential Equations
- Cauchy Random Features for Operator Learning in Sobolev Space
- PROSE-FD: A Multimodal PDE Foundation Model for Learning Multiple Operators for Forecasting Fluid Dynamics
- BCAT: A Block Causal Transformer for PDE Foundation Models for Fluid Dynamics
- Multiple Physics Pretraining for Physical Surrogate Models
- A Multimodal PDE Foundation Model for Prediction and Scientific Text Descriptions
- LeMON: Learning to Learn Multi-Operator Networks
- A Deep Learning Framework for Multi-Operator Learning: Architectures and Approximation Theory
- Fine-Tune Language Models as Multi-Modal Differential Equation Solvers
- PDEformer-2: A Versatile Foundation Model for Two-Dimensional Partial Differential Equations
- Probabilistic operator learning: generative modeling and uncertainty quantification for foundation models of differential equations
- PI-MFM: Physics-informed multimodal foundation model for solving partial differential equations
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks