The Error of Deep Operator Networks Is the Sum of Its Parts: Branch-Trunk and Mode Error Decompositions
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "The Error of Deep Operator Networks Is the Sum of Its Parts: Branch-Trunk and Mode Error Decompositions".
Jane: The paper was written by Alexander Heinlein and Johannes Taraz from Delft Institute of Applied Mathematics and Delft University of Technology.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Summary: Tom: So, we established that "The Error of Deep Operator Networks Is the Sum of Its Parts..." gives us a way to diagnose failure by decomposition. Now, let's talk about what the paper actually summarizes regarding these components.
Jane: The summary really drills down into *how* they define those different error types—the branch, the trunk, and the mode error. It's moving from abstract concept to concrete definition for us listeners.
Lu: They aren't just throwing terms around; they are providing rigorous mathematical frameworks to quantify exactly what 'mode error' means in practice, which is often the trickiest part of these networks.
Meng: When they talk about the specifics of these decompositions, I keep thinking about computational cost. Does isolating the mode error component add significant overhead that makes this approach impractical for real-time use?
Lalam: Thinking about it through the lens of usability, if we can break down failure modes, we can build diagnostic tools that don't just flag an issue but guide the user toward understanding *which* part of their input or process caused the deviation.
Tom: Right, so it’s not just saying "it failed"; it's pointing to the specific mechanism—the branch, the trunk, or the mode—that was responsible for pulling it down.
Jane: It helps us understand that these errors aren't just random noise; they are structured deviations rooted in specific parts of the network architecture or its mathematical assumptions.
Lu: The beauty is how it suggests a natural basis for decomposition within the operator space itself, treating the error not as residue but as a sum of known directional components.
Meng: From an implementation standpoint, if they can isolate these errors, it means we could potentially design modular retraining loops—we only need to retrain or constrain the specific component that's failing for a given task type.
Lalam: That modularity has massive implications for lifelong learning in AI; instead of having to retrain the whole massive system when it encounters a novel failure mode, you just patch up the faulty component.
Tom: So, if I'm hearing this correctly, they've moved us from general failure analysis to specific fault attribution based on these three error components.
Jane: It gives us diagnostic precision that was previously out of reach when dealing with such deep and
Paper discussion segment 2: Tom: So, if we can treat the error of these Deep Operator Networks as a sum of distinct parts—branch, trunk, and mode—it fundamentally changes how we think about model failure.
Jane: Exactly; instead of just having one big error number that tells us nothing useful, this decomposition breaks down *where* the model is failing. It separates the error sources into manageable components.
Lu: What's incredible about this separation is that it suggests a modular path toward achieving true physical consistency in AI models. By isolating the mode error, for instance, we can directly target violations of known physical laws, which has huge implications for scientific discovery using AI.
Meng: But Lu, if you’re talking about separating error sources based on theoretical modes and branches, how much does this increase the computational overhead during training? Are we just trading one complex optimization problem for three even more complex ones?
Tom: That's a really fair question, Meng; it forces us to think about whether this structural decomposition is computationally feasible on current hardware. It's not enough to just prove the math works, we need to build something that runs fast enough for real-time use.
Jane: Think of it like debugging code; instead of getting a general crash message, you get a report saying, "The memory allocation failed here," and another one saying, "You called this function too many times." It gives pinpoint accuracy on the failure mechanism.
Lu: And that pin-point accuracy is critical because it allows us to build corrective constraints directly into the loss function for each component. We aren't just optimizing for prediction; we're optimizing for *physical plausibility* across multiple, distinct axes.
Meng: So, practically speaking, this means that if I’m trying to model fluid dynamics and my current model fails because the mass isn't conserved—that’s one error component—but it also fails because the geometry is wrong—that's a separate component we can correct without messing up the conservation law correction.
Tom: Precisely, Meng; we get surgical control over what our AI learns and what it ignores. We are moving from black-box prediction to white-box understanding of model limitations.
Lu: This opens the door for genuinely multiscale modeling, where different parts of the system—like material properties versus large-scale flow—are handled by specialized network architectures simultaneously.
Meng: Does this require us to predefine every single mode or branch we want the AI to consider? Or is there a way for the model itself to discover those critical error modes from raw data?
Lalam: Looking at the broader impact, this shift fundamentally elevates AI from being a powerful predictor of patterns to becoming an accountable reasoning partner. By quantifying and separating failure modes, we are building trust into complex systems, making AI safer and more reliable across all cultures.
Tom: It truly feels like we're moving past just predicting what *is* to understanding what *must be*. This gives us a whole new framework for reliability. Next time, I want us to dig into how this concept of mode decomposition could revolutionize AI's interaction with real-world physical systems...
Paper discussion segment 3: Tom: So, we've established that error decomposition gives us diagnostic precision; now let's talk about what specific improvements the authors suggest to fix those errors and make our AI better at solving real-world PDEs.
Jane: The paper points out that the standard DeepONet architecture has a few practical issues—specifically, that its learned basis functions are often inferior to classic mathematical bases like SVD or trigonometric functions.
Lu: And they show, through the use of a modified DeepONet with a fixed SVD basis, exactly how much we are losing by replacing those learned coefficients with the true optimal coefficients from the Singular Value Decomposition.
Meng: That’s interesting, because it implies that if we can replace our neural network’s internal representation—the trunk functions—with something that is mathematically optimal right on the surface, we can dramatically reduce error. That would be a massive gain in accuracy for complex simulations.
Tom: But Meng, how do you scale that fixed SVD basis approach across a large input domain? It sounds very restrictive to replace the adaptive learning of the trunk network with something pre-calculated.
Jane: That's a valid concern, Tom; they acknowledge that while it' works well for small examples and is great for analysis, scaling up requires careful handling of interpolation and how we manage data points across the entire space.
Lu: The paper suggests that by using the SVD basis, we are essentially removing a large portion of the approximation error—the trunk error—allowing us to focus entirely on fixing the branch coefficients. This allows for much more targeted training and optimization in that specific part of the real-world problem.
Meng: Targeting is key; if we can eliminate the structural failure modes inherent in the learned basis, we can apply our resources to address just one error without risking collateral damage to other parts of solving the equation.
Jane: And Lu's point about focusing on mode loss is important; by looking at individual coefficients—each mode—we can see exactly which specific physical frequencies are poorly approximated.
Tom: It’s a systematic way of saying that when we can’t solve an entire problem, we start by fixing the smallest parts until the whole thing works.
Lu: We also see how using techniques like loss re-weighting or employing adaptive optimizers like Adam allows us to fix this systemic bias—the tendency of AI to only learn low-frequency, easily recognizable patterns.
Meng: So, instead of just guessing that a model is "bad," we can apply a specific optimization strategy tailored to overcome that inherent mode neglect, which is very practical.
Lalam: From a cultural perspective, this means our AI can become far more reliable in the physical sciences—it won't just guess the outcome; it will respect the underlying physics and correct its own deficiencies.
Tom: It sounds like we have several powerful tools now: structural fixes, focused optimization, and a clear understanding of where we are going to improve. What happens when we put all these ideas together in one single large-scale simulation?
Conclusion: Tom: We've seen how this paper gives us such detailed diagnostic tools for DeepONets, showing that their total error is really just a collection of branch, trunk, and mode errors.
Jane: And understanding these distinct failures is incredibly important because it tells us exactly where to look when our models start failing in the real world.
Lu: The ability to see the spectral bias and how it relates to the singular values of the training data opens up entirely new ways to think about what makes a neural network truly generalize across different physical parameters.
Meng: From a deployment standpoint, knowing that we can address these specific errors means we can design maintenance and retraining strategies that are surgical, rather than just throwing massive computational resources at a whole system.
Tom: That’s the goal, Meng—moving from black-box failure to white-box understanding of the performance limitations of DeepONets.
Jane: It feels like we've given ourselves a much better handle on this technology, making it feel less mysterious and more predictable in its behavior.
Lu: I think the whole scientific community will benefit from seeing that the errors are not just random noise but systematic artifacts, guiding the path to more robust operator learning.
Meng: It also helps us decide where to spend our engineering effort—we can' focus on fixing the branch network if that's where most of our loss is, or prioritize a wider network if we see coupling issues.
Lalam: This framework allows AI to become an inherently more accountable tool, ensuring its output reflects not just what it saw during training but also the underlying mathematical structure of our world.
Tom: It’s a powerful combination of theory and engineering that has made this research a real breakthrough.
Jane: I hope that when we look at future work, this clarity on the error decomposition helps accelerate how these models are actually deployed in critical fields.
Lu: Indeed, by focusing on the mode losses and spectral bias, we' are pushing the boundaries of what AI can handle in complex physical systems.
Meng: It's about making sure that our engineering efforts are aligned with this specific understanding of the fault modes.
Lalam: We must carry this knowledge forward so that when we use AI, we do so with a deep understanding of its strengths and the full title, "The Error of Deep Operator Networks Is the Sum of Its Parts: Branch-Trunk and Mode Error Decompositions."
Tom: It's certainly a major step in solving complex problems. We're going to see if this approach can handle even more challenging dynamics next time.
Alexander Heinlein, Johannes Taraz
Delft Institute of Applied Mathematics · Delft University of Technology
cs.LG, cs.NA, math.NA
Submitted: 2026-08-24
Updated: 2026-08-25
Comments: 36 pages, 15 figures
Code: https://github.com/jotaraz/ModeDecomposition-DeepONets
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 91/100
The gist: This paper provides a systematic analysis of the performance limitations and error distributions within Deep Operator Networks (DeepONets), a standard architecture for approximating solution
Key concepts
- Deep Operator Networks (DeepONets)
- These are a type of neural network architecture used to solve complex mathematical problems, such as partial differential equations. The paper focuses on diagnosing and improving the inherent error within these networks to make them more reliable.
- Error Decomposition
- This is a method that breaks down the total error of a DeepONet into three specific, manageable parts: branch error, trunk error, and mode error. This allows researchers to pinpoint exactly where a model is failing instead of just having one general failure metric.
- Mode Error
- A specific component of the network error that relates to violations of known physical laws or poorly approximated physical frequencies. Isolating this allows for targeted training to ensure the AI respects underlying physics and maintains accuracy.
- Singular Value Decomposition (SVD) Basis
- This is a mathematically optimal basis used to replace the learned internal representation of a neural network. By using SVD, researchers can reduce approximation error and focus resources on correcting specific structural failure modes.
Terminology
Summary
This paper provides a systematic analysis of the performance limitations and error distributions within Deep Operator Networks (DeepONets), a standard architecture for approximating solution operators of partial differential equations (PDEs). By decomposing total approximation error into distinct components, the authors aim to understand why these networks often exhibit limited accuracy and generalization in practice,
providing insights that are crucial for advancing operator learning in scientific computing.
Error Decomposition Framework
The researchers establish a formal mathematical framework to partition the total approximation error, denoted as ε, into two primary components: trunk error (εT) and branch error (εB). The trunk error is defined as the projection error,
which quantifies how well the target data can be reconstructed using the span of the learned basis functions. Conversely, the branch error measures how accurately the network predicts the coefficients for those basis functions.
The paper demonstrates that:
**) The approximation error is dominated by the branch network
when the internal dimension N is sufficiently large. **
**) The learned trunk basis can often be replaced by classical basis functions
(such as Legendre or Chebyshev polynomials) without a significant impact on performance. **
**) A small projection error is a necessary, but not sufficient, condition for low total error. **
Mode-Specific Analysis
To investigate the branch error more deeply, the authors construct a modified DeepONet where the trunk network is replaced by fixed left singular vectors derived from the training data's SVD. This allows for a mode loss decomposition,
where the branch error is split into individual errors for each mode. This modification reveals that:
**) The overall branch error is dominated by modes with intermediate singular values
rather than the smallest ones. **
**) A spectral bias
exists in the branch network, where coefficients of dominant, low-frequency modes are learned more effectively than high-frequency ones. **
**) The standard architecture's use of a shared branch network for all mode coefficients actually improves generalization of small modes
compared to a stacked architecture where coefficients are computed separately. **
Optimization and Coupling Effects
The study examines how different optimization schemes and architectural choices influence the distribution of these mode losses. The authors find that plain Gradient Descent (GD) suffers from gradient domination from large singular values,
causing it to systematically neglect smaller modes. While the Adam optimizer mitigates this through an implicit re-weighting scheme,
aggressive manual re-weighting can lead to overfitting.
Furthermore, the paper identifies two types of coupling that affect training:
**) **
Architectural Coupling:
The unstacked (shared) architecture acts as an implicit regularizer
through parameter sharing, which promotes generalization across modes, whereas stacked architectures tend to overfit individual modes.
**) **
Update-Based Coupling:
In the parameter space, there exists a detrimental coupling between modes
during gradient descent updates. This means that updating parameters to improve one mode can negatively impact others. The authors note that wider DeepONets exhibit significantly smaller negative coupling strengths,
suggesting that increasing network width is a key mechanism for reducing these detrimental interactions.
Conclusion and Implications
The research concludes that the multi-scale nature of the mode decomposition
makes standard loss functions poorly calibrated for DeepONets. By understanding that error is driven by intermediate modes and parameter-space coupling, the authors suggest that practitioners can improve performance by utilizing adaptive optimizers, increasing network width, or employing multi-task learning techniques to address gradient conflicts.
:---::---::---::---::---::---::---::---::---:
Improvements for AI systems
Based on the findings in this paper, I propose moving away from black-box
DeepONet training toward a specialized architecture and optimization strategy designed to solve the identified issues of gradient starvation, spectral bias, and detrimental mode coupling.
Here are the specific improvements to implement:
-
Implement an
SVD-Initialized/Fixed Trunk
Architecture with Adaptive Loss Re-weighting -
Transition from standard GD/Adam to a
Mode-Aware Adaptive Optimizer
-
Deploy
Architecturally Decoupled (Stacked) Branch Networks
for high-fidelity requirements
Sources
- Deep Operator Neural Network Model Predictive Control
- Learning to Control PDEs with Differentiable Predictive Control and Time-Integrated Neural Operators
- DeepOKAN: Deep Operator Network Based on Kolmogorov Arnold Networks for Mechanics Problems
- Multifidelity deep neural operators for efficient learning of partial differential equations with application to fast inverse design of nanoscale heat transport
- Learning Singularity-Encoded Green's Functions with Application to Iterative Methods
- Neural Green's Operators for Parametric Partial Differential Equations
- A Resolution Independent Neural Operator
- Deep Operator Network Approximation Rates for Lipschitz Operators
- HyPINO: Multi-Physics Neural Operators via HyperPINNs and the Method of Manufactured Solutions
- TI-DeepONet: Learnable Time Integration for Stable Long-Term Extrapolation
- Long-time integration of parametric evolution equations with physics-informed DeepONets
- Operator learning with PCA-Net: upper and lower complexity bounds
- Nonlinear model reduction for operator learning
- What do physics-informed DeepONets learn? Understanding and improving training for scientific computing applications
- Mitigating Spectral Bias in Neural Operators via High-Frequency Scaling for Physical Systems
- Reducing Frequency Bias of Fourier Neural Operators in 3D Seismic Wavefield Simulations Through Multi-Stage Training
- Multi-scale DeepOnet (Mscale-DeepOnet) for Mitigating Spectral Bias in Learning High Frequency Operators of Oscillatory Functions
- Blending Neural Operators and Relaxation Methods in PDE Numerical Solvers
- FEDONet : Fourier-Embedded DeepONet for Spectrally Accurate Operator Learning
- An Overview of Multi-Task Learning in Deep Neural Networks
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks