Inducing Permutation Invariant Priors in Bayesian Optimization for Carbon Capture and Storage Applications
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Inducing Permutation Invariant Priors in Bayesian Optimization for Carbon Capture and Storage Applications".
Jane: The paper was written by Sofianos Panagiotis Fotias and Vassilis Gaganisa from School of Mining and Metallurgical Engineering, National Technical University of Athens.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Summary: Tom: So, in "Inducing Permutation Invariant Priors in Bayesian Optimization for Carbon Capture and Storage Applications," the authors explain why standard optimization methods struggle with these complex well layouts.
Jane: They found that when you have groups of wells—say, injectors or producers—the simulator doesn't care about the order they are listed, but typical Gaussian Process surrogates do.
Tom: It’s like trying to find the best way to organize a stack of books where putting them in alphabetical order or by size is treated as two entirely different options for the AI.
Lu: The paper argues that this distinction is physically meaningless, and it leads to the surrogate wasting time exploring redundant solutions.
Meng: It's about recognizing that we need a method to aggregate or summarize these inputs so the optimization process doesn't get distracted by trivial differences in sequence.
Lalam: We are moving toward a more natural way of interacting with AI, where the machine understands the structure of the problem rather than just its arbitrary input format.
Tom: To solve this, they introduced GP-Perm, which is a Gaussian Process kernel designed specifically to respect that symmetry.
Jane: It does this by comparing sets using something called a stable divergence between their empirical representations, rather than just looking at simple distances.
Lu: I’m fascinated by the technical detail of the Sinkhorn divergence; it sounds like they are building a way to measure similarity while respecting that the elements within a set don't have an intrinsic ordering.
Meng: That is key for me; we need a metric that captures spatial relationships between sets without being sensitive to how many times we shuffle the individual components.
Lalam: It feels like this allows us to build AI models that understand the physical reality of the CCS operation, making our digital twins much more accurate.
Tom: And this approach is applied not just to one set, but across multiple sets—injectors and producers—and how they interact with each other.
Jane: We’re getting a summary of the problem and a very elegant solution for how to manage those complex, unordered inputs.
Improvements: Tom: Now we look at the improvements in "Inducing Permutation Invariant Priors in Bayesian Optimization for Carbon Capture and Storage Applications." The authors really put their methods through their paces.
Jane: They compared GP-Perm against several baselines, like non-invariant GPs, and also against other set-kernel approaches such as Double-Sum (DS) and Deep Embedding (DE).
Tom: And they also included a learned baseline using Deep Kernel Learning with the Deep Sets architecture, which is called DKL-DS.
Lu: The comparison reveals that while all these models have their strengths, GP-Perm seems to provide a very stable and direct way to encode the required physical symmetry.
Meng: For me, this stability is impressive because when we are running BO with only a few samples due to simulation cost, you can't afford any model instability or excessive variance.
Lalam: It’s encouraging that the performance gains aren’t just theoretical; they are showing up in real-world scenarios where climate action matters.
Tom: The results from the synthetic benchmarks show GP-Perm consistently achieving better sample efficiency than the non-invariant models, which is a huge win for minimizing costly simulation runs.
Jane: They aren't just looking at the final answer; they are measuring how fast the model gets to that answer through AUC of the best-so-far curve, which is a much more honest measure of performance.
Lu: The fact that GP-Perm maintains a moderate and predictable uncertainty profile in these synthetic tests speaks to its robustness against small, non-i.i.d datasets typical in BO loops.
Meng: That’s vital for my work; if the model is consistently performing well, it means we can trust the acquisition function to guide us toward better designs without unpredictable failures or poor recommendations.
Lalam: It suggests that AI can be deployed not just as a powerful tool, but as a truly reliable partner in optimizing our environmental strategies.
Tom: To transition into the next section, we need to see how this works when we apply it to an actual geological challenge: the Johansen formation.
Conclusion: Tom: So, we are wrapping up our discussion of "Inducing Permutation Invariant Priors in Bayesian Optimization for Carbon Capture and Storage Applications."
Jane: It’s clear that the authors, Panagiotis Fotias and Vassilis Gaganis, have provided a powerful solution to a long-standing problem with structured inputs.
Tom: The results from the Johansen case study are particularly compelling because they show GP-Perm maintaining its advantage even when we severely restrict the number of evaluations.
Lu: I think this proves that recognizing symmetry is not just an academic exercise; it’s a core element of optimizing complex, real-world physical systems.
Meng: For me, this validates the need for these explicit, geometry-aware kernels in industrial AI applications where resource constraints are severe.
Lalam: The ability to achieve better outcomes with fewer evaluations is exactly what we need if we want to scale up global carbon storage efforts.
Tom: It seems GP-Perm has shown itself to be the most robust and effective approach across various benchmarks and scenarios.
Jane: We’ve seen it outperform both the standard non-invariant models and even set-kernel baselines, which is a very strong result indeed.
Lu: It demonstrates that by respecting the underlying physics of group control, we can significantly enhance the performance of our AI systems.
Meng: The engineering takeaway is that this provides a reliable framework for making decisions under uncertainty in large geological formations.
Lalam: It gives us hope that we can use advanced AI to make some massive improvements in environmental stewardship.
Final Conclusion: Tom: Before we head off, let’s take one last look at the biggest implications of this paper, "Inducing Permutation Invariant Priors in Bayesian Optimization for Carbon Capture and Storage Applications."
Jane: It really shows that when we are dealing with systems where order is irrelevant—like well placements—we can boost AI efficiency dramatically.
Lu: I think the creativity here lies in building a stable divergence method to be able to capture those complex, non-local geometric relationships.
Meng: The practical impact is huge; it makes the "expensive" part of BO much more manageable, which means we can run these models on larger datasets and for longer periods.
Lalam: I hope this work shows that AI can reliably support critical infrastructure like CCS without forcing us to choose between computational power and environmental necessity.
Tom: It’s a powerful combination of elegant mathematical design meeting a very real-world climate challenge.
Jane: We’re looking at the future of how AI interacts with our physical world, making decisions that are both smart and contextually appropriate.
Lu: I just love the potential; if we' can apply these invariant kernels to other large-scale combinatorial problems, the possibilities are endless.
Meng: And from an engineering standpoint, it ensures that we’re designing robust systems that don't break down when the AI finds redundant paths.
Lalam: This work helps us build a culture where technological progress is aligned with global sustainability goals.
Tom: It’s a truly exciting intersection of cutting-edge math and real-world application, "Inducing Permutation Invariant Priors in Bayesian Optimization for Carbon Capture and Storage Applications."
Jane: We've seen how this method works, so we hope you have an amazing week!
School of Mining and Metallurgical Engineering, National Technical University of Athens
cs.LG
Submitted: 2026-05-04
Updated: 2026-09-04
Code: https://github.com/flammmes/Permutation-Invariant-Kernels
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 83/100
The gist: The following is a detailed, comprehensive summary of the scientific paper, "Inducing Permutation Invariant Priors in Bayesian Optimization for Carbon Capture and Storage Applications," utilizing
Key concepts
- Bayesian Optimization
- A method used for finding the best design or configuration within a system. It uses statistical models, like Gaussian Processes, to guide the search process. The paper focuses on improving this technique when dealing with complex structures where certain elements are interchangeable.
- Permutation Invariant Priors
- A method of designing AI models that understand symmetry. It means the model does not change its output based on the order in which inputs are listed. This is useful for physical systems, such as well groups, where the sequence of components is physically meaningless.
- GP-Perm
- A specific Gaussian Process kernel developed to respect permutation invariance. It measures similarity between sets using a stable divergence, allowing the AI to understand the structure of a problem rather than just its arbitrary input format.
Terminology
Summary
The following is a detailed, comprehensive summary of the scientific paper, Inducing Permutation Invariant Priors in Bayesian Optimization for Carbon Capture and Storage Applications,
utilizing only content extracted from the text.
Motivation and Problem Statement
Bayesian Optimization (BO) is a sample-efficient methodology used to optimize expensive black-box objective functions. A critical limitation arises when standard surrogate models, such as Gaussian Processes (GPs), are applied to inputs that possess inherent symmetries. The paper identifies this issue in the context of Carbon Capture and Storage (CCS) well placement, where the objective function is invariant under permutations within groups of injector wells (S m) and producer wells (S n). A non-invariant GP treats permuted representations of the same physical design as distinct inputs,
which can lead to unnecessary exploration and reduced sample efficiency
(Figure 1). The core challenge is that the simulator output depends on the well configurations only through the sets I x and P x, not on their internal ordering.
Proposed Methodologies
The authors introduce two primary approaches to encode this permutation symmetry:
-
GP-Perm (Kernel-Based Approach): A novel Gaussian Process kernel designed to enforce within-group exchangeability directly at the kernel level. GP-Perm is built around a composite distance D(x, x') that compares sets through a stable divergence between their induced empirical representations. This approach remains compatible with standard vector-valued kernels for auxiliary continuous controls.
-
DKL-DS (Learned Approach): A Deep Kernel Learning model using the Deep Sets architecture to learn a permutation-invariant embedding f theta(x).
** Detailed Methodology: GP-Perm Construction**
GP-Perm defines a composite squared distance D(x, x') that respects the structured input x = v x, I x, P x. The this distance is constructed as an ANOVA-style sum of contributions from four components (Figure 2):
D(x, x') = 1 over 2 [v - v' squared + sum I S epsilon I, I' + S epsilon P, P' + S epsilon R, R']
where R is the interaction set defined by R = p b - i a: i a in I, p b in P.
The set terms are computed using the Sinkhorn divergence S epsilon, and the debiasing step is applied via:
S epsilon(S, T) = W epsilon(S, T) - W epsilon(S, S) - W epsilon(T, T)
The resulting Matérn kernel is then constructed as:
k GP-Perm(x, x') = sigma squared M nu D(x, x')
Since S epsilon depends only on the induced empirical measures of the sets, D(x, x') and k GP-Perm(x, x') are invariant under S n times S m.
Detailed Methodology: DKL-DS Construction
The DKL-DS model uses a Deep Sets encoder to ensure invariance. The input components (I x, P x, and R x) are processed separately by the Deep Sets module, which applies a shared elementwise map phi(times) and then aggregates across the set using order-invariant pooling. The final latent representation z = f theta(x) is modeled by a GP kernel:
k(x, x') = k lat f theta(x), f theta(x')
** Comparison with Baselines**
The proposed surrogates are benchmarked against four baselines:
-
Non-invariant GP (Matérn kernel applied to the flattened input vector).
-
Non-invariant DKL (MLP maps the flattened input to a latent representation).
-
Double-Sum (DS) set-kernel GP (aggregates pairwise similarities across elements).
-
Deep-Embedding (DE) set-kernel GP (applies a radial kernel to the RKHS mean embedding distance).
** Experimental Evaluation**
The evaluation is conducted across three distinct settings:
-
Synthetic Benchmarks: A suite of 6 benchmarks (e.g., Particle Physics, Max Area Coverage, Distribution Matching) are used to verify that the surrogates
exploit permutation symmetry
under controlled conditions. -
CCS-like Two-Set Synthetic Benchmark: This setup mirrors the core geometric coupling of CCS well design, where a cost function f(x) is defined by cross-group matching (Equation 11) and within-set repulsion (Equation 12).
-
Johansen Case Study: A realistic CCS task using the OPM-Flow reservoir simulator to maximize a regularized Net Present Value (NPV), incorporating penalties for recycling and rate inconsistency.
** Key Results and Findings**
-
Synthetic Benchmarks: The results show that
no single surrogate dominates every benchmark,
indicating thatthe benefit of permutation invariance is task dependent.
-
CCS-like Benchmark: In this setting, GP-Perm achieved the best mean AUC (Table 5) and the best mean final objective (Table 6). This demonstrated that in a regime where the objective depends on relative geometry between two unordered sets,
GP-Perm consistently converts fewer evaluations into larger best-so-far gains.
-
Johansen Case Study: Despite having a significantly tighter BO budget (T=10, q=2) compared to the synthetic study (T=30, q=4), GP-Perm achieved the best mean AUC (7.717) and the best mean final objective (0.387),
outperforming both non-invariant baselines and competitive set-based alternatives.
** Discussion and Conclusion**
The findings suggest that GP-Perm provides the strongest overall performance among the tested approaches.
The study concludes that when inputs contain nuisance symmetries, encoding that symmetry reduces effective complexity by collapsing equivalent designs. In BO, this translates to a smoother posterior over the quotient space of designs and consequently, more reliable acquisition decisions under small evaluation budgets.
The authors provide usage guidelines: (i) if exchangeability is guaranteed by is the problem setup, enforcing permutation invariance is beneficial; (ii) explicit invariant kernels are a strong default when reliability under small data is critical; and (iii) learned invariant embeddings may require careful regularization and training protocols.
Improvements for AI systems
As a researcher, I have analyzed this methodology. The core breakthrough is not just an application to CCS; it is a fundamental solution to a class of structural bias problems in Bayesian Optimization (BO).
The improvement lies in transitioning from treating every permutation as a distinct input (the standard approach) to identifying and collapsing equivalent configurations (the proposed method). This drastically increases sample efficiency and ensures the posterior geometry accurately reflects physical reality.
Here are the specific improvements to AI systems, categorized by mechanism:
The primary improvement is the introduction of two distinct, highly effective methods for encoding symmetry directly into a surrogate model's kernel or latent space.
Mechanism: We implement a composite Gaussian Process kernel, GP-Perm, that is explicitly permutation invariant. This is achieved by defining the distance between two inputs x and x' using three distinct components:
-
Auxiliary Vector Distance: A standard (e.g., Matérn) kernel for auxiliary continuous parameters (v).
-
Set Divergence: A stable, entropically regularized Sinkhorn Divergence (S epsilon) between the injector sets (I and I' and P and P'). This measures similarity based on transport costs rather than simple pairwise averages, capturing fine geometric relationships.
-
Interaction Term: A distance measure over the derived interaction set (R), capturing cross-group spatial relationships (e injector–producer spacing).
Improvement: The resulting kernel, GP-Perm, is mathematically guaranteed to be invariant under any permutation of elements within the sets I and P.
What it allows the system to do: It correctly models optimization problems where the objective function depends on relative geometry (e.g., spacing) between two unordered groups, while being entirely agnostic to how those groups are ordered internally.
Mechanism: We integrate a permutation-invariant feature extractor, Deep Sets, within the Deep Kernel Learning (DKL) framework. The input is processed through:
-
MLP Encoding: Mapping auxiliary vector inputs (v).
-
Set Encoding (Deep Sets): Processing the unordered sets (I, P, R) using shared element-wise maps and order-invariant pooling (e.g., mean, max, standard deviation).
-
Concatenating these embeddings and passing them through a final fusion MLP to produce the latent representation z.
-
Applying a standard GP kernel on this learned latent space k lat(z, z').
The integration of these invariant surrogates into a BO loop offers critical improvements in decision-making:
A. Enhanced Sample Efficiency:
Because GP-Perm and DKL-DS recognize that multiple permutations represent the same physical design, the surrogate model does not waste capacity modeling distinctions that are physically meaningless. This allows the model to achieve a much better fit with fewer evaluations.
B. Robust Acquisition Strategy:
The improved surrogate model provides a smoother, more accurate posterior distribution over the quotient space of designs. This means the acquisition function (e.g., qLogEI) is based on a more reliable probability landscape for selecting the next batch of candidate points, rather than being perturbed by spurious distinctions caused by permutations.
The methodology is not limited to CCS; it applies universally to any complex physical system where the inputs are structured sets or unordered components.
What the improved system can do:
-
Optimize Complex Configurations: It excels at problems like well placement, sensor network design, component grouping in manufacturing, or even scheduling tasks where the order of execution is irrelevant but the combination matters.
-
Handle Cross-Group Dependencies: By explicitly modeling the interaction set R (as seen in GP-Perm), it can find optimal solutions where the success of one group depends on its specific geometric relationship to another group, a feature standard non-invariant models often miss.
Feature Standard BO Surrogate Improved System (GP-Perm / DKL-DS)
:---:---:---
View of Input (e.g, i 1, i 2 vs i 2, i 1) Two distinct inputs; requires separate modeling. One equivalent input; collapses the space.
Sample Efficiency (Fixed Budget) Lower—wasted evaluations on redundant permutations. Higher—focus on unique, physically meaningful configurations.
Posterior Geometry (Uncertainty Map) Distorted/fragmented due to over-representation of equivalent points. Smooth and accurate representation of the true design space.
Optimal Use Case Simple vector optimization problems where inputs are strictly ordered. Complex, structured systems with inherent symmetry and critical cross-group dependencies (e.g., geological placement).
Abstract
Bayesian Optimization is an iterative method, tailored to optimizing expensive black box objective functions. Surrogate models like Gaussian Processes, which are the gold standard in Bayesian Optimization, can be inefficient for inputs with permutation symmetries, as the most common kernels employed are better suited for vector inputs rather than unordered sets of items. Motivated by this issue, we turn to permutation invariant Bayesian Optimization for well placement in Carbon Capture and Storage projects. The high fidelity black box simulator is instructed to operate wells under group control, giving rise to permutation symmetries within injector and producer groups that cannot be exploited with standard GP kernels. In this work, our main contribution is a novel Gaussian Process kernel (GP-Perm) that encodes permutation invariance by comparing sets through a stable divergence between their induced empirical representations, and can be combined with standard kernels for additional vector-valued inputs. As a learned invariant baseline, we also consider a Deep Kernel Learning model (DKL-DS) using the Deep Sets architecture to learn a permutation-invariant embedding. We evaluate the proposed methodology across 8 use cases, comprising seven synthetic benchmarks and one realistic CCS case study (Johansen formation)
Sources
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks