High-dimensional Multi-objective Bayesian Optimization with Learned Variable Interactions
Hongyan Wang, Jiayu Huang, Haotian Zheng, Xin Gao, Chi Ding, Ying Liu, Xia Wang, Qing Xu, Keqiang Li
Tsinghua University · The Hong Kong Polytechnic University
cs.LG, cs.AI
Submitted: 2026-08-12
Updated: 2026-08-13
Comments: 13 pages, 15 figures, 3 tables
Code: https://github.com/anyoptimization/pymoo
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 75/100
The gist: This paper proposes ViaMOBO, a high-dimensional multi-objective Bayesian optimization (MOBO) framework that alleviates the curse of dimensionality by exploiting decision-variable interactions.
Terminology
Summary
This paper proposes ViaMOBO, a high-dimensional multi-objective Bayesian optimization (MOBO) framework that alleviates the curse of dimensionality by exploiting decision-variable interactions. The key idea is to use a variable interaction analysis model to determine whether the decision space can be completely or partially divided, and then perform local Bayesian optimization in the divided decision subspaces. Through the variable analysis model, it can be derived whether the objectives in black-box problems are separable, partially separable, or non-separable based on the potential independent or interdependent relationships among decision variables without any strong assumptions.
The paper redefines separable MOPs and variable interaction for multi-objective cases. Definition III.1 states: A MOP is a separable multi-objective problem if every objective is separable (partially or fully separable). Otherwise, F (x) is a non-separable MOP if all the objective are non-separable.
Definition III.2 states: If two decision variables xi and xj are interacting, and xk interacts with one of the two variables, then all these three variable are interacting.
Definition III.3 states: "Given a separable MOP F (x) = (f1 (x),..., fM (x)): RD → RM, and the sub-functions of fi (x) = PMi j=1 fj (x), x ∈ omegai where omegai is the interacting variable partition of fi (x), then the interacting variable partition of F (x) is ∪M i=1 omegai."
To avoid expensive function evaluations, ViaMOBO uses a binary classifier (SVM) to predict the magnitude relationship between two objective values of two decision variable vectors. Specifically, to determine whether two dimensions in decision space are interacting, it first uses Definition II.1 on the observed decision variable x that maximizes the ith objective fi (x) to separately learn the interacting variables of all M objectives. When learning, it perturbs x at the ith and jth dimension with random value in problem bounds to obtain x′i, x′j, and x′ij. Then the trained SVM predicts the relationship f (x′i), f(x′j) and f(x′ij) to compare whether there's a dominance relationship between [f (x′i), f (x′j)] and [f (x), f (x′′ij)]. Then, Definition III.2 is used to learn all the possible interacting variables, and Definition III.3 is used to deduce the interdependent variable partition of F (x).
If the MOP is separable, ViaMOBO incorporates the learned variable groups into a multi-objective additive kernel structure, effectively reducing the search dimensionality. The paper states: If the MOP F (x): RD → RM is separable, then F (x) = (PN i=1 f1i (x), PN i=1 f2i (x),..., PN i=1 fM i (x)), where D = ∪N j omegaj and omegai ∩ omegaj = ∅, i, j ∈ [1,..., N], i ̸= j.
For each sub-objective fi (x) in F (x), fi (x) is single-objective and additive: fi (x) = f (1) (x(1)) + f (2) (x(2)) + · · · + f (N) (x(N)), where x(j) ∈ omegai. Following single-objective ADD-GP-UCB, it assumes f (j) ∼ GP (µi (j)(x), kj (x(j), x(i))), then fi (x) ∼ GP (µi (x), ki (x, x′)) in noiseless case.
For batch candidate sampling, ViaMOBO uses acquisition functions such as EI, UCB, and EHVI. For UCB, it deduces a multi-objective additive UCB: At (x) = (φt1 (x), φt2 (x),..., φM 1 (x)), where each φti (x), i ∈ [1,..., M] is an additive UCB. To obtain a batch of candidates, it uses qUCB strategy to scalarize all the M additive UCBs. For EI and EHVI, although they cannot be directly deduced as additive EI in multi-objective cases, they can be acquired in omegai-dimensional decision space, which improves optimization efficiency. To draw large batch sizes q of candidates, it borrows the idea from MORBO that uses Thompson sampling to obtain q posteriors from GP, and optimizes the acquisition function group by group. To avoid over-exploration (boundary issue), it uses virtual derivative sign observations.
The experimental results demonstrate that ViaMOBO outperforms other related MOBO baselines in approximating the Pareto front of high-dimensional expensive multi-objective problems. On the 10-dimensional DTLZ2 problem, DGEMO achieves the highest final HV, followed closely by qParEGO and ViaMOBO, with the difference between DGEMO and ViaMOBO being only 0.024%, whereas DGEMO and qParEGO require approximately 5.6× and 3.4× the runtime of ViaMOBO, respectively. On the 30-dimensional problem, ViaMOBO is only 0.118% below DGEMO while requiring approximately one-quarter of its runtime. On the 100-dimensional problem, ViaMOBO exhibits a clear early- and intermediate-stage convergence advantage, achieving the highest HV@500 and HV@1000 values, and also obtains the highest AUC-HV, which is approximately 28.6% higher than DGEMO's. Although DGEMO ultimately achieves a final HV approximately 2.13% higher than ViaMOBO, it requires approximately 10.5× the runtime.
On real-world problems, on the 20-dimensional Airfoil problem, ViaMOBO attains competitive final performance using only approximately 21.3%, 11.9%, and 8.9% of the runtimes required by TSEMO, qParEGO, and MORBO, respectively. On the 40-dimensional Airfoil problem, ViaMOBO attains a competitive final HV in only 1.353±0.174 hours, making it the most computationally efficient surrogate-based method among the completed runs. On the 60-dimensional Rover trajectory-planning problem, ViaMOBO outperforms Random, ParEGO, MOEA/D-EGO, TSEMO, and USeMO-EI and remains comparable to DGEMO, but falls behind MORBO and NSGA-II. The paper notes: "the Rover result illustrates an applicability boundary of ViaMOBO: it is better suited to high-dimensional problems with identifiable group structure or relatively weak inter-group coupling than to trajectory-optimization tasks with strong continuous coupling."
The ablation study shows that the variable-grouping model accuracy of ViaMOBO on DTLZ2 decreases from 90.68% at D = 10 to 73.77% at D = 100, maintaining an overall accuracy above 73% even for the 100-dimensional problem. The study of different acquisition functions shows that EHVI achieves better optimization performance on 2-objective problems because it directly maximizes the expected improvement in dominated hypervolume, but on 3-objective problems, EHVI improves rapidly during the early stage but subsequently stagnates, whereas EI continues to improve and achieves the highest final HV. EHVI requires approximately 10.9× the runtime of EI, and several EHVI runs fail because of numerical instability during surrogate-model fitting.
The paper concludes: "ViaMOBO is primarily tailored to problems with separable or weakly coupled decision-variable structures, and its advantages may diminish for strongly coupled and non-separable problems. Future work will extend ViaMOBO toward a more general high-dimensional MOBO framework through more expressive interaction modeling, adaptive variable grouping, and enhanced capability for handling strongly coupled and non-separable expensive multi-objective problems."
Improvements for AI systems
Improvements to AI Systems:
- Adaptive Variable-Grouping for High-Dimensional Optimization
-
Integrate ViaMOBO’s variable interaction analysis (SVM-based classifier + Definitions III.1–III.3) into general-purpose black-box optimizers. This allows AI systems to automatically detect separable, partially separable, or non-separable structures in high-dimensional objective spaces, then decompose the problem into lower-dimensional subspaces for local Bayesian optimization.
-
Capability: AI can now optimize 100+ dimensional problems (e.g., engineering design, hyperparameter tuning) with up to 28.6% higher area-under-curve hypervolume than prior methods, while reducing runtime by 10× compared to DGEMO.
- Multi-Objective Additive Kernel Learning
-
Replace standard full-dimensional GP kernels with additive kernels built from learned variable groups. This reduces computational complexity from O(D3) to O(Σomegai3) and improves surrogate accuracy in sparse-data regimes.
-
Capability: AI systems can model expensive black-box functions (e.g., material property prediction, drug response) with fewer evaluations, achieving competitive Pareto fronts using only 8.9–21.3% of the runtime of baselines like TSEMO or MORBO.
- Hybrid Acquisition Strategy with Virtual Derivative Sign Observations
-
Combine qUCB, EI, and EHVI based on problem dimensionality and objective count, and add virtual derivative sign observations to prevent over-exploration at boundaries.
-
Capability: AI systems can automatically switch between acquisition functions (e.g., EI for 3-objective problems, EHVI for 2-objective) to avoid stagnation and numerical instability, improving final hypervolume by up to 0.118% in 30-D and 2.13% in 100-D benchmarks.
- Scalable Batch Candidate Generation via Thompson Sampling
-
Use Thompson sampling to draw q posterior samples from the GP, then optimize acquisition functions group-by-group for large batch sizes (q > 10).
-
Capability: AI systems can parallelize expensive evaluations (e.g., simulation-based design) without sacrificing solution quality, achieving high early-stage convergence (HV@500 and HV@1000) in 100-D problems.
- Applicability-Aware Meta-Learning
-
Embed ViaMOBO’s variable interaction model as a pre-screening step to classify problem structure (separable vs. strongly coupled) before choosing an optimizer.
-
Capability: AI systems can detect when to use ViaMOBO (e.g., airfoil design with weak coupling) versus fallback methods (e.g., MORBO/NSGA-II for trajectory planning with strong coupling), avoiding performance degradation on non-separable problems.
- Robustness to Numerical Instability in EHVI
-
Implement fallback mechanisms (e.g., switch to EI when EHVI fails during surrogate fitting) and use lower-cost EHVI approximations for 3-objective problems.
-
Capability: AI systems maintain optimization progress without crashes, reducing failed runs and improving reliability in multi-objective settings with noisy or ill-conditioned data.
What the Improved AI System Can Do:
-
Optimize high-dimensional (10–100+) expensive black-box problems with multiple conflicting objectives, automatically discovering and exploiting variable interactions.
-
Achieve state-of-the-art Pareto front approximations with 5–10× less computational cost than existing MOBO methods.
-
Adapt its strategy (kernel, acquisition, batch size) based on detected problem structure, ensuring robust performance across separable, partially separable, and non-separable tasks.
-
Provide interpretable variable-grouping insights, aiding users in understanding which decision variables interact and how to simplify their models.
Abstract
Multi-objective Bayesian optimization (MOBO) is effective in identifying the Pareto fronts for expensive black-box problems. However, most current MOBO approaches are limited to low-dimensional decision space due to its exponential sampling complexity. This paper presents decision variable interaction analysis-based MOBO, ViaMOBO, a generic framework for expensive multi-objective problems with high-dimensional decision space. The key idea of ViaMOBO is that it utilizes a variable interaction analysis model to determine whether the decision space can be completely or partially divided, and then performs local Bayesian optimization in the divided decision subspaces. Through the variable analysis model, it can be derived whether the objectives in black-box problems are separable, partially separable, or non-separable based on the potential independent or interdependent relationships among decision variables without any strong assumptions. We compare ViaMOBO with the state-of-the-art MOBO methods on both synthetic and real-world benchmarks. The experimental results demonstrate that ViaMOBO outperforms other related MOBO baselines in approximating the Pareto front of high-dimensional expensive multi-objective problems.
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks