Pareto-Optimal Offline Reinforcement Learning via Smooth Tchebycheff Scalarization
cs.LG, cs.AI, q-bio.BM, q-bio.QM
Submitted: 2026-04-14
Updated: 2026-09-11
Project page: https://multi-objective.github.io/moocore/r
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
The gist: Large language models can be aligned with human preferences through offline reinforcement learning (RL) on small labeled datasets.
Terminology
Abstract
Large language models can be aligned with human preferences through offline reinforcement learning (RL) on small labeled datasets. While single-objective alignment is well-studied, many real-world applications demand the simultaneous optimization of multiple conflicting rewards, e.g. activity and specificity for proteins, or helpfulness and harmlessness for chatbots. Prior work has largely relied on linear reward scalarization, which provably fails to recover non-convex regions of the Pareto front. In this paper, instead of scalarizing the rewards directly, we frame multi-objective RL itself as an optimization problem to be scalarized via smooth Tchebycheff scalarization, a recent technique that overcomes the drawbacks of linear scalarization. We use this formulation to derive Smooth Tchebycheff Optimization of Multi-Objective Preferences (STOMP), a novel offline RL algorithm that extends direct preference optimization to the multi-objective setting in a principled way by standardizing the individual rewards based on their observed distributions. We empirically validate STOMP using multiple base models on 3 protein engineering and 3 chatbot alignment datasets. Compared to state-of-the-art baselines, STOMP achieves or ties for the highest hypervolumes on 16/18 protein tasks and 5/6 natural language tasks. We thus demonstrate that STOMP is a powerful, robust multi-objective alignment algorithm that can meaningfully improve post-training in multiple domains.
Sources
- Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
- Training Deep Nets with Sublinear Memory Cost
- Adam: A Method for Stochastic Optimization
- How to make the most of your masked language model for protein engineering
- Proximal Policy Optimization Algorithms
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks