Objective-Behavior Alignment: Diagnostics for MORL Policy Selection
cs.LG
Submitted: 2026-06-19
Updated: 2026-09-02
Comments: 22 pages, 41 figures, Accepted to Transactions on Machine Learning Research (TMLR). OpenReview: openreview.net/forum?id=hfnMLNCCYz
Journal ref: Transactions on Machine Learning Research, 2026. ISSN 2835-8856
Code: https://github.com/ffelten/Behavior-vs-Objective-Space
License: http://creativecommons.org/licenses/by/4.0/
The gist: Real-world decision-making often requires optimizing multiple competing objectives simultaneously.
Terminology
Abstract
Real-world decision-making often requires optimizing multiple competing objectives simultaneously. In reinforcement learning (RL), this is typically addressed by combining reward signals into a single scalar objective via a scalarization function, which can be fragile: small changes in the weights can induce drastically different policies. Multi-objective reinforcement learning (MORL) instead produces sets of policies that explicitly represent trade-offs between objectives. However, these policies are typically presented to the decision maker only through their value vectors, which can obscure substantial behavioral variation: policies that induce distinct trajectories may appear indistinguishable when evaluated solely by expected returns. We propose an exploratory diagnostic workflow that automatically highlights behavioral variation along the Pareto front that objective values alone do not reveal, providing both quantitative and visual tools to support policy inspection. We validate our approach on simple grid examples and scale it to continuous control benchmarks, demonstrating that it remains effective as problem complexity increases.
Sources
- Is Conditional Generative Modeling all you need for Decision-Making?
- Scalable agent alignment via reward modeling: a research direction
- CoMI-IRL: Contrastive Multi-Intention Inverse Reinforcement Learning
- A Survey of Multi-Objective Sequential Decision-Making
- Contrast & Compress: Learning Lightweight Embeddings for Short Trajectories
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks