Policy Gradients for Cumulative Prospect Theory in Reinforcement Learning
cs.LG, cs.AI
Submitted: 2024-10-03
Updated: 2026-09-06
Comments: Published in Transactions on Machine Learning Research, camera-ready version includes an updated algorithm, new convergence results, an extended related work discussion and new future research directions
Journal ref: Transactions on Machine Learning Research 2026
License: http://creativecommons.org/licenses/by/4.0/
The gist: We derive a policy gradient theorem for Cumulative Prospect Theory (CPT) objectives in finite-horizon Reinforcement Learning (RL), generalizing the standard policy gradient theorem and encompassing
Terminology
Abstract
We derive a policy gradient theorem for Cumulative Prospect Theory (CPT) objectives in finite-horizon Reinforcement Learning (RL), generalizing the standard policy gradient theorem and encompassing distortion-based risk objectives as special cases. Motivated by behavioral economics, CPT combines an asymmetric utility transformation around a reference point with probability distortion. Building on our theorem, we design a first-order policy gradient algorithm for CPT-RL using a Monte Carlo gradient estimator based on order statistics. We establish statistical guarantees for the estimator and prove asymptotic convergence of the resulting algorithm to first-order stationary points of the (generally nonconvex) CPT objective. We complement our asymptotic analysis with a non-asymptotic total sample complexity analysis to reach an approximate first-order stationary policy. Simulations illustrate qualitative behaviors induced by CPT and compare our first-order approach to existing zeroth-order methods.
Sources
- A Short Note on Concentration Inequalities for Random Vectors with SubGaussian Norm
- Risk-Sensitive Reinforcement Learning with Exponential Criteria
- Policy Newton methods for Distortion Riskmetrics
- PyTorch: An Imperative Style, High-Performance Deep Learning Library
- Policy Gradient Methods for Distortion Risk Measures
- Risk-sensitive Markov Decision Process and Learning under General Utility Functions
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks