Cantelli Constrained Policy Optimization
cs.LG, stat.ML
Submitted: 2026-01-30
Updated: 2026-09-10
Code: https://github.com/ax-ml/jax
License: http://creativecommons.org/licenses/by/4.0/
The gist: We introduce Canary, a risk-averse method designed to optimize Value-at-Risk (VaR) constrained reinforcement learning (RL) problems.
Terminology
Abstract
We introduce Canary, a risk-averse method designed to optimize Value-at-Risk (VaR) constrained reinforcement learning (RL) problems. We employ Cantelli's inequality to obtain a tractable, conservative and smooth bound on the VaR constraint based on the first two moments of the cost return. This yields a constraint estimator that remains stable with tight violation thresholds in dense cost regimes. Extending the trust-region framework of the Constrained Policy Optimization (CPO) method, we further provide worst-case bounds for both policy improvement and constraint violation during the training process. Empirically during training, Canary is the only method that reliably satisfies the VaR constraint in every environment tested.
Sources
- Deep Reinforcement Learning for Autonomous Driving: A Survey
- Constrained Policy Optimization
- Hindsight Experience Replay
- Trust Region Policy Optimization
- Recent Advances in Reinforcement Learning in Finance
- Reward Constrained Policy Optimization
- Gymnasium: A Standard Interface for Reinforcement Learning Environments
- Towards Safe Reinforcement Learning via Constraining Conditional Value-at-Risk
- Sample Efficient Reinforcement Learning with REINFORCE
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks