Reinforcement learning with an expectile-based objective
cs.LG
Submitted: 2026-02-10
Updated: 2026-09-20
License: http://creativecommons.org/licenses/by/4.0/
The gist: We consider the policy evaluation and control in a finite horizon reinforcement learning (RL) setting under an expectile-based objective.
Terminology
Abstract
We consider the policy evaluation and control in a finite horizon reinforcement learning (RL) setting under an expectile-based objective. First, we derive the mean-squared error (MSE) and concentration bounds for the classic estimator of expectiles based on independent and identically distributed (i.i.d.) samples. To the best of our knowledge, expectiles have not been analyzed in the non-asymptotic regime and the bounds we derive may be of independent interest. Next, we analyze a Monte Carlo type estimator of expectile of the Markov chain underlying a given policy. We derive upper bounds that hold in expectation as well as with high probability for this estimator. Further, we show the order-optimality of our estimator by deriving a lower bound on expectile-based policy evaluation. For the problem of control, we adopt a policy gradient approach and derive a policy gradient theorem for expectiles. Using this result, we propose a gradient estimator with a O (1/m) mean-squared error bounds, where m is the number of trajectories. Further, under standard assumptions for policy gradient-type algorithms, we establish smoothness of the expectile-sensitive objective, in turn leading to stationary convergence rate bounds for the overall risk-sensitive policy gradient algorithm that we propose. Finally, we conduct numerical experiments to show the utility of expectiles on popular RL benchmarks.
Sources
- Bootstrapping Expectiles in Reinforcement Learning
- A Remark on the Structure of Expectiles
- Concentration Bounds for Optimized Certainty Equivalent Risk Estimation
- Distributional Reinforcement Learning with Dual Expectile-Quantile Regression
- Risk Estimation in a Markov Cost Process: Lower and Upper Bounds
- Policy Gradient Methods for Distortion Risk Measures
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks