Certified Safety Curation: Distribution-Free Guarantees for Safe Offline Reinforcement Learning
cs.LG
Submitted: 2026-09-10
Updated: 2026-09-26
Comments: 37 pages, 10 figures
License: http://creativecommons.org/licenses/by/4.0/
The gist: Safe offline reinforcement learning assumes a cost function on every transition.
Terminology
Abstract
Safe offline reinforcement learning assumes a cost function on every transition. We ask what remains possible when safety can be judged only by comparing short clips and occasionally asking whether an episode exceeded its budget. Certified safety curation answers with a filter-then-clone pipeline: a state-only value trained from segment comparisons scores whole trajectories, Learn-then-Test calibration certifies a selection threshold under a distribution-free (α, δ) bound on the unsafe fraction of the selection, and behavior cloning follows. We are not aware of prior work certifying the composition of a training set for offline RL or imitation. Oracle controls justify the design: reweighting individual transitions fails even with an exact value, so the value selects whole trajectories. The policies satisfy the cost budget on eleven of fifteen DSRL tasks, one short of cloning the ground-truth safe subset, which needs a label on every trajectory; the uncertified variant reaches twelve. Retrained on the certified selection, the strongest full-label method becomes safe where no setting of its own cost target rescues it. Refusal is predictable: the certificate's probability has a closed form in the purity the pool attains, which the calibration sample estimates and the scorer enters only through.
Sources
- Offline RLAIF: Piloting VLM Feedback for RL via SFO
- OSIL: Learning Offline Safe Imitation Policies with Safety Inferred from Non-preferred Trajectories
- RvS: What is Essential for Offline RL via Supervised Learning?
- COptiDICE: Offline Constrained Reinforcement Learning via Stationary Distribution Correction Estimation
- Robometer: Scaling General-Purpose Robotic Reward Models via Trajectory Comparisons
- Datasets and Benchmarks for Offline Safe Reinforcement Learning
- Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks