REVERSAL-BENCH: A Reversibility Axis and Reset Oracle for Measuring the Reset-Free RL Cliff
cs.LG, cs.AI, cs.RO
Submitted: 2026-09-15
Updated: 2026-09-15
Comments: 8 pages, 5 figures, 2 tables
License: http://creativecommons.org/licenses/by/4.0/
The gist: A central goal of autonomous reinforcement learning is continuous policy training without external resets.
Terminology
Abstract
A central goal of autonomous reinforcement learning is continuous policy training without external resets. However, existing paradigms largely depend on underlying environmental reversibility, a property absent in real world manipulation, where events such as pushing objects off tables or spilling granular substances cannot be undone. We introduce REVERSAL-BENCH, a benchmark that controls reversibility via a continuous parameter ρ in [0, 1] and provides a reset oracle, a ground-truth verification mechanism to test state recoverability across eight manipulation settings in five physics engines. Evaluating a broad spectrum of policy architectures, including standard actor-critic algorithms, safe RL, and specialized reset-free frameworks, reveals a sharp reversibility cliff: reset-free agents are consistently absorbed into irrecoverable states as ρ increases, whereas episodic agents maintain steady learning. We see this failure mode across autonomous reset-free baselines and constrained RL. Because reset-free agents lack external resets, any transition into an irrecoverable state results in permanent absorption, leaving the agent trapped where further learning halts. We show that this absorption phenomenon persists in full physics simulations under learned manipulation policies. By evaluating against geometrically identical reversible counterparts, we confirm that this breakdown is causally driven by irreversibility rather than obstacle complexity. We release the benchmark suite, a large multi-simulator dataset labeled with recoverability and a reset oracle. We also evaluate a safety shield that intervenes before irreversible failures occur, showing that while recoverability can be predicted accurately, active recovery primarily succeeds only when the agent can physically steer clear of the trap
Sources
- A State-Distribution Matching Approach to Non-Episodic Reinforcement Learning
- The Ingredients of Real-World Robotic Reinforcement Learning
- Autonomous Reinforcement Learning via Subgoal Curricula
- Autonomous Reinforcement Learning: Formalism and Benchmarking
- Intelligent Switching for Reset-Free RL
- Recovery RL: Safe Reinforcement Learning with Learned Recovery Zones
- There Is No Turning Back: A Self-Supervised Approach for Reversibility-Aware Reinforcement Learning
- Reset-free Reinforcement Learning with World Models
- Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor
- Responsive Safety in Reinforcement Learning by PID Lagrangian Methods
- Learning to Undo: Rollback-Augmented Reinforcement Learning with Reversibility Signals
- Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks