Continue, Abort, or Fall: Viability-Aware Policy Selection (VAPS) for Safe Humanoid Acrobatics
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.
Rosa: Today's paper: "Continue, Abort, or Fall".
Dev: Dynamic humanoid motions, such as flips, risk hardware damage due to suboptimal policies or disturbances, and this work presents Viability-Aware Policy Selection (VAPS), which treats safety as a policy-conditioned,
Rosa: First, who's behind it and why it matters.
Paper summary: Dev: So to wrap up what we've heard about "Continue, Abort, or Fall: Viability-Aware Policy Selection (VAPS) for Safe Humanoid Acrobatics," this paper essentially presents a structured way for humanoid robots to manage the risk of hardware damage during dynamic motions like flips.
Taro: It’s about moving beyond simple one-shot safety checks and instead implementing a continuous decision process where the robot constantly weighs what it can safely do next against potential future outcomes.
Rosa: The authors introduce Viability-Aware Policy Selection, or VAPS, which uses learned predictors to estimate the viability of different behaviors over a short time horizon before selecting between continuing, aborting for feet contact, or executing a protective fall.
Dev: This is significant because it shows that folding execution and recovery into one policy doesn't restore the necessary granularity for these fast maneuvers; VAPS offers a more sophisticated decision-making structure.
Taro: The impact here is that we can design autonomous systems that make trade-offs between achieving a difficult goal and maintaining a guaranteed level of physical safety through this structured policy selection.
Rosa: It proves that safety in complex dynamic tasks depends on what the robot does next and how much time remains to save it, which is something we need to consider as we deploy more complex robots.
Dev: The distinction between VAPS's behavior—keeping several named behaviors and retaining the most ambitious one that is still viable—and monolithic approaches really highlights the value of this explicit policy structure.
Taro: The future work they point toward is extending viability prediction to other tasks, suggesting this framework could be adapted for deciding when to abandon an entire task safely rather than just managing a single maneuver.
Rosa: Ultimately, VAPS gives us a concrete way to handle the complexity of dynamic motion safety by conditioning policy selection on predicted time-dependent risk.
Conclusion: Rosa: So, to wrap up our discussion on "Continue, Abort, or Fall: Viability-Aware Policy Selection (VAPS) for Safe Humanoid Acrobatics," we've seen how this work provides a systematic way for robots to decide whether to keep going with a flip, stop and land safely on their feet, or just take a controlled fall.
Dev: Exactly. The core idea is that the robot doesn't just react to the current state; it looks ahead using learned predictors to estimate how long any given action will actually last before it becomes dangerous. That predictive modeling is what gives this system its structure beyond simple reactive control.
Taro: I think what really stands out is that VAPS handles things when the world goes sideways. If something unexpected happens, the system doesn't just crash; it uses those viability predictions to choose the safest path forward, which is crucial for real-world deployment.
Rosa: And those implications are pretty huge because it moves safety from a fixed set of rules into a dynamic, time-dependent decision framework. It suggests that complex maneuvers can be managed by prioritizing the most ambitious safe path available at any given moment.
Dev: From an engineering standpoint, I'm impressed with how they manage the loop rate and latency within this predictive window; having those predictors run continuously allows for very fine control over the transition between policies without introducing noticeable lag.
Taro: I'm also interested in how this could affect autonomy in general, because if a robot can intelligently decide when to give up a high-risk task based on future risk assessment, that opens up possibilities for much more robust exploration.
Rosa: We should definitely keep thinking about where these kinds of viability predictors can be applied outside of acrobatics; I'm curious if this concept has value in other high-stakes physical tasks.
Dev: It definitely has potential, but the hardware validation they did on platforms like the LimX Oli is key to seeing if those theoretical predictions hold up when you're dealing with real torque limits and physical disturbances.
Taro: That’s a fair point; showing it works on different hardware setups is what moves this from a lab curiosity into something that could impact how we build reliable autonomous systems in the field.
Rosa: So, moving forward, we need to focus on those real-world testing scenarios where these policies are actually being tested under imperfect conditions.
Siwei Ju, Lu Liu, Jan Peters, Oleg Arenz
Department of Computer Science, Technical University of Darmstadt, Germany · Robotics Institute Germany (RIG) · LimX Dynamics · Hessian.AI · German Research Center for AI (DFKI), Research Department: Systems AI for Robot Learning
cs.RO
Submitted: 2026-10-01
Updated: 2026-10-01
Project page: https://vaps-r-al.github.io
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 91/100
The gist: Dynamic humanoid motions, such as flips, risk hardware damage due to suboptimal policies or disturbances, and this work presents Viability-Aware Policy Selection (VAPS), which treats safety as a
Key concepts
- Specialist Policies
- The system defines three distinct behaviors: motion success (tracking the goal), controlled abort (landing on feet), and protective falling. These policies are managed as a set where the robot chooses which one to execute based on real-time safety predictions.
- Policy-Conditioned Receding-Horizon Viability
- Instead of guessing the end of an episode, VAPS uses predictors to estimate the probability of no failure occurring over a short window (H steps). This allows the robot to continuously re-evaluate which action is safest and most viable at every single control step.
- Viability-Aware Policy Selection
- The final decision rule selects the policy based on predicted viability thresholds. This ensures the robot only proceeds as far as it is predicted to be safe, balancing the desire to complete a task with immediate safety constraints.
Terminology
Summary
Dynamic humanoid motions, such as flips, risk hardware damage due to suboptimal policies or disturbances, and this work presents Viability-Aware Policy Selection (VAPS), which treats safety as a policy-conditioned, receding-horizon decision framework to select between continuing the motion, aborting it for a controlled landing on feet, or executing a protective fall.
The gist
VAPS selects among the nominal tracking policy and backup policies—an abort policy seeking a controlled feet-first landing, and a protective-fall policy minimizing harm once upright recovery is no longer viable—by using learned predictors to estimate the finite-horizon viability of each behavior at every control step, preserving the most task-ambitious behavior that remains viable.
Specialist Policies
The system decomposes safe execution into three specialist behaviors: motion success, controlled abort, and protective falling. The policies are defined as a set: Π = Σπnom, πabort, πprotect (1).
The nominal policy (πnom) tracks the reference motion. The abort policy (πabort) minimizes damage by seeking an upright base with only the feet in ground contact, rewarding keeping the base’s vertical axis upright
and simultaneous contact of both feet.
The protective-fall policy (πprotect) is triggered when a fall is unavoidable and takes over to reduce its severity,
penalizing contact forces on the head and hands beyond static support load.
Policy-Conditioned Receding-Horizon Viability
VAPS utilizes predictors to estimate viability over a short, continuously re-evaluated window rather than by a single guess for the end of the episode. The Hstep viability of policy πi is defined as "the probability that no failure occurs within H steps, V H i(st) = Pr (∆i(st) > H st, πi). These predictors are trained using discounted distributions over future offsets and viability outcomes to minimize a cross-entropy loss function. This mechanism allows the system to
recover a mature policy from disturbances, supervises a new and unproven one, and handles a stop request issued by an operator or a workspace monitor."
Viability-Aware Policy Selection
The final selection rule at every control step is based on the predicted viability thresholds: "πt = f(πnom, V nom,t > τnom, πabort, V nom,t ≤ τnom AND V abort > τabort). This ensures that the robot
descend only as far as the predicted viability forces it to," prioritizing task ambition while maintaining safety. The thresholds (e.g., τnom = 0.2) are empirically selected to balance false alarms against lead time, with a required false-alarm rate of at most 1% on held-out episodes.
Monolithic Alternatives
The paper investigates single-network alternatives, including end-to-end RL policies where safety is folded into the task reward (e.g., re2e = rnom − λ (chead + chand + cjoint + ctorque)). It also examines students distilled from VAPS, using methods like DAgger to add student-visited states that the oracle routes to the protective fall.
The results show that VAPS dominates these monolithic alternatives on the Pareto frontier; for instance, at τ = 0.01, VAPS matches nominal success while achieving substantially lower head-contact rates than best end-to-end and distilled alternatives. The structure of VAPS keeps several named behaviors (continue, abort, fall) and retains the most ambitious one that is still viable,
a distinction that single networks lack.
Hardware Validation
VAPS was validated in simulation on the Unitree G1 and LimX Oli humanoid platforms, with transfer experiments conducted on the physical LimX Oli. In hardware tests involving manual switching, the abort policy returned the robot to a stance in 17 of 19 trials, while the protective-fall policy managed to protect the head and hands in 12/16 trials. The full-hierarchy VAPS achieved an overall success rate of 83.3% across four routing patterns on hardware, demonstrating its robustness against external disturbances. Furthermore, VAPS can protect the hardware during the early deployment of a new policy or a new robot
by only requiring retraining of the nominal predictor when a new nominal policy is introduced.
Conclusion and Outlook
VAPS successfully addresses the challenge that safety cannot be decided from the state alone; it depends on what the robot does next and on how much time remains to save it.
The explicit structure of VAPS proved superior to monolithic approaches because it preserves the decision of how much of the task to give up,
leading to better trade-offs between task success and safety. Future work involves extending viability prediction to other tasks, such as deciding when to abandon a task safely.
References
[1] T. He, J. Gao, W. Xiao, Y.
Improvements for AI systems
Based on the provided research paper, here are specific improvements for AI systems and what those improved systems can achieve:
) Improved System Capabilities: Viability-Aware Policy Selection (VAPS) Framework
The core improvement is moving from simple, reactive safety mechanisms to a proactive, policy-conditioned decision-making framework that treats safety as a receding-horizon selection problem.
-
A robot executing dynamic maneuvers (e.g., acrobatics like flips) can maintain high task ambition while ensuring hardware survival by selecting the
least-sacrificial
action among three hierarchical options: -
It can perform a controlled, safe landing on its feet if the maneuver is lost but recovery is still possible (Abort Policy).
-
It can execute a controlled, damage-mitigated fall if upright recovery is no longer viable (Protective Fall Policy).
-
The system will continuously monitor the viability of both the nominal tracking policy and backup policies over a short, continuously re-evaluated window using learned predictors based on recent proprioceptive history (Viability Predictors).
-
The system will dynamically select the most ambitious behavior that remains viable at every control step, preserving task success whenever possible while minimizing hardware risk.
-
The system can supervise partially trained or undertrained policies by retraining only the viability predictor when a new nominal policy is deployed, allowing for safe deployment of novel controllers without requiring full re-training of the primary motion policy.
-
The system can be compared against monolithic alternatives (End-to-End RL and distilled students) and is shown to dominate in terms of safety while maintaining high task success rates, demonstrating that explicit structural separation (specialist policies + predictors) is superior to folding objectives into a single network.
-
The system will exhibit superior performance in distinguishing between different failure modes: it can detect potential falls significantly earlier (30-60% lead time improvement over baseline final-outcome predictors) and reduce head contact forces substantially compared to monolithic approaches, resulting in lower peak impact forces on fragile components (head/hands).
-
The system is robust across different hardware platforms and deployment scenarios, successfully transferring safety guarantees from simulation to physical robots (LimX Oli) with minimal degradation in performance metrics.
Sources
- ASAP: Aligning Simulation and Real-World Physics for Learning Agile Humanoid Whole-Body Skills
- BeyondMimic: From Motion Tracking to Versatile Humanoid Control via Guided Diffusion
- Recovery RL: Safe Reinforcement Learning with Learned Recovery Zones
- Safe Reinforcement Learning via Shielding
- Safe Reinforcement Learning for Legged Locomotion
- SafeFall: Learning Protective Control for Humanoid Robots
- A Kung Fu Athlete Bot That Can Do It All Day: Highly Dynamic, Balance-Challenging Motion Dataset and Autonomous Fall-Resilient Tracking
- Unified Humanoid Fall-Safety Policy from a Few Demonstrations
- DeepMimic: Example-Guided Deep Reinforcement Learning of Physics-Based Character Skills
- OmniH2O: Universal and Dexterous Human-to-Humanoid Whole-Body Teleoperation and Learning
- Humanoid Parkour Learning
- Perceptive Humanoid Parkour: Chaining Dynamic Human Skills via Motion Matching
- KungfuBot: Physics-Based Humanoid Whole-Body Control for Learning Highly-Dynamic Skills
- Learning a Unified Control Policy for Safe Falling
- Discovering Self-Protective Falling Policy for Humanoid Robot via Deep Reinforcement Learning
- Robot Crash Course: Learning Soft and Stylized Falling
- Hamilton-Jacobi Reachability: A Brief Overview and Recent Advances
- A predictive safety filter for learning-based control of constrained nonlinear dynamical systems
- $\lambda$-Reachability: Geometric-Horizon Safety Bellman Equations for Humanoid Safety
- Agile But Safe: Learning Collision-Free High-Speed Legged Locomotion
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving