LIBERO-RECOVER: Beyond Task Success Towards Failure Recovery in Robotic Manipulation Models
summary
The gist
Vision-Language-Action (VLA) and World Action (WAM) models have achieved remarkable performance in robotic manipulation, but success under ideal conditions does not imply real-world robustness
In short
LIBERO-Recover is a benchmark testing how robotic models recover from failures during manipulation tasks, moving beyond simple success metrics. It collects real execution failures and creates 1,000+ scenarios across four difficulty levels—from simple action retries to complex environmental reasoning. The research shows that while models succeed in ideal tests, they often fail when faced with real-world errors.
Key concepts
- Failure Recovery
- This evaluates a robot's ability to continue a task after an unexpected error occurs during execution. It measures if the robot can recognize that its current plan is invalid and perform necessary steps—like redoing an action or fixing an object state—to get back on track toward the original goal.
- Failure Taxonomy
- This organizes failures into four progressive difficulty levels: Action Retry (L1), Action Adaptation (L2), Object State Recovery (L3), and Environmental Recovery (L4). These levels dictate how much reasoning and effort the robot needs to expend to fix the problem, ranging from simple replays to complex environmental analysis.
- Recovery Success Rate (RSR)
- This metric quantifies if a policy can successfully complete the original task after hitting a failure state. It is calculated by checking if the robot reaches its goal state after entering a failure condition. A higher RSR means better overall capability in handling execution errors.
- Recovery Degradation (RD)
- This measures how much the robot's performance drops after a failure occurs. It compares the task capability before and after the error, showing how much execution ability is lost. Lower RD indicates that the robot retains its original skill level more effectively following a mistake.
Terminology used across episodes
This episode discusses
- LIBERO-RECOVER: Beyond Task Success Towards Failure Recovery in Robotic Manipulation Models · Paper Radio
- Cosmos-Transfer1: Conditional World Generation with Adaptive Multimodal Control
- GR00T N1: An Open Foundation Model for Generalist Humanoid Robots
- pi 0: A Vision-Language-Action Flow Model for General Robot Control
- UniVLA: Learning to Act Anywhere with Task-centric Latent Actions
- RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation
- LIBERO-Safety: A Comprehensive Benchmark for Physical and Semantic Safety in Vision-Language-Action Models
- LIBERO-Plus: In-depth Robustness Analysis of Vision-Language-Action Models
- Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations
- pi 0.5: a Vision-Language-Action Model with Open-World Generalization
- OpenVLA: An Open-Source Vision-Language-Action Model
- Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning
- Evaluating Real-World Robot Manipulation Policies in Simulation
- RoboCasa: Large-Scale Simulation of Everyday Tasks for Generalist Robots
- FAST: Efficient Action Tokenization for Vision-Language-Action Models
- Wan: Open and Advanced Large-Scale Video Generative Models
- LIBERO-X: Robustness Litmus for Vision-Language-Action Models
- FLARE: Robot Learning with Implicit World Modeling
- LIBERO-PRO: Towards Robust and Fair Evaluation of Vision-Language-Action Models Beyond Memorization
- Unified World Models: Coupling Video and Action Diffusion for Pretraining on Large Robotic Datasets
The paper
LIBERO-RECOVER: Beyond Task Success Towards Failure Recovery in Robotic Manipulation Models · Read on arXiv
Lin Liu, Lu Zhang, Huchuan Lu, Wu Yang, Yuzheng Zhuang, Shuai Tao, Wulong Liu
School of Information and Communication Engineering, Dalian University of Technology · Beta Infinity
Transcript
Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.
Rosa: I'm Rosa, and with me are Dev and Taro, guest researcher.
Dev: Today's paper: "LIBERO-RECOVER: Beyond Task Success Towards Failure Recovery in Robotic Manipulation Models".
Rosa: Vision-Language-Action (VLA) and World Action (WAM) models have achieved remarkable performance in robotic manipulation,
Dev: First, who's behind it and why it matters.
Paper summary: Rosa: So, to build on what we just talked about, this paper "LIBERO-RECOVER: Beyond Task Success Towards Failure Recovery in Robotic Manipulation Models" argues that the current success rates seen on benchmarks are insufficient because they don't account for real-world failures. The core claim is that a robot must be evaluated not only on its ability to complete the task but also on its ability to recognize and recover from deviations when those deviations occur during execution.
Dev: That’s the main thesis, and what matters is that they are creating a comprehensive benchmark for this specific capability, focusing on collecting real execution failures from state-of-the-art embodied models.
Taro: They are systematically evaluating recovery capability by organizing these failures into four progressive difficulty levels: L1 Action Retry, L2 Action Adaptation, L3 Object State Recovery, and L4 Environmental Recovery.
Rosa: These levels define the complexity of the required recovery effort; for instance, a failure at Level three means the task-relevant object state has changed significantly and needs to be restored before proceeding.
Dev: And they claim that this structured evaluation allows them to measure things like Recovery Success Rate, Recovery Degradation, and Recovery Consistency across different failure states.
Taro: The methodology involves a three-stage pipeline: first executing the task multiple times to get trajectories; second using an AI to localize where the deviation happens in time; and third using another AI to determine the specific recovery level based on that failure's consequences.
Rosa: It really matters because this moves the conversation toward measuring robustness under realistic, messy conditions rather than just ideal execution paths.
Dev: And when you look at their findings, they found that models often excel at local corrections at L1 and L2 but struggle significantly with state recovery problems at L3 and L4.
Taro: That suggests that while models are getting better at simple reactive fixes, the real challenge for general autonomy lies in reasoning about complex state changes and environmental interactions.
Rosa: And they pointed out that even when models succeed in these predefined scenarios, their performance can drop by over fifty percent when exposed to failures that look like those encountered during standard task execution.
Dev: So essentially, the paper is establishing a new way to quantify failure recovery capability for VLA and WAM systems.
Taro: It provides a framework for understanding where these models need more training emphasis to become truly reliable partners in complex physical environments.
Conclusion: Rosa: Thinking about the title "LIBERO-RECOVER" itself, it really captures the essence of what this research is doing—moving past just measuring success to specifically focusing on recovery after a failure occurs during manipulation.
Dev: And I think the authors, Lin Liu, Lu Zhang, and Huchuan Lu are pointing toward a necessary shift in how we define robot competence in this domain.
Taro: The implication for the field is that we need to start demanding that models demonstrate genuine resilience when things inevitably go wrong in physical interactions.
Rosa: And if these models can handle those L3 and L4 recovery scenarios reliably, it means robots could be deployed in environments far more dynamic than what we've tested so far.
Dev: From my perspective, this research is a vital step because it provides a measurable way to assess how well the AI handles unexpected physical disturbances during operation.
Taro: I see this as laying the groundwork for future systems where autonomy isn't just about following pre-programmed steps, but about intelligently navigating and adapting when the world misbehaves.
Rosa: And what this means in simple terms is that we are moving from testing if a robot can follow a plan perfectly to testing if it can actually keep functioning when the world throws curveballs at it.
More episodes
- 2610.12154-Stochastic Distribution Network Reconfiguration under Load Uncertainty
- 2607.00148-3D Point World Models: Point Completion Enables More Accurate Dynamics Learning
- 2607.02403-ACID: Action Consistency via Inverse Dynamics for Planning with World Models
- 2510.26623-A Sliding-Window Filter for Online Continuous-Time Continuum Robot State Estimation
- 2406.13267-The Kinetics Observer: A Tightly Coupled Estimator for Legged Robots
- 2511.02147-Census-Based Population Autonomy For Distributed Robotic Teaming
- 2603.08260-Seed2Scale: A Self-Evolving Data Engine with Parallel Worlds Expansion for Scalable Robot Learning
- 2602.14032-RoboAug: One Annotation to Hundreds of Scenes via Region-Contrastive Data Augmentation for Robotic Manipulation
- 2602.15397-ActionCodec: What Makes for Good Action Tokenizers
- 2607.01819-Koopman operator theory: fundamentals, control, and applications