LIBERO-RECOVER: Beyond Task Success Towards Failure Recovery in Robotic Manipulation Models

summary

Video file (mp4)

The gist

Vision-Language-Action (VLA) and World Action (WAM) models have achieved remarkable performance in robotic manipulation, but success under ideal conditions does not imply real-world robustness

In short

LIBERO-Recover is a benchmark testing how robotic models recover from failures during manipulation tasks, moving beyond simple success metrics. It collects real execution failures and creates 1,000+ scenarios across four difficulty levels—from simple action retries to complex environmental reasoning. The research shows that while models succeed in ideal tests, they often fail when faced with real-world errors.

Key concepts

Failure Recovery
This evaluates a robot's ability to continue a task after an unexpected error occurs during execution. It measures if the robot can recognize that its current plan is invalid and perform necessary steps—like redoing an action or fixing an object state—to get back on track toward the original goal.
Failure Taxonomy
This organizes failures into four progressive difficulty levels: Action Retry (L1), Action Adaptation (L2), Object State Recovery (L3), and Environmental Recovery (L4). These levels dictate how much reasoning and effort the robot needs to expend to fix the problem, ranging from simple replays to complex environmental analysis.
Recovery Success Rate (RSR)
This metric quantifies if a policy can successfully complete the original task after hitting a failure state. It is calculated by checking if the robot reaches its goal state after entering a failure condition. A higher RSR means better overall capability in handling execution errors.
Recovery Degradation (RD)
This measures how much the robot's performance drops after a failure occurs. It compares the task capability before and after the error, showing how much execution ability is lost. Lower RD indicates that the robot retains its original skill level more effectively following a mistake.

Terminology used across episodes

This episode discusses

The paper

LIBERO-RECOVER: Beyond Task Success Towards Failure Recovery in Robotic Manipulation Models · Read on arXiv

Lin Liu, Lu Zhang, Huchuan Lu, Wu Yang, Yuzheng Zhuang, Shuai Tao, Wulong Liu

School of Information and Communication Engineering, Dalian University of Technology · Beta Infinity

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: I'm Rosa, and with me are Dev and Taro, guest researcher.

Dev: Today's paper: "LIBERO-RECOVER: Beyond Task Success Towards Failure Recovery in Robotic Manipulation Models".

Rosa: Vision-Language-Action (VLA) and World Action (WAM) models have achieved remarkable performance in robotic manipulation,

Dev: First, who's behind it and why it matters.

Paper summary: Rosa: So, to build on what we just talked about, this paper "LIBERO-RECOVER: Beyond Task Success Towards Failure Recovery in Robotic Manipulation Models" argues that the current success rates seen on benchmarks are insufficient because they don't account for real-world failures. The core claim is that a robot must be evaluated not only on its ability to complete the task but also on its ability to recognize and recover from deviations when those deviations occur during execution.

Dev: That’s the main thesis, and what matters is that they are creating a comprehensive benchmark for this specific capability, focusing on collecting real execution failures from state-of-the-art embodied models.

Taro: They are systematically evaluating recovery capability by organizing these failures into four progressive difficulty levels: L1 Action Retry, L2 Action Adaptation, L3 Object State Recovery, and L4 Environmental Recovery.

Rosa: These levels define the complexity of the required recovery effort; for instance, a failure at Level three means the task-relevant object state has changed significantly and needs to be restored before proceeding.

Dev: And they claim that this structured evaluation allows them to measure things like Recovery Success Rate, Recovery Degradation, and Recovery Consistency across different failure states.

Taro: The methodology involves a three-stage pipeline: first executing the task multiple times to get trajectories; second using an AI to localize where the deviation happens in time; and third using another AI to determine the specific recovery level based on that failure's consequences.

Rosa: It really matters because this moves the conversation toward measuring robustness under realistic, messy conditions rather than just ideal execution paths.

Dev: And when you look at their findings, they found that models often excel at local corrections at L1 and L2 but struggle significantly with state recovery problems at L3 and L4.

Taro: That suggests that while models are getting better at simple reactive fixes, the real challenge for general autonomy lies in reasoning about complex state changes and environmental interactions.

Rosa: And they pointed out that even when models succeed in these predefined scenarios, their performance can drop by over fifty percent when exposed to failures that look like those encountered during standard task execution.

Dev: So essentially, the paper is establishing a new way to quantify failure recovery capability for VLA and WAM systems.

Taro: It provides a framework for understanding where these models need more training emphasis to become truly reliable partners in complex physical environments.

Conclusion: Rosa: Thinking about the title "LIBERO-RECOVER" itself, it really captures the essence of what this research is doing—moving past just measuring success to specifically focusing on recovery after a failure occurs during manipulation.

Dev: And I think the authors, Lin Liu, Lu Zhang, and Huchuan Lu are pointing toward a necessary shift in how we define robot competence in this domain.

Taro: The implication for the field is that we need to start demanding that models demonstrate genuine resilience when things inevitably go wrong in physical interactions.

Rosa: And if these models can handle those L3 and L4 recovery scenarios reliably, it means robots could be deployed in environments far more dynamic than what we've tested so far.

Dev: From my perspective, this research is a vital step because it provides a measurable way to assess how well the AI handles unexpected physical disturbances during operation.

Taro: I see this as laying the groundwork for future systems where autonomy isn't just about following pre-programmed steps, but about intelligently navigating and adapting when the world misbehaves.

Rosa: And what this means in simple terms is that we are moving from testing if a robot can follow a plan perfectly to testing if it can actually keep functioning when the world throws curveballs at it.

More episodes

← Home