LIBERO-RECOVER: Beyond Task Success Towards Failure Recovery in Robotic Manipulation Models

arXiv:2609.05178 · cs.RO · Submitted 2026-09-04 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: I'm Rosa, and with me are Dev and Taro, guest researcher.

Dev: Today's paper: "LIBERO-RECOVER: Beyond Task Success Towards Failure Recovery in Robotic Manipulation Models".

Rosa: Vision-Language-Action (VLA) and World Action (WAM) models have achieved remarkable performance in robotic manipulation,

Dev: First, who's behind it and why it matters.

Paper summary: Rosa: So, to build on what we just talked about, this paper "LIBERO-RECOVER: Beyond Task Success Towards Failure Recovery in Robotic Manipulation Models" argues that the current success rates seen on benchmarks are insufficient because they don't account for real-world failures. The core claim is that a robot must be evaluated not only on its ability to complete the task but also on its ability to recognize and recover from deviations when those deviations occur during execution.

Dev: That’s the main thesis, and what matters is that they are creating a comprehensive benchmark for this specific capability, focusing on collecting real execution failures from state-of-the-art embodied models.

Taro: They are systematically evaluating recovery capability by organizing these failures into four progressive difficulty levels: L1 Action Retry, L2 Action Adaptation, L3 Object State Recovery, and L4 Environmental Recovery.

Rosa: These levels define the complexity of the required recovery effort; for instance, a failure at Level three means the task-relevant object state has changed significantly and needs to be restored before proceeding.

Dev: And they claim that this structured evaluation allows them to measure things like Recovery Success Rate, Recovery Degradation, and Recovery Consistency across different failure states.

Taro: The methodology involves a three-stage pipeline: first executing the task multiple times to get trajectories; second using an AI to localize where the deviation happens in time; and third using another AI to determine the specific recovery level based on that failure's consequences.

Rosa: It really matters because this moves the conversation toward measuring robustness under realistic, messy conditions rather than just ideal execution paths.

Dev: And when you look at their findings, they found that models often excel at local corrections at L1 and L2 but struggle significantly with state recovery problems at L3 and L4.

Taro: That suggests that while models are getting better at simple reactive fixes, the real challenge for general autonomy lies in reasoning about complex state changes and environmental interactions.

Rosa: And they pointed out that even when models succeed in these predefined scenarios, their performance can drop by over fifty percent when exposed to failures that look like those encountered during standard task execution.

Dev: So essentially, the paper is establishing a new way to quantify failure recovery capability for VLA and WAM systems.

Taro: It provides a framework for understanding where these models need more training emphasis to become truly reliable partners in complex physical environments.

Conclusion: Rosa: Thinking about the title "LIBERO-RECOVER" itself, it really captures the essence of what this research is doing—moving past just measuring success to specifically focusing on recovery after a failure occurs during manipulation.

Dev: And I think the authors, Lin Liu, Lu Zhang, and Huchuan Lu are pointing toward a necessary shift in how we define robot competence in this domain.

Taro: The implication for the field is that we need to start demanding that models demonstrate genuine resilience when things inevitably go wrong in physical interactions.

Rosa: And if these models can handle those L3 and L4 recovery scenarios reliably, it means robots could be deployed in environments far more dynamic than what we've tested so far.

Dev: From my perspective, this research is a vital step because it provides a measurable way to assess how well the AI handles unexpected physical disturbances during operation.

Taro: I see this as laying the groundwork for future systems where autonomy isn't just about following pre-programmed steps, but about intelligently navigating and adapting when the world misbehaves.

Rosa: And what this means in simple terms is that we are moving from testing if a robot can follow a plan perfectly to testing if it can actually keep functioning when the world throws curveballs at it.

Lin Liu, Lu Zhang, Huchuan Lu, Wu Yang, Yuzheng Zhuang, Shuai Tao, Wulong Liu

School of Information and Communication Engineering, Dalian University of Technology · Beta Infinity

cs.RO

Submitted: 2026-09-04

Updated: 2026-09-28

Project page: https://liulin815.github.io/LIBERO-Recovery

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 77/100

The gist: Vision-Language-Action (VLA) and World Action (WAM) models have achieved remarkable performance in robotic manipulation, but success under ideal conditions does not imply real-world robustness

Key concepts

Failure Recovery
This evaluates a robot's ability to continue a task after an unexpected error occurs during execution. It measures if the robot can recognize that its current plan is invalid and perform necessary steps—like redoing an action or fixing an object state—to get back on track toward the original goal.
Failure Taxonomy
This organizes failures into four progressive difficulty levels: Action Retry (L1), Action Adaptation (L2), Object State Recovery (L3), and Environmental Recovery (L4). These levels dictate how much reasoning and effort the robot needs to expend to fix the problem, ranging from simple replays to complex environmental analysis.
Recovery Success Rate (RSR)
This metric quantifies if a policy can successfully complete the original task after hitting a failure state. It is calculated by checking if the robot reaches its goal state after entering a failure condition. A higher RSR means better overall capability in handling execution errors.
Recovery Degradation (RD)
This measures how much the robot's performance drops after a failure occurs. It compares the task capability before and after the error, showing how much execution ability is lost. Lower RD indicates that the robot retains its original skill level more effectively following a mistake.

Terminology

Summary

Vision-Language-Action (VLA) and World Action (WAM) models have achieved remarkable performance in robotic manipulation, but success under ideal conditions does not imply real-world robustness because robots must recognize and recover from failures to continue tasks. This paper introduces LIBERO-Recover, a large-scale benchmark designed to evaluate the critical capability of failure recovery in embodied models by collecting real execution failures and constructing scenarios across four progressive difficulty levels.

The Gist

LIBERO-Recover is a large scale benchmark for failure recovery in robotic manipulation, shifting evaluation from Can the robot succeed? to Can the robot recover after failure?, by collecting real execution failures from SOTA embodied models and constructing 1,000+ scenarios across four recovery levels: (1) Action Retry, (2) Action Adaptation, (3) Object State Recovery, and (4) Environmental Recovery.

Problem Formulation

The paper contrasts conventional task execution with failure recovery. Conventional evaluation measures whether a policy can continuously execute the task from an initial state to a goal state. Failure recovery considers an execution that deviates from the intended progression and reaches an intermediate state where the original task goal is not satisfied, requiring subsequent interaction to restore necessary states before continuing toward the original goal. A failure state is defined as one where execution-induced changes invalidate the current plan, necessitating adaptation or recovery, while state changes that preserve task executability are treated as normal variations.

Failure Taxonomy and Difficulty Levels

The recoverable failures are organized into four progressive difficulty levels based on the required state reasoning and recovery effort:

  1. L1 Action Retry: The failure does not materially alter the task configuration; recovery only requires reattempting the failed action.

  2. L2 Action Adaptation: The failure causes a limited change, requiring adaptation of actions to the observed state rather than blindly repeating the failed action.

  3. L3 Object State Recovery: The failure substantially alters a task-relevant object, requiring first restoring the object’s required state before resuming the original task.

  4. L4 Environmental Recovery: The failure affects an object or state not directly manipulated by the original task but subsequently prevents execution, requiring reasoning about both task-relevant and task-irrelevant states and their interaction topology.

LIBERO-Recover Benchmark Construction

The benchmark is built upon LIBERO by collecting real execution failures from SOTA embodied models. The construction pipeline involves three stages:

  1. Stage 1: Task Execution, where multiple embodied policies are deployed to generate a complete failure trajectory, yielding videos of the execution.

  2. Stage 2: Failure Localization, utilizing Qwen3.5-27B-Instruct to analyze execution videos and temporally localize deviations into pre-failure, during-failure, and post-failure stages.

  3. Stage 3: Failure Characterization, where Qwen3.5 is prompted to assess the post-failure consequences and determine the corresponding recovery type and difficulty level according to the taxonomy defined in Section 3.3. This process results in 2178 recovery scenarios derived from 130 subtasks across four task categories, covering four recovery levels and 16 evaluation dimensions.

Evaluation Metrics

To assess model performance, three core metrics are proposed:

(Recovery Success Rate (RSR))

This measures whether a policy can successfully complete the original task after entering a failure state: RSR = 1/N X N i=1 G(s i T, l) = 1. Higher RSR indicates stronger failure recovery capability.

(Recovery Degradation (RD))

This quantifies execution capability lost after a failure: RD = 1 − Mpost + ε / Mpre + ε. Lower RD indicates better retention of execution capability.

(Recovery Consistency (RC))

This measures the stability of recovery across different failure states within the same task: RC = 1 − Stdk(Rk). Higher RC indicates more consistent recovery across failure states.

Key Findings

Experiments reveal that Success in predefined scenarios does not guarantee robust generalization to real failures, showing performance drops over 50% when exposed to naturally occurring failures. Furthermore, models excel at local correction (L1/L2) but struggle with state recovery (L3/L4). WAM policies exhibit more stable recovery, achieving higher Recovery Consistency (RC) than conventional VLAs due to their explicit modeling of action-conditioned state transitions. Finally, Smaller action chunks improve recovery by enabling more frequent state feedback, suggesting that effective failure recovery requires fine-grained closed-loop control. However, recovery training does not effectively transfer to failures encountered during standard task execution, indicating a failure-recovery transfer gap. Temporal context also improves performance, as providing the task initial frame as temporal context consistently increases recovery success rates.

Improvements for AI systems

Here are the specific improvements that can be made to AI systems based on the LIBERO-Recover benchmark, and what those improved systems will be able to do:

  1. The core improvement is shifting model evaluation from simple task completion success under ideal conditions (standard benchmarks) to robustness and resilience against real-world failures (LIBERO-Recover).

  2. Improved AI systems will move beyond achieving near 100% success rates in simulation to demonstrating the capability to recognize, diagnose, and recover from unexpected physical disruptions like failed grasps, collisions, and unintended object movements during live operation.

Specifically:

  1. Improved systems will exhibit superior performance across four distinct recovery levels:

  2. Recognize and execute simple retries (L1), such as reattempting a failed grasp when the target remains unchanged.

  3. Adapt their actions to limited state changes (L2), adjusting the next action based on observed, non-catastrophic deviations from the plan, rather than blindly repeating the failed action.

  4. Perform complex object state recovery (L3), such as restoring a displaced object's pose before continuing manipulation, requiring structured reasoning about object states and spatial relations.

  5. Execute sophisticated environmental recovery (L4), which involves reasoning about both task-relevant and irrelevant states, identifying environmental obstructions, and restoring necessary configurations to resume the original task.

Furthermore:

  1. Improved systems will demonstrate greater stability in their recovery behavior (higher Recovery Consistency) compared to current models, suggesting a transition from purely imitation-based learning to world-model-based reasoning that explicitly predicts action consequences and state transitions.

  2. The use of smaller action chunks in the policy execution loop will enable more frequent state feedback, allowing the system to track evolving post-failure states more effectively, leading to finer closed-loop control during recovery.

  3. Systems trained on failure scenarios will exhibit a reduced failure-recovery transfer gap, meaning they can reliably recover from failures that occur during standard, unperturbed task execution (i.e., improving overall success rates in real-world deployment).

  4. Incorporating temporal context (the initial frame) into the policy's decision-making process will allow systems to compare the current scene against the intended task configuration, significantly aiding in identifying state deviations caused by failures and inferring what has been disrupted.

Sources

Related papers