Recova: Agent-Guided Failure Recovery for Autonomous Robotic Manipulation

arXiv:2610.01178 · cs.RO · Submitted 2026-10-01 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: Today's paper: "Recova: Agent-Guided Failure Recovery for Autonomous Robotic Manipulation".

Dev: Manipulation failures can leave scenes from which a task policy cannot recover,

Rosa: First, who's behind it and why it matters.

Paper summary: Rosa: So, to recap, we're discussing "Recova: Agent-Guided Failure Recovery for Autonomous Robotic Manipulation," which essentially argues that standard manipulation policies struggle when mistakes happen in real life because they lack built-in ways to fix a scene after a failure. The paper claims the thesis is that by using an agent guided framework, you can jointly develop task execution and recovery skills within a digital twin, and then fine-tune those capabilities using real-world experience for better performance.

Dev: Exactly; the core claim is that this system decouples the task execution from scene recovery so they can learn in parallel on data tailored to their specific needs, which means each skill accumulates independently of any single task failure. The importance lies in how it addresses the long tail of unusual configurations that standard training data often misses, which otherwise leave policies unable to continue after a failed attempt.

Taro: I see how that separation is important for generalization; if the recovery skills can be learned robustly, they should apply across different types of tasks, not just one specific sequence of actions. The paper claims this architecture allows the system to develop a reusable skill library from failures rather than just learning a single successful path.

Rosa: That's what makes it matter for practical robotics; if we can create these robust recovery programs, robots can handle unforeseen environmental changes or simple slips without needing complete re-planning or human input every time. It suggests that failure isn't just an error to be debugged, but a source of new knowledge for the robot.

Dev: From my perspective as an engineer, the framework is significant because it formalizes how we can integrate simulation and reality; they use a coding agent in the digital twin to explore failures and then train recovery rollouts that form a dedicated dataset for the recovery policy. This structured approach gives us a way to systematically generate high-quality failure data for training.

Taro: That structured data generation is key, because it moves beyond just collecting successful trajectories; it's about explicitly creating the scenarios where the robot needs to perform complex recovery maneuvers, which is where autonomy really tests its limits.

Rosa: It’s exciting because they are showing how agent-guided learning can bridge the gap between perfect simulation and messy real-world execution through this iterative refinement process involving both task and recovery datasets.

Dev: And when you look at the results mentioned, they show that this method can significantly improve mean task success, citing a gain of fifty-three point seven percentage points to reach seventy-seven point five percent across their evaluations on two simulation benchmarks and four real-robot tasks.

Conclusion: Rosa: So, thinking about "Recova: Agent-Guided Failure Recovery for Autonomous Robotic Manipulation," the authors are Isabella Liu, An-Chieh Cheng, Johan Bjorck, Zhiding Yu, Hongxu Yin, Jan Kautz, Linxi Fan, Yuke Zhu and Sifei Liu from UC San Diego and UT Austin. The paper emphasizes that this is a system where failure recovery is not an afterthought but a core part of the manipulation process.

Dev: They are showing that by implementing this agent-guided framework, we can create robots that are much better at handling unexpected situations because they learn to restore a workable scene after execution stops. In simple terms, Recova means the robot learns how to fix itself when it gets stuck during a task.

Taro: The implication for the wider world is that this could mean deploying robots in environments far more complex than highly controlled labs where things are constantly changing and unpredictable, because they would have the capability to maintain operation autonomously.

Rosa: Precisely; it moves us closer to having robotic systems that are genuinely resilient, capable of continuing work even when things go wrong in a dynamic setting. It suggests we need to focus on building these complementary skills for manipulation tasks.

Dev: And from an engineering standpoint, the takeaway is that integrating simulation exploration with real-world verification allows us to create policies that are much more reliable when they hit the physical world, which is crucial for any practical deployment scenario.

Taro: I think this work suggests a direction where autonomy research needs to heavily focus on creating these agentic mechanisms for intelligent self-correction rather than just optimizing the initial success rate of a single attempt.

Rosa: It’s about making the robot smarter about its own failures, turning those moments into reusable skills, which is what makes this paper so compelling for anyone interested in field robotics.

Isabella Liu, An-Chieh Cheng, Johan Bjorck, Zhiding Yu, Hongxu Yin, Jan Kautz, Linxi Fan

University of California, San Diego 2 University of Texas at Austin 3 NVIDIA

cs.RO

Submitted: 2026-10-01

Updated: 2026-10-01

Comments: Project page: https://www.liuisabella.com/Recova

Project page: https://www.liuisabella.com/Recova

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 92/100

The gist: Manipulation failures can leave scenes from which a task policy cannot recover, and Recova presents an agent-guided framework that jointly develops task execution and recovery in a reconstructed

Key concepts

Digital Twin
A virtual replica of the physical workstation built from real robot recordings (like camera views and trajectories) using a simulation environment like MuJoCo. This twin is used for initial training and exploring failures before deployment.
Decoupled Policies
The framework separates task execution policies from scene recovery policies. This allows each policy to specialize in its specific role—completing the task or restoring a scene—leading to more focused and efficient learning for both capabilities.
DAgger-style Data Collection
A method of repeatedly collecting data by having the agent explore failures in the twin. During each round, policies are fixed while new trajectories from robot execution and human demonstrations are gathered, which is then used to fine-tune the respective policies.

Terminology

Summary

Manipulation failures can leave scenes from which a task policy cannot recover, and Recova presents an agent-guided framework that jointly develops task execution and recovery in a reconstructed digital twin, then verifies and refines both through real-world experience.

The gist

Recova is an agent-guided framework that develops task execution and scene recovery in a reconstructed digital twin, then continually improves both through real-world experience.

How it works

Recova decouples task execution from scene recovery and connects the two through an agent that coordinates simulation, deployment, and learning. This separation trains each capability on the data suited to its role and lets recovery skills accumulate independently of any single task. The framework involves several coordinated components:

  1. A coding agent builds a digital twin of the workstation from real-robot recordings (camera observations, robot trajectories, and calibration) using a MuJoCo scene.

  2. The agent explores failures in the twin; upon failure, it diagnoses the issue and develops corrective behaviors in two forms: a program for a skill library and recovery rollouts that form the recovery dataset.

  3. Successful task rollouts initialize separate task and recovery policies, while programs capture reusable recovery strategies.

  4. During real-world deployment, a vision-language monitor tracks progress through four steps: Monitor (checking completion), Detect (deciding if intervention is needed and naming a recovery instruction z), Decide (triggering autonomous or human intervention based on the situation), and Verify (checking scene restoration).

Learning Across DAgger Rounds

Recova refines the task and recovery policies through repeated rounds of DAgger-style data collection. During each round, both policies remain fixed while the agent collects trajectories from robot execution and human demonstrations. These collected trajectories are divided into four groups: Sk (Success), Ek (Expand, containing human task demonstrations), Rk (Recover, containing human recovery demonstrations labeled with a recovery instruction z), and Ak (Autonomous recoveries that pass the scene-restoration check). At the end of each round, successful task rollouts and human task demonstrations are added to the task dataset, and human recovery demonstrations and verified autonomous recoveries are added to the recovery dataset. Each policy is then fine-tuned on its updated dataset using its standard supervised objective.

Real-Robot Rollout Loop

The real-robot rollout loop closes the gap between simulation and reality. The task policy runs autonomously, Recova intervenes only when needed, and the resulting experience becomes training data that further improves both policies. The monitor implements four steps:

  1. Monitor: A requirement query converts the task instruction into an explicit completion condition (Cg).

  2. Detect: An intervention query compares recent frames with the initial scene to decide whether to intervene, reporting whether the scene is intact and naming a recovery instruction z if needed.

  3. Decide: If stalled progress occurs in an intact scene, a human operator provides a task demonstration; otherwise, the recovery policy executes z if available or requests human help.

  4. Verify: A restoration query uses initial images to check whether the scene is workable again, allowing the robot to resume execution after autonomous recovery.

Policy Refinement and Skill Expansion

The framework allows for specialized learning by keeping task and recovery datasets separate, enabling each policy to learn its distinct role: completing the task or restoring a scene from which task execution can resume. The recovery skill set also expands as the robot encounters new failures; when no known instruction applies, the intervention query proposes a new instruction based on human demonstration. This demonstration provides the first training example for the new skill, and it is added to K and guides the development of recovery programs in the digital twin. After training, this instruction is added to the recovery policy’s registry, making it available for autonomous execution in later rounds.

Performance Results

Recova achieves superior performance across benchmarks. On simulation benchmarks, Recova exceeds its strongest baseline by 7.1 percentage points on LIBERO-Pro and 26.9 points on MolmoSpaces. On real-robot tasks, DAgger fine-tuning more than triples the base task policy’s mean success (from 23.8% to 77.5%), and enabling recovery skills at deployment raises it further to 87.5%, the best result on every task across four real-robot settings. Across four collection rounds on one task, observed human intervention falls from 87.5% to 0%. This demonstrates how agent-guided recovery turns failures into reusable capabilities, improving robustness while progressively reducing human intervention.

Conclusion

Recova establishes scene recovery as a learnable capability that complements task execution and shows that agents can turn failures into reusable skills through an agent-guided real-to-sim-to-real framework.

Improvements for AI systems

Here are the specific improvements to AI systems based on the Recova framework, and what those improved systems can achieve:


) 1. Development of a Joint Task-and-Recovery Policy Architecture:

The core improvement is decoupling task execution from scene recovery by training them as separate policies, coordinated by an agent.

The improved system will feature a dedicated Recovery Policy trained specifically on restorative actions (e.g., pushing a stuck ring down) and a Task Policy trained on achieving the primary goal (e.g., stacking rings). This specialization allows each policy to become highly proficient in its domain, leading to higher mean success rates compared to monolithic VLA or code-as-policy models.

) 2. Agent-Guided Digital Twin Exploration (Sim-to-Real Bridging):

The system will incorporate a coding agent that uses a reconstructed digital twin (MuJoCo scene based on real camera/trajectory data) to proactively explore failure modes safely.

The improved system can autonomously discover novel, complex recovery skills in simulation without risking physical hardware damage or requiring frequent manual resets. The agent generates recovery rollouts and seeds a reusable skill library (like the set of learned recovery programs) that is ready for deployment.

) 3. A Robust, Parallel Deployment-and-Learning Loop:

The system will operate on a parallel DAgger data collection harness across multiple real-robot stations, coordinating the policies' refinement concurrently.

The improved system can scale data collection significantly by utilizing multiple physical robots simultaneously. This parallel loop allows the Task Policy and Recovery Policy to learn from diverse failure scenarios concurrently, leading to faster convergence and more generalized recovery skills that are robust across different task variations (e.g., across six LIBERO-Pro settings).

) 4. Dynamic Human-in-the-Loop Intervention:

The system will employ a sophisticated Vision-Language Monitor that intelligently decides when autonomous recovery fails or is insufficient, triggering targeted human intervention only when necessary.

The improved system can achieve near-zero human intervention rates on mastered tasks (falling to 0% in the study). When failure occurs, the monitor doesn't just call for help; it diagnoses the exact failure mode and proposes a specific recovery instruction from a learned registry, ensuring that human demonstrations are targeted and immediately contribute to expanding the autonomous skill set rather than just fixing a single mistake.

) 5. Scalable Skill Library Management:

The system will maintain a comprehensive, version-controlled code-as-policy library for recovery skills, allowing learned programs to be reused across different tasks.

The improved system can achieve true reusability of corrective behaviors. A skill like Realign a ring on the peg rim developed for one task can be instantly invoked by any other task that encounters that specific failure pattern, drastically reducing the need to relearn solutions for repetitive problems.

) 6. Continuous Skill Expansion Guided by Failure Feedback:

The system will incorporate a mechanism where newly encountered failures automatically seed the development of new recovery skills and guide human demonstrations.

The improved system gains resilience against unseen failure modes. When it encounters a configuration it doesn't know how to handle, the agent proposes a new recovery instruction, which is then refined by human input into a permanent skill for that policy, ensuring the robot’s recovery capabilities grow continuously throughout its operational life.

Sources

Related papers