FAR: Failure-Aware Retry for Test-Time Recovery and Continual Policy Improvement

arXiv:2607.01111 · cs.RO, cs.AI, cs.LG · Submitted 2026-07-01 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: Today's paper: "FAR: Failure-Aware Retry for Test-Time Recovery and Continual Policy Improvement".

Dev: Failure-Aware Retry (FAR) is a framework designed to enable robot policies to learn from previous failures at test time, adapt their behavior accordingly, and eventually complete tasks autonomously.

Rosa: First, who's behind it and why it matters.

Paper summary: Rosa: So we're diving into this paper now called "FAR: Failure-Aware Retry for Test-Time Recovery and Continual Policy Improvement." Basically, it tackles how robots can learn from their mistakes while they are actually doing a task in the real world, aiming to let them finish things on their own.

Dev: Yeah, I'm interested in what it claims regarding that self-learning aspect because for us on the engineering side, the reliability during those retries is everything. What is the main thesis of this framework?

Taro: The core idea seems to be combining two main things: using failure feedback to adapt immediately and then using small changes during a retry to explore new possibilities when things go wrong. It’s about making recovery smarter than just trying the same thing again over and over.

Rosa: Exactly, it claims that by doing this, robots can eventually complete tasks autonomously even when they run into unexpected failures in their environment. The authors propose this framework to move beyond naive retry methods that just repeat previous errors every time they fail.

Dev: From my standpoint as someone who deals with loop rates and latency, the focus on adapting behavior at test time is interesting because it suggests a faster way to correct a trajectory before things get completely out of hand. How does this adaptation actually happen in practice?

Taro: It starts by using something called Conservative Value Estimation to pinpoint which specific actions in a failed run were likely causing the failure, which is a big step toward understanding the failure-inducing behavior. Then they use that information to build preference data and update the policy based on those failures.

Rosa: That sounds like they're constructing targeted learning examples from those bad experiences, using Failure-Contrastive Preference Adaptation to steer the policy away from what didn't work before. It seems like a very focused way to teach the robot what *not* to do next time.

Dev: I see how that helps with stability, but what about exploring different routes when the first attempt fails? Does this paper suggest any way for the robot to try something structurally different during those retries?

Taro: Yes, they incorporate lightweight action perturbations during retries specifically for structural exploration to expand the policy's support for harder recovery cases. They sample a target perturbation and keep it fixed for a few steps, smoothing it out over time so it stays manageable on the actual robot.

Paper summary: Rosa: So, to summarize what we've heard about "FAR: Failure-Aware Retry for Test-Time Recovery and Continual Policy Improvement," the paper proposes a system that learns from failures during testing by using failure-aware adaptation and small action perturbations for exploration, while also building a loop to continuously improve the policy with successful recovery data.

Dev: And we've touched on how this aims to boost success rates, noting their results show average gains of seventeen point six percent in simulation and eleven point seven percent in the real world compared to standard diffusion policies <ref:2607.01111#pg0,in simulation and 11.7% in the real world>. That level of improvement over existing methods is pretty significant for us to consider on a deployment basis.

Taro: I think what really stands out is how they generate those informative recovery trajectories, which they then use as supervision for continual policy improvement, allowing the system to get better over time even with limited new interaction. The paper also shows that using a "Value-Percentile" for selecting negative samples in their failure attribution step gave them better overall performance than just a fixed threshold.

Rosa: That's fascinating because it shows that even the way they select which failures to learn from matters, suggesting there are nuanced ways to use this data. It really emphasizes that this isn't just one setting you can apply universally; the methodology itself is being tuned based on performance metrics.

Dev: When we think about the implications for actual deployment, Rosa, I have to ask how long this kind of continuous learning loop would realistically run in a real-world operational setting before needing significant human intervention to overhaul it. What are the practical limits on its sustained autonomy?

Taro: The paper focuses on making policies more robust and improving value estimation over time by leveraging successful recovery trajectories in their training loop, which is designed to gradually expand policy support while prioritizing high-value behaviors. They aim for sustained improvement by constantly feeding in these better recovery data points into the Critic Learning phase.

Rosa: That brings us right to the point of deployment, Dev; if we take this framework outside of a controlled simulation environment and put it on a physical robot, how long can we expect this continual improvement mechanism to keep working without constant re-calibration?

Paper summary: Dev: Well, because the system maintains three replay buffers—Dexp for offline demonstrations, Dsucc for successful online trajectories including recoveries, and Dfail for failure trajectories—it's designed to handle both known good data and recent bad data simultaneously. The latency of the action perturbation is kept low by using exponential smoothing on that target perturbation delta, which helps maintain stable execution even when exploring locally.

Taro: And those three buffers are key because they allow the Critic Learning phase to aggregate a diverse dataset D = Dexp Dsucc Dfail to optimize the critic models effectively, which is what drives the advantage-weighted policy update. This structured data management is what enables that long-term improvement potential.

Rosa: So, looking at the title of "FAR: Failure-Aware Retry for Test-Time Recovery and Continual Policy Improvement," it really encapsulates both aspects: fixing immediate failures during testing and building a mechanism to get better through experience over time. It's about creating a robot that isn't just reactive but actively learning from its struggles.

Dev: I agree, the title highlights the dual nature of the work; it’s not just about surviving one bad attempt, it’s about improving the whole policy incrementally across multiple sessions. The authors are clearly trying to address those limitations where standard data collection resets and doesn't utilize hard failure cases effectively.

Taro: I think a big implication here is that we might be able to deploy robots in more complex, messy real-world environments without needing constant manual reprogramming after every unexpected setback. Instead of stopping the task when it fails, the robot learns from that specific failure and adapts its strategy for the next attempt.

Rosa: That would have a huge impact on field robotics; imagine a delivery drone or a manipulator arm in a cluttered warehouse that encounters an obstacle it hasn't seen before during a retry and figures out how to navigate around it successfully. That level of autonomous resilience is what this suggests is achievable.

Dev: From an engineering standpoint, the fact that the perturbation mechanism allows for local exploration without completely destabilizing the system on real hardware is a crucial detail for us to focus on in terms of safety constraints and acceptable risk levels during those recovery steps. The exponential smoothing delta t = alpha delta t-one + (one - alpha) delta is what keeps that exploration controlled <ref:2607.01111#pg1>.

Paper summary: Taro: The paper itself points out that the authors are focusing on learning from failure states specifically to improve the policy from challenging situations while reducing costly resets and human effort, which is a direct response to the data-intensive nature of standard online improvement methods. They show how this targeted approach can be more efficient.

Rosa: It really shifts the focus from just getting a successful outcome in one go to building a robust system capable of handling inevitable failures gracefully through iterative learning. That's a significant shift in how we design these autonomous agents, isn't it?

Dev: It is, Rosa; it moves us toward policies that are inherently more resilient because they are actively incorporating failure knowledge into their future decision-making process. The paper presents a framework where the robot doesn't just recover from an error; it learns *why* the error happened and adjusts its approach accordingly for subsequent attempts.

Taro: And I think the authors' demonstration on both simulation and real-world tasks, achieving those gains of seventeen point six percent in one place and eleven point seven percent in another, validates that this combined approach actually translates into tangible performance improvements across different types of robot manipulation challenges <ref:2607.01111#pg0,both simulation and real-world>.

Rosa: So, to wrap up our discussion on "FAR: Failure-Aware Retry for Test-Time Recovery and Continual Policy Improvement," the main point is that this framework provides a systematic way for robots to learn from their mistakes during testing by combining failure adaptation and controlled exploration, leading to sustained performance gains in real-world tasks.

Dev: And we see its importance in how we approach deployment, because it suggests a path toward more robust autonomous agents that can handle unexpected issues with less reliance on constant human intervention for fine-tuning.

Taro: The implication is that we are moving toward systems where the learning process itself is more integrated into the execution loop, allowing for incremental policy refinement based on real operational data rather than just episodic successes.

Rosa: It’s exciting to think about what this means for complex tasks in robotics; it suggests a future where robots can tackle tasks with much greater inherent resilience and self-correction capabilities.

Conclusion: Rosa: So we've been talking about FAR: Failure-Aware Retry for Test-Time Recovery and Continual Policy Improvement, and now we get to wrap up by looking at what this whole thing really means for robots out there.

Dev: I agree, Rosa, it’s interesting how the authors framed the title—it really highlights that they're not just patching a single failure but building a system for long-term recovery.

Taro: Indeed, and thinking about those authors, I see why this approach is compelling; they’ve managed to integrate immediate adaptation with a mechanism for continuous learning.

Rosa: Exactly, and when you put it all together, the core of FAR is giving robots a smarter way to deal with unexpected problems during their actual work.

Dev: From my side as someone focused on the operational side, I think what's most important is understanding that this isn't just about getting one successful run; it's about building something that gets better over time.

Taro: That’s where the continual improvement loop comes in, which suggests that even if a robot encounters a new type of failure, it has a chance to learn from it and improve its performance for the next time.

Rosa: It really paints a picture where robots can handle messy real-world situations with much more inherent resilience than we see now.

Dev: And I'm curious about how long this kind of sustained improvement would actually last in a busy, industrial setting before it needs major reprogramming from a human operator.

Taro: That’s a valid question, Dev; the paper suggests that by constantly feeding successful recovery data back into the training process, the system is designed to gradually expand its abilities over time.

Rosa: So it moves us toward systems that can handle unexpected issues with much more grace than just stopping and restarting every single time.

Dev: And that grace comes from how they manage those failures during retries; the lightweight perturbations are a key detail for me, showing how they keep things stable while exploring new options.

Taro: The authors also showed that by using different methods to identify failure causes, like the Value-Percentile approach over a fixed threshold, you can get better results from that data.

Rosa: It really underscores how much nuance there is in designing these recovery systems; it’s not just one way to look at the problem.

Dev: I think the authors' focus on both test-time adaptation and offline replay buffers means they’re tackling the problem from multiple angles, which is impressive.

Taro: And that dual focus is what makes this framework so interesting for autonomy research; it bridges that gap between immediate response and long-term system growth.

Rosa: So, to sum up, FAR seems to be a very promising direction for creating robots that can learn from their struggles in a way that's both adaptive and cumulative.

Dev: And it really makes me think about the future of deployment; will we see this kind of self-improving resilience in our field robotics applications soon?

Carnegie Mellon University

cs.RO, cs.AI, cs.LG

Submitted: 2026-07-01

Updated: 2026-10-07

Comments: Accepted by CoRL 2026. Project Page: https://hoar012.github.io/FAR-Project

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 92/100

The gist: Failure-Aware Retry (FAR) is a framework designed to enable robot policies to learn from previous failures at test time, adapt their behavior accordingly, and eventually complete tasks autonomously.

Key concepts

Failure-Contrastive Preference Adaptation (FCPA)
This technique updates the robot's policy using feedback from failures. It identifies actions that led to poor outcomes and uses a preference objective, specifically a negative denoising loss, to train the policy so it avoids repeating those failed actions and prefers better ones.
Lightweight Action Perturbation
During retries, the system introduces small random changes (perturbations) to the robot's planned actions with some probability. This allows for structural exploration of different behaviors. These small changes help the policy explore new action possibilities to escape local failures while maintaining stable execution through temporal smoothing.
Failure Attribution
This process measures which specific parts of a failed trajectory caused the failure by observing changes in value estimates. Actions that cause a significant drop in value are flagged as failure-inducing behaviors, guiding the policy toward avoiding those specific unstable actions.

Terminology

Summary

Failure-Aware Retry (FAR) is a framework designed to enable robot policies to learn from previous failures at test time, adapt their behavior accordingly, and eventually complete tasks autonomously. The core contribution of this work lies in combining Failure-Contrastive Preference Adaptation with lightweight action perturbations during retries to steer the policy away from unsuccessful behaviors while encouraging local exploration.

Failure-Aware Retry Mechanism

The framework integrates two primary components to improve retry performance: (1) Failure-Contrastive Preference Adaptation (FCPA), which updates the policy using failure feedback, and (2) lightweight action perturbation, which facilitates structural exploration during retries. The process begins by using Conservative Value Estimation to identify failure-inducing actions from a failed trajectory. This involves training a Q-function and value function under an IQL-style loss to obtain more reliable estimates in out-of-distribution states. Subsequently, Failure Attribution is performed by measuring the effect of action chunks on the value change across the trajectory; chunks exhibiting a pronounced decrease in value indicate unstable or failure-inducing behavior. These identified chunks are then used to construct contrastive pairs by sampling candidates from the current policy and selecting those with high critic values as positive samples, paired with negative failure examples. Finally, Failure-Contrastive Preference Adaptation optimizes the policy using a preference objective, specifically employing the negative denoising loss for diffusion policies: "we define ldiff(a, s; θ) = Ek∈ϵ[∥ϵ − ϵθ(ak, s, k)∥2i], where ak is obtained by adding Gaussian noise to a at diffusion step k. We then optimize the pairwise preference objective Lpref = − log σβ ldiff(a−t, st; θ) − ldiff(a+t, st; θ), which encourages the model to assign lower denoising error to the preferred chunk than to the failed chunk."

Test-Time Adaptation and Exploration

The framework addresses limitations in exploration by incorporating lightweight action perturbations during retries. This component allows for structural exploration to expand the policy support for harder recovery cases. Specifically, with probability epsilonexplore, a target perturbation δ⋆ is sampled from N(0, Σ) and kept fixed for the next h steps; otherwise, it is set to zero. The applied perturbation is updated via exponential smoothing: δt = αδt−1 + (1 − α)δ⋆, where α ∈ [0, 1). This mechanism ensures that simple local exploration during retry occurs while the temporal smoothing helps maintain stable execution on real robots.

Continual Policy Improvement

The collected recovery trajectories are crucial for further policy refinement through a continual improvement loop. The system maintains three replay buffers: Dexp (offline demonstrations), Dsucc (successful online trajectories, including recovered ones), and Dfail (failure trajectories). During the Critic Learning phase, the critic models are optimized using Eq. (1) and Eq. (2) on the aggregated dataset D = Dexp ∪ Dsucc ∪ Dfail. Following critic training, an Advantage-weighted Policy Update is performed by computing the advantage A(s, a) = Qϕ(s, a) − Vψ(s), assigning weights w(s, a) = exp A(s,a)/η η to samples in Dexp ∪ Dsucc. The actor objective then optimizes the diffusion policy using an advantage-weighted denoising objective: Lπ = Ew(s, a0)∼Dexp∪Dsucc[k, ϵθ(ak, s, k)∥ϵ − ϵθ(ak, s, k)∥2i]. This loop is designed to gradually expand the policy support while prioritizing high-value behaviors, leveraging successful recovery trajectories to improve policy robustness and value estimation over time.

Experimental Validation and Ablations

Experiments on simulation and real-world tasks demonstrate that FAR substantially improves success rates. In simulations, FAR achieved average gains of 17.6% over the standard diffusion policy, while in the real world, it showed an improvement of 11.7%. The framework's effectiveness is further validated through ablation studies:

(a) Retry Algorithm:

FAR outperforms naive retry (DP-NR), with FCPA and perturbation both contributing to performance gains. The paper notes that FCPA helps the policy avoid repeating previous failures and make more effective retries, while perturbation encourages the policy to try novel actions, allowing it to reach new states and escape local optima.

(b) Failure Attribution:

Using a Value-Percentile for negative sample selection is empirically found to yield better overall performance than using a fixed threshold.

Improvements for AI systems

As a fastidious researcher, I have analyzed the proposed Failure-Aware Retry (FAR) framework. The core contribution lies in enabling robot policies to learn from failures at test time and continuously improve autonomously, significantly reducing reliance on costly human intervention and environment resets.

Here are the specific improvements that can be made to existing AI systems by implementing FAR, and what the resulting improved system can achieve:


The implementation of the Failure-Aware Retry (FAR) framework provides several high-impact improvements across three main vectors: Test-Time Robustness, Continual Policy Improvement, and Data Efficiency.

  1. Inference/Deployment Phase:

  2. Training/Learning Loop:

  3. Model Generalization & Adaptation:

  4. Inferences/Deployment Phase (Test-Time Recovery):

The improved system can perform Self-Correcting Autonomous Manipulation in real-world and simulated environments without requiring a full environment reset after every failure.

FAR enables the robot to execute a task, detect a failure state, immediately identify the action chunks responsible for that failure using conservative value estimation (IQL objective), and then update its policy within seconds using Failure-Contrastive Preference Adaptation (FCPA). This allows the system to:

  • Avoid repeating specific mistakes in subsequent retries.

  • Explore alternative actions in out-of-distribution states by injecting lightweight, temporally smoothed action perturbations during retries.

  1. Training/Learning Loop (Continual Improvement):

The improved system can transition from a static, offline policy to a dynamic, self-improving agent capable of long-term skill acquisition.

FAR integrates successful recovery trajectories into an online finetuning loop by maintaining specialized replay buffers (Dexp, Dsucc, Dfail). This allows the policy to:

  • Learn from hard negative examples (failures) and successful recovery trajectories, effectively expanding its capability boundary over time.

  • Achieve continual policy improvement with superior data efficiency compared to standard online RL methods that rely on costly environment resets.

  1. Model Generalization & Adaptation (Robustness):

The improved system can achieve robust performance across varied task complexities and initial conditions by leveraging failure-informed supervision.

FAR uses the generated recovery trajectories as informative supervision, providing a form of failure-aware regularization absent from offline training data. This allows the policy to:

  • Improve success rates substantially (e.g., 17.6% in simulation).

  • Adapt its behavior specifically around failure boundaries, leading to better generalization when encountering novel states that resemble previous failures.

  1. Data Efficiency (Resource Optimization):

The improved system can operate with significantly reduced data and computational overhead during continuous learning phases.

FAR improves data efficiency under both reset and timestep budgets by exploiting informative failure cases rather than wasting resources on redundant or easy state transitions. This means the robot can:

  • Learn from challenging failure cases efficiently, requiring fewer total environment interactions for a given level of policy improvement.

  • Maintain stable learning signals even when operating under tight environmental step constraints (e.g., in real-world deployment).

Sources

Related papers