RoboFAC: A Comprehensive Framework for Robotic Failure Analysis and Correction
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.
Rosa: I'm Rosa, and with me are Dev and Taro, guest researcher.
Dev: Today's paper: "RoboFAC: A Comprehensive Framework for Robotic Failure Analysis and Correction".
Rosa: Vision-Language-Action (VLA) models are advanced in robotic manipulation but lack structured supervision for failure diagnosis and recovery, which limits their robustness in open-world scenarios.
Dev: First, who's behind it and why it matters.
Title and authors: Dev: Moving onto the paper "RoboFAC: A Comprehensive Framework for Robotic Failure Analysis and Correction," the authors are Zewei Ye, Weifeng Lu, Minghao Ye, Tao Lin, Shuo Yang, and Junchi Yan. It’s a solid team of researchers from AI schools who seem to have put together something quite substantial here.
Rosa: They clearly have a strong foundation in the areas where these VLAs operate and where failure analysis is needed most right now. The title itself tells us that this isn't just another vision model; it’s focused on a whole framework for analyzing and correcting failures, which is much more specific than general task completion.
Taro: I noticed they mention the dataset has nine thousand four hundred forty erroneous manipulation trajectories and seventy-eight thousand six hundred twenty-three QA pairs across fifty-three scenes in both simulation and real-world settings; that sheer volume sounds like it’s going to give the model a very thorough education on what goes wrong <ref:2505.12224#pg0,9,440 erroneous manipulation trajectories and 78,623 QA pairs across 53>.
Dev: That scale is significant, especially since they intentionally varied the backgrounds, object configurations, and camera viewpoints to expose the models to realistic visual perturbations, which addresses that issue of domain shifts we were just discussing. It shows they aren't just testing in neat lab conditions.
Rosa: And what this means for us is that instead of training on perfect successes only, we're feeding the model examples of failure and telling it exactly where and how to fix those specific errors, which is a much more practical way to teach it robust behavior.
The paper's summary: Rosa: So, to summarize what the paper is really doing with "RoboFAC: A Comprehensive Framework for Robotic Failure Analysis and Correction," they are creating a large-scale dataset specifically designed around robotic failures across both simulated and real-world tasks. This dataset includes those nine thousand four hundred forty erroneous trajectories and the associated QA pairs that detail different types of errors <ref:2505.12224#pg0>.
Dev: That means they aren't just collecting data; they’re systematically categorizing what constitutes a failure into fundamental and atomic categories, breaking failures down by where they occur in the control hierarchy—task planning, motion planning, or execution control errors.
Taro: Decomposing failures into these distinct types is crucial because if the model can pinpoint whether it's a high-level planning error or just a low-level movement mistake, it can apply a much more targeted correction strategy when things go sideways.
Rosa: Right, and then they developed this specialized multimodal model based on Qwen2 point 5-VL that is specifically trained to understand these failures, analyze them, and provide corrections in natural language using all that rich supervision <ref:2505.12224#pg0>.
Dev: The training approach they took was interesting because they froze the visual encoder to keep those general visual representations intact but fully fine-tuned the merger and the LLM backbone; this allowed them to create a relatively lightweight model that could actually run with low latency.
The paper's improvements: Rosa: One of the major improvements they highlight is that their resulting RoboFAC model can perform four critical functions: failure detection, identification, locating the failure point, and providing an explanation for why it failed. It goes beyond just saying "it failed" to telling us precisely what went wrong and why.
Dev: And on top of diagnosis comes the correction part; they offer both high-level suggestions that map out the sequence of sub-tasks needed for recovery, and low-level commands which are precise movements to fix immediate physical errors. That dual approach is really smart for practical deployment.
Taro: The ability to locate the specific subtask stage where the error happened, as mentioned in their annotation design, gives us a very clear roadmap for debugging complex sequences rather than just seeing a final failed result.
Rosa: And when we look at how it compares to other models, they showed that this RoboFAC approach significantly improves failure reasoning compared to its base model, achieving an average score of seventy-nine point one zero on their benchmark, which is notably higher than GPT-4o's score of fifty-seven point four two and Gemini-two point zero's score of fifty-one point one one on the same task area.
Dev: That performance lift is impressive, but what’s really compelling for me as an engineer is that when integrated as an external supervisor in a real-world VLA control pipeline, it showed a twenty-nine point one percent relative improvement across four tasks while actually reducing latency compared to GPT4o.
Conclusion: Rosa: So, to wrap up our discussion on "RoboFAC: A Comprehensive Framework for Robotic Failure Analysis and Correction," the main point is that they've built a comprehensive system that uses a large, diverse dataset of failures to create a specialized multimodal model capable of diagnosing failure types across the control hierarchy and providing both strategic and tactical recovery suggestions.
Dev: Indeed, it seems they successfully bridge the gap between having powerful generalist models and needing reliable, fast, specialized tools for real-time robotic operation by achieving strong performance while keeping latency manageable for on-device deployment.
Taro: I think the implication here is that we are moving toward systems where a robot doesn't just execute a plan blindly but has an internal mechanism to self-diagnose and attempt recovery when the environment deviates from the expected path, which is essential for true autonomy.
Rosa: Exactly; it’s about making these complex robotic tasks more robust in open-world scenarios by giving them that structured understanding of failure. We've seen how this framework addresses visual perturbations through their dataset construction and how the model leverages that data to give actionable feedback on errors like position deviations or grasping issues.
Dev: And the low-level corrections being better than high-level ones suggests that most real-world failures are actually physical execution inaccuracies, which aligns perfectly with my concerns about loop rates and immediate control adjustments.
Taro: I think the future work should focus on extending this failure taxonomy further to cover more complex, emergent failures that aren't explicitly defined in the initial six types they used.
Rosa: That’s a good point for future research; expanding that taxonomy will definitely help make these systems even more resilient as they interact with increasingly unpredictable environments. So, that’s our take on the work presented by "RoboFAC: A Comprehensive Framework for Robotic Failure Analysis and Correction."
School of AI, Shanghai Jiao Tong University
cs.RO, cs.AI
Submitted: 2025-05-18
Updated: 2026-10-07
Code: https://github.com/MINT-SJTU/RoboFAC
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 92/100
The gist: Vision-Language-Action (VLA) models are advanced in robotic manipulation but lack structured supervision for failure diagnosis and recovery, which limits their robustness in open-world scenarios.
Key concepts
- RoboFAC Dataset
- This is a massive dataset containing 9,440 failed robotic manipulation trajectories from both simulation and real-world tasks. It is designed to expose the model to diverse visual perturbations by varying backgrounds and object configurations, ensuring it learns from a wide range of failure scenarios.
- Hierarchical Failure Taxonomy
- Failures are categorized into three levels: task planning errors, motion planning errors, and execution control errors. This structure allows the model to diagnose root causes at different levels of abstraction, moving from high-level strategy mistakes down to specific movement inaccuracies.
- RoboFAC Model
- This is a specialized multimodal model built on Qwen2.5-VL that is fine-tuned for robotic task understanding and failure reasoning. It uses a vision encoder and an LLM backbone to analyze robot videos, enabling it to detect, locate, explain failures in natural language, and suggest specific corrective actions.
- Low-level vs. High-level Correction
- The model suggests two types of fixes: high-level corrections provide the sequence of sub-tasks needed for recovery (useful for planning errors), while low-level corrections offer precise, immediate commands like specific movements or directions to fix execution errors.
Terminology
Summary
Vision-Language-Action (VLA) models are advanced in robotic manipulation but lack structured supervision for failure diagnosis and recovery, which limits their robustness in open-world scenarios. The RoboFAC framework addresses this by constructing a large-scale, failure-centric dataset and developing a specialized multimodal model to enable systematic failure diagnosis and correction.
RoboFAC Dataset Construction
The core of the framework is the RoboFAC dataset, which is described as a large-scale and diverse robotic failure analysis and correction dataset.
This dataset encompasses tasks of varying complexity in both simulated and real-world environments, intentionally varying backgrounds, object configurations, and camera viewpoints to expose the model to realistic visual perturbations.
The construction pipeline involves two stages: data collection and data annotation. For simulation data, the process involves defining an expert policy and then replacing it with a code snippet that generates an erroneous trajectory at the selected substage, causing the overall robotic task to fail.
Real-world data is collected via teleoperation using a robotic arm. The dataset includes 9,440 erroneous manipulation trajectories and 78,623 QA pairs across 53 scenes in both simulation and real-world environments.
Failure Taxonomy and Annotation
A key design principle of RoboFAC is to decompose robotic failures into a set of fundamental and atomic categories.
Failures are categorized into six types spanning different levels of the control hierarchy: task planning errors, motion planning errors, and execution control errors.
This hierarchical taxonomy captures root causes at multiple levels of abstraction. The dataset is annotated with rich, multi-dimensional supervision comprising eight question types and 78K video QA pairs. These question types are designed to evaluate a model’s ability in Task Understanding, Failure Analysis, and Failure Correction,
including:
-
Task identification
-
Task planning
-
Failure detection
-
Failure identification
-
Failure locating (identifying the subtask stage where the error occurred)
-
Failure explanation (providing a detailed reason for failure)
RoboFAC Model Development
Leveraging this dataset, the framework develops an MLLM termed the RoboFAC model, which is specialized for robotic task understanding, failure analysis, and corrective reasoning from robot videos.
The model is built on Qwen2.5-VL [39], consisting of an LLM backbone, a vision encoder, and an MLP-based vision-language merger. The training strategy involves freezing the visual encoder to preserve general visual representations while fully fine-tuning the merger and LLM backbone. This approach enables a relatively lightweight open-source model to match–and even surpass–general-purpose large models such as GPT-4o, while supporting low-latency and on-device deployment.
Failure Analysis and Correction Capabilities
The RoboFAC model performs four critical functions:
-
Failure detection: Determining whether the task was successfully completed.
-
Failure identification: Determining the type of failure that occurred (e.g.,
Position deviation,
Grasping error
). -
Failure locating: Identifying the specific subtask stage where the error happened.
-
Failure explanation: Providing a detailed explanation for why the task failed in natural language, using rich language instead of specific numerical values for reasons and high-level corrections.
Furthermore, the model provides two types of corrective suggestions:
-
High-level correction: Offering explicit guidance by specifying
the sequence of sub-tasks the model should execute to recover from the failure,
valuable for errors in task planning. -
Low-level correction: Providing
precise low-level commands for error correction and recovery,
such as suggesting specific movements or directions (Move the robot arm backward then move the robot arm to the left to align with the target object
).
Experimental Validation
Experiments demonstrate that RoboFAC significantly improves failure reasoning compared to its base model, achieving an average score of 79.10 on a benchmark, significantly surpassing GPT-4o (57.42) and Gemini-2.0 (51.11).
When integrated as an external supervisor in a real-world VLA control pipeline, RoboFAC yields a 29.1% relative improvement across four tasks while significantly reducing latency relative to GPT4o,
enabling more responsive and practical deployment in real robotic systems. Low-level corrections consistently outperform high-level corrections across correction rounds, indicating that most real-world failures arise from execution-level inaccuracies. The framework successfully demonstrates strong sim-to-real transfer capability, maintaining performance comparable to simulation across most task dimensions.
The gist: RoboFAC enables systematic failure diagnosis and recovery by constructing a large-scale hierarchical dataset and developing a specialized multimodal model that outperforms generalist models like GPT-4o in failure analysis accuracy and provides low-latency, effective correction for real-world VLA systems.
Improvements for AI systems
Based on the RoboFAC framework, here are specific improvements for AI systems and what those improved systems can achieve:
-
Improve robustness in open-world robotic manipulation by enabling systematic failure diagnosis and recovery.
-
Develop lightweight, deployable multimodal models (like the RoboFAC model) specialized for task understanding, failure analysis, and correction that can be used as external supervisors in real-world VLA control pipelines.
-
Enhance the ability of VLA models to generalize across diverse visual perturbations by training them on a large-scale dataset comprising 9,440 erroneous manipulation trajectories and 78,623 QA pairs across varied scenes and environments (simulated and real-world).
-
Enable fine-grained control over recovery strategies by providing both high-level (task sequence suggestion) and low-level (precise directional/magnitude commands) corrective feedback, which significantly improves the success rate of recovery in real-world manipulation tasks compared to generalist models like GPT-4o.
-
Achieve competitive performance with large proprietary models on failure reasoning tasks while maintaining low inference latency (approximately 3x faster than GPT-4o), making these critical systems suitable for real-time, on-device deployment.
-
Improve sim-to-real transfer capability by training the model on a diverse dataset that intentionally varies backgrounds, object configurations, and camera viewpoints to expose the model to realistic domain shifts.
These improved AI systems can:
-
Identify whether a robotic task was successfully completed (Failure Detection).
-
Diagnose the specific type of failure that occurred (Failure Identification), categorized into six atomic types: Task Planning Errors, Motion Planning Errors, and Execution Control Errors (e.g., Position Deviation, Grasping Error).
-
Locate the exact subtask stage where the error happened (Failure Locating), pinpointing the root cause within a complex sequence of actions.
-
Provide detailed explanations for task failure by analyzing temporal errors (Timing Error) or physical misalignments (Position Deviation).
-
Generate actionable, executable recovery plans:
List of sub-tasks to perform next (High-level Correction), or precise directional and movement commands for the end-effector to fix immediate misalignment (Low-level Correction).
Sources
- RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
- OpenVLA: An Open-Source Vision-Language-Action Model
- $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control
- ConLA: Contrastive Latent Action Learning from Human Videos for Robotic Manipulation
- Evo-1: Lightweight Vision-Language-Action Model with Preserved Semantic Alignment
- $\pi^{*}_{0.6}$: a VLA That Learns From Experience
- GR00T N1: An Open Foundation Model for Generalist Humanoid Robots
- EvoVLA: Self-Evolving Vision-Language-Action Model
- STI-Bench: Are MLLMs Ready for Precise Spatial-Temporal World Understanding?
- PALM: Progress-Aware Policy Learning via Affordance Reasoning for Long-Horizon Robotic Manipulation
- RaC: Robot Learning for Long-Horizon Tasks by Scaling Recovery and Correction
- OneTwoVLA: A Unified Vision-Language-Action Model with Adaptive Reasoning
- Phoenix: A Motion-based Self-Reflection Framework for Fine-grained Robotic Action Correction
- A Self-Correcting Vision-Language-Action Model for Fast and Slow System Manipulation
- REFLECT: Summarizing Robot Experiences for Failure Explanation and Correction
- AIC MLLM: Autonomous Interactive Correction MLLM for Robust Robotic Manipulation
- Automating Robot Failure Recovery Using Vision-Language Models With Optimized Prompts
- Large Language Models as General Pattern Machines
- Real-Time Anomaly Detection and Reactive Planning with Large Language Models
- AHA: A Vision-Language-Model for Detecting and Reasoning Over Failures in Robotic Manipulation
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving