iTeach: In the Wild Interactive Teaching for Failure-Driven Adaptation of Robot Perception

arXiv:2410.09072 · cs.RO · Submitted 2024-10-01 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: I'm Rosa, and with me are Dev and Taro, guest researcher.

Dev: Today's paper: "iTeach: In the Wild Interactive Teaching for Failure-Driven Adaptation of Robot Perception".

Rosa: iTeach, a failure-driven interactive teaching framework for deployment-time adaptation of robot perception, enables a co-located human to observe model predictions during deployment, identify failure cases,

Dev: First, who's behind it and why it matters.

Title and authors: Dev: Moving on from the concept, I want to talk about the core idea of "iTeach: In the Wild Interactive Teaching for Failure-Driven Adaptation of Robot Perception." It suggests that instead of just collecting tons of general data offline, we can get targeted supervision when a model actually fails during deployment.

Rosa: That’s a big shift in how we think about adapting these robots; it moves the learning process right into the moment of action.

Taro: What I find interesting is that they aren't just passively recording failures; they have this active feedback loop where the human physically interacts with the robot to expose those specific errors.

Dev: And that interaction is captured through an RGB-D sequence, which gives us a rich context for what caused the failure.

The paper's summary: Rosa: To sum up what the authors are proposing in "iTeach: In the Wild Interactive Teaching for Failure-Driven Adaptation of Robot Perception," they are building a framework where a robot's perception model can be updated on deployment by having a human observe its errors and provide targeted corrections.

Dev: They specifically use mixed reality headsets, like the HoloLens two to overlay the model's predictions directly in front of the user so they can see exactly where things are going wrong.

Taro: The method hinges on generating failure-driven samples—like when an object is missed or segmentation is off—and then using a specific labeling strategy called Few-Shot SemiSupervised, or FS3, to turn those short interactions into dense training supervision.

Rosa: That FS3 strategy is pretty clever because it only annotates the final frame of that interaction sequence using hands-free eye-gaze and voice commands, and then they propagate those labels across the whole video to get dense supervision.

Dev: So it minimizes annotation effort while still getting a lot of useful data from a very short human–object interaction.

The paper's improvements: Rosa: The paper points out several key advantages in their approach, starting with how it tackles the challenge of real-world adaptation that traditional methods struggle with.

Dev: They highlight that this method addresses the limitation where large-scale dataset efforts, like DROID or Open X-Embodiment, often lack the mechanism to resolve deploymentspecific failure modes when things go wrong.

Taro: I think their biggest contribution is using iterative fine-tuning: starting from a pretrained model M0, they refine it through iterations like M1 = Finetune(M0) and then redeploying it to collect new failure data, leading to M2, and so on.

Rosa: That iterative refinement process allows the perception model to progressively adapt its understanding of those specific deployment errors over time.

Dev: And they show that this method leads to improved performance on Unseen Object Instance Segmentation, starting from a pre-trained MSMFormer model, which is what we use for evaluation.

Conclusion: Rosa: So, wrapping up the discussion on "iTeach: In the Wild Interactive Teaching for Failure-Driven Adaptation of Robot Perception," the main implication is that we can bridge that gap between controlled training and real-world deployment by having the robot learn directly from its mistakes in situ.

Dev: It really shows how a failure-driven data collection strategy, coupled with an iterative fine-tuning paradigm, can lead to better performance when facing out-of-distribution conditions.

Taro: I think the idea that targeted supervision generated during deployment is more practical than just scaling up massive datasets for every single corner of the real world.

Rosa: Indeed, it means we don't have to wait around for months of data collection cycles before we can deploy a model that actually handles clutter and occlusion well.

Dev: And the speed of adaptation they mention, taking about fifteen to twenty-five minutes from failure observation to redeployment, is really something worth considering for high-stakes applications.

Taro: From my view, the closed-loop learning system itself is what’s important; it creates a self-correcting pipeline that learns from its own errors in the environment.

Rosa: Exactly, and I think we should keep an eye on how this translates to more robust grasping and pick-and-place success when we move these models off the tabletop.

Jishnu Jaykumar P*, Cole Salvato†, Vinaya Bomnale, Jikai Wang, Yu Xiang

The University of Texas at Dallas

cs.RO

Submitted: 2024-10-01

Updated: 2026-09-29

Code: https://github.com/jishnujayakumar/robokit

Project page: https://irvlutd.github.io/iTeach

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 86/100

The gist: iTeach, a failure-driven interactive teaching framework for deployment-time adaptation of robot perception, enables a co-located human to observe model predictions during deployment, identify failure

Key concepts

iTeach
A failure-driven interactive teaching framework designed for deployment-time adaptation of robot perception. It allows a human observer to identify and correct model errors while the robot is in use, shifting learning into the moment of action.
Few-Shot SemiSupervised (FS3)
A labeling strategy used in iTeach that turns short human–object interactions into dense training supervision. It primarily annotates only the final frame of an interaction sequence using eye-gaze and voice commands, then propagates those labels across the entire video to create dense data.
Iterative Fine-Tuning
The process where a robot model is refined through successive iterations (M0, M1, M2). The robot refines its perception model based on failure data collected during deployment and then redeploys to collect new failure data, progressively adapting to specific errors.
Unseen Object Instance Segmentation
A performance metric used to evaluate the improved perception models. The paper shows that the iterative fine-tuning method leads to better performance on this task, starting from a pre-trained MSMFormer model.

Terminology

Summary

iTeach, a failure-driven interactive teaching framework for deployment-time adaptation of robot perception, enables a co-located human to observe model predictions during deployment, identify failure cases, and perform short human–object interaction (HumanPlay) to expose informative object configurations while recording an RGB-D sequence. To minimize annotation effort, iTeach employs a Few-Shot SemiSupervised (FS3) labeling strategy where only the final frame of a short interaction sequence is annotated using hands-free eye-gaze and voice commands, and labels are propagated across the video to produce dense supervision. The collected failure-driven samples are used for iterative fine-tuning, enabling progressive deployment-time adaptation of the perception model.

The system consists of three components: (1) a robot (Fetch mobile manipulator) that captures RGB-D observations, (2) a mixed reality (MR) headset (Microsoft HoloLens 2) that overlays perception outputs in situ and enables hands-free annotation using eye gaze and voice commands, and (3) a compute node that performs inference, dataset aggregation, and model fine-tuning.

The iTeach loop begins with failure observation: "A perception model performs inference on RGB-D observations streamed from the robot. Predictions are rendered directly in the user’s field of view through the MR interface, allowing the human to assess perception performance in situ (Fig. 3). Failures such as missed objects, incorrect segmentation boundaries, or spurious detections are identified during operation. The co-located human actively navigates the robot across diverse environments to expose informative task spaces: The iTeach loop begins with failure observation. A perception model performs inference on RGB-D observations streamed from the robot. Predictions are rendered directly in the user’s field of view through the MR interface, allowing the human to assess perception performance in situ (Fig. 3). The co-located human actively navigates the robot across diverse environments to expose informative task spaces."

When perception failures are observed, a short human–object interaction, referred to as HumanPlay, is performed: "When perception failures are observed, the human performs short human–object interaction, referred to as HumanPlay, where the human rearranges objects to reduce occlusion and produce a clean final frame while recording a short RGB-D sequence. To generate dense supervision efficiently: To minimize annotation effort, iTeach employs a FewShot SemiSupervised (FS3) labeling strategy. Only the final frame of a short RGB-D HumanPlay sequence is annotated using hands-free eye–gaze and voice commands, as illustrated in Fig. 4. Sparse point prompts are converted into bounding box proposals, and labels are propagated across the recorded sequence using SAM2 [9] (from Robokit [44]) to produce dense supervision (Fig. 5)."

The iterative fine-tuning paradigm is implemented: "After collecting failure-driven samples, iTeach refines the perception model through iterative fine-tuning. Starting from a pretrained model M0, failure cases observed during deployment are collected using HumanPlay and used to train an updated model Mt. The better model is selected and redeployed for subsequent data collection, yielding RGB-D video frames collected via HumanPlay across timesteps. This process is formalized as: M 0 → M 1 → M 2 → … where best(·) selects the model with higher validation performance."

The evaluation focuses on Unseen Object Instance Segmentation (UOIS), starting from a pre-trained MSMFormer [11] model. The method demonstrates that Using a small number of failure-driven samples, our approach significantly improves segmentation performance across diverse real-world scenes, and These improvements directly translate to higher grasping and pick–and–place success on the SceneReplica benchmark and real robotic experiments. The results show that failure-driven, co-located interactive teaching enables efficient in–the–wild adaptation of robot perception and improves downstream manipulation performance. The adaptation time from failure observation to model redeployment is approximately 15–25 minutes, including data collection, mask propagation, dataset aggregation, and fine-tuning. The final results show that Replacing the perception module with iTeach-UOIS (Pipeline 4) improves both grasping and pick–and–place success while keeping the rest of the pipeline identical to Pipeline 3. The overall conclusion is that failure-driven data collection enables efficient adaptation to in–the–wild scenarios beyond static tabletop settings. The main contributions are: We propose iTeach, a failure-driven, co-located interactive teaching framework for deployment-time adaptation of robot perception in the wild, We introduce a FewShot Semi–Supervised (FS3) labeling strategy that converts sparse hands-free annotations into dense supervision, and We develop an iterative failure–driven fine–tuning paradigm. The paper also notes a limitation: "iterative fine-tuning does not guarantee monotonic improvement.

Improvements for AI systems

Here are the specific improvements that an AI system could gain by implementing the iTeach framework, along with what those improved systems can achieve:


  1. Enhance Robustness to Out-of-Distribution (OOD) Real-World Scenarios:

In a real-world deployment, perception models fail due to clutter, occlusion, and novel objects (out-of-distribution conditions). The iTeach system improves this by using failure cases observed in the wild for iterative fine-tuning.

The improved AI system can achieve:

A significant increase in performance on Unseen Object Instance Segmentation (UOIS) starting from a pre-trained model, as evidenced by the reported gains (e.g., +50.5% improvement on metric C over baseline). This translates directly to more accurate object masks and better scene understanding in cluttered, non-synthetic environments.

  1. Enable Deployment-Time Adaptation Without Slow Offline Retraining:

Current methods require hours or days for offline data collection and retraining, making deployment reactive rather than proactive. iTeach integrates human feedback directly into the deployment loop.

The improved AI system can achieve:

Progressive adaptation of the perception model in real-time during a single session. Instead of waiting for a massive offline dataset, the model can progressively refine its understanding of specific failure modes (e.g., a particular type of occlusion or novel object) within minutes (15–25 minutes total loop time), enabling continuous, on-the-fly improvement in deployment.

  1. Achieve Data Efficiency Through Few-Shot Semi-Supervised Learning:

Traditional data annotation is expensive and time-consuming. iTeach employs a Few-Shot SemiSupervised (FS3) labeling strategy using hands-free eye-gaze and voice commands to convert sparse interactions into dense supervision across an entire video sequence.

The improved AI system can achieve:

Massive reduction in the human effort required for data labeling. A short, interactive session can yield a fully labeled training sequence spanning seconds of video, drastically lowering the cost and time associated with generating high-quality, failure-driven training data compared to traditional frame-by-frame annotation.

  1. Bridge the Gap Between Perception Accuracy and Downstream Manipulation Success:

Improvements in perception quality often do not translate directly to successful manipulation tasks because planning modules remain unchanged. iTeach explicitly tests this link by evaluating manipulation success on the SceneReplica benchmark after upgrading only the perception module (MSMFormer).

The improved AI system can achieve:

Direct, measurable improvements in physical task performance, such as higher grasping and pick-and-place success rates on complex benchmarks like SceneReplica. This confirms that perception gains are not just academic but lead to tangible increases in real-world robotic utility.

  1. Create a Closed-Loop Learning Paradigm for Perception:

The system establishes a complete feedback cycle: Perception Failure Identification → Human Observation/Interaction (HumanPlay) → Dense Label Generation (FS3) → Model Refinement (Iterative FT).

The improved AI system can achieve:

A self-correcting, adaptive perception pipeline that learns from its own mistakes in the environment. It moves beyond static training by actively seeking out and learning from the specific hard cases encountered during actual deployment.

  1. Improve Contextual Integrity of Segmentation Masks:

The reliance on SAM2 for mask propagation across a sequence is mitigated by the HumanPlay interaction, which reduces occlusion before annotation, leading to cleaner final frames and more robust mask propagation.

The improved AI system can achieve:

More contextually coherent and precise object segmentation masks that are reliable across a video sequence. This robustness is crucial for downstream tasks like grasp proposal generation and collision-free planning, as it reduces errors caused by noisy or ambiguous boundary detections.

Abstract

We present iTeach, a deployable system that lets any co-located human fix a robot's perception failures on the spot without expertise, a workstation, or offline retraining. The operator wears a mixed reality (MR) headset, sees the robot's segmentation predictions overlaid on the real scene, and corrects failures hands-free: rearranging objects (HumanPlay), annotating via gaze and voice, and triggering SAM2 backward mask propagation. Each 20 s interaction yields 150-300 densely labeled training frames; the system fine-tunes the perception model onboard, keeps the better model, and redeploys, all without leaving the deployment site. The full loop requires only an RGB-D camera, onboard GPU, and an MR headset: any mobile robot, any environment. Starting from 26.1 on cluttered real-world scenes, 45 teaching interactions (13K frames) raise segmentation to 80.7 with no catastrophic forgetting; on three standard benchmarks the model never trained on, performance improves as well. Downstream pick-and-place on SceneReplica reaches 72/100, surpassing a model-based pipeline requiring CAD models. A 12-participant user study confirms non-experts match experts on annotation accuracy (95% box IoU), speed, and task load (NASA-TLX 21/100). The framework is architecture-agnostic: any fine-tunable perception model can serve as backbone.

Sources

Related papers