Instrumentation for Imitation Learning: Enhancing Training Datasets for Clothes Hanger Insertion
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.
Rosa: Today's paper: "Instrumentation for Imitation Learning".
Dev: Instrumentation for imitation learning provides invaluable state information that enables efficient learning for robotic manipulation, as demonstrated in this study focusing on clothes hanger insertion.
Rosa: First, who's behind it and why it matters.
Title and authors: Rosa: To kick things off, let's talk about the title and the people behind this work, which is "Instrumentation for Imitation Learning: Enhancing Training Datasets for Clothes Hanger Insertion." I want to explain in simple terms what that actually means for us.
Dev: From my perspective as someone who deals with control loops, I think the title immediately signals that they’re focusing on how incorporating sensor data into the learning process changes the outcome of a robotic task.
Taro: It sounds like they're proposing a way to enrich the training data for imitation learning by using sensors attached to objects, specifically for tasks involving manipulating clothes hangers.
Rosa: Precisely, Taro; it means they are looking at how providing an AI with direct readings from physical sensors on the object helps it learn better than just relying on what its cameras see.
Dev: So, rather than training the AI only on images and joint positions, they're including information about how those objects are physically interacting with each other through sensors.
Taro: That suggests they’re trying to give the model a richer understanding of the physical situation during the learning phase, which is crucial for complex movements like insertion.
Rosa: I think that’s right; it moves the focus from purely visual perception to a more comprehensive state representation that includes physical interaction data.
Dev: It's interesting because it builds on previous work where sensors have been used for state estimation in cloth, but this paper seems to be pushing that concept into the realm of garments without any sensors at all.
Taro: Indeed, the core idea here is using additional state information during learning and then deploying without it, relying only on what's available in the field.
Rosa: So it's about having a privileged piece of information during training that can help the policy figure out complex sequences more effectively than just raw visual input.
Dev: That makes sense; if you give the AI a hint about whether a sleeve is covering something, it saves it from trying to execute a move that will surely fail.
The paper's summary: Rosa: Moving on to the actual summary of "Instrumentation for Imitation Learning: Enhancing Training Datasets for Clothes Hanger Insertion," the authors explain their setup and what they did in detail.
Dev: They set up a specific scenario where a robot has to insert a hanger into an open T-shirt collar, lifting it, and then hanging it off both shoulders. This task was chosen because the rigid nature of the hanger works well with sensor integration.
Taro: So, they defined success very clearly: the autonomous execution of all these stages concluding with the T-shirt hanging only from the clothes hanger on both shoulders.
Rosa: They used one hundred eighty teleoperated demonstrations to train diffusion policies, and they trained two versions: one with access to instrumentation data and another without it <ref:2605.23847#pg0>.
Dev: The key result they reported is that the policies leveraging instrumentation outperform vision-only counterparts by fourteen to twenty-five percentage points and show greater task awareness <ref:2605.23847#pg0,policies leveraging instrumentation outperform vision-only counterparts by 14>.
Taro: And the paper notes something important about their learning process; a black-box imitation learning policy learns to prioritize those instrumentation signals without explicit guidance from the researchers.
Rosa: That’s a key takeaway: the AI figures out which sensor inputs are most valuable for achieving the goal just by observing how they affect performance.
Dev: It also details the hardware setup, mentioning four TCRT5000 reflective infrared sensors integrated into the hanger that detect when it's covered by cloth compared to when it's uncovered.
Taro: And in terms of methodology, they used a Diffusion Policy architecture with a ResNet18 vision backbone and a CNN-based noise prediction network to handle the action predictions.
Rosa: They trained all these models for one hundred thousand steps on an NVIDIA RTXfour thousand ninety GPU, which took about eleven hours just for training <ref:2605.23847#pg2,models for 100,000 steps>.
Dev: And they also mentioned that they created an enhanced dataset by taking successful rollouts from the instrumented policy to train a vision-only policy, which is quite a sophisticated training technique.
The paper's improvements: Rosa: Now let's discuss what the paper suggests for future work and how they think this approach can be improved in practice, moving beyond just the initial results.
Dev: They suggest developing a "soft sensor" model, which would use camera images to predict sensor values, which could then serve as learned instrumentation inputs for deployment where physical sensors aren't present.
Taro: I think that’s smart because it tackles the deployment hurdle head-on; if we can learn to simulate what the sensors are doing based on vision, we can make these policies more versatile.
Rosa: They also propose pretraining the ResNet18 model to predict those sensor values from camera images, essentially using vision as a backbone for that prediction task.
Dev: That would mean shifting the focus onto training a model that learns the relationship between what we see and what those physical sensors are reporting, which is a significant shift in how we think about sensory input.
Taro: I also noticed they flag that the performance gap between instrumented and vision-only policies might increase if there are less constraints on the task, like when dealing with more varied object types or different T-shirt positions.
Rosa: That limitation is important; it tells us we can't just assume this technique will work perfectly across every single real-world scenario without some kind of adaptation for those variations.
Dev: So, the implication is that while instrumentation helps a lot in controlled settings, generalizing that knowledge to unstructured environments will require more sophisticated modeling to handle those unexpected variations.
Conclusion: Rosa: Wrapping up our discussion on "Instrumentation for Imitation Learning: Enhancing Training Datasets for Clothes Hanger Insertion," we need to summarize the main implications before we move on.
Dev: In short, the paper demonstrates that by incorporating object-specific sensor information into imitation learning training data, we can significantly boost performance over purely vision-based methods in manipulation tasks.
Taro: It confirms that for complex physical manipulation, state information derived from interaction sensors gives the AI a much better understanding of what's happening than just looking at pixels.
Rosa: And they show that even a black-box learning policy can pick up on the importance of these sensor signals without being explicitly told to prioritize them during training.
Dev: From an engineering standpoint, this means we might see better performance in systems where sensors are physically limited, provided we can learn to use those learned representations effectively.
Taro: I think the biggest impact is that it paves the way for creating more capable manipulation policies that can handle real-world complexity by understanding the physical state of objects better.
Rosa: So, this study on "Instrumentation for Imitation Learning: Enhancing Training Datasets for Clothes Hanger Insertion" gives us a solid foundation to think about how to build smarter, more aware robotic systems.
Ghent University
cs.RO
Submitted: 2026-05-22
Updated: 2026-05-22
Comments: Accepted for presentation at ICRA2026
DOI: 10.1109/ICRA57385.2026.11696878
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
Importance score: 79/100
The gist: Instrumentation for imitation learning provides invaluable state information that enables efficient learning for robotic manipulation, as demonstrated in this study focusing on clothes hanger
Key concepts
- Instrumentation
- This refers to adding specific physical sensors (TCRT5000 IR sensors) to the clothes hanger. These sensors detect changes in infrared light reflection when the hanger is covered by a T-shirt, providing crucial state information about the object's position that cameras alone might miss.
- Vision-Only Policies
- These are imitation learning models trained only on visual data from cameras (like RealSense and Zed 2i). They attempt to learn the task solely through what they see. The research found these policies often fail because they lack awareness of subtle physical states, like a T-shirt being folded.
- Diffusion Policy Architecture
- This is the specific type of deep learning model used for training. It uses a ResNet18 vision backbone and a separate network to predict noise. The policy takes camera images and a state vector as input to determine the next set of robot actions, predicting 16 actions at 10 Hz.
- Task Awareness
- This measures how well the AI understands the overall goal of inserting a hanger. Instrumenting policies demonstrated higher task awareness because they used sensor data to recognize when a step failed or succeeded, leading them to adjust their behavior more intelligently than vision-only models.
Terminology
Summary
Instrumentation for imitation learning provides invaluable state information that enables efficient learning for robotic manipulation, as demonstrated in this study focusing on clothes hanger insertion. The core finding is that policies leveraging instrumentation outperform vision-only counterparts by 14–25 %pt and exhibit greater task awareness, with a black-box imitation learning policy learning to prioritize these signals without explicit guidance.
Task Description and Setup
The research focuses on the challenging task of clothes hanger insertion, chosen as a testbed because the rigid clothes hanger lends itself well to sensor integration. The task involves inserting the first leg of the hanger into an open collar of a T-shirt, lifting it until it is in position for the second insertion, and finally hanging only from the clothes hanger. To alleviate data requirements, constraints were applied: both objects are held by robot arms, only one T-shirt and one hanger are used at a time, and task success is defined as the autonomous execution of these stages concluding with the T-shirt hanging only from the clothes hanger on both shoulders.
Instrumentation Hardware
The instrumented clothes hanger is equipped with four TCRT5000 reflective infrared (IR) range sensors. These sensors function by detecting a large change in photocurrent when the IR light, emitted by an integrated LED, reflects off the T-shirt covering the hanger compared to when it is uncovered. Data from these sensors are communicated to a workstation via Bluetooth Low Energy (BLE). The robot setup involves two UR5e collaborative robot arms; one holds the T-shirt with a RealSense D405 camera for close-up views, and an externally mounted Zed 2i camera provides a view of the entire scene.
Policy Architecture and Training
The policies trained in this work utilize a Diffusion Policy architecture featuring a ResNet18 vision backbone and a CNN-based noise prediction network. The model predicts the joint positions of the right UR5e and the gripper opening of the left UR5e, given an observation consisting of 720x500x3 images from both wrist and scene cameras, along with a state vector. Policies are denoted as “instrumented” policies (Πt instr) if they include the clothes hanger sensor data in the state vector, and “vision-only” policies (Πt vis) otherwise. The models predict a chunk of 16 actions at a control frequency of 10 Hz.
Data Collection Protocols
Three types of datasets were collected:
-
Human demonstrations via teleoperation: A total of 180 human demonstrations were collected, divided into subsets defined by the number of Type I (full task), Type II (missed first insertion), and Type III (first insertion only) episodes.
-
Policy rollouts for evaluation metrics: All policy rollouts used for success rates are of Type I, given at least one minute to complete an insertion.
-
Policy rollouts for enhanced dataset: To improve vision-only policies, successful evaluation rollouts from an instrumented policy (Π180 instr) were added to the teleoperation data, forming the enhanced set of 237 demonstrations for training a better vision-only policy (Π237 vis+).
Key Findings and Failure Modes
The experimental end-to-end success rates show that instrumented policies consistently outperform their vision-only counterparts by 14 to 25 %pt.
Furthermore, the study demonstrated that rollouts from such an “expert” policy can be used as extra training data to enhance the performance of a “student” policy that cannot rely on instrumentation.
Failure mode analysis revealed distinct behaviors:
A vision-only policy often executes robot movements without concern for the state of the T-shirt, resulting in a dropped T-shirt.
An instrumented policy relies more strongly on the sensor data than the visual feed. After a failed first insertion, the policy realises it cannot continue.
Crucially, for Π50 instr policies, the Drop failures of Π50 instr were caused by the outer sensor of the second hanger leg being covered by a folded sleeve of the T-shirt,
illustrating how strongly they rely on instrumentation data. This showed that a black-box IL policy can recognise the importance of instrumentation signals without explicit guidance.
Future Work Directions
The paper suggests several avenues for future research, including developing a soft sensor
model to predict sensor values from camera images for use as instrumentation inputs, pretraining a ResNet18 model to predict sensor values given camera images to serve as the vision backbone, and training the network to explicitly predict both robot actions and sensor data. These approaches aim to capture essential information about the goal and task progression more effectively. Additionally, it is noted that the performance gap between instrumented and vision-only policies may increase with less constraints on the task,
such as increased variation in object types or T-shirt positions.
Improvements for AI systems
Here are the specific improvements that can be made to AI systems based on this research, along with what those improved systems can achieve:
-
The implementation of a state-of-the-art Diffusion Policy architecture for robotic manipulation that incorporates heterogeneous sensor modalities (visual input and dedicated instrumented signals).
-
The development of a novel training paradigm where the policy learns to intrinsically prioritize and leverage privileged instrumentation data signals without explicit human guidance or reward shaping for that data stream.
-
The creation of an
Expert Policy Rollout Augmentation
technique: using successful rollouts from an instrumented expert policy to significantly boost the performance of a vision-only student policy, enabling vision-only models to achieve near-expert levels by learning from high-quality, instrumented demonstrations without requiring the hardware itself for deployment. -
The implementation of a
Soft Sensor
model: training an additional specialized model to predict sensor values from camera images, which can then be used as learned instrumentation inputs for the primary action policy during inference, enabling deployment in environments lacking physical sensors.
This improved AI system (a Diffusion Policy) can achieve the following specific capabilities:
-
It can perform complex, multi-stage manipulation tasks like clothes hanger insertion with significantly higher success rates (up to 25% improvement over vision-only methods) by effectively fusing visual perception with internal state information derived from object interaction sensors.
-
It can exhibit superior task awareness compared to purely vision-based policies, demonstrating a better understanding of the physical state of the object (e.g., recognizing when a sleeve covers a sensor or determining readiness for the next insertion step).
-
The resulting system can be deployed in real-world, uninstrumented settings (like any household) while maintaining performance comparable to systems trained with full instrumentation, by using learned representations and augmented datasets derived from instrumented data.
-
It can rapidly learn complex manipulation skills by leveraging high-quality demonstrations from instrumented experts to create robust vision-only models, effectively bridging the gap between controlled lab environments and unstructured real-world deployment.
Abstract
Large behaviour models have transformed the field of robotic manipulation, but prohibitive data requirements have thus far prevented a revolution similar to vision language models. We believe that instrumentation, i.e. sensor integration in objects, can provide invaluable state information and enable efficient learning for robotic manipulation. In this paper, we present instrumented imitation learning of clothes hanger insertion. Using 180 teleoperated demonstrations, we train diffusion policies with and without access to instrumentation data. Results show that policies leveraging instrumentation outperform vision-only counterparts by 14-25 %pt and exhibit greater task awareness. Crucially, a black-box imitation learning policy learns to prioritise instrumentation signals without explicit guidance. In addition, enhancing the teleoperation dataset with rollouts from an instrumented expert policy, enables a vision-only student policy to achieve performance comparable to the instrumented expert, thereby surpassing the original vision-only policy. These findings establish instrumentation as a promising strategy to enhance imitation learning for robotic manipulation. Datasets are available on Zenodo.
Sources
- A Careful Examination of Large Behavior Models for Multitask Dexterous Manipulation
- $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control
- DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset
- BridgeData V2: A Dataset for Robot Learning at Scale
- Interactive Imitation Learning in Robotics: A Survey
- Real-Time Operator Takeover for Visuomotor Diffusion Policy Training
- HG-DAgger: Interactive Imitation Learning with Human Experts
- RACER: Rich Language-Guided Failure Recovery Policies for Imitation Learning
- Robot Learning as an Empirical Science: Best Practices for Policy Evaluation
- Instrumentation for Better Demonstrations: A Case Study
- VITaL Pretraining: Visuo-Tactile Pretraining for Tactile and Non-Tactile Manipulation Policies
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving