VE2VF: Vision-Enabled to Vision-Free Distillation via Real-world Reinforcement Learning for Robust Contact-Rich Manipulation

summary

Video file (mp4)

The gist

When using reinforcement learning for contact-rich robotic manipulation, vision can provide task-relevant information that accelerates learning beyond what proprioception alone can achieve, but

In short

The VE2VF framework uses a vision-enabled teacher policy trained with human input to distill its knowledge into a vision-free student policy. This allows the student to perform robust contact-rich manipulation using only proprioceptive data like pose and force, achieving strong generalization without needing extensive visual training or augmentation.

Key concepts

Teacher-Student Distillation
This technique transfers skills from a complex, vision-enabled policy (teacher) to a simpler, vision-free policy (student). The student learns by mimicking the teacher's performance on the same data, effectively distilling the visual knowledge into a format usable without images.
Vision-Free Policy
This is the resulting agent that makes decisions using only physical measurements like end-effector pose, velocity, and force/torque. It relies entirely on proprioceptive sensing rather than visual input to successfully complete manipulation tasks.
Knowledge Distillation Loss
A specific mathematical term used during training where the student's loss function is modified to match the teacher's critic. This forces the student to learn not just from the reward, but also from how the teacher evaluates states and actions.
Human-in-the-Loop RL
A training setup where a human supervisor interacts with an RL agent, either by providing demonstrations or correcting its actions in real-time. This allows the agent to learn effectively from both expert examples and its own generated experiences.

Terminology used across episodes

This episode discusses

The paper

VE2VF: Vision-Enabled to Vision-Free Distillation via Real-world Reinforcement Learning for Robust Contact-Rich Manipulation · Read on arXiv

Autonomous Systems, Technische Universitaet Wien (TU Wien) · Institute of Robotics and Mechatronics (DLR), German Aerospace Center

When using reinforcement learning (RL) for contact-rich robotic manipulation, vision can provide task-relevant information that complements robot proprioception. However, vision-enabled policies tend to overfit to the visual conditions seen during training, limiting their robustness and transferability. We present a human-in-the-loop RL framework that employs teacher-student distillation to achieve robust performance across multiple task variants, trained entirely in the real world without requiring domain randomization or data augmentation. A vision-enabled teacher distills its knowledge into a vision-free student that relies solely on pose, twist, and wrench sensing, combining fast training with strong task generalization. On the real-world NIST assembly benchmark board, our approach achieves 95% overall success after approximately 50 minutes of training on 3 representative tasks, including robust generalization to 8 unseen task variants. Fine-tuning with distillation achieves full success on the most challenging task. We demonstrate that the resulting policies outperform baselines in both robustness and adaptability. Page: https://tuwien-asl.github.io/VE2VF/.

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: I'm Rosa, and with me are Dev and Taro, guest researcher.

Dev: Today's paper: "VE2VF: Vision-Enabled to Vision-Free Distillation via Real-world Reinforcement Learning for Robust Contact-Rich Manipulation".

Rosa: When using reinforcement learning for contact-rich robotic manipulation, vision can provide task-relevant information that accelerates learning beyond what proprioception alone can achieve,

Dev: First, who's behind it and why it matters.

Paper summary: Rosa: So, looking at "VE2VF: Vision-Enabled to Vision-Free Distillation via Real-world Reinforcement Learning for Robust Contact-Rich Manipulation," the main thesis is that vision provides an advantage in learning contact-rich manipulation tasks, but this visual reliance leads to policies that overfit to the specific visual conditions they were trained in.

Dev: They propose a solution called VE2VF, which is a two-stage approach where you first train a vision-enabled teacher policy that benefits from rich perceptual feedback. Then, they use knowledge distillation to transfer those skills into a vision-free student policy that operates solely on pose, twist, and wrench sensing.

Rosa: The paper claims this combination allows them to achieve robust performance across multiple task variants while training entirely in the real world without needing any domain randomization or data augmentation techniques. That part is quite compelling for practical applications because it simplifies the training pipeline significantly.

Taro: It matters because contact-rich tasks are often complex, and relying solely on visual input can be a distraction from the fundamental force and geometric relationships that actually determine task success; this framework aims to isolate those core mechanics.

Dev: Essentially, they are using the vision as a temporary guide to learn the skill efficiently, and then distilling it down to a controller that is less susceptible to sensory noise or visual occlusions during operation.

Rosa: It’s about creating a policy that retains the exploration benefits of seeing things while shedding the unreliability of vision for long-term deployment in physical environments.

Taro: And they are using a human-in-the-loop approach, specifically HILSERL, to guide this process in the real world, which gives them a strong foundation based on actual physical interaction rather than just simulation data.

Dev: That human feedback loop is important because it helps define what success means in contact manipulation without needing to painstakingly design intricate reward functions from scratch for every single task variant.

Rosa: So, the core claim is that this distillation technique effectively transfers the skills learned with visual input into a vision-free system that generalizes well, which addresses a key weakness in current vision-based RL approaches.

Taro: The impact could be significant because it suggests we can build manipulators that are more adaptable to unexpected physical variations because they aren't overly dependent on perfect visual cues during execution.

Dev: I just hope the resulting policy is fast enough for real-time control; if the distillation process adds too much overhead, that loop rate could become a problem in a high-speed contact scenario.

Rosa: That’s definitely something we need to watch closely as we move toward deploying this on physical hardware; the speed of the inference on that vision-free policy is critical for its success in dynamic situations.

Conclusion: Rosa: Wrapping up the discussion on "VE2VF: Vision-Enabled to Vision-Free Distillation via Real-world Reinforcement Learning for Robust Contact-Rich Manipulation," this paper by Kowalski, Li, and Lee explores a novel way to build robust robotic manipulators. The main implication is that we can move toward having controllers that are not overly reliant on visual input during the execution of complex physical tasks.

Dev: It really boils down to taking the strengths of visual RL for initial learning and distilling them into a more deterministic, proprioception-based controller that handles unexpected physical situations better in the field.

Taro: For autonomy research, this suggests that if we can distill skills from vision into pose and wrench sensing, our autonomous systems could be much more resilient when visual sensors fail or provide ambiguous data during critical contact phases.

Rosa: And I think this means we are closer to having manipulators that can operate effectively across a wider range of real-world scenarios without needing extensive pre-training with massive datasets for every single new condition.

Dev: The long-term impact hinges on whether this distillation method scales efficiently enough to handle the complexity of industrial applications where we need reliability over sheer raw visual fidelity.

Taro: I believe the real world implication is that we can deploy robots in environments where perfect visual tracking isn't guaranteed, relying instead on the learned physical relationships encoded in force and motion feedback.

Rosa: So, in simple terms, it’s about making robotic manipulation skills more reliable by replacing vision with a distilled representation of those essential physical dynamics.

Dev: It’s a solid contribution because it shows how to leverage existing learning paradigms—like teacher-student distillation—to create systems that are more focused on the underlying physics rather than just the superficial visual appearance.

Taro: I'm optimistic about its potential for future work, seeing how this vision-free controller handles tasks that require high levels of fine motor precision under challenging, unpredictable physical conditions.

More episodes

← Home