VE2VF: Vision-Enabled to Vision-Free Distillation via Real-world Reinforcement Learning for Robust Contact-Rich Manipulation

arXiv:2605.29564 · cs.RO · Submitted 2026-05-28 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: I'm Rosa, and with me are Dev and Taro, guest researcher.

Dev: Today's paper: "VE2VF: Vision-Enabled to Vision-Free Distillation via Real-world Reinforcement Learning for Robust Contact-Rich Manipulation".

Rosa: When using reinforcement learning for contact-rich robotic manipulation, vision can provide task-relevant information that accelerates learning beyond what proprioception alone can achieve,

Dev: First, who's behind it and why it matters.

Paper summary: Rosa: So, looking at "VE2VF: Vision-Enabled to Vision-Free Distillation via Real-world Reinforcement Learning for Robust Contact-Rich Manipulation," the main thesis is that vision provides an advantage in learning contact-rich manipulation tasks, but this visual reliance leads to policies that overfit to the specific visual conditions they were trained in.

Dev: They propose a solution called VE2VF, which is a two-stage approach where you first train a vision-enabled teacher policy that benefits from rich perceptual feedback. Then, they use knowledge distillation to transfer those skills into a vision-free student policy that operates solely on pose, twist, and wrench sensing.

Rosa: The paper claims this combination allows them to achieve robust performance across multiple task variants while training entirely in the real world without needing any domain randomization or data augmentation techniques. That part is quite compelling for practical applications because it simplifies the training pipeline significantly.

Taro: It matters because contact-rich tasks are often complex, and relying solely on visual input can be a distraction from the fundamental force and geometric relationships that actually determine task success; this framework aims to isolate those core mechanics.

Dev: Essentially, they are using the vision as a temporary guide to learn the skill efficiently, and then distilling it down to a controller that is less susceptible to sensory noise or visual occlusions during operation.

Rosa: It’s about creating a policy that retains the exploration benefits of seeing things while shedding the unreliability of vision for long-term deployment in physical environments.

Taro: And they are using a human-in-the-loop approach, specifically HILSERL, to guide this process in the real world, which gives them a strong foundation based on actual physical interaction rather than just simulation data.

Dev: That human feedback loop is important because it helps define what success means in contact manipulation without needing to painstakingly design intricate reward functions from scratch for every single task variant.

Rosa: So, the core claim is that this distillation technique effectively transfers the skills learned with visual input into a vision-free system that generalizes well, which addresses a key weakness in current vision-based RL approaches.

Taro: The impact could be significant because it suggests we can build manipulators that are more adaptable to unexpected physical variations because they aren't overly dependent on perfect visual cues during execution.

Dev: I just hope the resulting policy is fast enough for real-time control; if the distillation process adds too much overhead, that loop rate could become a problem in a high-speed contact scenario.

Rosa: That’s definitely something we need to watch closely as we move toward deploying this on physical hardware; the speed of the inference on that vision-free policy is critical for its success in dynamic situations.

Conclusion: Rosa: Wrapping up the discussion on "VE2VF: Vision-Enabled to Vision-Free Distillation via Real-world Reinforcement Learning for Robust Contact-Rich Manipulation," this paper by Kowalski, Li, and Lee explores a novel way to build robust robotic manipulators. The main implication is that we can move toward having controllers that are not overly reliant on visual input during the execution of complex physical tasks.

Dev: It really boils down to taking the strengths of visual RL for initial learning and distilling them into a more deterministic, proprioception-based controller that handles unexpected physical situations better in the field.

Taro: For autonomy research, this suggests that if we can distill skills from vision into pose and wrench sensing, our autonomous systems could be much more resilient when visual sensors fail or provide ambiguous data during critical contact phases.

Rosa: And I think this means we are closer to having manipulators that can operate effectively across a wider range of real-world scenarios without needing extensive pre-training with massive datasets for every single new condition.

Dev: The long-term impact hinges on whether this distillation method scales efficiently enough to handle the complexity of industrial applications where we need reliability over sheer raw visual fidelity.

Taro: I believe the real world implication is that we can deploy robots in environments where perfect visual tracking isn't guaranteed, relying instead on the learned physical relationships encoded in force and motion feedback.

Rosa: So, in simple terms, it’s about making robotic manipulation skills more reliable by replacing vision with a distilled representation of those essential physical dynamics.

Dev: It’s a solid contribution because it shows how to leverage existing learning paradigms—like teacher-student distillation—to create systems that are more focused on the underlying physics rather than just the superficial visual appearance.

Taro: I'm optimistic about its potential for future work, seeing how this vision-free controller handles tasks that require high levels of fine motor precision under challenging, unpredictable physical conditions.

Autonomous Systems, Technische Universitaet Wien (TU Wien) · Institute of Robotics and Mechatronics (DLR), German Aerospace Center

cs.RO

Submitted: 2026-05-28

Updated: 2026-10-05

Project page: https://tuwien-asl.github.io/VE2VF

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 83/100

The gist: When using reinforcement learning for contact-rich robotic manipulation, vision can provide task-relevant information that accelerates learning beyond what proprioception alone can achieve, but

Key concepts

Teacher-Student Distillation
This technique transfers skills from a complex, vision-enabled policy (teacher) to a simpler, vision-free policy (student). The student learns by mimicking the teacher's performance on the same data, effectively distilling the visual knowledge into a format usable without images.
Vision-Free Policy
This is the resulting agent that makes decisions using only physical measurements like end-effector pose, velocity, and force/torque. It relies entirely on proprioceptive sensing rather than visual input to successfully complete manipulation tasks.
Knowledge Distillation Loss
A specific mathematical term used during training where the student's loss function is modified to match the teacher's critic. This forces the student to learn not just from the reward, but also from how the teacher evaluates states and actions.
Human-in-the-Loop RL
A training setup where a human supervisor interacts with an RL agent, either by providing demonstrations or correcting its actions in real-time. This allows the agent to learn effectively from both expert examples and its own generated experiences.

Terminology

Summary

When using reinforcement learning for contact-rich robotic manipulation, vision can provide task-relevant information that accelerates learning beyond what proprioception alone can achieve, but vision-enabled policies often overfit to visual conditions seen during training. This paper presents a human-in-the-loop RL framework employing teacher-student distillation to achieve robust performance across multiple task variants, trained entirely in the real world without requiring domain randomization or data augmentation.

The gist: A vision-enabled teacher policy distills its knowledge into a vision-free student policy that relies solely on pose, twist, and wrench sensing, combining fast training with strong task generalization.

How it works

The proposed framework, VE2VF (Vision-Enabled to VisionFree), follows a two-stage approach: first, training a vision-enabled teacher policy to benefit from rich perceptual feedback; second, employing knowledge distillation to transfer these skills into a vision-free student policy. This method is designed specifically for contact-rich manipulation where visual input can act as a distractor from the essential force and geometric relationships governing task success.

The process involves several key components:

  1. A vision-enabled teacher policy trained via human-in-the-loop RL on representative tasks, utilizing all input modalities including images.

  2. Knowledge distillation to transfer the acquired skills to a vision-free student policy that observes only proprioceptive inputs (pose, twist, and wrench). This is achieved by mapping sensed poses, velocities, and force-torques (wrenches) to actions.

  3. An optional third stage involving fine-tuning with distillation from a new vision-enabled teacher for adapting to challenging, unseen tasks.

Policy Optimization and Learning Paradigm

The policies are learned using the off-policy max-entropy RL algorithm Soft Actor-Critic (SAC), which leads to stochastic policies that converge after a reasonable number of environment transitions. The critic learns the Q-function by minimizing the temporal difference error over past transitions stored in a buffer D, incorporating both reward and soft value of the next state. The actor optimizes the policy to maximize the expected Q-value while maintaining sufficient entropy for exploration, balancing exploitation with exploration through entropy regularization.

The training is conducted in a human-in-the-loop setting following HIL-SERL [5], which uses the RLPD variant to symmetrically sample transitions from policy-generated data and human demonstrations. The human interaction can involve both offline demonstrations and online teleoperation where the human supervisor can overwrite the RL agent's actions, enabling continuous improvement through self-generated experience.

State, Action, and Reward Formulation

The MDP is formulated with a state observation space consisting of images from multiple cameras and a proprioceptive state vector composed of end-effector pose, velocity, force, and torque. The action space is defined as a pose displacement expressed in the local coordinate frame of the end-effector.

A sparse reward function is used where the reward is defined as:

(1) r t = 1, if C(s t) = 1 (task success), 0, otherwise.

This binary reward signal, combined with human demonstrations and corrections, provides a direct and effective scheme for training across multiple tasks without the need for task-specific reward shaping. The initial state distribution is set to a random pose within the action space.

Distillation Mechanism

The distillation process is crucial for achieving robustness and generalization. To distill from a vision-enabled teacher policy (π t) to a vision-free student policy (π s), the critic loss function is modified to match the teacher's critic by adding a mean-square error term:

(13) L s Q(ϕ) = E(o,a,r,o′)∼Ds HIL [(Q s ϕ(o s, a) − y) squared + Q t(o t, a) − Q s(o s, a)] squared.

Furthermore, the vision-free student policy is regularized to match the teacher's actions by adding a KL divergence term to its actor loss function:

(16) L s π(θ) = E o s∼Ds HIL [L s s(θ)] α log π s θ(ao s) − Q s ϕ(o s, a) + DKL(π t(ao))∥ π s(ao s).

Experimental Validation and Results

The approach was validated on the NIST Assembly Board I benchmark board across training tasks, disturbed versions of training tasks (introducing visual distractors and target pose uncertainty), and out-of-distribution tasks.

Improvements for AI systems

Here are the specific improvements that can be made to existing AI systems based on the VE2VF framework, and what those improved systems can achieve:


  1. The core improvement is a shift from purely vision-dependent policies to robust, proprioception-grounded policies via cross-modal distillation.

  2. An AI system trained using the VE2VF method can perform complex, contact-rich assembly tasks (like those on the NIST Assembly Board) with high success rates (up to 95%) even when operating in novel environments or under visual disturbance, by relying primarily on sensed pose, twist, and wrench data.

  3. The improved system exhibits superior robustness compared to vision-only policies, as it avoids overfitting to specific visual textures or lighting conditions that plague traditional vision-based RL.

  4. The system demonstrates strong generalization capability; it can achieve zero-shot success or full task completion on unseen assembly variants simply by mapping the new task's pose into a task-relative proprioceptive coordinate system and using the distilled policy.

  5. The system achieves rapid adaptability through an efficient two-stage training process (Teacher Training followed by Distillation), requiring minimal total real-world interaction time (approximately 50 minutes for complex skills), significantly reducing the need for massive, time-consuming simulation infrastructure or extensive human labeling during initial learning phases.

  6. The policy leverages a vision-free student architecture that is inherently more transferable across different visual contexts (lighting, background) because its decision-making process is decoupled from raw image processing and instead relies on geometrically invariant physical relationships (pose, twist, wrench).

Abstract

When using reinforcement learning (RL) for contact-rich robotic manipulation, vision can provide task-relevant information that complements robot proprioception. However, vision-enabled policies tend to overfit to the visual conditions seen during training, limiting their robustness and transferability. We present a human-in-the-loop RL framework that employs teacher-student distillation to achieve robust performance across multiple task variants, trained entirely in the real world without requiring domain randomization or data augmentation. A vision-enabled teacher distills its knowledge into a vision-free student that relies solely on pose, twist, and wrench sensing, combining fast training with strong task generalization. On the real-world NIST assembly benchmark board, our approach achieves 95% overall success after approximately 50 minutes of training on 3 representative tasks, including robust generalization to 8 unseen task variants. Fine-tuning with distillation achieves full success on the most challenging task. We demonstrate that the resulting policies outperform baselines in both robustness and adaptability. Page: https://tuwien-asl.github.io/VE2VF/.

Sources

Related papers