GIFT: Glove-Inferred Force Transfer: Force-Aware Human-to-Robot Skill Transfer from a Wearable Sensing Glove to a Robot Hand Without Tactile Sensors
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.
Rosa: I'm Rosa, and with me are Dev and Taro, guest researcher.
Dev: Today's paper: "GIFT: Glove-Inferred Force Transfer: Force-Aware Human-to-Robot Skill Transfer from a Wearable Sensing Glove to a Robot Hand Without Tactile Sensors".
Rosa: GIFT (Glove-Inferred Force Transfer) is an end-to-end pipeline for human-to-robot skill transfer of force without tactile hardware on the robot,
Dev: First, who's behind it and why it matters.
Title and authors: Rosa: So we're diving into the paper "GIFT: Glove-Inferred Force Transfer: Force-Aware Human-to-Robot Skill Transfer from a Wearable Sensing Glove to a Robot Hand Without Tactile Sensors." This sounds like something that really tackles the problem of teaching robots how to handle delicate tasks, like grasping things, using just what a human does.
Dev: Yeah, it's interesting because the whole concept revolves around treating the human hand and robot hand as one physical unit instead of relying on some shared sensor between them. It seems they're proposing a way to transfer force information directly without needing tactile hardware on the robot at all.
Taro: I wonder how this system handles situations where things go wrong, like when the environment suddenly changes during the grasp, or if there's unexpected resistance? I need to know what happens when the world misbehaves during deployment.
Rosa: That’s a fair concern, Taro; it definitely sounds like a safety consideration we need to think about. The paper suggests that the system uses force feedback from human demonstrations to guide the robot's grip strength, which could be really helpful in learning how to handle varying object properties.
Dev: From an engineering standpoint, I'm thinking about the latency involved here. If this whole pipeline is meant for real-time operation on a deployment hand, we need to make sure that force estimation from actuator current residuals doesn't introduce any significant delays that could cause instability in the control loop.
Taro: Exactly; if there's a delay, and the world misbehaves, the robot might react too late or incorrectly based on outdated force information. I hope their estimation method is robust enough to handle those kinds of unexpected dynamics.
Rosa: Well, what the paper describes is that they capture finger flexions and fingertip forces from a calibrated glove while filming a demonstration with no robot involved during the data collection phase. That’s a huge piece of context for understanding how they get this information in the first place.
Dev: That setup sounds complex to manage, especially getting all those sensors—five flex sensors, five FSRs, and wrist orientation—to sync up properly with a camera feed on that ESP32-S3 microcontroller during capture. I'm curious about the reliability of that hardware setup in a real lab environment.
Title and authors: Taro: The fact that they use a calibrated reference scale to turn those raw FSR readings into newtons is important; it confirms they are dealing with physical quantities, which gives the policy something concrete to learn from.
Rosa: Right, and then they train this policy using a glove-space state representation that includes all those flexes, forces, and IMU angles from the demonstration data to predict where the fingers should go next. That’s how the AI learns the desired movements.
Dev: The state vector itself is quite rich; having five flex, five force readings, and three IMU angles means the policy has a lot of input to work with when deciding what action to take next on that hand. I have to wonder about the computational load this representation puts on whatever hardware runs the policy during actual deployment.
Taro: It’s interesting how they show that vision-only policies fail completely, which really solidifies the idea that combining those physical force inputs with visual data is necessary for acquiring a stable grasp. That points toward how AI needs to incorporate physical constraints when learning manipulation skills.
Rosa: And then they retarget those actions from the glove space into commands for a robot hand using something called spline-based retargeting, which even has a specific mechanism where the thumb's single flex channel drives three joints and opposition rides a spline to mimic the demonstrator’s motion.
Dev: Spline-based mapping is interesting because it suggests they are trying to model the kinematics of the human hand quite closely in that translation step, which is crucial for maintaining fidelity when moving from one physical system to another. I'm concerned about how sensitive that spline mapping is to errors in the initial glove state readings.
Taro: If the mapping isn't robust, any small error in measuring finger flexion or force could translate into a big error in the actual joint commands sent to the robot hand, which would definitely cause trouble when things get dynamic.
Rosa: The deployment side is where things get really clever because they estimate force from actuator current residuals compared to a freespace baseline, effectively bypassing any need for tactile sensors on the robot itself. That's a significant design choice.
Title and authors: Dev: That estimation relies on comparing the robot hand’s actuator currents against a freespace baseline and mapping those residuals into the same newton range the policy saw in training, which means it's essentially inferring what force is being exerted based on how much current is needed to maintain position. I need to understand if that mapping remains accurate across different deployment conditions.
Taro: That reliance on actuator currents as a proxy for force seems like a clever way to achieve the goal without adding new hardware, but it introduces its own set of uncertainties we have to account for when we think about real-world scenarios.
Rosa: So, looking at the overall result from that evaluation, they tested two different action-chunking policies trained on these demonstrations and found that the policy retaining force inputs performed significantly better in terms of grip strength, showing a median per-rollout hold-phase grip-force estimate of one point two zero newtons compared to two point five five newtons when force inputs were included.
Dev: That difference between those two values is pretty substantial; a fifty percent reduction in the estimated grip force when the policy utilized those force inputs really speaks to their effectiveness in learning how to be delicate rather than just brute-force grasping. It’s a strong data point for loop rate considerations, though.
Taro: That result confirms what we suspected from the ablation study; it clearly shows that vision alone isn't enough to get the right grip strength, and those force inputs are actually determining how hard the policy holds onto the object.
Rosa: It really highlights that this paper is showing how you can share information by focusing on a physical unit, where human fingertip force in newtons becomes a shared quantity that informs both sides of the transfer pipeline. This seems like a practical way to bridge the gap between perception and physical action.
Dev: If we look at the limitations they mention, they flag that shear forces aren't captured by those FSRs, and there's also drift and hysteresis in those force-sensitive resistors which could affect accuracy over time during extended use. Those are classic hardware challenges we have to consider when moving this from a controlled capture rig to a long-term deployment.
Taro: So, while the pipeline is powerful for cup grasp tasks under ideal conditions, we need to keep an eye on those limitations when applying it to more complex or messy environments where shear forces might be involved. That’s where real autonomy gets tricky.
Title and authors: Rosa: It’s exciting because they prove that you don't need a shared tactile sensor between the human and robot; you can just share a physical measurement, like force in newtons, and the robot can infer its own force through its own actuation feedback. This opens up possibilities for much more generalized skill transfer.
Dev: And if we consider the practical implications, having this capability means we could equip many different types of position-controlled hands with sophisticated grasping skills without needing to integrate expensive tactile sensor arrays into every single robot arm. That flexibility is valuable for deployment platforms.
Taro: I think the biggest implication is that AI systems can learn nuanced physical interactions not just through visual imitation, but by learning the underlying mechanics of force application, which could be really important for robots interacting with fragile objects in complex settings.
Rosa: So to wrap up this discussion on "GIFT: Glove-Inferred Force Transfer: Force-Aware Human-to-Robot Skill Transfer from a Wearable Sensing Glove to a Robot Hand Without Tactile Sensors," it really shows how we can establish a shared physical unit—measured force in newtons—to transfer skill knowledge between human and robot systems.
Dev: It’s an end-to-end pipeline that requires no robot during the data collection phase and no tactile sensors on the deployment hand, instead estimating force through actuator current residuals relative to a freespace baseline.
Taro: The evaluation showed that policies using force inputs performed significantly better in terms of grip strength, with the policy with force inputs holding the cup at less than half of its estimated grip force compared to one without those inputs.
Rosa: Overall, this work suggests that by focusing on how physical forces are measured and mapped consistently between systems, we can build more capable robotic manipulation skills directly from human demonstrations.
Dev: We've seen how it works in a controlled setting for a cup grasp-and-hold task, but the real challenge now is making sure those force estimations hold up when the system moves outside that specific test scenario.
Taro: I think the next step involves testing this on more challenging tasks where unexpected dynamics are common, to see how robust those force-informed policies actually are in unpredictable environments.
The paper's summary: Rosa: So, to recap, this paper introduces GIFT, which is an end-to-end pipeline designed to let robots learn how to grasp objects by observing human demonstrations while completely avoiding the need for any tactile sensors on the robot hand itself.
Dev: That's right; it treats the physical contact force measured from a calibrated human glove as a shared unit of information, and then estimates that same force on the robot side using only actuator current readings.
Rosa: What I find most striking is how they achieved this without having to have a robot present during the data collection phase, which really opens up possibilities for collecting demonstrations in labs or even just from videos.
Dev: And from an engineering standpoint, the method of estimating force through those current residuals against a freespace baseline seems clever, but I'm still looking at how stable that mapping is when you move to something more dynamic than a static cup grasp.
Rosa: Taro, you mentioned earlier that this system learns how hard to hold based on fingertip force input; does the paper show any examples of it handling those situations where the object changes shape during the hold?
Taro: It does, but their ablation study was pretty telling; they showed that a vision-only policy couldn't manage a single grasp because it didn't have that crucial feedback on grip strength, whereas policies including force inputs actually acquired the grasp and determined how firm it held.
Dev: That fifty-three percent reduction in estimated grip force they found during their evaluation really suggests that this isn't just about getting a position right; it’s fundamentally about learning the correct physical pressure required for the task. I need to keep thinking about the loop rate here, though; if we want this running on a fast deployment platform, we need to make sure that current-residual estimation doesn't introduce any unacceptable latency into the control loop.
Rosa: That makes sense; if there's a delay in estimating force from the robot hand’s currents, it could lead to instability when interacting with something fragile. This paper really shows how you can transfer skill knowledge by focusing on how physical forces are measured and mapped consistently between systems.
Taro: The real impact here is that we can train these sophisticated manipulation skills just by filming a human performing the task, without needing a robot arm in the capture setup, which makes collecting data much more accessible for researchers across various fields.
Dev: From an engineering standpoint, if this framework proves robust outside of a controlled lab setting for long periods—say, if we can get those drift and hysteresis issues they mentioned under control—then the flexibility to deploy it on any position-controlled hand without specialized tactile hardware is a big deal for robot design.
Rosa: It seems like the next logical step is testing how this performs in more complex scenarios where things aren't perfectly rigid or where unexpected resistance occurs during deployment. That’s what we need to know when we think about real-world application.
The paper's improvements: Rosa: So, we’ve looked at how GIFT works now, and I want to talk about what they suggest as improvements for making this system even better for the field.
Dev: What I'm looking forward to is seeing if these suggested enhancements actually translate into a more reliable control loop, because if the estimation method gets messier, that latency could become a major issue on deployment.
Rosa: They propose moving toward higher fidelity force-aware manipulation tasks where the robot can execute grips at a fraction of its maximum capability, which sounds really practical for handling fragile materials.
Dev: That's something I can get behind; if we can reliably estimate that grip force and know it's only using a small percentage of the motor capacity, it significantly reduces the risk of crushing an object during manipulation.
Rosa: What’s another big improvement they point to? I heard something about how this framework allows AI systems to learn nuanced physical interactions by treating force in newtons as a primary shared unit.
Taro: That's significant because it means the robot isn't just mimicking visual positions; it’s learning the underlying mechanics of how pressure translates into grip strength, which is a much deeper level of skill transfer.
Dev: If we can treat force as that primary shared unit, then we have more concrete data to work with for training, even if the hardware itself isn't sharing sensors. It shifts the burden from perfect sensor alignment to accurate physical modeling in the policy.
Rosa: And they also noted that this approach allows us to collect demonstrations using only visual data and hand-command states, meaning we don't always need a robot present for every single data collection session.
Taro: That scalability is huge for researchers; it means we can gather more diverse demonstration datasets without needing a fully equipped robotic lab setup every time, which should speed up the development of new manipulation policies.
Dev: I do wonder if that reliance on hand-command states to acquire the grasp creates any dependency on the initial state acquisition being perfectly accurate; if that handshake isn't flawless, the whole force estimation chain could become unreliable quickly.
Rosa: That’s a fair point about state acquisition accuracy; it shows they are thinking about the entire pipeline, not just the training phase. This system seems designed to be highly flexible for deployment because it doesn't mandate any specific tactile hardware on the robot itself.
Taro: The fact that we can deploy this on any position-controlled hand reporting motor current is a huge win for adaptability; it means the skill transfer capability isn't locked into one specific robot platform.
Dev: If we consider the limitations they flagged—like shear forces not being captured by FSRs—then their improvements must be focused on mitigating those known weaknesses, or else we’re just moving errors around in a different way.
Rosa: Exactly, so the future work seems focused on making this system robust enough to handle those messy real-world dynamics where perfect force measurement isn't always possible. That's where the real challenge lies for field application.
Conclusion: Rosa: So we’ve reached the end of our discussion on GIFT: Glove-Inferred Force Transfer: Force-Aware Human-to-Robot Skill Transfer from a Wearable Sensing Glove to a Robot Hand Without Tactile Sensors, and I want to wrap up by summarizing its main implications.
Dev: Essentially, this work shows that we can transfer complex manipulation skills from human hands to robot hands by sharing a calibrated physical unit—measured force in newtons—instead of relying on specialized tactile sensors on the robot itself.
Rosa: It’s pretty exciting because it opens up ways for us to build much more capable robotic systems without having to integrate expensive tactile sensor arrays into every single arm we deploy.
Dev: That flexibility is key; if we can estimate force through actuator current residuals, then any position-controlled hand can potentially be equipped with these force-aware grasping skills.
Taro: From an autonomy research angle, the real impact here is proving that AI systems can learn nuanced physical interactions not just through visual imitation, but by learning the underlying mechanics of force application, which is important for robots interacting with fragile objects in complex settings.
Rosa: And I think that capability to learn grip strength dynamically based on force feedback means we move beyond simple positional control toward genuine dexterity in grasping tasks.
Dev: I'm still thinking about those deployment scenarios; how long do you think this pipeline can maintain its accuracy outside of a controlled lab environment before the inherent limitations start showing up?
Taro: That’s the question for future work; we need to rigorously test how robust these force-informed policies are when they encounter unexpected dynamics or situations where shear forces are involved, as those limitations were explicitly mentioned.
Rosa: So, in summary, GIFT gives us a powerful framework for force-aware skill transfer by focusing on shared physical measurements rather than shared hardware sensors.
Dev: It’s a solid piece of engineering because it addresses the need for robust manipulation without adding complex new hardware to the robot's end effector.
Taro: I just want to see this framework applied in environments that are less predictable, where the world doesn't always behave according to our training data.
Rosa: And that’s exactly what we need to look at next—how we can push this capability into more challenging, unpredictable field conditions.
cs.RO
Submitted: 2026-09-12
Updated: 2026-09-28
Comments: 7 pages, 5 figures, 1 table. Project page: https://tzahsarusi.github.io
Code: https://github.com/huggingface/lerobot
Project page: https://tzahsarusi.github.io
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 84/100
The gist: GIFT (Glove-Inferred Force Transfer) is an end-to-end pipeline for human-to-robot skill transfer of force without tactile hardware on the robot, where fingertip force is measured in newtons on the
Key concepts
- GIFT (Glove-Inferred Force Transfer)
- An end-to-end pipeline for transferring human skill demonstrations to a robot hand by treating the physical contact force from a calibrated human glove as a shared unit of information. This allows the robot to learn how to grasp objects by observing human demonstrations without needing tactile hardware on the robot.
- Force Estimation via Actuator Current Residuals
- The method used for the robot to infer its own grip force. It compares the current drawn by the robot hand's actuators against a freespace baseline and maps those differences (residuals) into newton units, effectively bypassing the need for tactile sensors on the robot.
- Spline-Based Retargeting
- The mechanism used to translate actions from glove space to robot hand commands. This involves modeling the kinematics of the human hand closely, with specific mappings like using a thumb's flex channel to drive multiple joints and opposition riding a spline to mimic human motion.
Terminology
Summary
GIFT (Glove-Inferred Force Transfer) is an end-to-end pipeline for human-to-robot skill transfer of force without tactile hardware on the robot, where fingertip force is measured in newtons on the human side and estimated in newtons on the robot side. The interface between human and robot is treated as a physical unit rather than a shared sensor.
The system involves:
-
A calibrated sensing glove that records
finger flexions, calibrated fingertip force, and wrist orientation.
Specifically, it mountsfive flex sensors, one per finger along the dorsal side
to measure finger curl andfive force-sensitive resistors at the fingertips measure normal contact force.
These are read by an ESP32-S3 microcontroller. Fingertip force is calibrated to newtons against a digital reference scale with a fitted curve, ensuring the recorded channel is a physical quantity. -
A head-mounted camera that records the demonstration, with
no robot present during data collection.
-
A policy trained on this data that uses a
glove-space state representation and predicts finger-position targets.
The observation for this policy includes an egocentric image plus a 13-dimensional state vector:five flex, five fingertip forces, and three IMU angles.
-
A retargeting decoder that maps the policy's actions from glove space to robot actuators at runtime. This involves
spline-based retargeting from glove space to hand actuators,
where the thumb's single flex channel drives three joints, andthumb opposition rides the spline, so as the hand closes the thumb sweeps into opposition as the demonstrator’s did.
-
A force estimator at deployment that estimates force from actuator current residuals relative to a freespace baseline. The robot hand has
no tactile sensors,
and its force observation is estimated through this process:the robot hand’s actuator currents are compared against a freespace baseline and mapped into the same newton range the policy saw in training.
This estimation relies on the residual principle of force estimation from actuation, whereicontact = imeasured −ˆifree(q).
The pipeline is designed around four design principles: (1) Force is a physical unit: the captured channel is expressed in newtons,
(2) No robot during capture,
(3) No tactile sensor at deployment,
and (4) One instrument scores both policies.
In evaluation, GIFT was tested on a cup grasp-and-hold task with two action-chunking policies trained on the same demonstrations, differing only in whether the fingertip-force inputs were retained or zeroed. In a 50-rollout evaluation:
both policies succeeded in all 25 rollouts
The results showed that the policy with force inputs performed significantly better regarding grip strength. Specifically, "the median of the per-rollout hold-phase grip-force estimates was 53% lower with force inputs: 1.20 N versus 2.55 N (one-sided Mann–Whitney U, p < 0.0001)."
An observation ablation demonstrated the contribution of each channel: a vision-only policy achieved 0/15 grasps, policies given hand-command state acquired the grasp, and the force inputs determined how hard the policy held.
This indicates that hand-command state acquires the grasp, and the fingertip-force inputs determine how hard it holds.
In summary, GIFT transfers human demonstrations to a robot hand without tactile sensors by sharing a physical unit—calibrated fingertip force captured on a human hand—and estimating force on the robot side via actuator current residuals. The pipeline requires no robot during capture and no tactile sensor at deployment, with the final comparison being between policies that utilize this shared physical unit versus those that do not. The mechanism is described as force-informed imitation: observing force shifts which demonstrated positions the policy reproduces.
The primary contributions are: (1) GIFT, an end-to-end pipeline for human-to-robot skill transfer of force without tactile hardware on the robot; (2) A 50-rollout result on real hardware showing a 53% reduction in grip force when force inputs are used; and (3) An observation ablation isolating the channels, confirming that the fingertip-force inputs determine how hard the policy holds.
The pipeline requires no robot during capture and no tactile sensor at deployment, and the human and robot share no sensor.
Limitations noted include: shear forces are not captured by FSRs; FSRs exhibit drift and hysteresis; the current-residual estimator is a relative instrument that saturates near 5 N; and the evaluation measures force regulation, not reliability. The paper concludes that "the pipeline carries no dependency on the sensing or actuation hardware used here.
Improvements for AI systems
Here are the specific improvements and capabilities that can be derived from the GIFT framework for AI systems:
-
The system gains the ability to perform high-fidelity, force-aware manipulation tasks on robots without requiring tactile sensors on those robots.
-
The robot hand can execute grip forces at a fraction of its maximum capability (specifically, 53% lower estimated grip force in the cup grasp example), ensuring delicate handling and reducing the risk of crushing objects.
-
The AI system achieves robust skill transfer from human demonstrations to robot hands by treating physical contact force (measured in Newtons) as a primary, shared unit of information, rather than relying on shared hardware or learned sensor alignment.
-
The AI can be trained using only visual data (camera input) and hand-command states, allowing for scalable demonstration collection without needing a robot present during the capture phase.
-
The system can dynamically regulate the strength of its grasp based on force feedback; it learns that
how hard
to hold is determined by the fingertip force input, not just the visual state or position commands. -
The robot's deployment platform is highly flexible; any position-controlled hand that reports motor current can serve as a deployment site, eliminating the need for specialized tactile hardware on the robot itself.
-
The AI architecture benefits from a
glove-space state representation,
which allows for more intuitive and effective policy learning by mapping physical contact dynamics directly into the policy's observation space (flex, force, IMU).
Sources
- OSMO: Open-Source Tactile Glove for Human-to-Robot Skill Transfer
- Feel the Force: Contact-Driven Learning from Humans
- TactAlign: Human-to-Robot Policy Transfer via Tactile Alignment
- FACTR 2: Learning External Force Sensing for Commodity Robot Arms Improves Policy Learning
- AnySkin: Plug-and-play Skin Sensing for Robotic Touch
- DexUMI: Using Human Hand as the Universal Manipulation Interface for Dexterous Manipulation
- RealDexUMI: A Wearable Universal Manipulation Interface for Dexterous Robot Learning
- DEX-Mouse: A Low-cost Portable and Universal Interface with Force Feedback for Data Collection of Dexterous Robotic Hands
- EgoScale: Scaling Dexterous Manipulation with Diverse Egocentric Human Data
- FlexiTac: A Low-Cost, Open-Source, Scalable Tactile Sensing Solution for Robotic Systems
- Tactile-VLA: Unlocking Vision-Language-Action Model's Physical Knowledge for Tactile Generalization
- Imagining the Sense of Touch: Touch-Informed Manipulation via Imagined Tactile Representations
- FACTR: Force-Attending Curriculum Training for Contact-Rich Policy Learning
- LeRobot: An Open-Source Library for End-to-End Robot Learning
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving