TACROSS: An Efficient and Low-Cost Scalable Human Touch System Across Heterogeneous Tactile Sensors for Dexterous Robot Learning

arXiv:2610.11945 · cs.RO, cs.LG · Submitted 2026-10-08 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: I'm Rosa, and with me are Dev and Taro, guest researcher.

Dev: Today's paper: "TACROSS: An Efficient and Low-Cost Scalable Human Touch System Across Heterogeneous Tactile Sensors for Dexterous Robot Learning".

Rosa: The gist:

Dev: First, who's behind it and why it matters.

Title and authors: Rosa: So we're looking at this paper called TACROSS: An Efficient and Low-Cost Scalable Human Touch System Across Heterogeneous Tactile Sensors for Dexterous Robot Learning. It sounds like they're tackling the problem of making robots learn how to touch things based on human data, but they focus on making it cheap and scalable across different types of sensors.

Dev: Yeah, that's what the title suggests. The core idea is bridging that gap between human touch and robot sensors in a way that doesn't require super expensive setups or massive datasets. It’s about efficiency and cost here, which is always a big deal when you’re building systems for real-world deployment.

Taro: I'm curious about the scope since we see other papers here focusing on things like navigation or world models. Does this paper focus strictly on the tactile interface, or does it have broader implications for how robots acquire skills?

Rosa: Well, it focuses heavily on that touch interface, specifically using a piezoresistive glove as a low-cost human acquisition tool and then mapping that information to whatever robot sensor you have. They are trying to make the transfer work even when the sensors are fundamentally different.

Dev: Exactly, they're addressing those differences in how raw sensor values represent contact because the hardware is so varied, which is a huge hurdle for any cross-sensor learning system.

Taro: So, if I boil it down for someone just listening, it seems like they're proposing a way to use cheap human touch data to train robots without needing those incredibly expensive robot-only demonstrations that take forever to collect.

Rosa: That’s right. It’s about using the human demonstration as a starting point, but then making sure that this information translates reliably into something the robot can actually use for its own actions.

Dev: We'll see how robust that translation is when we get into the specific technical details of how they align those signals.

The paper's summary: Rosa: Okay, so diving into what TACROSS actually does, it integrates a human tactile system, a contact-semantic encoder with both human and robot input branches, and then an ACT policy that uses the robot's own observations at deployment. It’s a whole pipeline.

Dev: They use the human touch to learn the representation part—that shared tactile latent space—and then they ground that latent with actual robot actions for supervision during training. That’s a key part of how they manage the learning process, right?

Taro: I see them using human touch for representation learning and robot actions for ground-truth supervision; that sounds like a smart way to get the representation right without just relying on imitation.

Rosa: They use something called contact-semantic targets and contrastive alignment to learn those shared representations, which helps them avoid having a representation that’s just constant or stuck in one spot.

Dev: And they use this loss function, Lrep, which combines contrastive loss on the hand latents with contact-prediction loss and domain adversarial loss to keep things balanced during training. That sounds like a lot of machinery to manage for stability.

Taro: It sounds like they are very careful about how the representation evolves, trying not to get stuck in a fixed state while still learning meaningful contact dynamics from both human and robot inputs.

Rosa: And they use valid retargeted hand targets as pseudo-labels for auxiliary supervision, which lets them train effectively using both the actual robot ground truth and this human help.

The paper's improvements: Dev: Let's talk about what they suggest is better than just standard approaches. They introduce canonicalizers with validity masks and residual adapters alongside a shared causal temporal encoder with attention across fingers to handle the alignment between human and robot signals.

Rosa: That alignment mechanism is crucial because it converts those raw, different signals into a common masked anatomical contact field, capturing things like contact probability or loading dynamics with explicit validity masks.

Taro: So they are explicitly modeling not just where things are touching, but the physics of that touch—like force and duration—in a standardized way across the different sensors.

Dev: They achieve this by using these canonicalizers to convert raw signals into that common contact semantic field, which then feeds into a shared causal temporal encoder with attention across fingers to generate that two hundred fifty-six-dimensional shared tactile latent space <ref:2610.11945#pg1>.

Rosa: So they are distilling complex, heterogeneous information down into this single latent space of two hundred fifty-six dimensions that captures the coordination among fingers across different robot embodiments <ref:2610.11945#pg1>.

Taro: That sounds like a way to generalize the concept of 'touch' itself so that it’s not tied to one specific robot hardware configuration.

Dev: The result is they claim this system achieves a mean success rate of ninety-one point nine percent compared to seventy-three point one percent without human tactile input, even when only using thirty robot and one hundred fifty human demonstrations for the training setup described in TACROSS: An Efficient and Low-Cost Scalable Human Touch System Across Heterogeneous Tactile Sensors for Dexterous Robot Learning #pg2.

Conclusion: Rosa: So, to wrap up on TACROSS, it’s a system that lets you scale tactile data collection without needing a robot running constantly by combining synchronized human visual, motion, and tactile demonstrations with robot supervision. It’s practical because the hardware component itself is relatively inexpensive at USD ten point eight six for the glove <ref:2610.11945#pg2>.

Dev: The main implication is that you can get high success rates in complex manipulation tasks—like fixed-point dispensing or liquid transfer—by leveraging human data efficiently, and they showed efficiency improvements of three point five times compared to conventional teleoperation systems <ref:2610.11945#pg2>.

Taro: What I see is that the biggest limitation they point out is still the reliance on robot demonstrations and the need for retargeting and calibration specific to each embodiment, meaning it’s not entirely plug-and-play across every setup yet.

Rosa: Right, so while it's very powerful for learning from human touch, you still have that calibration step where you have to adapt the system to a new robot hand configuration. The next steps they mentioned are extending TACROSS to more robotic hands and exploring shared tactile representations across different embodiments in that latent space.

Dev: From an engineering standpoint, it’s a solid framework for getting good results with relatively low-cost hardware, but the dependency on those specific calibrations means deployment still requires some tailored work.

Taro: I think the future work into a shared latent space is what really matters because that would mean the representation learned from one robot hand could actually benefit another hand.

Rosa: That’s a good thought. So we've looked at how TACROSS uses contact semantics and alignment to make heterogeneous tactile data useful for robot learning, and it’s an open-source dataset of over one hundred fifty hours of recordings is coming soon <ref:2610.11945#pg2>.

Dev: Indeed, it sets a clear path forward by showing how to decouple the acquisition cost from the learning quality.

Taro: We'll be watching how they extend this concept beyond just one glove and one robot setup next.

Bo Chen, Huanzhang Hu, Junyang Ma, Bo Yue, Fangdi Yu, Haijier Chen, Xianxin Lai, Shuyu Pan, Zhen Yang, Xiaoquan Sun

The University of Hong Kong

cs.RO, cs.LG

Submitted: 2026-10-08

Updated: 2026-10-08

Project page: https://tacross-touch-project.github.io

The gist: The gist: TACROSS presents an efficient and low-cost system for learning from human touch and transferring it to robots by aligning tactile streams at the level of contact events rather than raw

Key concepts

Tactile Alignment Across Sensors
This addresses the 'Sensing Gap' where different sensors give different readings for the same touch. TACROSS uses canonicalizers and residual adapters to convert raw human and robot signals into a common contact field, capturing key details like force, area, and duration.
Shared Tactile Latent
A 256-dimensional representation derived from the shared encoder captures the dynamics of contact events across fingers in different embodiments. This latent space allows a robot policy to understand coordinated finger movements regardless of the specific sensor hardware used.
Policy Learning from Robot Demonstrations
The system learns a robot's movement policy primarily from demonstrations provided by a robot itself. Human touch data is used to improve tactile representation learning through confidence-weighted supervision, allowing the robot to combine its own sensory inputs with learned human interaction patterns.

Terminology

Summary

The gist: TACROSS presents an efficient and low-cost system for learning from human touch and transferring it to robots by aligning tactile streams at the level of contact events rather than raw sensor values.

System Overview

TACROSS is an integrated hardware–software system designed for human tactile data acquisition and cross-sensor transfer for robot learning The hardware component of TACROSS integrates a piezoresistive glove with five layers and a cost of USD 10.86 with 285 sensing points. This system bridges the heterogeneity between human and robot tactile sensors by aligning them at the level of contact events. The core idea is to map heterogeneous signals into a shared tactile latent with 256 dimensions via a temporal Transformer with attention across fingers.

Tactile Alignment Across Sensors

The system addresses the Sensing Gap where raw sensor values do not denote equal contact because of differences in transduction principle, layout, and dynamic response. To solve this, TACROSS employs canonicalizers with validity masks together with residual adapters and a shared causal temporal encoder with attention across fingers. These canonicalizers convert raw human and robot signals into a common masked anatomical contact field, capturing active region, contact probability, coarse force, contact area, duration, loading dynamics, phase, instability with explicit validity masks The shared encoder represents contact event dynamics and coordination among fingers across embodiments, yielding a shared tactile latent with 256 dimensions.

Policy Learning from Robot Demonstrations

The system utilizes a robot-grounded policy learning scheme where robot demonstrations provide the sole source of ground-truth action supervision. Human demonstrations support tactile representation learning and provide confidence-weighted auxiliary supervision through valid retargeted hand targets. The robot policy combines the tactile latent with RGB and proprioception to predict an action chunk.

Training and Objectives

Training involves using masked contact-semantic targets and contrastive alignment to learn shared representations. The training objective includes terms like contrastive attraction/repulsion with phase/force hard negatives to discourage a constant representation The combined objective function Lrep incorporates several loss terms, including contrastive loss on hand latents and contact-prediction loss.

Evaluation and Results

TACROSS is evaluated on four contact-rich manipulation tasks: T1 Fixed-Point Dispensing, T2 Beaker-to-Beaker Liquid Transfer, T3 Precision Pipetting, and T4 Sequential Fruit Pick-and-Place The system demonstrates significant efficiency improvements compared to conventional teleoperation, achieving a 3.5× efficiency improvement while reducing demonstration acquisition equipment cost by 95.7%. For example, Vision-IK Ego achieves similar success (91.5% vs. 92.2%) at 3.5× the average collection rate and 95.7% lower equipment cost than Robot-only. The system is open-source, and a tactile dataset comprising over 150 hours of recordings will be publicly released.

Conclusion

TACROSS provides a practical way to scale tactile data collection without requiring continuous robot operation by combining synchronized human visual, motion, and tactile demonstrations with robot supervision The limitation noted is that the transfer still relies on robot demonstrations and retargeting and calibration specific to each embodiment. The future work includes extending TACROSS to additional robotic hands and exploring tactile representations across embodiments in a shared latent space.

--- Page 8 ---

The full-hand glove has a component BOM of USD 10.86, with 285 physical sensing locations and an output rate of approximately 150 Hz The current Robot-only, Ego-R, and Ego-V acquisition configurations cost approximately USD 12,000, USD 2,150, and USD 520 respectively. Relative to Robot-only, the listed Ego-R and EgoV equipment investments are 82.1% and 95.7% lower. The full TACROSS with Contact-Semantic Alignment achieves 91.9% mean success when compared with 73.1% without human tactile input and 49.4% without contact-semantic alignment, using 30 robot and 150 human demonstrations. Removing human tactile input or contact-semantic alignment reduces success rates in both training settings. The full TACROSS with Contact-Semantic Alignment achieves 91.9% mean success when compared with 73.1% without human tactile input and 49.4% without contact-semantic alignment, using 30 robot and 150 human demonstrations.

--- Page 9 ---

The system is open-source, and a tactile dataset comprising over 150 hours of recordings will be publicly released. The limitation noted is that the transfer still relies on robot demonstrations and retargeting and calibration specific to each embodiment. The future work includes extending TACROSS to additional robotic hands and exploring tactile representations across embodiments in a shared latent space.

--- Page 7 ---

The full-hand glove has a component BOM of USD 10.86, with 285 physical sensing locations and an output rate of approximately 150 Hz. The current Robot-only, Ego-R, and Ego-V acquisition configurations cost approximately USD 12,000, USD 2,150, and USD 520 respectively. Relative to Robot-only, the listed Ego-R and EgoV equipment investments are 82.1% and 95.7% lower. The full TACROSS with Contact-Semantic Alignment achieves 91.9% mean success when compared with 73.1% without human tactile input and 49.4% without contact-semantic alignment, using 30 robot and 150 human demonstrations. Removing human tactile input or contact-semantic alignment reduces success rates in both training settings.

--- Page 6 ---

The full TACROSS with Contact-Semantic Alignment achieves 91.9% mean success when compared with 73.1% without human tactile input and 49.4% without contact-semantic alignment, using 30 robot and 150 human demonstrations. Removing human tactile input or contact-semantic alignment reduces success rates in both training settings.

--- Page 5 ---

The system is open-source, and a tactile dataset comprising over 150 hours of recordings will be publicly released. The limitation noted is that the transfer still relies on robot demonstrations and retargeting and calibration specific to each embodiment. The future work includes extending TACROSS to additional robotic hands and exploring tactile representations across embodiments in a shared latent space.

--- Page 6 ---

The full TACROSS with Contact-Semantic Alignment achieves 91.9% mean success when compared with 73.1% without human tactile input and 49.

Improvements for AI systems

  1. This system can bridge sensor heterogeneity by aligning heterogeneous human and robot touch in a shared contact-semantic space with validity masks. This allows for learning from low-cost human tactile demonstrations (like piezoresistive gloves) and transferring that knowledge to robots equipped with different sensors (like capacitive Revo2).

  2. The system can achieve Efficient and Low-Cost Data Acquisition System by utilizing a hardware component, the TACROSS glove, which has a BOM cost USD 10.86/glove and offers high frame rates (150 Hz), making scalable data collection feasible at a fraction of the cost of traditional methods.

  3. The system can improve policy learning by grounding policy learning in executed robot actions, which provide the sole source of ground-truth action supervision, ensuring that learned policies are grounded in physically executable movements rather than just retargeted human hand motions.

  4. The system can enhance representation quality by using a shared causal temporal encoder with attention across fingers to produce a shared tactile latent with 256 dimensions, which effectively captures contact event dynamics and coordination among fingers across embodiments.

  5. The system can improve training stability and generalization by employing the loss function: Lrep = X Σ λj Lj, where (λj) = (1, 1, 0.5, 0.5, 1, 0.2, 0.1), which combines contrastive learning on hand latents with contact-prediction loss and domain-adversarial loss to discourage a constant representation.

  6. The system can support robust deployment by using valid retargeted hand targets serve as pseudolabels for auxiliary supervision, weighted by confidence and masked by validity, allowing the ACT policy to be trained effectively using both robot ground truth and human auxiliary supervision.

Sources

Related papers