Adaptive Visuo-Tactile Fusion with Predictive Force Attention for Dexterous Manipulation

arXiv:2505.13982 · cs.RO · Submitted 2025-05-20 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: Today's paper: "Adaptive Visuo-Tactile Fusion with Predictive Force Attention for Dexterous Manipulation".

Dev: Effectively utilizing multi-sensory data is important for robots to generalize across diverse tasks,

Rosa: First, who's behind it and why it matters.

Title and authors: Rosa: So, to summarize what this paper, "Adaptive Visuo-Tactile Fusion with Predictive Force Attention for Dexterous Manipulation," actually suggests, it introduces a force-guided attention fusion module that uses the measured force signals as queries to adaptively adjust the importance of visual and tactile features.

Dev: So, simply put, it means the robot doesn't just blend its vision and touch data equally; instead, it uses the actual physical interaction force to make a decision on whether to lean more on what its camera sees or what its tactile sensors are feeling.

Taro: That’s a significant departure from previous work that might have just glued features together without any dynamic guidance based on the actual physics happening at that specific moment.

Rosa: Precisely; they introduce this module where force signals act as queries, and visual and tactile features serve as keys and values in an attention framework to calculate those fusion weights dynamically.

Dev: So, if we measure a high force during a push, the system learns to give more weight to the tactile input because that's what directly reflects that strong interaction happening right now.

Taro: That sounds like it introduces some implicit knowledge about physical interaction dynamics that was previously hidden in how we thought we should combine these sensory streams.

Rosa: They also supplement this with a self-supervised future force prediction auxiliary task during training, which is designed to actively reinforce the tactile modality and help fix data imbalance issues.

Dev: During training, the system is essentially forced to get better at predicting what the next force will be so that its tactile sense gets stronger without needing perfect labels for every single contact scenario.

Taro: That predictive guidance seems like a very clever way to inject structure into the learning process, specifically targeting those areas where tactile data might be sparse or noisy.

Rosa: So, the summary boils down to an adaptive system that learns to prioritize visual or tactile features based on the current force context, while using future force predictions as a training aid for getting better tactile representations.

The paper's summary: Dev: Looking at what this paper suggests as its main improvements, one major thing is that this approach avoids relying on manual labeling or fixed assumptions about which sense should be dominant.

Rosa: That’s a big deal because it means we don't have to spend all our time creating task-specific labels just to tune those modality weights; the system learns those adjustments automatically based on the data.

Taro: That automatic tuning capability really speaks to the scalability of the solution; it suggests that this method can handle a wider variety of manipulation tasks without needing bespoke tuning for every single one.

Dev: It also addresses a major problem with simpler concatenation methods, like three deeTacDex-P, which often suffer from tactile overfitting where the system just learns to rely too much on touch without truly understanding the underlying visual context <ref:2505.13982#pg2>.

Rosa: By incorporating that future force prediction guidance during training, they are directly tackling data imbalance by reinforcing the tactile modality precisely when it needs more attention.

Taro: And those results show a significant performance jump, with success rates reaching ninety-three percent on contact-rich tasks, which is substantial when you consider the baseline policies struggled with those specific tasks <ref:2505.13982#pg0>.

Dev: I agree, and I'm also paying attention to how this method shifts attention from visual dominance during reaching phases to tactile dominance during precise contact and manipulation phases.

Rosa: That stage-based modulation of focus based on the force context is what allows the policy to be more context-aware about when it needs which sense for optimal performance.

The paper's improvements: Taro: To wrap up our discussion on "Adaptive Visuo-Tactile Fusion with Predictive Force Attention for Dexterous Manipulation," the main implication is that we’ve seen a method where robots can dynamically manage their sensory inputs based on physical interaction forces during manipulation.

Dev: It means we have a system that doesn't rely on fixed sensor assumptions, but rather learns the correct balance between vision and touch at every point in the sequence.

Rosa: This adaptability, combined with the self-supervised future force prediction task for tactile reinforcement during training, points toward a more flexible control architecture that handles complex physical tasks much better than methods relying on static rules.

Taro: I still want to emphasize how this system can be used to build agents that are more capable of handling unstructured environments by anticipating physical consequences through force prediction.

Dev: From an engineering standpoint, the latency is what we need to focus on, so ensuring these adaptive weight adjustments happen fast enough for real-time operation is the main hurdle we've identified.

Rosa: It’s certainly an exciting direction for field robotics because it shows how physics can be used not just as a passive measurement but as an active control signal.

Conclusion: Rosa: So, to wrap up our discussion on "Adaptive Visuo-Tactile Fusion with Predictive Force Attention for Dexterous Manipulation," we’ve seen how this work introduces a force-guided attention fusion module that uses measured forces to dynamically adjust how much visual and tactile features the robot prioritizes.

Dev: It really means we have a system that doesn't rely on fixed sensor assumptions, but rather learns the correct balance between vision and touch at every point in the manipulation sequence.

Taro: This adaptability, combined with the self-supervised future force prediction task for tactile reinforcement during training, points toward a more flexible control architecture that handles complex physical tasks much better than methods relying on static rules.

Rosa: I still want to emphasize how this system can be used to build agents that are more capable of handling unstructured environments by anticipating physical consequences through force prediction.

Dev: From an engineering standpoint, the latency is what we need to focus on, so ensuring these adaptive weight adjustments happen fast enough for real-time operation is the main hurdle we've identified.

Taro: I think this paper provides a strong framework for future research into embodied AI that needs to focus on making the interaction dynamics between perception and action more inherently coupled.

Rosa: It’s certainly an exciting direction for field robotics because it shows how physics can be used not just as a passive measurement but as an active control signal.

Dev: Indeed, the way they handle the uncertainty through that auxiliary task is a promising avenue to explore for making these loops more stable in real-world deployment.

Taro: I think we need to keep watching this area closely because coupling perception and action based on actual physical feedback seems like a necessary step toward truly intelligent embodied systems.

cs.RO

Submitted: 2025-05-20

Updated: 2025-07-21

Journal ref: 2025 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pp. 3232-3239 (2025)

DOI: 10.1109/IROS60139.2025.11246476

Project page: https://adaptac-dex.github.io

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 90/100

The gist: Effectively utilizing multi-sensory data is important for robots to generalize across diverse tasks, but fusing heterogeneous modalities like vision and touch remains challenging because each

Key concepts

Force-Guided Attention Fusion
This core mechanism uses the actual force signal as a query to decide how much attention to pay to visual and tactile features. It allows the robot to dynamically prioritize touch information when manipulation requires it, instead of always relying on vision.
Self-Supervised Future Force Prediction
During training, a separate neural network is tasked with predicting the net force that will occur in several future steps. This prediction helps reinforce the tactile modality and addresses data imbalance by providing context about upcoming contact forces.
Adaptive Weighting
The system learns to adjust the blending weights between visual features and tactile features dynamically. This means if a task stage demands fine motor control, the robot increases reliance on touch; if it requires broad spatial awareness, it leans more on vision.
Sparse Encoders
These are specialized neural network components used to efficiently extract meaningful information from raw sensory data like point clouds (vision) and tactile readings. They compress the complex input into a manageable set of features for the subsequent fusion process.

Terminology

Summary

Effectively utilizing multi-sensory data is important for robots to generalize across diverse tasks, but fusing heterogeneous modalities like vision and touch remains challenging because each requires different levels of attention at different manipulation stages. This work proposes a force-guided attention fusion module that adaptively adjusts the weights of visual and tactile features without human labeling, complemented by a self-supervised future force prediction auxiliary task to reinforce the tactile modality and improve data imbalance, achieving an average success rate of 93% across three fine-grained, contact-rich tasks.

The gist

Our policy leverages visual and tactile information to predict future force and combines it with the observed force to adaptively adjust the attention of different modalities at different stages of dexterous manipulation.

How it works: Force-Guided Attention Fusion

The core mechanism introduces a force-guided attention module where force signals serve as queries, and visual/tactile features as keys and values. The process involves several steps:

  1. Sparse encoders are used to extract features: We use a sparse encoder [41] to extract point cloud features Zpc = ϕpc(Opc) and a pretrained tactile encoder [3] to extract tactile features Ztac = ϕtac(Otac).

  2. Feature alignment is performed using separate MLPs: To align the feature dimensions, we apply separate MLPs to the point cloud and tactile features: e pc = gpc(Z pc) ∈ R 512, etac = gtac(Z tac) ∈ R 512.

  3. Query, Key, and Value representations are constructed: Q F = e FWQ, K = [e pc, etac]WK, V = [e pc, etac]WV, where WQ, WK, and WV are learnable projection matrices.

  4. Attention weights are computed between the force-guided query and the keys: The attention mechanism first computes fusion weights between the force-guided query and the keys from visual and tactile features: α pc, αtac = σ(QFKT / √dk).

  5. The fused representation is produced by applying these weights to the values: The resulting weights are then applied to the corresponding values to produce the fused representation: Z fuse = α pcV pc + α tacV tac. This enables adaptive weighting of visual and tactile features based on force, allowing the policy to prioritize touch when needed rather than always relying on vision [13].

How it works: Future Force Prediction and Guidance

To further guide attention modulation, a self-supervised auxiliary task is introduced. This involves:

  1. Training a diffusion head: We design a self-supervised future force prediction task during training, where we introduce a transformer-based [52] diffusion head [18] for future net force prediction. This head takes visual and tactile features as input to predict the future net force F np at the next n steps.

  2. Creating the guide force: We then concatenate the observed net force F nO and predicted future forces F np to form the guide force F n g, and project it as the query: Q F = gF ([F nO, F np])WQ, where gF is an MLP that projects the force vectors into the query space.

  3. Guiding fusion: This query guides the attention module using both current and future contact information. This ensures context-aware modality weighting and efficient attention adjustment.

How it works: Visuo-Tactile Policy Learning

The fused features are then integrated into an imitation learning framework based on a 3D diffusion model policy (RISE [22]). The training objective is defined as: L = Lπ + αLffp, where Lπ is the policy loss and Lffp is the future force prediction loss, with α being a hyperparameter. This allows the policy to learn from both observed and predicted forces.

Experimental Validation

The method was tested on three dexterous, contact-rich manipulation tasks: (i) Open Box, (ii) Reorientation, and (iii) Flipping a dish sponge. The results showed that Ours outperforms all baselines across all tasks. Specifically, the vision-only baseline (RISE) struggled with the Flip task. Ablation studies confirmed the necessity of both components: without FFP and FFG, success rates were significantly lower (e.g., 50% for Flip), whereas incorporating observed force prediction and guidance increased success to 70%. Analysis showed that during reaching stages, more attention is given to visual input, shifting towards tactile features during contact. Furthermore, the policy demonstrated strong generalization on unseen objects with a 75% success rate.

Improvements for AI systems

Here are the specific improvements that can be made to current AI systems by adopting the methodology described in this paper, and what those improved systems will be capable of:


  1. The proposed architecture (AdapTac) enables robots to perform complex, contact-rich dexterous manipulation tasks by dynamically weighting visual and tactile inputs based on the current physical interaction state (force).

  2. The system can execute fine-grained tasks such as flipping a dish sponge or precisely reorienting a cup by effectively prioritizing tactile feedback during critical force application phases, overcoming limitations of vision-only policies (like RISE) that fail when precise force control is required.

  3. The integration of the self-supervised future force prediction auxiliary task allows the robot to anticipate physical consequences of its actions, leading to significantly improved generalization and robustness against data imbalance in tactile modalities.

  4. The resulting system can achieve high success rates (up to 93%) on complex, real-world tasks by learning appropriate attention adjustments at every stage—shifting from visual dominance during reaching phases to tactile dominance during precise contact and manipulation phases—without requiring manual labeling of task-specific attention weights.

  5. The improved AI system will exhibit superior performance in tasks where force application is paramount (e.g., maintaining a firm grip, applying the correct torque for reorientation), as it avoids the instability caused by manually defined contact thresholds (like FoAR) and mitigates tactile overfitting that plagues simple concatenation methods (like 3DTacDex-P).

  6. The system demonstrates strong generalization on unseen objects with varying geometries, meaning it can successfully perform novel manipulation tasks or handle objects not explicitly seen during training, as evidenced by achieving a 75% success rate on unseen items.

Sources

Related papers