TacDyn-WAM: Learning Implicit Tactile Dynamics in a Heterogeneous Visuo-Tactile World Action Model

arXiv:2610.00638 · cs.RO · Submitted 2026-09-30 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: I'm Rosa, and with me are Dev and Taro, guest researcher.

Dev: Today's paper: "TacDyn-WAM: Learning Implicit Tactile Dynamics in a Heterogeneous Visuo-Tactile World Action Model".

Rosa: World action models improve robotic manipulation by conditioning actions on predicted futures, yet existing tactile variants largely inherit video-generation pipelines that reconstruct future tactile observations through iterative denoising.

Dev: First, who's behind it and why it matters.

Paper summary: Rosa: We started by looking at this paper, "TacDyn-WAM: Learning Implicit Tactile Dynamics in a Heterogeneous Visuo-Tactile World Action Model," which essentially tackles the problem that existing world action models rely on reconstructing future tactile observations through iterative denoising pipelines.

Dev: It claims that this approach is problematic because deployment drift can cause small shifts in contact position or force to substantially alter tactile pixels, which can mislead policies even when the actual contact evolution remains predictable.

Taro: So the paper argues for a shift toward modeling implicit tactile dynamics: predicting how contact itself changes, rather than reconstructing the visual appearance of that change.

Rosa: TacDyn-WAM proposes a heterogeneous visuo-tactile world action model that predicts these implicit tactile dynamics within a representation space called TacRep, which is designed to capture this evolution.

Dev: This means the core claim is that predicting contact evolution in this dedicated space is more robust than trying to predict future tactile pixel observations directly.

Taro: The paper argues that by doing this, they are moving away from methods where latent prediction isn't suited for forecasting tactile evolution, contrasting it with other models like N0-VTLA or V-JEPA two point one which have shown limitations in modeling frame-to-frame dynamics <ref:2610.00638#pg0>.

Rosa: By focusing on predicting the evolution in a representation space tailored for dynamics, they aim to create a system that handles unpredictable real-world tactile changes much better than methods based on static image encoders like DINOv2 or single-frame predictors.

Dev: This distinction is important because it addresses the issue of how different representation spaces suit different forecasting tasks; TacDyn-WAM’s design makes it specifically fit for capturing the temporal changes in touch.

Taro: The necessity of this modeling comes from autonomy research where we need systems that can handle unexpected world behavior, and this paper provides a mechanism for predicting that unexpected interaction without relying on perfect pixel fidelity.

Rosa: It really boils down to the idea that if we can predict the underlying physical interaction—the sliding or deformation—we get a more reliable action model than if we are just guessing what the next set of pixels will look like.

Conclusion: Rosa: So, looking at "TacDyn-WAM: Learning Implicit Tactile Dynamics in a Heterogeneous Visuo-Tactile World Action Model," the authors are Yilun Chen and the rest of the team, and they've shown how to model implicit tactile dynamics.

Dev: The implications are that this moves robotic action models toward systems that don't get confused by minor, unpredictable environmental changes during deployment because they focus on contact physics instead of just pixels.

Taro: This suggests that for future autonomy, we might see a trend where learning systems prioritize predicting the physical interaction over visual fidelity when dealing with unpredictable touch.

Rosa: In simple terms, this means robots become better at handling situations where things get bumped or deformed because they are learning the dynamics of the contact itself.

Dev: It’s about building in stability by making sure that even if the visual input shifts slightly, the underlying predictive model for contact remains sound.

Taro: From an autonomy viewpoint, this could mean a significant step toward more reliable systems when faced with novel physical interactions that weren't perfectly covered in training data.

Rosa: So, ultimately, this work offers a way to achieve more stable robotic manipulation by predicting the dynamics of contact evolution rather than just reconstructing the visual outcome.

Enyi Wang, Mingxin Wang, Quan Shi, Hetian Guo, Hongyu Wang, Xi Wang, Bin Qian, Yupeng Zheng

Institute for AI Industry Research (AIR), Tsinghua University Institute of Automation, Chinese Academy of Sciences, The Hong Kong University of Science and Technology (Guangzhou), Tsinghua University, Fudan University, TARS Robotics

cs.RO

Submitted: 2026-09-30

Updated: 2026-09-30

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 91/100

The gist: World action models improve robotic manipulation by conditioning actions on predicted futures, yet existing tactile variants largely inherit video-generation pipelines that reconstruct future tactile

Key concepts

TacRep
This is a specialized representation space designed to capture how contact changes over time rather than exact pixel details. It is trained using methods that focus on the trend of contact and regularized by relating local patch structures to ensure the representation is robust against deployment drift.
Implicit Tactile Dynamics Expert (ITDE)
This expert predicts how tactile sensations will evolve in TacRep space over multiple future time steps in a single calculation. It uses context tokens from TacRep and future queries to generate predictions about the change in tactile features, focusing learning on areas where contact is expanding or sliding.
Heterogeneous Visuo-Tactile World Action Model
This architecture combines two separate pathways: a Visual Generation Expert that handles vision and an Implicit Tactile Dynamics Expert that handles touch. These experts interact through joint attention, allowing the model to process both visual and tactile information simultaneously for action prediction.
Joint Attention
This mechanism allows the different expert pathways (visual and tactile) to communicate effectively at every layer of the model. It enables them to coordinate their predictions based on each other's current state, ensuring that the visual and tactile predictions are mutually informed during training.

Terminology

Summary

World action models improve robotic manipulation by conditioning actions on predicted futures, yet existing tactile variants largely inherit video-generation pipelines that reconstruct future tactile observations through iterative denoising. TacDyn-WAM introduces a heterogeneous visuo-tactile world action model that predicts implicit tactile dynamics rather than reconstructing future tactile observations, which is crucial because pixel-level changes under deployment drift can mislead policies.

The gist: TacDyn-WAM learns to predict contact evolution in a dynamics-aware representation space called TacRep, enabling one-pass multi-horizon prediction without iterative denoising or pixel reconstruction.

TacDyn-WAM Architecture and Experts

TacDyn-WAM is a heterogeneous visuo-tactile world action model featuring two distinct world expert pathways coupled through joint attention. The Visual Generation Expert predicts future visual latents, while the Implicit Tactile Dynamics Expert (ITDE) predicts tactile evolution within the TacRep space. This architecture allows for separate parameters and target spaces for each expert, enabling them to interact via joint attention at every layer of the model. A separate Tactile Understanding Memory provides direct access to the current tactile state as a read-only component, avoiding the need for a fifth expert.

TacRep: Dynamics-Aware Tactile Representation Space

TacRep is introduced as a representation space designed to capture contact evolution rather than exact pixel details, which mitigates issues arising from deployment drift. It is trained through Dynamics Target Encoder Ebar and regularized by Relational Structure Distillation. The training involves:

  1. Tactile Dynamics Prediction (TDP): A light predictor learns to recover the target-encoder features of masked tokens, ensuring features respond to the trend of contact rather than static pixel patterns.

  2. Relational Structure Distillation (RSD): A frozen DINOv2 ViT-B serves as a Relational Structure Teacher to regularize local patch relations, aligning them without forcing feature values to match.

Implicit Tactile Dynamics Expert (ITDE)

The ITDE operates on the tactile block T encoded by the frozen TacRep Encoder Ebar and predicts future representations and their changes at multiple horizons in a single forward pass. It utilizes Context Tokens derived from the TacRep Encoder, which are combined with Future Token Queries. The ITDE predicts both future representations and their changes (Future Readout and Delta Readout) by applying the frozen TacRep Encoder to single frames to generate targets for prediction. The loss function employs a Patch-Adaptive Loss Reweighting scheme, where patch weights are proportional to the magnitude of ground-truth change, focusing learning on regions undergoing contact expansion or sliding.

Staged Training and Integration

The model is trained using a four-stage procedure to progressively integrate the tactile modules with pretrained experts. Stage 1 focuses on TacRep Training to learn the representation space. Stage 2 trains the ITDE and its readouts with Tactile World Grounding. Stage 3 involves Tactile–Action Alignment, training the tactile memory and Action Expert using a loss function called Ltac, allowing the action block to read both tactile pathways. Stage 4 is Joint Training, optimizing all trainable modules with a combined loss: L = Lact + λvisLvis + λtacLtac.

Performance and Real-World Validation

TacDyn-WAM achieves state-of-the-art performance on UniVTAC using only provided demonstrations, reaching an average success rate of 81.5%. On five real-robot tasks, it reaches 71.0% average success with pertask demonstrations alone, and modest-scale pretraining raises this to 85.0%. Ablations confirm the benefits of both tactile pathways and TacRep over pixel-reconstruction methods, showing that tactile prediction provides the largest gains on tasks involving contact rotation or sustained deformation. Furthermore, inference latency is low; TacDyn-WAM computes its predictions in compact latent spaces without iterative denoising, resulting in a constant cost over an episode.

Contributions Summary

The work proposes TacDyn-WAM to jointly predict future visual latents and tactile evolution in distinct target spaces via joint attention. It introduces TacRep for one-pass multi-horizon prediction, and it demonstrates state-of-the-art performance on UniVTAC using only demonstrations, validating the design through ablation studies showing the necessity of both tactile pathways and memory. The method successfully outperforms vision-only and existing tactile baselines on real robot tasks when modest pretraining is applied.

Limitations

A limitation noted is that TacRep and the ITDE are trained for a single type of visuo-tactile sensor; generalizing them across different modalities, such as force or taxel arrays, remains an open challenge. The model's performance on certain tasks, like Insert HDMI, is less sensitive to tactile prediction because their demonstrations contain little tactile variation.

Improvements for AI systems

Here are specific improvements that can be made to AI systems based on the TacDyn-WAM framework:

  1. Improve robustness under deployment drift by shifting tactile world model prediction from pixel-level reconstruction to dynamics-aware representation learning (TacRep). The improved system will predict physical contact evolution (e.g., imprint deepening, sliding, rotation) in a target space that tolerates minor tactile pixel shifts, making predictions reliable even when contact position or force slightly deviates from demonstrations.

  2. Enhance predictive foresight across multiple horizons by training the Implicit Tactile Dynamics Expert (ITDE) to predict both the future state and the change in state (delta representation). This will allow for more nuanced action planning that anticipates not just an endpoint, but the entire contact evolution sequence, leading to more stable manipulation during complex tasks like Pull-out Key or Unscrew Cup Lid.

  3. Achieve high performance on contact-rich manipulation tasks using only demonstration data (zero-shot generalization). The improved system will achieve state-of-the-art performance on benchmarks like UniVTAC, demonstrating that implicit dynamics modeling is sufficient for robust policy learning without requiring massive amounts of pretraining data for every new task.

  4. Improve generalization to unseen contact modalities by establishing a modular, expert-structured fusion mechanism rather than relying on a single learned sparse router. The improved system will be capable of integrating future tactile sensors, including force or taxel arrays, by adapting the appropriate modality-specific experts without needing to retrain the entire world model from scratch.

  5. Reduce inference latency significantly compared to iterative video denoising methods (like LingBot-VA). The improved system will perform predictions in a single forward pass using compact latent spaces (TacRep) and reuse learned blocks across action chunks, resulting in constant, low-latency prediction suitable for real-time robotic deployment.

  6. Enable rapid adaptation to new contact situations through a read-only Tactile Understanding Memory. The improved system will maintain a precise current state representation that the Action Expert can directly access, allowing for immediate policy adjustments based on the most recent tactile feedback without waiting for complex re-computation or full model retraining.

  7. Improve learning efficiency by employing staged training to progressively integrate components. The improved system will leverage pretrained experts (like those from large-scale visuo-tactile pretraining) and only train newly introduced modules (like the ITDE) on task-specific data, preventing catastrophic forgetting and allowing for faster fine-tuning on real-world tasks.

Sources

Related papers