Agile-WAM: An Agile Tactile World Action Model for Contact-Rich Robot Control
cs.RO, cs.LG
Submitted: 2026-09-17
Updated: 2026-09-18
Project page: https://hanchuzhou.github.io/TARO
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: World Action Models (WAMs) advance beyond conventional visuomotor policies by jointly predicting future world states and robot actions, enabling the policy to learn physical dynamics that support
Terminology
Abstract
World Action Models (WAMs) advance beyond conventional visuomotor policies by jointly predicting future world states and robot actions, enabling the policy to learn physical dynamics that support effective control. However, recent tactile WAMs often rely on large-scale pretrained generative backbones to capture contact-rich physical dynamics, which limit their inference efficiency and flexible deployment. In this paper, we present, an agile tactile World Action Model for contact-rich robot control. encodes visual and tactile observations into a shared latent that serves as the source of a direct vision-tactile-to-action flow-matching process, which can jointly generate latent representations of action chunks and future visual/tactile latents. A key observation is that vision and tactile signals evolve at inherently different timescales: adjacent visual frames are often highly similar, whereas tactile signals can change abruptly upon contact. We therefore introduce multi-horizon multimodal prediction in, which provides supervision for visual latent at a larger temporal offset while predicting the tactile latent in the next frame to capture fine-grained contact dynamics. Across nine simulated and five real-world contact-rich manipulation tasks, demonstrates strong and robust performance, outperforming the strongest baseline in success rate while maintaining low inference latency. In particular, in five real-world experiments, yields a relative gain of 29.4% in overall success rates while achieving inference latency of 11.9 ms. These results demonstrate that multimodal WAM can be achieved with an agile architecture suitable for precise and high-frequency robot control. More details are available on our project page: https://hanchuzhou.github.io/TARO project page/.
Sources
- IN-RIL: Interleaved Reinforcement and Imitation Learning for Policy Fine-Tuning
- World Action Models are Zero-shot Policies
- Reactive Diffusion Policy: Slow-Fast Visual-Tactile Policy Learning for Contact-Rich Manipulation
- FACTR: Force-Attending Curriculum Training for Contact-Rich Policy Learning
- VTAM: Video-Tactile-Action Models for Complex Physical Interaction Beyond VLAs
- Tactile-WAM: Touch-Aware World Action Model with Tactile Asymmetric Attention
- Dream-Tac: A Unified Tactile World Action Model for Contact-Rich Robot Manipulation
- AnyTouch 2: General Optical Tactile Representation Learning For Dynamic Tactile Perception
- Flow Matching for Generative Modeling
- Affordance-based Robot Manipulation with Flow Matching
- Learning Robotic Manipulation Policies from Point Clouds with Conditional Flow Matching
- $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control
- ManiFeel: Benchmarking and Understanding Visuotactile Manipulation Policy Learning
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving