ContactWorld: What Representations Matter for Vision-Tactile Latent World Models in Contact-Rich Manipulation
cs.RO
Submitted: 2026-06-11
Updated: 2026-09-24
Project page: https://contact-world.github.io
Terminology
Sources
- Learning Precise, Contact-Rich Manipulation through Uncalibrated Tactile Skins
- Reactive Diffusion Policy: Slow-Fast Visual-Tactile Policy Learning for Contact-Rich Manipulation
- Touch begins where vision ends: Generalizable policies for contact-rich manipulation
- V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
- LeWorldModel: Stable End-to-End Joint-Embedding Predictive Architecture from Pixels
- What Drives Success in Physical Planning with Joint-Embedding Predictive World Models?
- Being-H0.7: A Latent World-Action Model from Egocentric Videos
- FLARE: Robot Learning with Implicit World Modeling
- Reconstruction or Semantics? What Makes a Latent Space Useful for Robotic World Models
- DINO-WM: World Models on Pre-trained Visual Features enable Zero-shot Planning
- Visuo-Tactile World Models
- ManiFeel: Benchmarking and Understanding Visuotactile Manipulation Policy Learning
- 3D-ViTac: Learning Fine-Grained Manipulation with Visuo-Tactile Sensing
- OmniVTA: Visuo-Tactile World Modeling for Contact-Rich Robotic Manipulation
- World Guidance: World Modeling in Condition Space for Action Generation
- World Models
- Learning Latent Dynamics for Planning from Pixels
- Mastering Diverse Domains through World Models
- MC-JEPA: A Joint-Embedding Predictive Architecture for Self-Supervised Learning of Motion and Content Features
- Graph-level Representation Learning with Joint-Embedding Predictive Architectures
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving