UMI-Bridge: Action-Anchored Latent Alignment across Human and Robot Manipulation Data
cs.RO
Submitted: 2026-09-16
Updated: 2026-09-28
Terminology
Sources
- OpenVLA: An Open-Source Vision-Language-Action Model
- $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization
- DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset
- EgoDex: Learning Dexterous Manipulation from Large-Scale Egocentric Video
- UniVLA: Learning to Act Anywhere with Task-centric Latent Actions
- MVP-LAM: Learning Action-Centric Latent Action via Cross-Viewpoint Reconstruction
- Causally Debiased Latent Action Model for Embodied Action-Conditioned World Models
- HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data Alone
- RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
- $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control
- Learning to Act without Actions
- Genie: Generative Interactive Environments
- LARA: Latent Action Representation Alignment for Vision-Language-Action Models
- WALA Learning Executable Latent Actions from Action-Labeled Demonstrations and Action-Free Videos
- EgoMimic: Scaling Imitation Learning via Egocentric Video
- Translation as a Bridging Action: Transferring Manipulation Skills from Humans to Robots
- ACE-Ego-0: Unifying Egocentric Human and Robotic Data for VLA Pretraining
- EgoGuide: Egocentric Guidance for Efficient Robot-Free Demonstration Collection and Learning
- Bridging Handheld and Teleoperated Supervision for Contact-Rich Manipulation via State-Gated Experts
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving