RoboEdit: Turning Human Manipulation Videos into Scalable Robot Experience
cs.RO
Submitted: 2026-08-19
Updated: 2026-09-27
Project page: https://roboedit.github.io
Terminology
Sources
- Dexterous Imitation Made Easy: A Learning-Based Framework for Efficient Dexterous Manipulation
- HOT3D: Hand and Object Tracking in 3D from Egocentric Multi-View Videos
- $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control
- VINO: A Unified Visual Generator with Interleaved OmniModal Context
- H2R-Grounder: A Paired-Data-Free Paradigm for Translating Human Interaction Videos into Physically Grounded Robot Videos
- Project Aria: A New Tool for Egocentric Multi-Modal AI Research
- GigaHands: A Massive Annotated Dataset of Bimanual Hand Activities
- DexPilot: Vision Based Teleoperation of Dexterous Robotic Hand-Arm System
- OpenVE-3M: A Large-Scale High-Quality Dataset for Instruction-Guided Video Editing
- LoRA: Low-Rank Adaptation of Large Language Models
- Hand Pose Estimation via Latent 2.5D Heatmap Regression
- AnyV2V: A Tuning-Free Framework For Any Video-to-Video Editing Tasks
- Kinematic Motion Retargeting for Contact-Rich Anthropomorphic Manipulations
- H2R: A Human-to-Robot Data Augmentation for Robot Pre-training from Videos
- OmniV2V: Versatile Video Generation and Editing via Dynamic Content Manipulation
- Kiwi-Edit: Versatile Video Editing via Instruction and Reference Guidance
- EditCtrl: Disentangled Local and Global Control for Real-Time Generative Video Editing
- TACO: Benchmarking Generalizable Bimanual Tool-ACtion-Object Understanding
- LIV: Language-Image Representations and Rewards for Robotic Control
- DexVIP: Learning Dexterous Grasping with Human Hand Pose Priors from Video
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving