Counterfactual Video Generation Enables Scalable Humanoid Loco-Manipulation
cs.RO, cs.CV, cs.GR
Submitted: 2026-09-29
Updated: 2026-09-29
Terminology
Sources
- Visual Imitation Enables Contextual Humanoid Control
- World Action Models are Zero-shot Policies
- HDMI: Learning Interactive Humanoid Whole-Body Control from Human Videos
- OmniRetarget: Interaction-Preserving Data Generation for Humanoid Whole-Body Loco-Manipulation and Scene Interaction
- VLK: Learning Humanoid Loco-Manipulation from Synthetic Interactions in Reconstructed Scenes
- LessMimic: Long-Horizon Humanoid Interaction with Unified Distance Field Representations
- ULTRA: Unified Multimodal Control for Autonomous Humanoid Whole-Body Loco-Manipulation
- Generalizing from References using a Multi-Task Reference and Goal-Driven RL Framework
- Learning from Massive Human Videos for Universal Humanoid Pose Control
- HumanX: Toward Agile and Generalizable Humanoid Interaction Skills from Human Videos
- Large Video Planner Enables Generalizable Robot Control
- Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation
- Robotic Manipulation by Imitating Generated Videos Without Physical Demonstrations
- DreamGen: Unlocking Generalization in Robot Learning through Video World Models
- Seedance 2.0: Advancing Video Generation for World Complexity
- CRISP: Contact-Guided Real2Sim from Monocular Video with Planar Scene Primitives
- SAM 2: Segment Anything in Images and Videos
- SAM 3D: 3Dfy Anything in Images
- FoundationPose: Unified 6D Pose Estimation and Tracking of Novel Objects
- Perceptive Humanoid Parkour: Chaining Dynamic Human Skills via Motion Matching
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving