HeteroGenManip: Generalizable Manipulation For Heterogeneous Object Interactions
cs.RO, cs.AI
Submitted: 2026-05-11
Updated: 2026-09-06
Project page: https://yzmyalier.github.io/HeteroGenManip
License: http://creativecommons.org/licenses/by/4.0/
The gist: Generalizable manipulation involving cross-type object interactions is a critical yet challenging capability in robotics.
Terminology
Abstract
Generalizable manipulation involving cross-type object interactions is a critical yet challenging capability in robotics. To reliably accomplish such tasks, robots must address two fundamental challenges: "where to manipulate" (contact point localization) and "how to manipulate" (subsequent interaction trajectory planning). Existing foundation-model-based approaches often adopt end-to-end learning that obscures the distinction between these stages, exacerbating error accumulation in long-horizon tasks. Furthermore, they typically rely on a single uniform model, which fails to capture the diverse, category-specific features required for heterogeneous objects. To overcome these limitations, we propose HeteroGenManip, a task-conditioned, two-stage framework designed to decouple initial grasp from complex interaction execution. First, Foundation-Correspondence-Guided Grasp module leverages structural priors to align the initial contact state, thereby significantly reducing the pose uncertainty of grasping. Subsequently, Multi-Foundation-Model Diffusion Policy (MFMDP) routes objects to category-specialized foundation models, integrating fine-grained geometric information with highly-variable part features via a dual-stream cross-attention mechanism. Experimental evaluations demonstrate that HeteroGenManip achieves robust intra-category shape and pose generalization. The framework achieves an average 31% performance improvement in simulation tasks with broad type setting, alongside a 36.7% gain across four real-world tasks with different interaction types.
Sources
- RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation
- Segment Anything
- RAM: Retrieval-Based Affordance Transfer for Generalizable Zero-Shot Robotic Manipulation
- BEHAVIOR-1K: A Human-Centered, Embodied AI Benchmark with 1,000 Everyday Activities and Realistic Simulation
- ManiCM: Real-time 3D Diffusion Policy via Consistency Model for Robotic Manipulation
- RISE: 3D Perception Makes Real-World Robot Imitation Simple and Effective
- SKIL: Semantic Keypoint Imitation Learning for Generalizable Data-efficient Manipulation
- AffordDP: Generalizable Diffusion Policy with Transferable Affordance
- Leveraging Locality to Boost Sample Efficiency in Robotic Manipulation
- robosuite: A Modular Simulation Framework and Benchmark for Robot Learning
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving