From Solo to Ensemble: A Hierarchical Framework for Composable Multi-Agent Human-Object Interaction
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.
Rosa: I'm Rosa, and with me are Dev and Taro, guest researcher.
Dev: Today's paper: "From Solo to Ensemble".
Rosa: The gist: This paper proposes a hierarchical framework that converts single-agent human-object interaction policies into reusable Object-oriented Motion Skills,
Dev: First, who's behind it and why it matters.
Paper summary: Rosa: So we're looking at this paper called "From Solo to Ensemble: A Hierarchical Framework for Composable Multi-Agent Human-Object Interaction." Basically, they are tackling the problem of taking skills that work for one person interacting with an object and making them useful when you have multiple people involved in a task.
Dev: That sounds like moving from single agents to a whole team coordination, Rosa. I'm wondering what their main idea is here—is it just about making the learning process easier, or are they changing how the control loop actually works?
Rosa: The core thesis of this paper is that you can take those existing single-agent interaction policies and convert them into something reusable called an Object-oriented Motion Skill. This skill lets you describe the desired motion of an object proxy instead of just controlling every joint for a specific task.
Taro: So, if I'm thinking about autonomy, this sounds like decoupling the high-level planning from the messy, low-level physics execution. It suggests that instead of one giant policy trying to figure everything out at once, you break it down into smaller pieces that can be mixed and matched.
Dev: Exactly. They propose a hierarchical framework where you have this low-level skill, the Object-oriented Motion Skill, which handles the contact-rich execution part, and then a high-level policy that uses these skills to coordinate multiple agents by focusing on those object-level proxy motions.
Rosa: The key mechanism they use is reinterpreting teacher rollouts—those pre-recorded demonstrations from single human interactions—as supervision for this new skill by extracting short-horizon object-proxy motions from those executed trajectories. This distillation process creates a low-level policy that operates within an Object-oriented Action Space.
Taro: I see what you mean, so they are distilling the specific task knowledge out of the teacher's full motion and putting it into this compact space defined by anchor displacements and the desired near-future motion of that object proxy. That lets a high-level policy just reason in terms of these object actions without getting bogged down in every single joint control detail.
Dev: And they do two training stages for this low-level skill. The first stage uses teacher guidance with PPO, conditioned on the future proxy motion, along with some online behavior cloning loss to get it started. Then there's a post-training stage where they optimize it just against randomly generated feasible object-proxy motions using only the object-proxy tracking reward.
Rosa: That second stage is interesting because they remove the teacher supervision and only train it to realize diverse proxy motions through physically plausible interaction, which prevents it from just becoming a generic full-body motion tracker.
Paper summary: Taro: So, if you're looking at a situation where the world misbehaves—like an object slips or you need a different kind of grip—this framework suggests that once the skill is frozen after those two stages, the high-level policy can still direct it to handle it using these object-oriented actions.
Dev: That’s the promise for multi-agent coordination. The high-level policy generates region-wise object-oriented actions conditioned on things like the shared object, the task goal, and where each agent is located, outputting one proxy motion condition for each humanoid in an N humanoids scenario.
Rosa: The experimental validation shows that this hierarchy transfers across different object geometries, interaction modes—like carrying versus pushing—and team sizes. They found that this method outperformed CooHOI on its supported carrying tasks and was the only one reported to test generalization beyond source carrying when testing pushing tasks.
Taro: That’s a strong result if you look at the complexity of cooperative manipulation. It suggests that this object-level coordination approach isn't just neat for simple setups; it actually scales well when you introduce more agents and different physical challenges.
Dev: But I do have to point out what they don't cover yet, which is a limitation they flag: the framework doesn't explicitly handle composition across different skill types or long-horizon temporal task structures, like multi-stage task decomposition where you change roles mid-task.
Rosa: So, while it’s great for coordinating multiple agents using reusable skills for short-horizon motions, it seems like integrating these skills into very complex plans that require changing the entire plan structure over many steps is still something they haven't fully addressed.
Taro: For someone listening who is just trying to understand the immediate impact, this means we can build systems where a team of robots or humans can cooperate on manipulating an object, and instead of needing a brand new policy for every single combination, we reuse the learned physics skill.
Dev: It shifts the focus from fine-tuning a single human to coordinating multiple agents through compact object-level proxy motions. That's the big shift in how we think about multi-agent interaction control.
Rosa: So, to wrap up this paper "From Solo to Ensemble: A Hierarchical Framework for Composable Multi-Agent Human-Object Interaction," it proposes converting single human policies into reusable Object-oriented Motion Skills by reinterpreting teacher rollouts as object-oriented action supervision.
Taro: The implication is that high-level planners don't need to deal with the low-level contact physics directly; they just need to manage these compact object proxy motions, which makes scaling up the team much more feasible.
Dev: It seems like a solid way to move forward when we want robust execution across different environments and team sizes, even if it doesn't solve every possible long-term planning problem yet.
Rosa: That’s the gist of what they achieved with this hierarchical framework for composable physics-based HOI.
Conclusion: Rosa: So we're wrapping up this look at "From Solo to Ensemble," which is about taking single human movements and turning them into reusable skills for coordinating multiple agents together, and who wrote this stuff is Peng, Abbeel, Levine, and a few others.
Dev: It’s interesting because it moves the focus away from just fine-tuning one person's robot to actually building a system where you can mix and match those skills for a whole team.
Taro: The authors are using teacher rollouts as supervision to distill these object-oriented motions, so the idea is that you don't need to learn every single joint action from scratch for every new task.
Rosa: It’s about this hierarchy where the high-level policy just manages these compact object actions and lets the low-level skill handle all the messy physical contact details.
Dev: And they show that this transfer works across different objects, different types of interaction, and even when you have a bigger team of agents involved.
Taro: The main thing is that it makes scaling up multi-agent interaction feasible because you don't need a whole new policy from scratch every time you want to do something new.
Rosa: It really changes how we think about building robotic teams by suggesting that composition through these reusable skills is a way to handle the complexity of real-world manipulation.
Dev: But they did flag something important, which is that this framework doesn't explicitly cover long-horizon tasks or changing roles mid-way through a complicated plan.
Rosa: So it’s a strong foundation for short-term coordination, but we still need to figure out how to make these skills work together over much longer sequences of actions.
Dev: That leaves us wondering how robust this system is when the task itself requires complex, multi-stage planning instead of just one smooth sequence.
Zekai Deng, Kangyi Chen, Ye Shi, Jingya Wang
ShanghaiTech University
cs.RO
Submitted: 2026-10-08
Updated: 2026-10-08
The gist: The gist: This paper proposes a hierarchical framework that converts single-agent human-object interaction policies into reusable Object-oriented Motion Skills, enabling high-level policies to
Key concepts
- Object-oriented Motion Skill
- A reusable low-level policy designed to execute short, physically feasible motions for an object proxy. It operates within a compact action space that bridges high-level reasoning and physical execution. This skill is distilled from task rollouts and can be frozen for later use.
- Object-oriented Action Space
- The interface between high-level task planning and low-level execution, defined by a sequence of anchor displacements. It specifies the desired near-future motion of the object proxy, providing a structured way for high-level policies to command object movement.
- Teacher Rollout Distillation
- A process where task-specific single-agent rollouts are reinterpreted as supervision for an Object-oriented Action Space. This extracts short-horizon object-proxy motions from executed trajectories, effectively converting task teachers into a reusable low-level skill.
- High-Level Policy Coordination
- The top layer of the framework that coordinates multiple agents by generating region-wise object actions. It conditions its decisions on the shared object, the overall task goal, and local agent states to manage multi-agent interactions through compact proxy motions.
Terminology
Summary
The gist: This paper proposes a hierarchical framework that converts single-agent human-object interaction policies into reusable Object-oriented Motion Skills, enabling high-level policies to coordinate multiple agents through compact object-level proxy motions.
How it works
-
The framework establishes a reusable low-level policy called the Object-oriented Motion Skill, which realizes short-horizon motions of object proxies through physically feasible humanoid control. This skill operates within an Object-oriented Action Space defined as a compact interface between high-level task reasoning and contact-rich execution.
-
The low-level policy is conditioned on the humanoid state, the object proxy, and an object-oriented action to output humanoid control actions. The Object-oriented Action Space specifies a short-horizon sequence of anchor displacements, where mt specifies the desired near-future motion of the object proxy.
-
A unified skill is distilled from task-specific single-human HOI rollouts by reinterpreting teacher rollouts as object-oriented action supervision by extracting short-horizon object-proxy motions from executed trajectories. This converts task-specific teachers into an Object-oriented Motion Skill that can be frozen and reused as a low-level executor.
-
The high-level policy coordinates multiple agents by generating region-wise object-oriented actions conditioned on the shared object, task goal, agent states, and local manipulation regions. This shifts multi-agent HOI learning from direct contact-rich full-body control to compact object-level proxy-motion coordination.
Key Components and Training Stages
** Low-Level Policy Training:**
The Object-oriented Motion Skill is trained in two consecutive stages. The first stage involves teacher-guided distillation where the student is optimized with PPO conditioned on the future proxy-motion condition, together with an auxiliary online behavior cloning loss. This stage uses a mixed rollout policy π(p)mix that exposes the student to feasible interaction states during early training. The core reward in this stage is object-proxy trajectory tracking, defined as robjt = exp − 1/K Σ k=0 ∆dptk − ∆p⋆t,k.
** Post-Training for Robustness:**
In the second stage of skill learning, the policy is post-trained with randomly generated feasible object-proxy motions. In this stage, teacher action replacement and behavior cloning loss are removed, and the low-level skill is optimized only with the object-proxy tracking reward under randomly generated proxy-motion conditions. This prevents the policy from becoming a full-body motion tracker and instead encourages it to realize diverse proxy motions through physically plausible humanoid-object interaction.
** High-Level Policy Training:**
The high-level policy is trained on top of the frozen skill, generating region-wise object-oriented actions conditioned on the shared object, task goal, agent states, and local manipulation regions. The high-level policy is optimized with PPO using task-specific rewards. For multi-agent tasks with N humanoids, the policy outputs one proxy-motion condition mi t for each humanoid.
Experimental Validation
The framework was validated on cooperative HOI tasks, showing that an Object-oriented Motion Skill supports robust proxy-motion execution and enables highlevel policy learning across different interaction types, object geometries, and team sizes. Experiments showed that the hierarchy transfers across object geometries, interaction modes, and team sizes. The method outperformed CooHOI on its supported carrying tasks; for pushing tasks outside its released carrying-focused setting, only the method was reported to test generalization beyond source carrying. The Object-oriented Motion Skill achieved the best tracking performance on teacher-rollout targets and showed improved target-reaching and trajectory success rates compared to baselines.
Limitations
The framework does not explicitly address composition across heterogeneous skill types or long-horizon temporal task structures, such as multi-stage task decomposition and dynamic role reassignment.
Conclusion
The framework presents a hierarchical framework for composable physics-based HOI that converts taskspecific single-human interaction policies into a reusable Object-oriented Motion Skill. By relabeling teacher rollouts with short-horizon object-proxy motions, the method learns a low-level skill under the Object-oriented Action Space, instead of directly fine-tuning single-human policies into task-specific cooperative controllers. The frozen skill realizes region-wise object-oriented actions through contact-rich full-body control, allowing downstream policies to transfer across different object geometries, interaction types, and team sizes. The framework successfully enables high-level policies to coordinate multiple humans through short-horizon object-proxy motions.
Acknowledgments.
References [1] X. B. Peng, P. Abbeel, S. Levine, and M.
Improvements for AI systems
-
The system can achieve robust multi-agent cooperative manipulation across diverse object geometries, interaction types, and team sizes because it
reinterprets teacher rollouts as object-oriented action supervision by extracting short-horizon object-proxy motions from executed trajectories.
This allows the high-level policy tocoordinate multiple agents by generating region-wise object-oriented actions conditioned on the shared object, task goal, agent states, and local manipulation regions,
shifting learning fromdirect contact-rich full-body control to compact object-level proxy-motion coordination.
-
The AI system will exhibit improved generalization across different interaction modes by utilizing a transferable low-level skill. The paper demonstrates that the
Objectoriented Motion Skill distilled from a single-human skill supports robust proxy-motion execution and enables highlevel policy learning across different interaction types, object geometries, and team sizes.
-
The framework will overcome the limitation of task-specific fine-tuning by decoupling low-level execution from high-level coordination. The system implements a "hierarchical and decoupled control framework, where the low-level Object-oriented Motion Skill handles contact-rich execution and the high-level policy performs compositional object-level coordination through proxy-motion planning."
-
The system will demonstrate enhanced robustness to out-of-distribution states through staged training. The post-training stage
exposes the policy to a broader range of object motions and forces it to realize them through physical humanoid-object interaction rather than relying on teacher-guided trajectories,
ensuring the skill isreusable
by optimizing withrandomly generated feasible object-proxy motions.
-
The system will prevent coordination deadlocks common in existing methods like CooHOI by explicitly planning short-horizon future states. The high-level policy
explicitly predicts short-horizon local future trajectories,
which, combined with a low-level skill post-trained with RL, allows the system tokeep producing meaningful corrective actions instead of becoming inactive
even fromimperfect contact states or near-static configurations.
Sources
- PhysHOI: Physics-Based Imitation of Dynamic Human-Object Interaction
- SkillMimic: Learning Basketball Interaction Skills from Demonstrations
- CLoSD: Closing the Loop between Simulation and Diffusion for multi-task character control
- CooHOI: Learning Cooperative Human-Object Interaction with Manipulated Object Dynamics
- Full-Body Articulated Human-Object Interaction
- CORE4D: A 4D Human-Object-Human Interaction Dataset for Collaborative Object REarrangement
- DartControl: A Diffusion-Based Autoregressive Motion Model for Real-Time Text-Driven Motion Control
- Controllable Human-Object Interaction Synthesis
- HOI-Diff: Text-Driven Synthesis of 3D Human-Object Interactions using Diffusion Models
- InterAgent: Physics-based Multi-agent Command Execution via Diffusion on Interaction Graphs
- SMPLOlympics: Sports Environments for Physically Simulated Humanoids
- Multi-Quadruped Cooperative Object Transport: Learning Decentralized Pinch-Lift-Move
- Isaac Gym: High Performance GPU-Based Physics Simulation For Robot Learning
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving