MASkillBlender: Decentralized Whole-Body Coordination for Multi-Humanoid Loco-Manipulation via Skill Blending
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.
Rosa: I'm Rosa, and with me are Dev and Taro, guest researcher.
Dev: Today's paper: "MASkillBlender: Decentralized Whole-Body Coordination for Multi-Humanoid Loco-Manipulation via Skill Blending".
Rosa: MASkillBlender proposes a general multi-agent reinforcement learning framework that enables decentralized whole-body coordination for multiple humanoids by learning a shared high-level policy over reusable pre-trained single-humanoid skills,
Dev: First, who's behind it and why it matters.
Title and authors: Rosa: Well, we’re diving into MASkillBlender today. It looks like this paper addresses the big hurdle in multi-humanoid coordination: getting them to move together without needing specific motion references for every single task.
Dev: That’s exactly what caught my attention; relying on task-specific motion data feels so brittle when you're dealing with real-world latency and sensor noise, Rosa. I'm curious how this framework manages the coordination aspect decentralization you mentioned in the title.
Taro: From an autonomy standpoint, it sounds promising because if the system learns a high-level policy over reusable skills, it should handle unexpected situations better when the world misbehaves. I want to know how robust this approach is when things go wrong outside of a perfectly controlled simulation setup.
Rosa: Exactly, Taro, that’s my main question for Dev: where does this framework actually work once we take it out of the lab? Can we expect it to operate reliably in a messy environment for an extended period?
Dev: I'm thinking about the loop rate and failure modes right away. If the system is relying on blending actions from multiple skills, what’s the expected latency impact on achieving smooth, coordinated motion? We need to know if this framework can keep up with real-time demands.
Taro: That ties into what I was saying earlier about misbehavior; if it’s decentralized, does each humanoid have enough local intelligence to recover gracefully when its partner deviates from the intended path?
Rosa: The paper suggests that by learning a shared decentralized high-level policy, it can manage this complexity without needing constant external guidance. It seems to be moving away from brittle motion tracking toward a more general coordination strategy.
Dev: Moving away from motion tracking is great for data efficiency, but the reliance on task-level rewards means we still have to engineer those rewards carefully so the policy learns the right behavior, doesn't it? That’s where my engineering concerns kick in.
Taro: If we can decouple the high-level coordination from low-level motor control using reusable skills like walking or reaching, then when things go wrong, the system might be able to fall back onto a known stable primitive skill to maintain some level of function.
Rosa: That’s a good point about fallback mechanisms. The paper emphasizes learning over reusable pre-trained single-humanoid skills, which implies that even if the high-level policy makes a suboptimal choice, the underlying motor primitives are already robust.
Dev: But how much control does that high-level policy actually have over those primitive skills? If it’s just blending actions, we need to make sure there isn't some hidden instability introduced by summing up those weighted skill actions.
Title and authors: Taro: The paper seems to tackle this by proposing a specific factorization for the total policy, which suggests a structured way the high-level intent maps onto the low-level execution units. I’m interested in how that structure handles dynamic changes in the task requirements.
Rosa: It appears they are using a structure where a decentralized high-level policy outputs goal vectors and weights, which are then clipped before feeding into deterministic local actions for each primitive skill. That clipping step seems important for ensuring the resulting motion is physically plausible.
Dev: The clipping function sounds like a crucial safety mechanism, but I wonder how sensitive the system is to errors in those raw goal vectors if they come from noisy local observations, Rosa? That input quality really dictates the output quality here.
Taro: If we consider real-world deployment, localization errors can corrupt those local observations. How does MASkillBlender account for that uncertainty when generating those goal vectors before they get clipped into the skill actions?
Rosa: The discussion around permutation-based data augmentation is interesting because it theoretically guarantees that augmenting samples doesn't change the policy gradient direction under the Homogeneous Markov Game formulation. This suggests a strong foundation for training efficiency.
Dev: That theoretical guarantee sounds powerful for sample efficiency, but in practice, does that augmentation actually help when we run into asynchronous control delays or significant communication lags during actual execution? I worry about translating that theoretical invariance into real-world stability.
Taro: If the system can handle those delays and errors by learning a policy that is inherently permutation invariant, it means the learned coordination logic itself is quite resilient to how information arrives from different agents at slightly different times.
Rosa: It seems they are pushing this framework across different embodiments, showing generalization between the nineteen-DoF Unitree H1 and the twenty-one-DoF Unitree G1 without much retraining, which speaks to the generality of skill blending.
Dev: That cross-embodiment transfer is impressive, but I need to see how it handles the inherent differences in joint degrees of freedom between those two robot types during complex maneuvers. Are there any specific kinematic constraints that cause issues?
Taro: The paper mentions that they tested this across three distinct tasks: Carry, Push, and Move, showing coordination across different physical goals. That variety is what really tests the generality of the learned skill compositions.
Rosa: Those three tasks cover a good range of physical interactions, from collaborative transport to moving toward targets while avoiding collisions. It shows that the system isn't just good at one isolated movement but can compose behaviors for complex scenarios.
Title and authors: Dev: The evaluation across those specific tasks is helpful, but I need more detail on the performance metrics when things fail during those tasks. For instance, what happens when a collision avoidance maneuver fails because of an unforeseen external perturbation?
Taro: That points to the limitations they acknowledge: while it generalizes well across embodiments, the robustness against severe physical perturbations needs further testing in deployment scenarios that are more chaotic than simulation.
Rosa: So we have this framework that learns coordination through skill blending, uses task rewards, and has theoretical backing for data augmentation efficiency. It’s certainly a substantial piece of work for multi-humanoid control research.
Dev: It is certainly substantial, Rosa, but the practical deployment hurdles—latency management and handling unpredictable environmental disturbances—still need rigorous attention before we can say it’s ready for real-world use at a high frequency.
Taro: I agree that the theoretical guarantees are strong foundations, but for true autonomy in dynamic settings, we still need to see how this system handles the kind of unpredictable failures that occur when interacting with an uncontrolled environment.
Rosa: So, to wrap up this discussion on MASkillBlender: it’s a framework focused on decoupling high-level coordination from low-level motion tracking by learning over reusable skills.
Dev: The core mechanism relies on blending actions from selected primitive skills using a decentralized high-level policy that outputs goal vectors and weights.
Taro: It offers a strong path toward scalable coordination because it leverages pre-trained skills instead of requiring extensive task-specific motion reference data for every new challenge.
Rosa: We’ve seen how the permutation augmentation provides theoretical backing for training efficiency, which is a key aspect of making the learning process more effective.
Dev: However, we still have questions about the practical performance under real-world latency and localization errors when running at high loop rates.
Taro: The paper does point out that while it generalizes well across different humanoid designs like the H1 and G1, testing its resilience against severe physical disturbances in unstructured settings is still an open area for future work.
Rosa: Overall, MASkillBlender provides a general framework for multi-humanoid locomotion by learning coordination through skill blending driven by task rewards.
Dev: I think the next step is to focus on hardening the execution pipeline to ensure that these learned high-level goals translate into reliable, low-latency physical movements under stress.
Taro: And from an autonomy view, we need more demonstrations of this system operating successfully when it encounters unexpected physical obstacles or control delays in a less controlled setting.
The paper's summary: Rosa: So, to recap, MASkillBlender is proposing a framework where multiple humanoids coordinate their whole-body movements by learning a shared high-level policy over pre-trained single-humanoid skills, all driven by simple task rewards without needing specific motion data for each move.
Dev: That’s the core idea summarized; it’s about building a flexible system that can compose complex actions just by knowing which basic skills to use and how to blend them together locally.
Taro: What I find really interesting is how they manage that coordination without needing explicit task-specific motion references, which suggests a much more general approach than what we usually see in these systems.
Rosa: Exactly, Taro, it means the system isn't brittle; it learns the *how* of cooperation through the reward signal rather than being explicitly programmed with every possible joint movement.
Dev: From an engineering standpoint, that decoupling of high-level planning from low-level motor control is smart because it should make the policy much more robust when dealing with unexpected noise or small deviations in local observations.
Taro: And I think that reliance on reusable skills like walking and reaching offers a good safety net; if the high-level policy makes a poor choice, the underlying primitives are already trained for stability.
Rosa: It really does feel like they’re moving toward systems that can handle a wider variety of physical challenges because they aren't tied down to one specific motion reference for every scenario.
Dev: But my concern remains about the execution speed; if this blending happens in real-time, we need to make sure that the overhead of selecting and blending those skills doesn't introduce unacceptable latency into the control loop.
Taro: That’s a fair point, Dev; I think they address that by focusing on a decentralized policy factorization, which suggests the system can handle that complexity locally rather than waiting for a centralized bottleneck.
Rosa: It certainly sounds like it has potential for real-world deployment because it shows generalization across different humanoid bodies and even different task types like carrying versus pushing.
Dev: The cross-embodiment transfer capability they showed between the H1 and G1 is compelling, but I still need to see how well that skill blending holds up when the kinematic differences between those two robots become more pronounced during complex maneuvers.
Taro: I think their evaluation across those varied tasks—carry, push, and move—gives us a good idea of its scope; it’s not just for simple navigation but for actual physical interaction in dynamic settings.
Rosa: And the theoretical backing with the permutation-based data augmentation is significant; it suggests that the training process itself is much more efficient than we might think in high-dimensional reinforcement learning.
Dev: That efficiency is crucial, Rosa, because training on massive datasets takes a long time; if their method really speeds up convergence without sacrificing performance on the actual coordination task, that's a big win for deployment timelines.
Taro: The implication here is that we could start moving away from painstakingly hand-engineering motion trajectories for every novel manipulation task and toward these more general, skill-based coordination models.
Rosa: That’s the big picture I’m excited about; it opens up the possibility of creating multi-agent systems that can tackle a much broader range of physical interaction problems in unstructured environments.
Dev: But we still have to figure out how to guarantee that the blending function itself doesn't introduce oscillations or instability when multiple agents try to execute conflicting skill blends simultaneously under stress.
Taro: That’s the next big question for autonomy research; it moves us from "can it do it" to "how reliably and safely can we trust its decisions when things get chaotic?"
Rosa: Exactly, so while the framework is impressive on paper and in simulation, we need to see sustained performance under real-world latency and physical disturbances before we can call this ready for any kind of deployment.
The paper's improvements: Rosa: So, to summarize the improvements in MASkillBlender, the authors focus on several key enhancements that make this framework more practical for real applications.
Dev: They've introduced a permutation-based data augmentation strategy, which they claim is theoretically sound under their Homogeneous Markov Game formulation and should significantly boost training efficiency.
Taro: That theoretical guarantee is important because it means we can train the system faster by reusing existing data in a smart way, which is something I’ve been thinking about for improving sample efficiency in my autonomy work.
Rosa: It really does, Taro; that augmentation helps the learning process stay focused on the policy gradient direction even when we're feeding it augmented samples from existing rollouts.
Dev: But as a controls engineer, I need to ask if that theoretical boost translates into actual stability during execution; does this augmentation strategy introduce any new types of failure modes or timing issues when we run it at high frequencies?
Taro: The authors seem to be addressing the robustness of the system against those things by proving that the policy gradient direction remains unchanged, which implies a more stable learning trajectory overall.
Rosa: It’s also worth mentioning their focus on generalization; they demonstrated that this skill blending logic works across different humanoid embodiments like the Unitree H1 and G1 without needing a complete overhaul of the model.
Dev: That cross-embodiment transfer is impressive, Rosa, but I still need to see how well those learned skills adapt when the underlying physical constraints—like joint limits or center of mass differences—vary significantly between robot designs.
Taro: The paper tackles this by learning a shared high-level policy that abstracts away some of those specific physical details, which is exactly what we want for true generalization in embodied autonomy.
Rosa: It sounds like the authors are pushing the system toward being more versatile, handling tasks that require different physical setups without needing completely separate training runs for each robot type.
Dev: My main concern with versatility is performance degradation; if the generalized policy has to compromise on precision because it’s trying to work across too many embodiments, we could lose the fine control needed for manipulation.
Taro: I think they balance that trade-off by using a hierarchical structure where low-level skills handle the specific motor details and high-level policy handles the task sequencing, which keeps things organized.
Rosa: So they’re trying to get the best of both worlds: broad applicability across different robots while maintaining precise execution through those reusable, pre-trained skill modules.
Dev: It’s a solid approach for deployment feasibility in simulations, but we still have to worry about the gap between simulation and reality when dealing with real-world factors like localization errors or asynchronous control delays.
Taro: That is the next big hurdle for autonomy; we need to see how this system handles those real-world uncertainties without needing constant, explicit error correction from an external system.
Rosa: Ultimately, MASkillBlender seems positioned to be a tool that moves us closer to building multi-agent systems capable of tackling complex physical tasks in unstructured settings by learning coordination through reusable skills.
Conclusion: Rosa: So we’ve got to wrap up our discussion on MASkillBlender, which is essentially this framework for decentralized whole-body coordination in multi-humanoid locomotion using skill blending.
Dev: It really is a piece of work that tackles the coordination challenge by learning a shared high-level policy over reusable single-humanoid skills, which cuts down on needing task-specific motion references.
Taro: I think the impact here is huge because it suggests we can move away from painstakingly programming every single movement sequence for every scenario toward a system that learns how to combine existing skills effectively.
Rosa: That’s the core implication: more general, adaptable coordination systems that don't get stuck on specific movements.
Dev: From a controls engineering standpoint, the decentralized execution looks promising because it keeps decision-making local and fast, but I still need to see if that blending mechanism introduces any hidden control instability when agents are operating at high loop rates.
Taro: I agree with Dev on the stability question; autonomy hinges on those moments where the world misbehaves, and we can’t afford a coordination system that becomes erratic under unexpected physical forces.
Rosa: The way they handle generalization across different robot bodies is another major takeaway, suggesting this isn't just a simulation trick but something that could be applied to real-world deployment across various hardware.
Dev: If it performs well outside of the lab, we need concrete data on how long it can maintain reliable coordination under fluctuating network conditions or when localization estimates drift.
Taro: I’m curious about the long-term vision here; if this skill blending works reliably, could we start seeing multi-agent systems that perform complex physical tasks in unstructured environments for extended periods?
Rosa: That's what we're hoping for—systems that are robust enough to handle the messy reality of physical interaction.
Dev: Before we wrap up, I just want to reiterate my focus on the execution pipeline; getting those high-level goals translated into smooth, low-latency physical actions under stress is the real engineering challenge here.
Taro: I think that’s exactly where future work needs to go—proving that this coordination logic can withstand the kind of chaotic, unpredictable physical interactions we see in actual field robotics.
Rosa: So, to sum it up, MASkillBlender offers a powerful new way for AI systems to achieve complex multi-humanoid locomotion by learning coordinated behavior through skill composition rather than explicit motion programming.
Dev: It’s a solid framework for tackling coordination complexity with good theoretical backing on data efficiency.
Taro: It certainly opens up possibilities for much more flexible and general embodied autonomy in the near future.
Yifan Hu, Luhang Hong, Mingkang Long, Danning Wang
Nanyang Technological University · Southeast University · Purple Mountain Laboratories
cs.RO, cs.LG, cs.MA
Submitted: 2026-10-01
Updated: 2026-10-01
Code: https://github.com/Humanoid-SkillBlender/SkillBlender
Project page: https://maskillblender.github.io
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 87/100
The gist: MASkillBlender proposes a general multi-agent reinforcement learning framework that enables decentralized whole-body coordination for multiple humanoids by learning a shared high-level policy over
Key concepts
- Skill Blending
- This involves combining actions from multiple pre-trained primitive skills to create a final, coordinated movement. The framework selects relevant skills based on the task and then blends their outputs using learned weights to produce the robot's total action, allowing for flexible and adaptive whole-body control.
- Homogeneous Markov Games (H-MG)
- This is a mathematical framework used to model multi-agent systems where all agents share the same reward structure and observations. It provides the theoretical foundation for learning decentralized policies in scenarios where agents need to coordinate their actions based on local information while optimizing a common goal.
- Permutation-based Data Augmentation
- This technique enhances training by creating synthetic data by permuting the order of collected samples. The theory proves that this augmentation does not change the direction of the policy gradient, meaning it improves learning efficiency without compromising the policy's convergence.
Terminology
Summary
MASkillBlender proposes a general multi-agent reinforcement learning framework that enables decentralized whole-body coordination for multiple humanoids by learning a shared high-level policy over reusable pre-trained single-humanoid skills, which is significant because it achieves this coordination using only task-level rewards without requiring task-specific motion references.
The gist
MASkillBlender proposes a general multi-agent reinforcement learning framework to achieve decentralized multi-humanoid whole-body coordination by learning a shared decentralized high-level policy over reusable pre-trained single-humanoid skills, enabling coordinated behaviors using only task-level rewards without requiring task-specific motion references.
How it works
The framework is built upon the concept of skill blending within a Multi-Agent Reinforcement Learning (MARL) setting, specifically formulated under the Homogeneous Markov Games (H-MG) framework. The core idea is to learn a shared decentralized high-level policy that coordinates the composition of reusable pre-trained primitive skills across humanoids using local observations.
The decentralized control policy factorizes as:
π tot(as t) = QN i=1 π(a i,to i(s t))
This policy is constructed hierarchically:
-
A reusable skill library is constructed for each humanoid embodiment, including
Walking, Reaching, and Squatting,
which are pre-trained using dense rewards to capture low-level control objectives likewhole-body stability and energy efficiency.
-
For a given task, the system first selects relevant primitive skills from the library based on task requirements.
-
A decentralized high-level policy, denoted as µhigh, takes the local observation as input and outputs a raw joint goal vector (g̃ i,t) and a raw joint weight vector (w˜ i,t).
-
A
skill-specific clipping function
is applied to g̃ i,t to produce the final goal vector g c i,t used by each primitive skill µ c low to generate its deterministic local action u c i,t. -
The final local action ui,t is generated by blending the actions from all selected primitive skills:
u i,t = Σ C c=1 w c i,t · u c i,t for all k = 1...d.
Data Efficiency through Augmentation
To improve learning efficiency and data utilization in simulation-heavy settings, MASkillBlender incorporates a permutation-based data augmentation strategy. This strategy exploits the permutation symmetry of homogeneous multi-humanoid systems.
The theoretical justification is established through Lemma 1 and Theorem 2, which prove that under the H-MG formulation where all agents share a team reward and observations are permutation-preserving, any collected sample (st, ot, at) and its permuted counterpart (Mst, Mot, M at) yield the same policy gradient.
This means augmenting a minibatch B with its permuted counterpart leaves the policy-gradient direction unchanged,
providing theoretical backing for this technique in on-policy MARL.
Training and Evaluation
The framework is trained using MAPPO (Multi-Agent PPO). A centralized critic Vϕ(s t), parameterized by ϕ, takes the global state as input to approximate the value function Vθ(s t) associated with the current joint policy π totθ. The decentralized execution allows each humanoid to operate independently using its local observation.
The evaluation is conducted across three representative tasks:
-
Carry (two humanoids collaboratively transport a large box).
-
Push (two humanoids collaboratively push a heavy box toward a target position on the ground plane).
-
Move (three humanoids move toward their assigned targets while avoiding inter-humanoid collisions).
The framework is evaluated across two humanoid embodiments: the 19-DoF Unitree H1 and the 21-DoF Unitree G1, demonstrating that it generalizes across different humanoid embodiments without substantial modifications.
Robustness is tested against localization errors, asynchronous control delays, and inter-robot communication delays,
maintaining effective coordination in these challenging scenarios.
Key Contributions
The main contributions of this work are:
We propose MASkillBlender, a general skill-based MARL framework that learns decentralized multi-humanoid coordination over reusable pre-trained single-humanoid skills using only task-level rewards, without requiring task-specific motion references.
**We incorporate a permutation-based data augmentation strategy into MASkillBlender and theoretically show that the augmented samples preserve the policy-gradient direction of the original samples.
Improvements for AI systems
Here are the specific improvements and capabilities derived from the MASkillBlender framework:
) Improved AI System Capabilities: Multi-Humanoid Loco-Manipulation via Skill Blending (MASkillBlender)
The MASkillBlender framework enables a generalized, decentralized system for coordinating multiple humanoid robots to perform complex, task-specific physical manipulation and locomotion. This system moves beyond brittle, task-specific motion tracking by learning to compose reusable motor skills into coordinated behaviors using only high-level task rewards.
Here are the specific improvements and what the improved AI system can do:
-
-
Decentralized Skill Composition for Scalable Coordination:
MASkillBlender learns a shared, decentralized high-level policy that dynamically composes a library of pre-trained, task-agnostic primitive skills (e.g., Walking, Reaching, Squatting). This allows the system to solve complex coordination problems by selecting and blending only the necessary skills for a given task (e.g., selecting 'Walking' and 'Reaching' for box-pushing) rather than learning low-level motor control from scratch for every scenario.
-
-
Task-Level Reward Driven Learning:
The system is trained using only task-level rewards, completely eliminating the need for extensive, costly reward engineering specific to every single coordination scenario. The policy focuses on high-level skill composition and inter-agent coordination, making the training process significantly more generalizable across different tasks and humanoid embodiments.
-
-
Theoretical Guarantee of Permutation Invariance:
The framework incorporates a permutation-based data augmentation strategy that is theoretically proven to preserve the policy-gradient direction of the original samples under the Homogeneous Markov Game (H-MG) formulation. This allows for more efficient on-policy training by generating augmented data from existing rollouts, significantly improving sample efficiency in high-dimensional MARL settings.
-
-
Robustness to Deployment Variations and Asymmetry:
The system demonstrates strong robustness under realistic deployment conditions:
-
It maintains effective coordination despite introduced control delays (0–20 ms) and localization errors (position/orientation perturbations).
-
It can execute long-horizon tasks from asymmetric initial configurations by decomposing the task into sequential phases (e.g., moving to an interaction region, then cooperative execution), reusing pre-trained primitive skills for navigation during the approach phase.
-
-
Cross-Simulator Transferability:
The learned policies exhibit strong Sim2Sim transfer capability, successfully deploying from simulation environments (like Isaac Gym) to real physics engines (like MuJoCo) with minimal retraining, proving the generalization of the skill-blended coordination logic across different simulators and humanoid morphologies (e.g., Unitree H1 to Unitree G1).
) What the Improved AI System Can Do:
The improved MASkillBlender system can perform complex, multi-humanoid physical tasks in unstructured environments with high reliability. Specifically, it can:
-
Perform collaborative heavy object transportation (Carry and Push) of various sizes.
-
Execute complex coordinated locomotion sequences (Move) across multiple agents while maintaining collision avoidance.
-
Adapt to different humanoid robot designs (embodiments) without needing a complete retraining cycle, simply by leveraging the generalized skill library.
-
Handle real-world deployment challenges such as communication latency and localization errors effectively due to its task decomposition strategy and inherent robustness.
Sources
- Humanoid Locomotion and Manipulation: Current Progress and Challenges in Control, Planning, and Learning
- ASAP: Aligning Simulation and Real-World Physics for Learning Agile Humanoid Whole-Body Skills
- Learning Context-Aware Motion Priors for Humanoid Control
- HumanoidBench: Simulated Humanoid Benchmark for Whole-Body Locomotion and Manipulation
- SkillBlender: Towards Versatile Humanoid Whole-Body Loco-Manipulation via Skill Blending
- Humanoid Locomotion as Next Token Prediction
- Advancing Humanoid Locomotion: Mastering Challenging Terrains with Denoising World Model Learning
- Learning Smooth Humanoid Locomotion through Lipschitz-Constrained Policies
- Physics-Guided Biomechanical Gait Adaptation for Humanoid Locomotion on Extreme Sloped Terrains
- Sim-to-Real Reinforcement Learning for Vision-Based Dexterous Manipulation on Humanoids
- Expressive Whole-Body Control for Humanoid Robots
- ExBody2: Advanced Expressive Humanoid Whole-Body Control
- OmniRetarget: Interaction-Preserving Data Generation for Humanoid Whole-Body Loco-Manipulation and Scene Interaction
- HumanPlus: Humanoid Shadowing and Imitation from Humans
- Learning Human-to-Humanoid Real-Time Whole-Body Teleoperation
- OmniH2O: Universal and Dexterous Human-to-Humanoid Whole-Body Teleoperation and Learning
- ULC: A Unified and Fine-Grained Controller for Humanoid Loco-Manipulation
- TWIST: Teleoperated Whole-Body Imitation System
- HOMIE: Humanoid Loco-Manipulation with Isomorphic Exoskeleton Cockpit
- WoCoCo: Learning Whole-Body Humanoid Control with Sequential Contacts
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving