ExpertGen: Scalable Sim-to-Real Expert Policy Learning from Imperfect Behavior Priors
summary
The gist
EXPERTGEN is a framework designed to automate expert policy learning in simulation to enable scalable sim-to-real transfer by learning generalizable and robust behavior cloning policies from
In short
EXPERTGEN automates learning robust robotic policies for real-world deployment by using imperfect demonstrations as a starting point. It uses diffusion models and reinforcement learning to refine these flawed behaviors into high-success expert policies, enabling sim-to-real transfer without extensive manual reward engineering.
Key concepts
- Learning Imperfect Behavior Priors
- The process starts by creating an initial policy from imperfect demonstrations, such as those from LLMs or humans. These demonstrations are assumed to have errors like poor state coverage or dynamics mismatches, but they provide a structured starting point for learning.
- Steering Prior Model in Massively Parallel Simulation
- Reinforcement learning is used to improve the diffusion model's initial noise, steering it toward successful task completion. This is done efficiently using Diffusion Steering Reinforcement Learning (DSRL) and FastTD3, allowing the system to learn without needing complex reward engineering.
- Visuomotor Policy Distillation via DAgger
- The refined expert policies are then used as teachers in a DAgger distillation stage. This involves simulating trajectories, collecting data under the expert's guidance, and training a final policy that mimics the expert's behavior while remaining robust to visual uncertainty.
Terminology used across episodes
This episode discusses
- ExpertGen: Scalable Sim-to-Real Expert Policy Learning from Imperfect Behavior Priors · Paper Radio
- OpenMathInstruct-2: Accelerating AI for Math with Massive Open-Source Instruction Data
- GPT-4 Technical Report
- DeepSeek-V3 Technical Report
- LLaMA: Open and Efficient Foundation Language Models
- Qwen Technical Report
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- Theia: Distilling Diverse Vision Foundation Models for Robot Learning
- SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features
- pi 0: A Vision-Language-Action Flow Model for General Robot Control
- pi 0.5: a Vision-Language-Action Model with Open-World Generalization
- OpenVLA: An Open-Source Vision-Language-Action Model
- MolmoAct: Action Reasoning Models that can Reason in Space
- GR00T N1: An Open Foundation Model for Generalist Humanoid Robots
- AnyTask: an Automated Task and Data Generation Framework for Advancing Sim-to-Real Policy Learning
- Self-Improving Vision-Language-Action Models with Data Generation via Residual RL
- VIRAL: Visual Sim-to-Real at Scale for Humanoid Loco-Manipulation
- Opening the Sim-to-Real Door for Humanoid Pixel-to-Action Policy Transfer
- Isaac Lab: A GPU-Accelerated Simulation Framework for Multi-Modal Robot Learning
- Steering Your Diffusion Policy with Latent Space Reinforcement Learning
- FastTD3: Simple, Fast, and Capable Reinforcement Learning for Humanoid Control
The paper
ExpertGen: Scalable Sim-to-Real Expert Policy Learning from Imperfect Behavior Priors · Read on arXiv
Zifan Xu, Ran Gong, Maria Vittoria Minniti, Kausik Sivakumar, Ahmet Salih Gundogdu, Eric Rosen, Riedana Yan, Tushar Kusnur
Robotics and AI Institute, University of Texas at Austin, Sony AI
Learning generalizable and robust behavior cloning policies requires large volumes of high-quality robotics data. While human demonstrations (e.g., through teleoperation) serve as the standard source for expert behaviors, acquiring such data at scale in the real world is prohibitively expensive. This paper introduces ExpertGen, a framework that automates expert policy learning in simulation to enable scalable sim-to-real transfer. ExpertGen first initializes a behavior prior using a diffusion policy trained on imperfect demonstrations, which may be synthesized by large language models or provided by humans. Reinforcement learning is then used to steer this prior toward high task success by optimizing the diffusion model's initial noise while keep original policy frozen. By keeping the pretrained diffusion policy frozen, ExpertGen regularizes exploration to remain within safe, human-like behavior manifolds, while also enabling effective learning with only sparse rewards. Empirical evaluations on challenging manipulation benchmarks demonstrate that ExpertGen reliably produces high-quality expert policies with no reward engineering. On industrial assembly tasks, ExpertGen achieves a 90.5% overall success rate, while on long-horizon manipulation tasks it attains 85% overall success, outperforming all baseline methods. The resulting policies exhibit dexterous control and remain robust across diverse initial configurations and failure states. To validate sim-to-real transfer, the learned state-based expert policies are further distilled into visuomotor policies via DAgger and successfully deployed on real robotic hardware.
Transcript
Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.
Rosa: Today's paper: "ExpertGen: Scalable Sim-to-Real Expert Policy Learning from Imperfect Behavior Priors".
Dev: EXPERTGEN is a framework designed to automate expert policy learning in simulation to enable scalable sim-to-real transfer by learning generalizable and robust behavior cloning policies from imperfect demonstrations.
Rosa: First, who's behind it and why it matters.
Title and authors: Rosa: So, we're looking at ExpertGen today, which tackles that big problem of needing massive amounts of real-world data to train good policies for robots. The authors introduce this framework specifically to bridge that sim-to-real gap by learning generalizable and robust behavior cloning policies from imperfect demonstrations.
Dev: That sounds like it hits on a huge pain point in robotics, Rosa, because getting those high-quality real-world datasets is just too costly for most researchers to manage effectively.
Taro: It seems like the core idea here is using diffusion models to capture the distribution of what's possible from imperfect data, which is a clever way to get a starting point before trying to learn anything else.
Rosa: Exactly, and they focus on learning policies that can generalize well and recover gracefully even when things go wrong in deployment.
Dev: But I always wonder how they handle the practical realities of running this stuff; if we're aiming for real-world deployment, we need to worry about the loop rate and any potential latency issues introduced by these complex models.
Taro: The paper suggests that one of the major hurdles with existing methods is that scripted policies from LLMs or human teleoperation often lack diversity and are brittle when things deviate from what they saw during training.
Rosa: That's a key point they address with ExpertGen, because their initial approach involves training a state-space diffusion policy using those imperfect demonstrations, whether they come from LLM plans or human teleoperation.
Dev: And then what happens next? Since we can't just rely on the initial prior, I'm curious about the second stage where they refine that prior model to achieve actual task success.
Taro: The framework moves into a steering mechanism using Diffusion Steering Reinforcement Learning, or DSRL, which is designed to optimize the initial noise of that diffusion model while keeping the original policy frozen.
Rosa: It’s interesting how they use DSRL in massively parallel simulation to steer this prior toward high task success by optimizing only the initial noise without needing any reward engineering.
Dev: That addresses one of my main concerns about manual reward design, which is a huge win for scalability across different tasks and environments.
Taro: And the paper claims this steering process preserves the natural motion data manifold specified by the data, which keeps the resulting motions within a realistic range compared to unconstrained reinforcement learning methods.
Rosa: That preservation of the motion manifold is important because it suggests we won't get completely bizarre or unrealistic movements when deploying these policies in physical hardware.
Title and authors: Dev: But if we're talking about scaling this up, how do they manage the computational load of running this diffusion steering process across a large number of simulation instances?
Taro: They use FastTD3 for this steering process specifically because it’s efficient for massively parallel simulation training and helps amortize the cost associated with long-horizon action chunk rollouts.
Rosa: It sounds like they've thought about the computational efficiency needed to make this framework practical for real-world deployment scenarios, which is something I always look at when considering field testing.
Dev: And then we get to the final stage where they distill these state-based expert policies into something actually deployable on hardware, and that leads us to DAgger distillation.
Taro: The DAgger stage involves rolling out the expert policy in simulation, collecting trajectories under the learner's induced state distribution, and then iteratively refining the policy with corrective actions at each step.
Rosa: That iterative refinement process sounds like a solid way to make sure the final policy doesn't just mimic the expert but actually handles visual uncertainty when it hits reality.
Dev: So, by combining these diffusion priors, RL steering in parallel sims, and DAgger distillation, they're building a very layered approach to get that sim-to-real transfer working.
Taro: The main result they highlight is that this entire EXPERTGEN pipeline transforms a small number of imperfect demonstrations into sim-to-real ready expert policies with zero reward engineering required.
Rosa: That’s quite an accomplishment, Taro; moving from just having a few examples to something that works without manual reward tuning seems like a significant step for the field.
Dev: I'm still thinking about the practical implications for latency; if this entire pipeline is running in simulation to generate data, we need to ensure that when we actually run it on hardware, the resulting control loop doesn't suffer from excessive lag.
Taro: The paper does mention that they've shown robust zero-shot sim-to-real transfer of visuomotor policies through this large-scale DAgger distillation process.
Rosa: And that robustness is what really excites me; if a policy can handle visual noise and uncertainty during deployment, that opens up so many possibilities for deploying complex behaviors outside of a clean lab setting.
Dev: I agree about the robustness, but I want to know what happens when the world misbehaves in a way we didn't anticipate—for example, if an external force pushes the robot unexpectedly.
Taro: The results show that these policies exhibit strong failure recovery capabilities under targeted perturbations, with performance drops of only zero point five percent and twenty-eight point six percent compared to the evaluation without any perturbation.
Title and authors: Rosa: That level of recovery suggests they’ve captured some of the underlying physical principles rather than just memorizing trajectories, which is a big deal for field robotics applications.
Dev: It's promising, but I still want more details on how long these policies might remain reliable in a continuously operating system where drift or unexpected changes are constant factors.
Taro: The authors also noted that using human motions as behavior priors actually improves performance substantially, showing that those more diverse and adaptable patterns help both the diffusion policy and the EXPERTGEN trained with SkillMimicGen data outperform their scripted counterparts.
Rosa: So, it’s not just about scaling up to use more data; incorporating richer behavioral diversity from human demonstrations seems to give the system a better foundation to build upon.
Dev: That makes sense; if the initial prior is already more adaptable, the subsequent refinement steps should be much smoother and less prone to catastrophic failures in simulation.
Taro: The overall implication I see is that we can move away from painstakingly engineering rewards for every single task, which opens up a lot of time for researchers to focus on designing the underlying learning architecture itself.
Rosa: I think that’s a huge shift; it lets us focus on the general method rather than getting bogged down in task-specific reward hacking.
Dev: It does sound like this paper, ExpertGen: Scalable Sim-to-Real Expert Policy Learning from Imperfect Behavior Priors, offers a concrete path toward deploying complex visuomotor skills in physical systems without needing endless hours of hand-tuned reward functions.
Taro: Indeed, and the fact that it shows strong performance on industrial assembly tasks achieving a ninety point five percent overall success rate is compelling evidence for its capability when applied to real-world robotics challenges.
Rosa: Well, it seems like this paper gives us a powerful tool to start building more capable robots using less expensive real-world data collection, which is what we're all hoping for in the long run.
Dev: It’s definitely something worth keeping an eye on as we try to integrate these types of learned policies into actual robotic hardware that needs to operate reliably over long periods.
Taro: I think the way they combine diffusion steering with DAgger distillation is a really sophisticated way to ensure that the final output policy is both generalizable and physically sound for deployment.
Rosa: Well, that’s what we’ve got today on ExpertGen; it seems like a really solid piece of work pushing the boundaries of how we learn skills for robots.
The paper's summary: Rosa: So, ExpertGen is essentially proposing a way to take just a few imperfect examples of how a robot should move and use them to build something that can actually work in the real world without needing millions of hours of expensive real-world data.
Dev: That sounds like it aims to drastically cut down the time and cost associated with getting expert data, which is exactly what we need for scalable deployment.
Taro: The core mechanism involves using a diffusion policy trained on these imperfect demonstrations as a starting point, and then using reinforcement learning in massive parallel simulation to steer that prior toward actually succeeding at the task without needing any reward engineering.
Rosa: That steering process is clever because it keeps the natural motion patterns of the diffusion model intact while boosting success under sparse rewards, which is a big deal for making sure the robot doesn't learn weird, unrealistic movements.
Dev: I’m still focused on those real-world deployment questions; if this system is running in simulation to generate that data and steering signals, how reliable are we talking about its loop rate when it finally hits physical hardware?
Taro: The paper suggests they use FastTD3 for the steering phase because it handles massively parallel simulation efficiently, which helps manage those long-horizon action rollouts without bogging down the training process.
Rosa: And then after that, they distill these state-based policies into something deployable using a DAgger technique, which means they're refining the policy by having it correct itself based on data collected under its own simulated distribution.
Dev: That iterative refinement sounds like a solid safety mechanism; minimizing imitation loss over an aggregated dataset helps ensure that the final policy is robust to visual uncertainty when it leaves the sim environment.
Taro: What’s really compelling about this framework is its ability to achieve this zero-shot transfer of visuomotor policies by using these techniques, which means we don't need task-specific reward engineering to get a functional skill working from imperfect initial data.
Rosa: That capability to bypass manual reward engineering for complex manipulation tasks is what makes me really optimistic about the potential impact on how quickly we can deploy advanced robotic skills.
Dev: It’s certainly a strong technical achievement, but I wonder if these policies maintain that high level of performance when faced with unexpected physical disturbances or changes in the environment that aren't covered in their training set.
Taro: The authors did address failure recovery, showing that the resulting policies exhibit quite good behavior under targeted perturbations, meaning they don't just fail catastrophically when things go wrong.
Rosa: That kind of inherent robustness is exactly what we need for field applications; if a robot can handle a little unexpected push or visual noise without completely breaking down, that’s huge.
Dev: So, to wrap up this summary, ExpertGen provides an end-to-end pipeline that uses diffusion priors and RL steering to generate sim-to-real policies ready for deployment via DAgger distillation without requiring manual reward tuning.
Taro: Exactly; it shows how you can leverage a small amount of imperfect input data to create a high-quality, generalizable policy that performs well across different scenarios.
Rosa: This research really opens up avenues for deploying complex behaviors in physical systems much faster than before, and I’m eager to see how this translates into actual hardware deployment scenarios soon.
The paper's improvements: Taro: So, we've seen how ExpertGen uses diffusion steering and DAgger distillation to create these policies, and now I want to talk about what they suggest as improvements for this system.
Rosa: What are the key enhancements they propose? Are they focusing on making it work better in the real world or just improving its performance in simulation?
Dev: I’m interested if these improvements address the loop rate concerns we had earlier, like how fast it can actually react when a control signal comes through.
Taro: They focus on scaling this diffusion steering to massively parallel robotics simulation, which demonstrates that it keeps the natural motion manifold of diffusion models while significantly boosting success rates even with sparse rewards.
Rosa: That’s important because it means we can get better results without needing to manually design reward functions for every single task we want the robot to do.
Dev: But what about robustness? I need to know if these suggested improvements help the resulting policies handle unexpected physical failures or sudden changes in the environment when they are deployed in a real setting.
Taro: They also introduce alternatives like Residual RL, which adds a learnable residual policy on top of the fixed prior to predict corrective actions, and Score-Matching Motion Priors that use distillation sampling to provide a reusable behavior regularizer.
Rosa: It sounds like they are pushing for more adaptive behaviors by incorporating richer motion diversity from human demonstrations into the initial priors, showing that those more varied patterns help both the diffusion policy and the EXPERTGEN trained with SkillMimicGen data perform better.
Dev: Incorporating human motion priors is interesting; it suggests that starting with a prior that already has diverse behavioral patterns makes the subsequent reinforcement learning refinement stage much smoother than starting from a very rigid, scripted plan.
Taro: The framework also shows strong failure recovery capabilities under targeted perturbations, meaning these improvements aim to make the policies more resilient when things deviate from the expected path in deployment.
Rosa: So, what’s the big picture implication of all this refinement? Does it mean we can deploy complex skills much sooner than we thought possible?
Dev: If these improvements allow for a more reliable and faster learning process, it could mean we can get robots into complex manipulation tasks in real environments much quicker than relying on massive amounts of meticulously collected data.
Taro: The implication is that the architecture itself, rather than just collecting more data or tweaking rewards, is capable of generating high-quality behaviors robust enough for real-world use.
Rosa: I’m really excited about this potential because it could fundamentally change how we approach sim-to-real transfer in robotics by making it more data efficient and less reliant on perfect initial conditions.
Conclusion: Rosa: So, we’ve covered how ExpertGen uses diffusion steering and DAgger distillation to create sim-to-real policies from imperfect data, and now we’re wrapping up with some final thoughts on its impact.
Dev: It really seems like this framework offers a much more practical way to get robust skills into the hands of physical robots without needing mountains of real-world testing data.
Taro: I think the biggest implication is that we can move away from painstakingly engineering rewards for every single task and instead focus on designing a learning architecture that naturally handles generalization.
Rosa: That’s huge; it shifts the focus from task-specific optimization to building a foundational capability, which opens up so much room for new types of robotic skills.
Dev: I'm still thinking about the practical side; if this works in simulation and then transfers well, how long do you think these policies will reliably hold up when they’re running continuously in a real industrial setting?
Taro: The robustness against perturbations mentioned suggests they’ve built something that handles unexpected physical issues better than standard imitation learning methods, which is a major step toward deployability.
Rosa: It does sound like ExpertGen provides a solid path forward for deploying complex visuomotor skills in physical systems much faster than before, and I’m eager to see how this translates into actual hardware deployment scenarios soon.
Dev: I agree; the efficiency gains from using methods like FastTD3 for steering and the structured approach to distillation make it computationally viable for large-scale robotics applications.
Taro: To finish up, ExpertGen is a very strong piece of work because it shows that combining diffusion priors with RL steering can produce high-quality policies with minimal manual reward tuning.
Rosa: Indeed, this paper on ExpertGen offers a powerful tool to start building more capable robots using less expensive real-world data collection, which is what we're all hoping for in the long run.
Dev: It’s definitely something worth keeping an eye on as we try to integrate these types of learned policies into actual robotic hardware that needs to operate reliably over long periods.
Taro: I think the way they combine diffusion steering with DAgger distillation is a really sophisticated way to ensure that the final output policy is both generalizable and physically sound for deployment.
More episodes
- 2610.11768-Narrow and Deep: An Ontology Tower as the Knowledge of an LLM Agent for an Industrial Equipment System
- 2610.11904-Large-Scale Partition-Based RIS Beamforming For Uplink RIS-Equipped Multi-User Systems: Asymptotic Analysis
- 2610.11885-Redefining fuel poverty: Introducing the temporal equity framework (TEF)
- 2610.11900-Reach-Stabilize Control of Control-Affine Systems with Unknown Affine Parameters
- 2610.11964-From Asymptotic to Designer-Assigned-Time Control: A Review of Stability Notions, Design Mechanisms, and Controller Architectures
- 2610.12226-Stabilization of Unidirectional First-Order PDE-ODE Coupled Systems with Boundary and Distributed Input Delays
- 2610.12028-Policy Synthesis for Finite Populations of MDP Agents under Aggregate Reach-Avoid Chance Constraints
- 2610.12103-Predefined-Time Integral Reinforcement Learning for Unknown Nonlinear Systems via Inverse-Optimal Design
- 2610.12110-Adaptive dynamic programming using Lyapunov function constraints
- 2610.12324-Convex Safety Filtering via Spectral Selection for Nonconvex Safe Sets