Reconstruct, Practice, Go Real: Guided Self-Improvement for Embodied Agents
summary
The gist
Building reliable robot capabilities across diverse tasks requires substantial human effort to develop and maintain skills, design rewards, and integrate perception with control.
In short
The RPG framework develops reusable robot skills through autonomous practice without retraining model weights. It uses specialized agents to diagnose failures during simulation, leading to iterative refinement of symbolic skills and system prompts. This process transforms implicit assumptions into explicit procedural checks, enabling robots to improve performance across diverse tasks for physical deployment.
Key concepts
- Agent-as-Policy
- This approach treats the multimodal Runtime Agent as a policy that follows a system prompt. It allows the agent to dynamically select and combine reusable skills from a library during task execution. This structure enables flexible behavior composition based on learned skills.
- Constructor
- The Constructor agent is responsible for selecting initial source tasks from an offline dataset. It then builds related practice tasks in simulation while keeping the task definitions and evaluators fixed. This ensures that the self-improvement process focuses on skill development rather than changing the core problem definition.
- Self-Improvement via Diagnostics
- The system uses a feedback loop where agents diagnose failures during practice. The Video Analyzer and Privileged Agent provide diagnostic inputs, which inform the Implementor to develop or revise skills and prompts. This iterative refinement process systematically improves robot capabilities based on observed execution errors.
- Skill Transfer
- Learned corrections are not limited to specific tasks that generated feedback. The framework shows that procedural changes made during practice can be reused in new, unrelated tasks. This means the robot learns general, reusable manipulation patterns rather than just task-specific solutions.
Terminology used across episodes
This episode discusses
- Reconstruct, Practice, Go Real: Guided Self-Improvement for Embodied Agents · Paper Radio
- ASPIRE: Agentic /Skills Discovery for Robotics
- RHO: Your Coding Agent is Secretly a Roboticist
- Self-Evolving Embodied Agents via Skill-Harness Evolution · Paper Radio
- Practice Makes Policies: Bootstrapping and Consolidating Robotic Capabilities from Zero Human Demonstrations
- GRASP: Gated Regression-Aware Skill Proposer for Self-Improving LLM Agents · Paper Radio
The paper
Reconstruct, Practice, Go Real: Guided Self-Improvement for Embodied Agents · Read on arXiv
Yen-Jen Wang, Haozhe Jiang, Shuying Deng, Haoru Xue, Weirui Ye, Rocky Duan, Nika Haghtalab, S. Shankar Sastry
UC Berkeley Department of Computer Science and Engineering (implied by 1UC Berkeley) · Amazon Federal Research (FAR) · MIT
Building reliable robot capabilities across diverse tasks requires substantial human effort to develop and maintain skills, design rewards, and integrate perception with control. We present Reconstruct, Practice, Go Real (RPG), a framework for autonomous improvement of robot execution systems without updating model weights. RPG identifies manipulation capabilities in an offline dataset and constructs related practice tasks in simulation. During practice, RPG uses execution feedback, privileged simulator state, and available dataset videos to diagnose failures. It develops new reusable symbolic skills, refines existing skills, and revises the system prompt based on these diagnoses. Cross-task evaluation tests individual candidate changes and merged revisions before they are retained for reuse. At test time, a multimodal LLM uses the resulting system prompt and skill library to coordinate perception and robot control. On held-out initializations of 22 manipulation tasks, RPG improves task success from 28.6% after the first practice round to 95.0% after 15 rounds, outperforming all evaluated baselines, including ASPIRE (75.5%) and CaP-Agent0 powered by GPT-6 Astra Pro (60.0%). After a common calibration and hardware-adaptation procedure, the frozen system succeeds in all 30 physical trials, with ten trials on each of three tasks. Project Website: https://rpg-robot.github.io/
Transcript
Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.
Rosa: Today's paper: "Reconstruct, Practice, Go Real".
Dev: Building reliable robot capabilities across diverse tasks requires substantial human effort to develop and maintain skills, design rewards, and integrate perception with control.
Rosa: First, who's behind it and why it matters.
Paper summary: Rosa: So we’re looking at this paper titled "Reconstruct, Practice, Go Real: Guided Self-Improvement for Embodied Agents," which seems to propose a way for robots to develop skills autonomously through practice without needing constant human reprogramming of the core model. What’s the main idea here?
Dev: Essentially, the thesis is that by using an offline dataset to build simulation tasks and then letting an agent practice those tasks, it can use feedback from execution to diagnose its own failures and create or refine reusable symbolic skills and system prompts. It claims this process allows agents to improve across different tasks by revising shared skills and the overall system prompt before finally deploying the improved version onto a physical robot.
Taro: I’m interested in what this means when things go wrong in real-time, Rosa; does the system have a way to handle unexpected situations outside of those pre-defined practice scenarios?
Rosa: That's exactly where I want to ask, Taro; does it work outside the lab for extended periods, and how long can we expect these improved skills to hold up when the robot encounters something completely novel in a real environment?
Dev: From an engineering standpoint, the paper outlines a framework with six specialized agents—the Constructor, Runtime Agent, Privileged Agent, Video Analyzer, Implementor, and Merger—all coordinating through an agent-as-policy execution system where the Runtime Agent selects skills from a library guided by a system prompt. This whole loop is designed to be iterative for skill refinement.
Taro: The paper mentions that this self-improvement process converts implicit assumptions into explicit procedural checks, like verifying preconditions before acting; how robust is this when the environment misbehaves in unpredictable ways?
Rosa: That’s a big point, Taro; the paper suggests that these learned corrections can transfer to tasks without direct task-specific improvement feedback, which means those learned procedural patterns might be useful even when the specific object changes. However, it does flag that deformable-object manipulation, like folding a towel, still presents a bottleneck because current skill compositions aren't fully capturing those complex states.
Dev: The system architecture relies on the Video Analyzer comparing executions from both agents with available dataset videos to diagnose failures and recommend concrete changes to the Implementor; I'm curious about the latency of that diagnostic feedback loop when we’re running these practices in simulation.
Paper summary: Taro: If we look at what this RPG system does, it identifies manipulation capabilities from an offline dataset and builds practice tasks in simulation while keeping task definitions and evaluators fixed during self-improvement; how does that constrain the agent's ability to truly generalize?
Rosa: The paper claims that after fifteen rounds of practice, task success goes from twenty-eight point six percent up to ninety-five point zero percent, significantly outperforming baselines like ASPIRE at seventy-five point five percent; it shows a substantial gain in performance over time within the simulation environment.
Dev: The results on the quantitative evaluation focus on task success across held-out initializations of twenty-two manipulation tasks, showing that this method can reach very high success rates when given enough practice rounds; we also see it compared against CaP-Agent0 powered by GPT-six Astra Pro, which scored sixty point zero percent.
Taro: So, if we consider the overall trajectory described in "Reconstruct, Practice, Go Real: Guided Self-Improvement for Embodied Agents," the real impact seems to be in how it handles the iterative refinement of skills and system prompts based on execution feedback. What are your thoughts on what this means for long-term autonomous skill acquisition?
Rosa: I think the core implication is that we might move away from needing massive amounts of task-specific human engineering, allowing agents to build a baseline of reusable skills through practice and self-correction, which could speed up deployment considerably.
Dev: From my side, the architecture shows that the combination of the Privileged Agent providing simulator state and the Video Analyzer using that with traces and videos provides complementary information for failure diagnosis; it’s a layered approach to debugging.
Taro: And looking at how RPG handles failures, it systematically converts implicit assumptions into explicit procedural checks, which is a very practical way for an agent to learn robust behavior in complex physical interactions.
Rosa: So, we have this self-improving system that gets better through practice guided by feedback from simulation and video analysis; the paper's authors are really showing how these iterative loops can lead to high success rates on held-out seeds during the final deployment phase.
Dev: I’m still thinking about the practical deployment aspect, Rosa; it states that after a common calibration and hardware adaptation procedure, the frozen system succeeds in all thirty physical trials across three evaluated tasks. That suggests a good level of generalization once it hits the real world.
Paper summary: Taro: If we take this paper's findings seriously, what do you see as the biggest potential impact this has on how we approach creating embodied agents that can operate reliably in messy, unpredictable environments?
Rosa: I think it points toward a future where agents don't just follow a fixed sequence of commands but actively learn the necessary procedural checks to maintain success across varied tasks, which is what this paper demonstrates through its practice rounds.
Dev: The system evolution shows that over fifteen rounds, the skill library grew from fifteen to thirty-eight entries, with "twenty-three new skills and sixty-six modifications to existing skills," indicating a lot of actual learning happening in the system's knowledge base.
Taro: That expansion of the skill library is significant; it shows that the agent isn't just patching one thing; it’s building a richer repertoire of actions based on what it learned during its practice phase.
Rosa: It really does show how this iterative self-improvement process can lead to substantial gains, like that jump from forty-three point two percent to seventy-three point six percent success between rounds three and four, driven by a system prompt revision that introduced a bounded perception–action loop.
Dev: That specific revision is interesting; it suggests that the quality of the high-level instruction given to the agent is just as important as the low-level skill composition itself when boosting performance.
Taro: And when we look at how this self-improvement works, it seems to be a very structured way for an agent to evolve its behavior, moving from guessing what to do implicitly to having explicit rules it checks before acting.
Rosa: That transition from implicit assumptions into explicit procedural checks is a key finding because it suggests that learned behaviors can transfer between tasks without needing specific retraining for every single one.
Dev: But we also have to consider the limitations mentioned; the paper points out that deformable-object manipulation, specifically folding a towel, remains a bottleneck, suggesting that richer representations of deformable state are needed beyond what’s currently in place.
Taro: So while it shows massive improvement in structured tasks, it also highlights where current representation methods still fall short when dealing with highly complex physical states.
Rosa: Exactly; the work on "Reconstruct, Practice, Go Real: Guided Self-Improvement for Embodied Agents" gives us a very concrete framework for how embodied agents can improve their manipulation capabilities through guided self-improvement in simulation before real world deployment.
Conclusion: Rosa: So, this paper is titled "Reconstruct, Practice, Go Real: Guided Self-Improvement for Embodied Agents," and it’s authored by a team that has been pushing hard in the robotics space lately. What are the real-world implications of this approach for how we build these physical machines?
Dev: I think the authors are showing us a system that doesn't just rely on pre-programmed code; they've built an iterative loop where agents improve their own skills through practice and feedback, which is something we’ve always wanted to see implemented more robustly.
Taro: From an autonomy standpoint, the paper suggests that this method allows agents to convert those implicit assumptions into explicit procedural checks, which means they can handle uncertainty in the world better than just following a fixed script.
Rosa: Exactly, Taro; it's about giving the agent a way to learn how to navigate messy situations on its own by constantly checking its steps against what actually works.
Dev: And from an engineering side, I'm really interested in how this self-improvement loop is structured; they’ve got this whole coordination of agents—Constructor, Runtime Agent, and others—which suggests a lot of careful design around latency and execution flow.
Taro: That iterative nature is key because it means the system evolves its skill library over time, which implies that the agent isn't stuck with one set of behaviors but can adapt as it gains experience.
Rosa: It really shows a path toward creating agents that are more flexible and less brittle when things go off-script in physical deployment.
Dev: But my main concern remains, Rosa; I need to know how long this self-improvement process can sustain itself outside of the controlled simulation environment before we deploy it onto a real robot.
Taro: That’s a fair question, Dev; if the learning happens in simulation, how do we ensure those learned procedural checks actually transfer reliably when the physical dynamics or environment are slightly different?
Rosa: That leads us right into the next big discussion: whether these refined skills are truly portable across different real-world scenarios, and how long we can trust this self-tuning process to keep improving without constant human intervention.
More episodes
- 2610.10846-Cross-Embodiment Robot Foundation World Models with Latent Actions
- 2610.10601-Teaching a Robot Dog New Tricks: Diverse Quadruped Skills via Combined Reinforcement and Imitation Learning with Adversarial Task Selection
- 2610.10637-TacHair: Tactile Contact-Distribution Guided Online Correction for Robotic Hair Stroking and Perception
- 2610.10646-Masked Generative Motion Planning with Geometry-Guided Token Search
- 2610.10812-Skill-SLM: Agent Skill-driven Small Language Models for Reliable Robot Operation
- 2610.10801-Same Action, Different Outcome: Variability in Dynamic Cloth Manipulation
- 2610.10810-Diagnosing and Recovering from Observation-Space Shift at Long-Horizon Skill Seams
- 2610.10748-TAPNAV: Humanoid Navigation through Tactile Active Perception
- 2610.10855-OmniHOI: Dexterous Hand-Object Interaction from Monocular Human Video
- 2610.11003-ActiveReg: Information-Driven Active Regional Probing for Partial-to-Full Bone Registration