Reconstruct, Practice, Go Real: Guided Self-Improvement for Embodied Agents

arXiv:2610.02204 · cs.RO, cs.AI, cs.SY, eess.SY · Submitted 2026-10-01 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: Today's paper: "Reconstruct, Practice, Go Real".

Dev: Building reliable robot capabilities across diverse tasks requires substantial human effort to develop and maintain skills, design rewards, and integrate perception with control.

Rosa: First, who's behind it and why it matters.

Paper summary: Rosa: So we’re looking at this paper titled "Reconstruct, Practice, Go Real: Guided Self-Improvement for Embodied Agents," which seems to propose a way for robots to develop skills autonomously through practice without needing constant human reprogramming of the core model. What’s the main idea here?

Dev: Essentially, the thesis is that by using an offline dataset to build simulation tasks and then letting an agent practice those tasks, it can use feedback from execution to diagnose its own failures and create or refine reusable symbolic skills and system prompts. It claims this process allows agents to improve across different tasks by revising shared skills and the overall system prompt before finally deploying the improved version onto a physical robot.

Taro: I’m interested in what this means when things go wrong in real-time, Rosa; does the system have a way to handle unexpected situations outside of those pre-defined practice scenarios?

Rosa: That's exactly where I want to ask, Taro; does it work outside the lab for extended periods, and how long can we expect these improved skills to hold up when the robot encounters something completely novel in a real environment?

Dev: From an engineering standpoint, the paper outlines a framework with six specialized agents—the Constructor, Runtime Agent, Privileged Agent, Video Analyzer, Implementor, and Merger—all coordinating through an agent-as-policy execution system where the Runtime Agent selects skills from a library guided by a system prompt. This whole loop is designed to be iterative for skill refinement.

Taro: The paper mentions that this self-improvement process converts implicit assumptions into explicit procedural checks, like verifying preconditions before acting; how robust is this when the environment misbehaves in unpredictable ways?

Rosa: That’s a big point, Taro; the paper suggests that these learned corrections can transfer to tasks without direct task-specific improvement feedback, which means those learned procedural patterns might be useful even when the specific object changes. However, it does flag that deformable-object manipulation, like folding a towel, still presents a bottleneck because current skill compositions aren't fully capturing those complex states.

Dev: The system architecture relies on the Video Analyzer comparing executions from both agents with available dataset videos to diagnose failures and recommend concrete changes to the Implementor; I'm curious about the latency of that diagnostic feedback loop when we’re running these practices in simulation.

Paper summary: Taro: If we look at what this RPG system does, it identifies manipulation capabilities from an offline dataset and builds practice tasks in simulation while keeping task definitions and evaluators fixed during self-improvement; how does that constrain the agent's ability to truly generalize?

Rosa: The paper claims that after fifteen rounds of practice, task success goes from twenty-eight point six percent up to ninety-five point zero percent, significantly outperforming baselines like ASPIRE at seventy-five point five percent; it shows a substantial gain in performance over time within the simulation environment.

Dev: The results on the quantitative evaluation focus on task success across held-out initializations of twenty-two manipulation tasks, showing that this method can reach very high success rates when given enough practice rounds; we also see it compared against CaP-Agent0 powered by GPT-six Astra Pro, which scored sixty point zero percent.

Taro: So, if we consider the overall trajectory described in "Reconstruct, Practice, Go Real: Guided Self-Improvement for Embodied Agents," the real impact seems to be in how it handles the iterative refinement of skills and system prompts based on execution feedback. What are your thoughts on what this means for long-term autonomous skill acquisition?

Rosa: I think the core implication is that we might move away from needing massive amounts of task-specific human engineering, allowing agents to build a baseline of reusable skills through practice and self-correction, which could speed up deployment considerably.

Dev: From my side, the architecture shows that the combination of the Privileged Agent providing simulator state and the Video Analyzer using that with traces and videos provides complementary information for failure diagnosis; it’s a layered approach to debugging.

Taro: And looking at how RPG handles failures, it systematically converts implicit assumptions into explicit procedural checks, which is a very practical way for an agent to learn robust behavior in complex physical interactions.

Rosa: So, we have this self-improving system that gets better through practice guided by feedback from simulation and video analysis; the paper's authors are really showing how these iterative loops can lead to high success rates on held-out seeds during the final deployment phase.

Dev: I’m still thinking about the practical deployment aspect, Rosa; it states that after a common calibration and hardware adaptation procedure, the frozen system succeeds in all thirty physical trials across three evaluated tasks. That suggests a good level of generalization once it hits the real world.

Paper summary: Taro: If we take this paper's findings seriously, what do you see as the biggest potential impact this has on how we approach creating embodied agents that can operate reliably in messy, unpredictable environments?

Rosa: I think it points toward a future where agents don't just follow a fixed sequence of commands but actively learn the necessary procedural checks to maintain success across varied tasks, which is what this paper demonstrates through its practice rounds.

Dev: The system evolution shows that over fifteen rounds, the skill library grew from fifteen to thirty-eight entries, with "twenty-three new skills and sixty-six modifications to existing skills," indicating a lot of actual learning happening in the system's knowledge base.

Taro: That expansion of the skill library is significant; it shows that the agent isn't just patching one thing; it’s building a richer repertoire of actions based on what it learned during its practice phase.

Rosa: It really does show how this iterative self-improvement process can lead to substantial gains, like that jump from forty-three point two percent to seventy-three point six percent success between rounds three and four, driven by a system prompt revision that introduced a bounded perception–action loop.

Dev: That specific revision is interesting; it suggests that the quality of the high-level instruction given to the agent is just as important as the low-level skill composition itself when boosting performance.

Taro: And when we look at how this self-improvement works, it seems to be a very structured way for an agent to evolve its behavior, moving from guessing what to do implicitly to having explicit rules it checks before acting.

Rosa: That transition from implicit assumptions into explicit procedural checks is a key finding because it suggests that learned behaviors can transfer between tasks without needing specific retraining for every single one.

Dev: But we also have to consider the limitations mentioned; the paper points out that deformable-object manipulation, specifically folding a towel, remains a bottleneck, suggesting that richer representations of deformable state are needed beyond what’s currently in place.

Taro: So while it shows massive improvement in structured tasks, it also highlights where current representation methods still fall short when dealing with highly complex physical states.

Rosa: Exactly; the work on "Reconstruct, Practice, Go Real: Guided Self-Improvement for Embodied Agents" gives us a very concrete framework for how embodied agents can improve their manipulation capabilities through guided self-improvement in simulation before real world deployment.

Conclusion: Rosa: So, this paper is titled "Reconstruct, Practice, Go Real: Guided Self-Improvement for Embodied Agents," and it’s authored by a team that has been pushing hard in the robotics space lately. What are the real-world implications of this approach for how we build these physical machines?

Dev: I think the authors are showing us a system that doesn't just rely on pre-programmed code; they've built an iterative loop where agents improve their own skills through practice and feedback, which is something we’ve always wanted to see implemented more robustly.

Taro: From an autonomy standpoint, the paper suggests that this method allows agents to convert those implicit assumptions into explicit procedural checks, which means they can handle uncertainty in the world better than just following a fixed script.

Rosa: Exactly, Taro; it's about giving the agent a way to learn how to navigate messy situations on its own by constantly checking its steps against what actually works.

Dev: And from an engineering side, I'm really interested in how this self-improvement loop is structured; they’ve got this whole coordination of agents—Constructor, Runtime Agent, and others—which suggests a lot of careful design around latency and execution flow.

Taro: That iterative nature is key because it means the system evolves its skill library over time, which implies that the agent isn't stuck with one set of behaviors but can adapt as it gains experience.

Rosa: It really shows a path toward creating agents that are more flexible and less brittle when things go off-script in physical deployment.

Dev: But my main concern remains, Rosa; I need to know how long this self-improvement process can sustain itself outside of the controlled simulation environment before we deploy it onto a real robot.

Taro: That’s a fair question, Dev; if the learning happens in simulation, how do we ensure those learned procedural checks actually transfer reliably when the physical dynamics or environment are slightly different?

Rosa: That leads us right into the next big discussion: whether these refined skills are truly portable across different real-world scenarios, and how long we can trust this self-tuning process to keep improving without constant human intervention.

Yen-Jen Wang, Haozhe Jiang, Shuying Deng, Haoru Xue, Weirui Ye, Rocky Duan, Nika Haghtalab, S. Shankar Sastry

UC Berkeley Department of Computer Science and Engineering (implied by 1UC Berkeley) · Amazon Federal Research (FAR) · MIT

cs.RO, cs.AI, cs.SY, eess.SY

Submitted: 2026-10-01

Updated: 2026-10-01

Comments: 17 pages, 6 figures, 10 tables

Project page: https://rpg-robot.github.io

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 89/100

The gist: Building reliable robot capabilities across diverse tasks requires substantial human effort to develop and maintain skills, design rewards, and integrate perception with control.

Key concepts

Agent-as-Policy
This approach treats the multimodal Runtime Agent as a policy that follows a system prompt. It allows the agent to dynamically select and combine reusable skills from a library during task execution. This structure enables flexible behavior composition based on learned skills.
Constructor
The Constructor agent is responsible for selecting initial source tasks from an offline dataset. It then builds related practice tasks in simulation while keeping the task definitions and evaluators fixed. This ensures that the self-improvement process focuses on skill development rather than changing the core problem definition.
Self-Improvement via Diagnostics
The system uses a feedback loop where agents diagnose failures during practice. The Video Analyzer and Privileged Agent provide diagnostic inputs, which inform the Implementor to develop or revise skills and prompts. This iterative refinement process systematically improves robot capabilities based on observed execution errors.
Skill Transfer
Learned corrections are not limited to specific tasks that generated feedback. The framework shows that procedural changes made during practice can be reused in new, unrelated tasks. This means the robot learns general, reusable manipulation patterns rather than just task-specific solutions.

Terminology

Summary

Building reliable robot capabilities across diverse tasks requires substantial human effort to develop and maintain skills, design rewards, and integrate perception with control. The framework presented develops and refines reusable robot skills through autonomous practice without updating model weights.

The gist

RPG identifies manipulation capabilities in an offline dataset and constructs related practice tasks in simulation; during practice, agents use execution feedback to diagnose failures, develop new reusable symbolic skills, refine existing skills, and revise the system prompt based on these diagnoses; finally, the improved system is calibrated and frozen for deployment on the physical robot.

How it works

The framework coordinates six specialized agents across three stages: Constructor, Runtime Agent, Privileged Agent, Video Analyzer, Implementor, and Merger. The execution system is formulated as Agent-as-Policy: a multimodal agent (the Runtime Agent) follows a system prompt to complete tasks by selecting and composing reusable skills from a shared skill library.

  1. The Constructor selects source tasks from an offline dataset (e.g., ABC [10]) to identify manipulation capabilities and constructs related practice tasks in simulation, ensuring task definitions and evaluators remain fixed during self-improvement.

  2. During Practice, the Runtime Agent attempts tasks using a shared skill library, while the Privileged Agent runs separate episodes receiving simulator state for diagnosis. The Video Analyzer compares executions from both agents with available dataset videos to diagnose failures and recommend concrete changes to the Implementor.

  3. The Implementor develops new skills, refines existing ones, and revises the system prompt based on failure diagnoses. Candidate changes are evaluated across practice tasks before being merged; a candidate is eligible for merging only if it meets specific criteria: mean task success increases while no individual task drops by more than 20 percentage points.

  4. The Merger combines selected revisions, and the resulting system is retested before retention. The final RPG system is then calibrated and frozen for physical deployment, where the Runtime Agent coordinates perception and control through the shared skill library.

Key Evaluation Metrics

Quantitative evaluation focuses on task success across held-out initializations of 22 manipulation tasks. RPG improves task success from 28.6% after the first practice round to 95.0% after 15 rounds, outperforming baselines like ASPIRE (75.5%) and CaP-Agent0 powered by GPT-6 Astra Pro (60.0%).

Diagnostic Inputs and Skill Revisions

The framework demonstrates that specific diagnostic inputs drive improvement:

Removing the Video Analyzer reduces success by 26.4 percentage points, while removing the Privileged Agent reduces it by 27.3 points.

The analysis shows that the two provide complementary information, with the Privileged Agent supplying reference executions using privileged state and the Video Analyzer using those executions together with Runtime Agent traces and videos from the offline dataset to diagnose failures.

System Evolution and Limitations

Over 15 practice rounds, the skill library grows from 15 to 38 entries, with 23 new skills and 66 modifications to existing skills. The largest improvement occurred between rounds 3 and 4 when success increased from 43.2% to 73.6%, driven by a system prompt revision that introduced a bounded perception–action loop. A key finding is that self-improvement converts implicit assumptions into explicit procedural checks, such as verifying preconditions before acting, which allows learned behaviors to transfer between tasks without direct task-specific improvement feedback. However, the paper notes that deformable-object manipulation, such as Fold towel (hard), remains a bottleneck, suggesting that richer representations of deformable state are needed beyond existing rigid-object skill compositions.

Physical Deployment Success

The frozen RPG system successfully completes all 30 physical trials across three evaluated tasks after common calibration and hardware adaptation. The results show that shared skill code and the system prompt can serve as practical targets for persistent robot improvement, with the system succeeding in all trials on ten held-out seeds per task. Qualitatively, the frozen execution system is shown to apply to new physical instructions without task-specific real-world tuning, demonstrating its utility for zero-shot deployment.

Transfer of Learned Behaviors

The accumulated changes are not restricted to tasks that generate direct feedback. The results on tasks without direct feedback show evidence that reusable manipulation changes can transfer to tasks without direct task-specific improvement feedback, as seen in examples where Two-arm handover improves from 3/10 in Round 1 to 10/10 in Round 15. This suggests that the learned corrections are reusable procedural patterns rather than being tied to a single object instance.

Failure Taxonomy

The analysis of failures reveals a common trend: self-improvement progressively converts implicit assumptions into explicit procedural checks. Recurring failure modes include:

Localization and geometry uncertainty

Grasp identity and retention

**"

Improvements for AI systems

As a diligent researcher, I have analyzed the RPG (Reconstruct, Practice, Go Real) framework described in this paper. The core innovation is developing and refining reusable robot skills autonomously through simulation practice guided by failure diagnosis from privileged execution data, without updating model weights.

Here are the specific improvements and capabilities that can be derived from implementing the RPG framework:


) 1. Autonomous Skill Refinement via Failure Diagnosis

The system can autonomously identify brittle or implicit assumptions in its existing skill library (e.g., insufficient path validation, end-effector-centric geometry).

[Improvement Detail] The Video Analyzer compares Runtime Agent failures with Privileged Agent executions and offline dataset videos to diagnose the root cause (e.g., failure to clear a rim during bottle placement).

[Resulting Capability] The system develops new, explicit procedural checks (e.g., clear rim before lowering, verify containment after release) that transform brittle open-loop primitives into robust, observation-driven procedures.

) 2. Cross-Task Validation for Robustness

The system ensures that skill revisions do not create negative externalities by testing candidates across the entire suite of practice tasks.

[Improvement Detail] Candidate skill revisions are evaluated on complete Runtime Agent episodes across all 22 tasks using a strict validation gate: mean task success must increase, and no individual task success can drop by more than 20 percentage points.

[Resulting Capability] The system maintains high reliability (e.g., achieving 95% success on held-out initializations) even after significant library changes, guaranteeing that improvements are generalizable across diverse manipulation goals (e.g., a fix for bottle placement helps in other tasks).

) 3. System Prompt Evolution for Behavioral Guidance

The system can autonomously adjust the high-level instructions (system prompt) guiding the multimodal LLM agent to better leverage its skill library.

[Improvement Detail] The Implementor revises the system prompt based on diagnosis, leading to changes like introducing a bounded perception–action loop or tracking remaining interaction budgets.

[Resulting Capability] The agent becomes more efficient by dynamically adjusting its reliance on visual observation vs. internal estimates, preventing slow decision-making loops while ensuring adherence to safety and task constraints.

) 4. Knowledge Transfer Beyond Direct Feedback (Generalization)

The framework demonstrates that learned corrections are not confined to the tasks that generated them, suggesting true procedural knowledge acquisition rather than rote memorization of trajectories.

[Improvement Detail] The system successfully improves performance on no-feedback held-out tasks (e.g., Transfer object, Close sunglasses case) after 15 rounds of practice, indicating that learned verification patterns are transferable.

[Resulting Capability] The robot acquires reusable manipulation patterns—such as test-lift verification or object-centric alignment—that can be composed and applied to entirely new physical tasks it has never seen before (zero-shot deployment).

) 5. Persistent Execution System for Physical Deployment

The final output is a frozen system ready for real-world use, requiring only calibration and hardware adaptation, not continuous online learning.

[Improvement Detail] The RPG process culminates in Go Real, where the refined shared skill library and system prompt are calibrated once and frozen. The resulting system succeeds on all physical trials without further model weight updates.

[Resulting Capability] Deployment of highly reliable robot systems that require minimal real-time inference cost or complex online reasoning, making the agent practically deployable in industrial or service environments where latency is critical.

Abstract

Building reliable robot capabilities across diverse tasks requires substantial human effort to develop and maintain skills, design rewards, and integrate perception with control. We present Reconstruct, Practice, Go Real (RPG), a framework for autonomous improvement of robot execution systems without updating model weights. RPG identifies manipulation capabilities in an offline dataset and constructs related practice tasks in simulation. During practice, RPG uses execution feedback, privileged simulator state, and available dataset videos to diagnose failures. It develops new reusable symbolic skills, refines existing skills, and revises the system prompt based on these diagnoses. Cross-task evaluation tests individual candidate changes and merged revisions before they are retained for reuse. At test time, a multimodal LLM uses the resulting system prompt and skill library to coordinate perception and robot control. On held-out initializations of 22 manipulation tasks, RPG improves task success from 28.6% after the first practice round to 95.0% after 15 rounds, outperforming all evaluated baselines, including ASPIRE (75.5%) and CaP-Agent0 powered by GPT-6 Astra Pro (60.0%). After a common calibration and hardware-adaptation procedure, the frozen system succeeds in all 30 physical trials, with ten trials on each of three tasks. Project Website: https://rpg-robot.github.io/

Sources

Related papers