RoboCoach: World Models as Active Coaches for Compositional Robot Skills

arXiv:2609.39685 · cs.RO, cs.AI · Submitted 2026-09-30 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: Today's paper: "RoboCoach: World Models as Active Coaches for Compositional Robot Skills".

Dev: Long-horizon robot manipulation reuses skills across many task compositions, but improving these compositions with additional end-to-end demonstrations is costly.

Rosa: First, who's behind it and why it matters.

Paper summary: Rosa: So, we're looking at the paper "RoboCoach: World Models as Active Coaches for Compositional Robot Skills." The main idea seems to be tackling the expense of improving skills in long-horizon robot manipulation by figuring out exactly which demonstrations are most valuable to collect next.

Dev: Exactly, Rosa. The authors address the fact that adding more end-to-end demonstrations is really costly when you're dealing with sequences of reusable skills that change the state for the next step. They want a system that decides both which skill to teach and where to apply those updates.

Taro: That cost issue is huge, especially when you consider physical robots where every trajectory takes time and effort. So, what's the core claim they are making about how this works?

Rosa: The core of RoboCoach is using a world model to guide coaching through a loop called RIDI, which stands for Route, Imagine, Diagnose, and Improve. It claims that imagined failures can be used to guide demonstration requests and expert updates.

Dev: The RIDI loop sounds like a structured way to manage the coaching process, which is something we need when we're dealing with complex tasks. It’s not just about gathering data; it’s about being proactive in how you get that data.

Taro: I'm curious about the Route phase specifically, since it translates a long-horizon instruction into a sequence of subtask-conditioned expert calls. How does the system decide which expert to call next in that route?

Rosa: The Route phase is all about translating that big goal into smaller, manageable steps by querying the current active expert and generating a chunk of CoachWorld. It commits those observations and then checks with a progress judge to see if the subtask is done or if it's time for a timeout.

Dev: That sounds like they're building this active coaching cycle inside CoachWorld, which is that shared action-conditioned world model that lets them interact in a closed loop. I need to know how fast this loop can run; the latency matters for real-time feedback.

Taro: When things go wrong in that loop, what happens during the Diagnose phase when the system encounters a bottleneck or a timeout? That’s where I want to know how it handles unexpected world misbehavior.

Rosa: The Diagnose phase uses a progress judge to estimate subtask completion based on calibrated thresholds and time limits for each subtask–embodiment pair. It records the first unresolved subtask–expert pair when a subtask stays incomplete until its time limit, which helps identify those recurring bottlenecks.

Dev: So, if you hit a timeout, what does that mean for the overall system's stability? The paper mentions recording the "first unresolved subtask–expert pair along the executed route" when a timeout occurs. That suggests some kind of attribution mechanism for failures.

Paper summary: Taro: And how does that information flow into the Improve phase to actually select what to do next? We need to see how the imagined failures translate into actual learning actions.

Rosa: The Improve phase aggregates all these records into a coaching scorecard to choose which subtask demonstrations and which expert adapters need updates. They calculate a task-balanced first-timeout mass for candidate pairs, picking the top M pairs to acquire new demonstrations for and update the experts with LoRA fine-tuning using a smaller replay sample.

Dev: That sounds like they're balancing the data acquisition effort against the expected improvement, which is smart given our constraints. But what about generalization? If we teach one expert with targeted demonstrations, does it stick across different task compositions?

Taro: That’s a big question for autonomy systems; if the skill learned generalizes well, that’s where the real utility lies. The results mentioned show coached experts can generalize to four held-out compositions, achieving an average success of thirty-five point zero percent compared with zero percent for a shared-policy baseline updated with uniformly acquired demonstrations.

Rosa: That correlation between imagined and deployed success is quite strong, showing a Spearman correlation of ρ = zero point eight four zero across two simulation suites and two real-robot platforms. It also showed that applying the same targeted demonstrations to corresponding skill experts, rather than a shared global adapter, improved final complete-task success by three point four percent on LIBERO and thirteen point two percent on RoboTwin.

Dev: I'm still concerned about deployment outside of the lab environment where we face occlusion and out-of-view interactions. The paper acknowledges that the framework relies on action-conditioned predictions, which can be tricky when you can't see everything.

Taro: That points to a real limitation they flag: the model's ability to predict trajectories remains challenging under occlusion and out-of-view interaction. Also, their expert library is predefined by skill semantics, and figuring out how to learn the structure of that library itself is still an open direction.

Rosa: So, to wrap up this part of the paper, RoboCoach uses world models as active coaches to turn imagined failures into targeted supervision for modular policy improvement. It's a very structured way to approach skill improvement in complex manipulation tasks.

Dev: It’s certainly a sophisticated framework, and I think the idea of using imagined failures as a signal for targeted updates is compelling for reducing the data collection burden. We'll need to see how robust this RIDI loop is when we push it into more dynamic, real-world scenarios where those world model predictions might fail.

Taro: It suggests that for complex tasks, the future of skill improvement might not just be about massive data collection, but about using sophisticated models to intelligently guide the learning process. We need to keep pushing on how these world models can handle uncertainty in real-time environments.

Rosa: That’s what we'll be looking at next, as we move into the conclusion where we discuss the broader implications of RoboCoach and this research.

Conclusion: Rosa: So, to wrap up, RoboCoach uses world models as active coaches to turn imagined failures into targeted supervision for modular policy improvement across skill compositions.

Dev: That's a pretty concise way to put it, Rosa; I’m still thinking about how this entire RIDI loop manages the flow of information and decisions under real-time constraints.

Taro: Yeah, the structure itself is interesting because it moves beyond just collecting more data and starts using imagination to guide that collection process.

Rosa: Exactly, Taro; we're looking at how they've framed skill improvement not as a data-gathering treadmill but as an intelligent coaching cycle driven by prediction.

Dev: I wonder how stable this system is when we take it off the simulator and onto a physical robot; will those world model predictions hold up under real-world unpredictability?

Taro: That’s a big question for me, Dev; I'm interested in what happens when the world misbehaves during that imagined interaction phase.

Rosa: It seems like they're tackling that uncertainty head-on by making the world model an active participant in deciding what to learn next.

Dev: I think it’s promising because it directly addresses the cost of end-to-end demonstrations, which is a major hurdle for scaling robot learning systems.

Taro: And their findings on generalization across different task compositions are really telling about how robust these learned skills actually are in practice.

Rosa: It certainly suggests that for complex manipulation, the future of skill improvement might be less about brute force data collection and more about using sophisticated models to intelligently guide the learning process.

Jiajun Liu, *Yifan Chen*, *Yichao Liu*, *Jiayi Zhang*, Ruoqu Chen, Shaoxuan Xie, Guocai Yao, Mengdi Xu†, Sen Cui†‡, Changshui Zhang†

Renmin University of China · Tsinghua University · University of Nottingham · Beijing Academy of Artificial Intelligence 5

cs.RO, cs.AI

Submitted: 2026-09-30

Updated: 2026-09-30

Comments: https://robocoach-ai.github.io/

Project page: https://robocoach-ai.github.io

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 91/100

The gist: Long-horizon robot manipulation reuses skills across many task compositions, but improving these compositions with additional end-to-end demonstrations is costly.

Key concepts

Route
This phase translates a complex, long-horizon instruction into a sequence of calls to specific subtask experts. The router decides which skill comes next by checking the current expert's status and generating world model chunks to determine the path forward.
Imagine
The active expert is run inside CoachWorld in a closed loop. It predicts future end-effector trajectories based on sparse visual history and task instructions, creating 'future-video predictions' that allow for policy evaluation through imagined interaction.
Diagnose
This phase uses a progress judge to estimate how complete each subtask is. It identifies bottlenecks by recording the first unresolved expert pair when a subtask hits its time limit, pinpointing where coaching is most needed.
Improve
Records from the RIDI loop are aggregated into a scorecard. This phase selects which specific demonstrations and which expert adapters to update using LoRA fine-tuning, ensuring targeted improvements based on task-balanced first-timeout mass.

Terminology

Summary

Long-horizon robot manipulation reuses skills across many task compositions, but improving these compositions with additional end-to-end demonstrations is costly. RoboCoach presents a world-model-guided coaching framework that uses imagined failures to guide demonstration requests and expert updates.

How it works

RoboCoach utilizes a Route–Imagine–Diagnose–Improve (RIDI) loop to manage the coaching process. The core idea is to decompose long-horizon execution into reusable atomic skills, each handled by a corresponding expert, linked via a subtask–expert pair. The router selects the next unfinished subtask and dispatches its expert. This active coaching cycle operates inside CoachWorld, a shared action-conditioned world model that allows for imagined interaction in closed loop.

Route: from a long-horizon instruction to an atomic expert

The Route phase translates a long-horizon instruction into a sequence of subtask-conditioned expert calls. The router determines the next skill by querying the current active expert and generating one complete CoachWorld chunk, committing observations, and querying the progress judge. If the active subtask reaches its completion threshold, it is appended to the completed prefix before routing to the next subtask; otherwise, if it reaches its time limit, a TIMEOUT is recorded. This process ensures that additional demonstrations improve skills that transfer across tasks.

Imagine: policy-in-the-loop rollout with CoachWorld

The Imagine phase involves rolling out the active expert inside CoachWorld in closed loop. At time t, the active expert predicts an end-effector trajectory, which is converted into a future video chunk via a lightweight adapter. This prediction is conditioned on sparse visual history and task instructions to generate future-video prediction. The resulting chunk is committed before the next policy or judge query, forming a closed-loop rollout that allows for policy evaluation through imagined interaction.

Diagnose: progress-based switching and first-timeout attribution

The Diagnose phase uses a progress judge to estimate subtask completion, defined by calibrated completion thresholds and time limits for each subtask–embodiment pair. The system records the first unresolved subtask–expert pair when a subtask remains incomplete until its time limit. This mechanism identifies recurring subtask–expert bottlenecks, and when a timeout occurs, it records the first unresolved subtask–expert pair along the executed route.

Improve: scorecard-guided data acquisition

The Improve phase aggregates records into a coaching scorecard to select which subtask demonstrations to acquire and which expert adapters to update. The task-balanced first-timeout mass is calculated for candidate pairs, and the top M pairs are selected. New demonstrations are acquired for these pairs, and the corresponding experts are updated using LoRA fine-tuning with a smaller replay sample from existing data. This ensures that the coaching method outperforms matched baselines under matched data budgets and update schedules. The process then closes the RIDI loop, where updated experts return to the route.

Key Findings

Experiments across two simulation suites and two real-robot platforms show that imagined and deployed success correlate over 22 task–policy pairs with a Spearman correlation of ρ = 0.840. Controlled comparisons demonstrate that applying the same targeted demonstrations to corresponding skill experts rather than a shared global adapter improves final complete-task success by 3.4 and 13.2 pp on LIBERO and RoboTwin, respectively. Furthermore, coached experts can generalize to four held-out compositions, achieving an average success of 35.0% compared with 0% for a shared-policy baseline updated with uniformly acquired demonstrations. The results show that world models can serve as active coaches, turning imagined failures into targeted supervision for modular policy improvement.

Limitations

The framework relies on action-conditioned predictions, which remains challenging under occlusion and out-of-view interaction. Agreement across generation seeds reduces sensitivity to stochastic variation but cannot rule out systematic world-model bias. The expert library is predefined by skill semantics, and learning its structure remains an open direction.


The gist

RoboCoach uses imagined execution to decide which demonstrations to collect and which reusable skill experts should learn from them, connecting world-model prediction, progress-based diagnosis, and targeted supervision to iterative policy improvement.

Route: from a long-horizon instruction to an atomic expert

The Route phase translates a long-horizon instruction into a sequence of subtask-conditioned expert calls.

Improvements for AI systems

Here are the specific improvements and capabilities that an AI system, guided by RoboCoach, can achieve:

  1. A self-improving agent capable of autonomously deciding which skills to teach next and which policy components to update based on its own imagined failures.

  2. The ability to improve long-horizon robot manipulation tasks (like complex cooking or table setting) using only a limited budget of real-world demonstrations (e.g., 150 additional subtask demos instead of requiring full end-to-end task demonstrations).

  3. Targeted policy improvement: Instead of updating a shared global policy, the system can selectively update specific, reusable skill experts (via LoRA adapters) that are directly responsible for the failed subtask. This prevents catastrophic forgetting and ensures modularity.

  4. Data efficiency gains: The system achieves significantly higher success rates on real-world platforms (e.g., 75% on Franka, 83.8% on AgileX) using a fraction of the data compared to uniform acquisition strategies (e.g., from 13.3% to 75.0%).

  5. Generalization across unseen compositions: The improved system learns reusable skills that can be successfully recombined into novel task sequences or sequences with reordered steps that were not explicitly used during the initial coaching phase, demonstrating true compositional generalization (achieving 35% success on held-out compositions).

  6. Robust decision-making under uncertainty: By using a progress judge based on world model predictions, the system can reliably distinguish between a subtask that is genuinely complete and one that has simply timed out or failed to progress, leading to targeted supervision instead of random data collection.

  7. Improved policy fidelity: The system ensures that imagined successes within the world model align strongly (Spearman correlation of ρ = 0.840) with actual deployed policy outcomes, meaning the internal simulation accurately reflects real-world performance differences.

  8. Enhanced diagnostic capability: The system can pinpoint exactly where a long-horizon failure occurs (e.g., the bread placement failed), allowing for highly specific interventions rather than broad policy tuning.

Sources

Related papers