Deliberate Practice: Learning Robot Skills under a Budget

arXiv:2608.13415 · cs.RO, cs.AI · Submitted 2026-08-13 · Read on arXiv

Shivam Vats, Sudarshan Harithas, Mete Tuluhan Akbulut, Arvind Raghunathan, George Konidaris

Brown University · Mitsubishi Electric Research Laboratories

cs.RO, cs.AI

Submitted: 2026-08-13

Updated: 2026-08-14

Comments: 16 pages including appendices

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 75/100

The gist: Learning Robot Skills under a Budget" by Shivam Vats, Sudarshan Harithas, Mete Tuluhan Akbulut, Arvind Raghunathan, and George Konidaris addresses the problem of autonomously learning robot skills

Terminology

Summary

Summary

The paper Deliberate Practice: Learning Robot Skills under a Budget by Shivam Vats, Sudarshan Harithas, Mete Tuluhan Akbulut, Arvind Raghunathan, and George Konidaris addresses the problem of autonomously learning robot skills for sequential tasks under a limited practice budget. The authors propose an active skill learning algorithm called Deliberate Practice (DP) that computes a provably budget-optimal allocation—practicing skills that maximize expected cumulative reward while being learnable within the budget.

The key insight is that downtime is often known in advance, and robots should leverage this information to adapt their learning process to the resulting practice budget. The paper formalizes this as budgeted skill learning and considers the standard Task and Motion Planning (TAMP) setting where the robot is given high-level skill specifications including preconditions, termination conditions, and effects. The robot must learn parameterized control policies that ground these skills into low-level actions through practice.

The problem is formulated as a bilevel optimization problem: the outer problem computes a budget allocation for practicing policies, while the inner problem solves the resulting task MDP to compute expected planning performance. The authors note that computing an optimal allocation requires solving a challenging bilevel optimization problem and that existing active learning algorithms approximate this problem greedily, leading to myopic and suboptimal learning.

The paper demonstrates that greedy active learning methods converge to local minima using an illustrative example with three skills (π1, π2, π3) where greedy methods practice only π1, missing the optimal solution π2, π3 which provides higher reward under the full budget. The authors explain that identifying the optimal budget allocation requires reasoning about deliberate practice across π2 and π3 over multiple rounds.

The main technical contribution is an exact single-level reformulation of the bilevel problem using the linear programming (LP) formulation of Markov decision processes. This reformulation yields a structured single-level bilinear program that enables efficient optimization using standard nonlinear programming solvers. The bilinear program incorporates practice budget constraints directly into the standard dual linear program of an MDP, with the competence improvement function modeled as a piece-wise linear function: fimprov(u, b) = min(1, pu + Δu b), where Δu is the rate of improvement estimated online.

The authors prove Theorem 1: Deliberate Practice is budget-optimal, i.e, it computes a globally optimal allocation of the practice budget. The proof relies on showing that the bilinear program is an exact reformulation of the bilevel program by replacing the inner optimization with its LP formulation, utilizing LP duality, and solving to global optimality using spatial branch and bound methods with Piecewise McCormick Envelopes.

The Deliberate Practice algorithm comprises three steps: (1) Competence Prediction, which models competence improvement as a function of practice budget; (2) Budget Allocation, which computes a budget-optimal task plan and corresponding budget allocation; and (3) Skill Practice, where the robot sequentially masters skills that are reachable from the start and uses them to plan to reach and practice downstream skills.

Experiments are conducted in three long-horizon table-top manipulation tasks: Cleanup (simulated), Cleanup-Multi (simulated), and Breakfast (real-robot). The Cleanup task has 47 abstract states and 10 skills, while Cleanup-Multi has 5000 abstract states and 22 skills. The Breakfast task requires learning forceful manipulation skills to interact with novel articulated objects.

Results show that "DP intelligently chooses which skills to learn based on the available budget: conservatively placing items in the top drawer when the budget is 100 episodes, placing items in the middle drawer under a budget of 150 episodes, and maximizing reward by placing items in the bottom drawer under a budget of 250 episodes. By contrast, baselines cannot adapt their behavior to the budget and greedily practice the easiest (and lowest-reward) task plan, irrespective of the budget."

The paper reports that DP and EES perform similarly under a low budget, but DP significantly outperforms all baselines in medium (150 episodes) and high budget (250 episodes) settings. In the more complex Cleanup-Multi task, random practice outperforms EES in this setting because EES lacks an exploration mechanism when the myopic task improvement measure ΔJtask is zero for all skills.

For real-world validation on the Breakfast task, DP correctly chooses the conservative toast-bread plan under the smaller budget, but learns to microwave oatmeal under the larger budget, thereby achieving higher task reward. With a budget of 30 episodes, DP practices to toast bread (reward 1), while under a budget of 60 episodes, it practices to microwave oatmeal (reward 2).

The paper acknowledges limitations: "One key assumption of our approach is access to approximate priors over skill competence. If these priors are overly optimistic, the robot may allocate practice to task plans that are actually infeasible within the available budget. Additionally, solving the bilinear program to global optimality may become challenging for very large problems," suggesting future work on bounded-suboptimal optimization strategies and extending to generalizable skills across multiple tasks on mobile manipulators.

Improvements for AI systems

Improvements to AI Systems:

  1. Budget-Aware Sequential Decision-Making: AI systems can now explicitly incorporate a known or estimated practice/learning budget into their optimization objectives. Instead of greedily maximizing immediate reward or learning the easiest skills, the system can compute a globally optimal allocation of resources across multiple skills, adapting its strategy based on whether the budget is low, medium, or high.

  2. Non-Myopic Active Learning via Exact Reformulation: The system replaces greedy, myopic active learning heuristics with an exact single-level bilinear program derived from the LP formulation of MDPs. This allows the AI to reason about multi-round, coordinated practice across multiple skills (e.g., practicing π2 and π3 together) rather than getting stuck in local minima (e.g., always practicing only π1).

  3. Online Competence Modeling with Piecewise-Linear Priors: The system can estimate its own rate of skill improvement (Δu) online and model competence as a piecewise-linear function of practice budget. This enables the AI to predict which skills are learnable within the remaining budget and to avoid over-committing to skills that are too difficult or too easy relative to the reward they yield.

  4. Hierarchical Skill Grounding under Constraints: The system can now plan at the abstract Task and Motion Planning (TAMP) level while simultaneously deciding how many practice episodes to allocate to each low-level control policy. This bridges high-level symbolic planning with low-level continuous control, ensuring that learned policies are actually reachable and executable within the budget.

  5. Adaptive Task Selection Based on Budget: The AI can dynamically choose different task plans (e.g., placing items in top vs. middle vs. bottom drawer) depending on the available budget, maximizing expected cumulative reward without exceeding resource limits. This is a shift from fixed-policy learning to budget-conditional policy selection.

  6. Robustness to Exploration Gaps: By incorporating a global optimization over budget allocation, the system avoids the failure mode where random practice outperforms active learning (as seen with EES in Cleanup-Multi). The improved system can detect when myopic improvement signals are zero and still allocate practice to unexplored but potentially high-reward skills.

  7. Provably Optimal Resource Allocation: The system now has a theoretical guarantee (Theorem 1) that its practice allocation is globally optimal under the given budget and competence model. This provides reliability for safety-critical or resource-constrained robotic applications where suboptimal learning could lead to task failure.


What the Improved AI System Can Do:

  • In Robotics: A robot can autonomously decide, given a fixed number of practice episodes (e.g., 30 vs. 60), whether to learn a simple low-reward skill (e.g., toasting bread) or a complex high-reward skill (e.g., microwaving oatmeal), and then execute the corresponding task plan optimally.

  • In Reinforcement Learning: An agent can allocate a limited training compute budget across multiple sub-policies in a hierarchical task, avoiding wasted effort on unreachable or low-impact skills, and achieving higher cumulative reward than greedy or random baselines.

  • In Curriculum Learning: The system can design a practice curriculum that is provably optimal for a given budget, selecting which skills to master first and how much time to spend on each, rather than relying on hand-crafted or heuristic schedules.

  • In Multi-Task Learning: The AI can jointly optimize which tasks to practice and how to share practice across them, ensuring that the overall learning budget is spent on the combination of skills that maximizes expected performance on the final objective.

  • In Real-Time Adaptation: When the budget changes mid-task (e.g., due to unexpected downtime), the system can re-solve the bilinear program to reallocate remaining practice, gracefully degrading or upgrading its task plan accordingly.

Abstract

We consider the problem of autonomously learning robot skills under a limited practice budget for sequential tasks. We propose an active skill learning algorithm, Deliberate Practice (DP), that computes a provably budget-optimal allocation---practicing skills that maximize expected cumulative reward while being learnable within the budget. DP estimates both the time needed to master skills and the cumulative reward of the task plans that the skills unlock. Computing a budget-optimal allocation is challenging as it requires reasoning about combinatorially many skill plans over a large practice budget. Our key contribution is a bilinear program that can compute this exactly using off-the-shelf solvers. Through simulated and real-world experiments on long-horizon manipulation tasks, we show that our approach allows robots to optimally use limited practice time to acquire useful policies and improve long-horizon planning.

Sources

Related papers