Adaptation of Generalist Robot Policies with Minimal Data

arXiv:2608.11363 · cs.RO, cs.LG · Submitted 2026-08-11 · Read on arXiv

Shreyas Kowshik, Sreyas Venkataraman, Leo Wang, Niharika Pant, Max Simchowitz, Aviral Kumar

Carnegie Mellon University

cs.RO, cs.LG

Submitted: 2026-08-11

Updated: 2026-08-13

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 95/100

The gist: The paper introduces minimal-data adaptation (MDA), a regime in which "a pre-trained robot policy must learn a new task from as little as one demonstration followed by autonomous online interaction."

Terminology

Summary

The paper introduces minimal-data adaptation (MDA), a regime in which a pre-trained robot policy must learn a new task from as little as one demonstration followed by autonomous online interaction. The authors note that fully autonomous learning remains difficult with current policies: sparse rewards and weak zero-shot exploration make it unlikely that a robot will discover successful behavior from scratch. MDA serves as the closest tractable proxy for fully autonomous improvement, allowing us to study whether minimal human guidance can bootstrap autonomous learning and what algorithmic ingredients make it feasible.

The paper formalizes the setting: "Given a pre-trained vision-language-action policy πbase, K successful demonstrations Ddemo collected in a target environment (denoted by Mtarget), and access to autonomous interaction with Mtarget, minimal-data adaptation asks whether we can learn an adapted policy that achieves high return in Mtarget using only this limited supervision. The case K=1 is described as the sharpest version of this setting."

The authors build MiDAS (Minimal-Data Adaptation Strategy), described as "a simple offline-to-online RL recipe that first anchors a pre-trained VLA to the target task with behavior cloning on single/few demonstrations, then improves it through value-based online RL on a residual policy parameterization."

The method consists of two stages:

Stage I: Behavior cloning on minimal demonstrations. Stage I fine-tunes πbase on the K demonstrations with behavior cloning, producing a policy that can attempt the correct task but remains unreliable. The authors adapt the VLM backbone using a low-rank adapter (LoRA) and fine-tune the complete action head using the standard flow matching objective.

Stage II: Online improvement via value-based residual RL. Stage II freezes the base policy and trains a lightweight residual actor-critic on top of its representations using sparse-reward online RL. The residual policy is parameterized as: πθres(·st, abase) = tanh(N(μθ(st, abase), σ2θ(st, abase))), which directly predicts the edited action rather than an additive offset. The authors explain: "Conditioning the actor on abase exposes it directly to the action mode selected by the base policy, so online RL refines a task-relevant VLA proposal rather than learning to represent a complex, multimodal action distribution from scratch."

Three key design choices make Stage II work:

  1. Offline warmup: "A randomly initialized residual actor would produce actions unrelated to the base policy, making early exploration unstable. We therefore construct a 'warmup' buffer from the K demonstrations together with a small number of autonomous rollouts" and train the residual actor to reproduce base proposals.

  2. Success balancing: "Since the base policy may only succeed 2-5% in some cases after Stage I, we employ success balancing to amplify the sparse reward signal... Critic minibatches are sampled from a mixture of these buffers, oversampling the success buffer to prevent the value function from being dominated by failed rollouts."

  3. PA-RL residual policy training: We implement this by sampling a single proposal action from the base policy, and then applying PA-RL style Best-of-N distillation objective with an action gradient to a batch of N actions sampled from the residual.

Simulation benchmarks. The authors evaluate on LIBERO-Long, which studies long-horizon language-conditioned manipulation, and RoboCasa-365, which studies long-horizon household manipulation in diverse kitchen scenes. Using π0.5 as the pre-trained base policy with K=1 demonstration:

  • On LIBERO-Long, MiDAS achieves 91.2% average success versus 33.5% for BC, 33.5% for DSRL, and 39.0% for Filtered BC.

  • On RoboCasa, MiDAS achieves 89.3% average success versus 22.2% for BC, 19.1% for DSRL, and 34.0% for Filtered BC.

Real-world transfer. We finally evaluate whether minimal-data adaptation is tractable in the real world using a bimanual YAM platform. On two tasks—placing colored blocks in containers and placing a knife and donut on a plate—MiDAS improves success from 40% to 67% and from 27% to 80% respectively, after 5-6 hours of autonomous interaction.

1. Pretraining and demonstrations enable coarse task-specific behaviors. The authors find that behavior cloning on minimal demonstrations instills rough task solving behavior into the pre-trained policy. They show that alternatives fail: a flow-matching policy trained on pre-trained ResNet features... a flow-matching policy trained on privileged state inputs... and the zero-shot pre-trained policy all perform poorly. The privileged-state policy and πbase obtain 0% success, suggesting that MDA requires both task grounding from the demonstration and behavior-aware representations from pretraining.

2. Pre-trained representations simplify online adaptation. Representations from the pretrained π0.5 backbone enable rapid improvement throughout training. Without visual pretraining, MiDAS fails to converge at K=1. Generic visual pretraining with frozen DINO substantially improves learning, but still requires considerably more interaction than VLA features.

3. Value-based RL learns actions beyond the base policy's support. The authors demonstrate that the fine-grained control missing from the base policy in this case does not lie within the base policy's action support; sharpening or steering within this support cannot reach it. They show via UMAP visualization and distance metrics that MiDAS selects an action in a disjoint region of the action space and succeeds, with minimum distance to the base action cloud of 4.5371 versus 0.1122 for Filtered BC and 0.0244 for DSRL.

Observation shifts. MiDAS generalizes to observation shifts via the frozen VLM backbone. Under visual shifts (color, texture changes) and language paraphrases, the adapted policy shows a similar retention pattern to the base policy, with average success dropping only from 88.7% to 82.0% under visual perturbation and from 91.2% to 90.2% under language perturbation.

State shifts. Generalization to state shifts depends on whether new behavior is required. Under object swap perturbations, both base and adapted policies drop to 0%. Under shape changes, performance preserves when grasp affordances transfer, but degrades sharply when the required contact geometry changes. Under object changes (replacing with categorically different objects), both policies fail.

Curriculum for position generalization. The authors show that progressively widening the reset distribution during online RL recovers most of the position-generalization gap with the 50-demonstration policy without additional demos. On Both Moka Pots task, curriculum training raises 10 cm success from 22.7% to 75.3%, nearly matching the 76.0% of the 50-demonstration policy.

The authors note several limitations: learning qualitatively new behaviors or reaching distinct behavioral modes remains substantially harder, and reset expansion cannot recover behaviors that lie outside the support of the initial policy. They also acknowledge that our experiments primarily consider pick-and-place tasks that are plausibly well represented in the pretraining distribution of π0.5, and that it is therefore unclear whether the same interaction between pretraining, minimal demonstrations, and autonomous RL will hold when adaptation requires greater policy capacity or more substantial behavioral change. The LoadDishwasher task in RoboCasa shows the limitation of our proposed approach and provides motivation for future work to extend MiDAS to long-horizon, sparse-reward settings.

Improvements for AI systems

Improvements to AI Systems:

  1. Residual policy architecture with base-policy conditioning: Instead of training RL from scratch or fine-tuning the entire policy, the AI system maintains a frozen base policy and learns a lightweight residual actor that conditions on the base policy's proposed action. This allows the system to refine expert proposals rather than rediscover action distributions, dramatically reducing the data needed for adaptation (from thousands of demonstrations to 1).

  2. Two-stage offline-to-online adaptation pipeline: The system first anchors behavior via behavior cloning on minimal demonstrations (using LoRA for the vision-language backbone and full fine-tuning of the action head), then switches to sparse-reward online RL with a residual actor-critic. This enables the system to bootstrap from coarse task knowledge and refine it through autonomous interaction, achieving 91.2% success on LIBERO-Long with just one demonstration.

  3. Offline warmup buffer for stable RL initialization: Before online RL, the system trains the residual actor to reproduce base policy proposals using a buffer constructed from demonstrations plus a few autonomous rollouts. This prevents the instability of randomly initialized residual actors and ensures early exploration stays in task-relevant regions of action space.

  4. Success-balanced replay sampling: The system maintains separate buffers for successful and failed rollouts and oversamples successes during critic training. This prevents the value function from being dominated by failed trajectories when the base policy succeeds only 2-5% of the time, enabling effective credit assignment in sparse-reward settings.

  5. PA-RL style Best-of-N distillation with action gradients: During online RL, the system samples multiple residual actions per state, evaluates them via the critic, and distills the best ones using action gradients. This improves sample efficiency by leveraging parallel action proposals and provides denser learning signal than single-action policy gradient methods.

  6. Curriculum-based reset distribution expansion: The system progressively widens the initial state distribution during online RL (e.g., starting near the goal and moving farther away). This recovers position generalization gaps without additional demonstrations, improving success from 22.7% to 75.3% on a long-horizon task, nearly matching a policy trained with 50 demonstrations.

What the Improved AI System Can Do:

  • Learn new manipulation tasks from a single human demonstration followed by autonomous practice, achieving over 90% success in simulation and 67-80% in real-world bimanual tasks after 5-6 hours of self-interaction.

  • Adapt to new environments without catastrophic forgetting of the base policy's general capabilities, maintaining robustness to visual perturbations (color/texture changes) and language paraphrases.

  • Discover actions outside the base policy's support when fine-grained control requires novel motor commands, as demonstrated by selecting actions in disjoint regions of action space that succeed where sharpening or steering the base policy fails.

  • Generalize to position shifts through curriculum-based self-training, eliminating the need for dense human demonstrations across varied starting configurations.

  • Operate with sparse rewards by amplifying rare success signals through balanced replay and leveraging pre-trained vision-language representations that provide behavior-aware features for rapid value learning.

  • Scale to long-horizon tasks (e.g., LIBERO-Long with multi-step manipulation) while remaining sample-efficient, though the system still struggles with very long-horizon sparse-reward tasks like dishwasher loading, indicating a clear direction for future improvement.

Abstract

A central goal in robot learning is to move beyond task-specific human data collection toward robots that improve through autonomous interaction. Yet fully autonomous learning remains difficult with current policies: sparse rewards and weak zero-shot exploration make it unlikely that a robot will discover successful behavior from scratch. We study minimal-data adaptation, a regime in which a pre-trained robot policy must learn a new task from as little as one demonstration followed by autonomous online interaction. This setting serves as the closest tractable proxy for fully autonomous improvement, allowing us to study whether minimal human guidance can bootstrap autonomous learning and what algorithmic ingredients make it feasible. We build MiDAS, a simple offline-to-online RL recipe that first anchors a pre-trained VLA to the target task with behavior cloning on single/few demonstrations, then improves it through value-based online RL on a residual policy parameterization. Across LIBERO and RoboCasa, MiDAS recovers strong task performance from as little as one demonstration, substantially outperforming baselines and generalizing beyond demonstrated conditions. We further evaluate MiDAS on a bimanual YAM platform. Starting from a fragile low-success policy obtained from a single demonstration, MiDAS improves its robustness and learns new successful behaviors over 6 hours of online interaction. To the best of our knowledge, this is the first demonstration of reliable robot policy adaptation from a single task demonstration.

Sources

Related papers