SAIL: Test-Time Scaling for In-Context Imitation Learning with VLM
cs.RO, cs.AI
Submitted: 2026-03-09
Updated: 2026-09-19
Comments: Accepted to IROS 2026
License: http://creativecommons.org/licenses/by/4.0/
The gist: In-context imitation learning allows robots to acquire skills from demonstrations, yet one-shot trajectory generation remains fragile under environmental variation.
Terminology
Abstract
In-context imitation learning allows robots to acquire skills from demonstrations, yet one-shot trajectory generation remains fragile under environmental variation. We propose SAIL, a framework that reframes robot imitation as an iterative refinement problem capable of scaling with test-time compute. SAIL utilizes Monte Carlo Tree Search, where each node is a complete trajectory and edges correspond to trajectory refinements. The process is guided by three core components: an automated archive of successful trajectories for contextually relevant retrieval, a vision language model-based scoring mechanism for trajectory evaluation, and a step-level feedback that provides trajectory-aligned scores for iterative refinement. Experiments across six diverse manipulation tasks in simulation and real-world validation clearly demonstrate that increasing test-time compute consistently improves success rates, achieving up to 95% on complex tasks. Our results suggest that trajectory-level test-time scaling is a robust path toward more generalizable robotic agents.
Sources
- MimicDroid: In-Context Learning for Humanoid Robot Manipulation from Human Play Videos
- In-Context Imitation Learning via Next-Token Prediction
- RT-Trajectory: Robotic Task Generalization via Hindsight Trajectory Sketches
- RoboMP$^2$: A Robotic Multimodal Perception-Planning Framework with Multimodal Large Language Models
- RAM: Retrieval-Based Affordance Transfer for Generalizable Zero-Shot Robotic Manipulation
- Trajectory Adaptation using Large Language Models
- GELATO: Multi-Instruction Trajectory Reshaping via Geometry-Aware Multiagent-based Orchestration
- LLM3:Large Language Model-based Task and Motion Planning with Motion Failure Reasoning
- Code-as-Symbolic-Planner: Foundation Model-Based Robot Planning via Symbolic Code Generation
- Do As I Can, Not As I Say: Grounding Language in Robotic Affordances
- Prime the search: Using large language models for guiding geometric task and motion planning by warm-starting tree search
- SayPlan: Grounding Large Language Models using 3D Scene Graphs for Scalable Robot Task Planning
- Gemini Robotics: Bringing AI into the Physical World
- Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection
- SAM 2: Segment Anything in Images and Videos
- Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware
- GSWorld: Closed-Loop Photo-Realistic Simulation Suite for Robotic Manipulation
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving