InSight: Self-Guided Skill Acquisition via Steerable VLAs
cs.RO, cs.AI, cs.LG
Submitted: 2026-06-23
Updated: 2026-09-21
Comments: Project website: https://insight-vla.github.io
Project page: https://insight-vla.github.io
License: http://creativecommons.org/licenses/by/4.0/
The gist: Vision-language-action (VLA) models excel at robot manipulation via imitation learning, but adapting them to new tasks often requires additional human demonstrations, which can be costly or
Terminology
Abstract
Vision-language-action (VLA) models excel at robot manipulation via imitation learning, but adapting them to new tasks often requires additional human demonstrations, which can be costly or infeasible. Meanwhile, vision-language models (VLMs) offer semantic task understanding but lack the physical grounding required for execution. To bridge this gap, we present InSight, a framework for self-guided skill acquisition that uses a VLM to identify primitives missing from a VLA's repertoire, grounds the VLM's proposals through robot execution, and distills new primitives from successful rollouts into the VLA. Primitive steerability, the ability to execute and terminate primitives on command, enables the robot to reuse known primitives while collecting training data for missing primitives without requiring full-task human demonstrations for each new task. InSight has two stages: (1) a VLM automatically segments existing demonstrations into primitive-labeled trajectories to fine-tune a primitive-steerable VLA, and (2) the VLM plans a sequence of known primitives executed by the VLA and new primitives attempted by VLM-parameterized low-level controllers. New-primitive segments from successful task rollouts are added to the training data, and the VLA is retrained. The adapted VLA can then reliably execute new skills using the acquired primitives, without per-primitive VLM calls. We evaluate InSight on six simulated and real-world tasks with no human demonstrations of target skills, including block flipping, drawer closing, sweeping, twisting, and pouring. On hardware, acquired twisting and pouring skills achieve 92% and 96% success, versus 32% and 16% for a zero-shot CaP-X baseline. Composing both skills into a 14-primitive task achieves 80% success with no combined-task demonstrations. Project website: https://insight-vla.github.io/.
Sources
- OpenVLA: An Open-Source Vision-Language-Action Model
- $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization
- GR00T N1: An Open Foundation Model for Generalist Humanoid Robots
- QT-Opt: Scalable Deep Reinforcement Learning for Vision-Based Robotic Manipulation
- Steering Your Diffusion Policy with Latent Space Reinforcement Learning
- Learning Diffusion Policy from Primitive Skills for Robot Manipulation
- Integrated Task and Motion Planning
- Steerable Vision-Language-Action Policies for Embodied Reasoning and Hierarchical Control
- Code as Policies: Language Model Programs for Embodied Control
- Do As I Can, Not As I Say: Grounding Language in Robotic Affordances
- VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models
- CaP-X: A Framework for Benchmarking and Improving Coding Agents for Robot Manipulation
- ${\pi}_{0.7}$: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities
- VLS: Steering Pretrained Robot Policies via Vision-Language Models
- Movement Primitives in Robotics: A Comprehensive Survey
- Learning Compositional Behaviors from Demonstration and Language
- Bottom-Up Skill Discovery from Unsegmented Demonstrations for Long-Horizon Robot Manipulation
- Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models
- RLDG: Robotic Generalist Policy Distillation via Reinforcement Learning
- ReWiND: Language-Guided Rewards Teach Robot Policies without New Demonstrations
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving