REFACTOR-VLA: Unsupervised Library Learning of Typed Motor Programs
cs.LG, cs.AI, cs.RO
Submitted: 2026-09-01
Updated: 2026-09-01
Comments: 30 pages, 5 figures
License: http://creativecommons.org/licenses/by/4.0/
The gist: Most vision-language-action (VLA) models -- OpenVLA, π 0, RT-2, RDT-1B -- are monolithic: they emit raw motor commands or short action chunks without organizing behavior into reusable abstractions,
Terminology
Abstract
Most vision-language-action (VLA) models -- OpenVLA, π 0, RT-2, RDT-1B -- are monolithic: they emit raw motor commands or short action chunks without organizing behavior into reusable abstractions, so they degrade on long-horizon tasks and resist interpretation. Existing skill-discovery methods sidestep the core question of when two action sequences are behaviorally equivalent, either clustering contrastive embeddings or delegating the judgment to a language model uncalibrated to the robot's dynamics. We introduce REFACTOR-VLA, a wake/sleep system for learning reusable skills. Its sleep phase clusters motor-program fragments under a Behavioral-Equivalence Kernel (BEK) computed from rollouts of a learned latent world model M ϕ; its wake phase emits typed lambda terms over a Hindley--Milner-inspired vocabulary, consumed by a library-conditioned rectified-flow action decoder. Abstractions are admitted only if they pass Minimum Description Length and return-preservation gates. On LIBERO we report two findings. First, enlarging the world model from 188M to 430M parameters worsened performance on 4 of 4 suites, so capacity alone does not help. Second, the training objective matters far more: adding an auxiliary supervised contrastive (InfoNCE) loss during world-model warmup substantially improves sleep-phase clustering, giving Normalized Mutual Information at n=3 seeds of 0.462 plus or minus 0.021 (object), 0.867 plus or minus 0.025 (spatial), 0.915 plus or minus 0.013 (goal) and 0.754 plus or minus 0.010 (LIBERO-10), and beating the strongest published baseline on all 4 suites by a mean Δ= +0.184. Across providers (n=12) the 95% bootstrap confidence interval for mean pairwise NMI is [0.683, 0.729] (mean 0.705). The sleep phase also yields the first real-LIBERO task-language library: the decoder uses 2 of 3 admitted abstractions and rewrites all 256 sampled demonstrations.
Sources
- The Option-Critic Architecture
- Hierarchical State Space Models for Continuous Sequence-to-Sequence Modeling
- Top-Down Synthesis for Library Learning
- Scalable methods for computing state similarity in deterministic Markov Decision Processes
- Data-driven model predictive control: closed-loop guarantees and experimental results
- A Simple Framework for Contrastive Learning of Visual Representations
- Diffusion Policy: Visuomotor Policy Learning via Action Diffusion
- DreamCoder: Growing generalizable, interpretable knowledge with wake-sleep Bayesian program learning
- LILO: Learning Interpretable Libraries by Compressing and Documenting Code
- Relay Policy Learning: Solving Long-Horizon Tasks via Imitation and Reinforcement Learning
- Mastering Diverse Domains through World Models
- Supervised Contrastive Learning
- OpenVLA: An Open-Source Vision-Language-Action Model
- CompILE: Compositional Imitation Learning and Execution
- Revisiting k-means: New Algorithms via Bayesian Nonparametrics
- LIBERO: Benchmarking Knowledge Transfer for Lifelong Robot Learning
- Learning Compositional Behaviors from Demonstration and Language
- Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow
- RoboCasa: Large-Scale Simulation of Everyday Tasks for Generalist Robots
- DINOv2: Learning Robust Visual Features without Supervision
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks