HABILIS Brain 0: Geometry-Change Supervision for Vision-Language-Action and Residual Flow Recovery
cs.RO, cs.LG
Submitted: 2026-09-22
Updated: 2026-09-22
License: http://creativecommons.org/licenses/by/4.0/
The gist: Vision-language-action policies benefit from geometric supervision, but current-frame geometry alone does not explicitly describe the changes associated with manipulation.
Terminology
Abstract
Vision-language-action policies benefit from geometric supervision, but current-frame geometry alone does not explicitly describe the changes associated with manipulation. This design is motivated by the goal of learning an embodiment-agnostic visual interface that can be pretrained across robot and egocentric video before robot-specific action alignment. We introduce Geometry-Change VLA (GC-VLA), which learns to predict multiview future-current geometry-change tokens from current observations. Offline frame pairs define a nominal 0.5-second prediction horizon; future observations are used only to construct training targets. Stage 1 trains a geometry-change vision-language model (GC-VLM). Stage 2 introduces a continuous ActionExpert and aligns it with robot actions while stopping action-flow gradients at the VLM interface. Stage 3 enables these gradients to update the trainable VLM components jointly with the ActionExpert. Stage 4 freezes GC-VLA and applies Geometry-Conditioned Residual Flow (GCRF), using a binary intervention router and a single bounded residual velocity policy learned from closed-loop feedback. GC-VLA achieves 95.20% success on LIBERO, and GC-VLA with GCRF achieves 99.55%. Inference uses current observations and the learned GC representation without executing the offline target encoders.
Sources
- $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control
- RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
- RT-1: Robotics Transformer for Real-World Control at Scale
- ABot-M0.5: Unified Mobility-and-Manipulation World Action Model
- SRPO: Self-Referential Policy Optimization for Vision-Language-Action Models
- DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset
- Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success
- Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning
- OpenVLA: An Open-Source Vision-Language-Action Model
- SimpleVLA-RL: Scaling VLA Training via Reinforcement Learning
- PhysBrain 1.0 Technical Report
- Being-H0.5: Scaling Human-Centric Robot Learning for Cross-Embodiment Generalization
- SimVLA: A Simple VLA Baseline for Robotic Manipulation
- $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization
- HABIT: Human-Aware Behavior and Interaction Training Dataset for Robot Manipulation
- RFS: Reinforcement Learning with Residual Flow Steering for Dexterous Manipulation
- VLA-Adapter: An Effective Paradigm for Tiny-Scale Vision-Language-Action Model
- Xiaomi-Robotics-0: An Open-Sourced Vision-Language-Action Model with Real-Time Execution
- RLinf-VLA: A Unified and Efficient Framework for Reinforcement Learning of Vision-Language-Action Models
- 3D Diffusion Policy: Generalizable Visuomotor Policy Learning via Simple 3D Representations
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving