GALA: Geometry-Aware Latent Action Modeling for Vision-Language-Action Model Pretraining across Embodiments
cs.RO, cs.CV
Submitted: 2026-09-18
Updated: 2026-09-18
Project page: https://puzhenyuan.github.io/GALA-website
Terminology
Sources
- OpenVLA: An Open-Source Vision-Language-Action Model
- Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success
- GR00T N1: An Open Foundation Model for Generalist Humanoid Robots
- $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control
- $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization
- UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent
- Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations
- METIS: Multi-Source Egocentric Training for Integrated Dexterous Vision-Language-Action Model
- Latent Action Diffusion for Cross-Embodiment Manipulation
- One-Policy-Fits-All: Geometry-Aware Action Latents for Cross-Embodiment Manipulation
- Cross-Hand Latent Representation for Vision-Language-Action Models
- UniVLA: Learning to Act Anywhere with Task-centric Latent Actions
- HARP-VLA: Human-Robot Aligned Representation Learning for Vision-Language-Action Model
- ConLA: Contrastive Latent Action Learning from Human Videos for Robotic Manipulation
- RoboCasa: Large-Scale Simulation of Everyday Tasks for Generalist Robots
- DexVLA: Vision-Language Model with Plug-In Diffusion Expert for General Robot Control
- UniSkill: Imitating Human Videos via Cross-Embodiment Skill Representations
- FLARE: Robot Learning with Implicit World Modeling
- DiT4DiT: Jointly Modeling Video Dynamics and Actions for Generalizable Robot Control
- JoyAI-RA 0.1: A Foundation Model for Robotic Autonomy
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving