Grounded Action Model: 3D Grounding as a Foundation for Robotics
cs.RO
Submitted: 2026-09-20
Updated: 2026-09-25
Project page: https://grounded-action-model.github.io
Terminology
Sources
- MolmoAct2: Action Reasoning Models for Real-world Deployment
- World Action Models are Zero-shot Policies
- Fast-WAM: Do World Action Models Need Test-time Future Imagination?
- Unified World Models: Coupling Video and Action Diffusion for Pretraining on Large Robotic Datasets
- Unified 4D World Action Modeling from Video Priors with Asynchronous Denoising
- LIBERO-PRO: Towards Robust and Fair Evaluation of Vision-Language-Action Models Beyond Memorization
- RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation
- WildDet3D: Scaling Promptable 3D Detection in the Wild
- Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding
- Knowledge Insulating Vision-Language-Action Models: Train Fast, Run Fast, Generalize Better
- MolmoAct: Action Reasoning Models that can Reason in Space
- ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing?
- CoTracker3: Simpler and Better Point Tracking by Pseudo-Labelling Real Videos
- StarVLA: A Lego-like Codebase for Vision-Language-Action Model Developing
- Galaxea Open-World Dataset and G0 Dual-System VLA Model
- Xiaomi-Robotics-0: An Open-Sourced Vision-Language-Action Model with Real-Time Execution
- EventVLA: Event-Driven Visual Evidence Memory for Long-Horizon Vision-Language-Action Policies
- ABot-M0: VLA Foundation Model for Robotic Manipulation with Action Manifold Learning
- X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action Model
- Spatial Forcing: Implicit Spatial Representation Alignment for Vision-language-action Model
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving