See like a Robot: Robot-Centric Pointmaps for VLA Models
cs.RO, cs.AI
Submitted: 2026-07-13
Updated: 2026-09-21
Comments: Project page: https://davian-robotics.github.io/pointmap/
Project page: https://davian-robotics.github.io/pointmap
License: http://creativecommons.org/licenses/by/4.0/
The gist: Vision-language-action (VLA) models require 3D spatial reasoning, yet RGB observations encode robot-object geometry only implicitly.
Terminology
Abstract
Vision-language-action (VLA) models require 3D spatial reasoning, yet RGB observations encode robot-object geometry only implicitly. Lifting depth with camera intrinsics makes this geometry explicit as dense, image-aligned pointmaps, but their camera-frame coordinates depend on camera placement. We propose SeeR-VLA, which transforms pointmaps into a robot-centric frame with an end-effector origin and robot-base-aligned axes. An encoder initialized from pretrained RGB weights extracts pointmap features, which are added to corresponding RGB tokens without increasing the token count. Across 24 RoboCasa tasks and four real-world tasks, SeeR-VLA improves average success over RGB-only π 0.5 by 6.4 and 32.5 percentage points, respectively. It also exceeds the strongest evaluated 3D-augmented baseline, PointVLA, by 3.5 and 23.7 percentage points, respectively. Beyond these gains, our ablations clarify how coordinate choices affect VLA performance, showing that end-effector centering is most effective with robot-base-aligned axes. The benefits grow as training viewpoints diversify, highlighting the importance of using robot-frame pointmaps when learning from diverse camera configurations.
Sources
- $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control
- GR00T N1: An Open Foundation Model for Generalist Humanoid Robots
- DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset
- LIBERO-Plus: In-depth Robustness Analysis of Vision-Language-Action Models
- Colosseum V2: Benchmarking Generalization for Vision-Language-Action Models
- SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model
- RoboUniView: Visual-Language Model with Unified View Representation for Robotic Manipulation
- GeoVLA: Empowering 3D Representations in Vision-Language-Action Models
- 3D CAVLA: Leveraging Depth and 3D Context to Generalize Vision Language Action Models for Unseen Tasks
- DepthVLA: Enhancing Vision-Language-Action Models with Depth-Aware Spatial Reasoning
- Any3D-VLA: Enhancing VLA Robustness via Diverse Point Clouds
- PointACT: Vision-Language-Action Models with Multi-Scale Point-Action Interaction
- OG-VLA: Orthographic Image Generation for 3D-Aware Vision-Language Action Model
- Evo-0: Vision-Language-Action Model with Implicit Spatial Understanding
- PaliGemma: A versatile 3B VLM for transfer
- Canonical Policy: Learning Canonical 3D Representation for SE(3)-Equivariant Policy
- SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics
- FP3: A 3D Foundation Policy for Robotic Manipulation
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving