VersaCamVLA: Camera-Configurable VLA Policies for Robotic Manipulation
cs.CV, cs.RO
Submitted: 2026-10-08
Updated: 2026-10-08
Project page: https://boyaohan.github.io/VersaCamVLA.github.io/Abstract
Terminology
Sources
- $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization
- $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control
- Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success
- RT-1: Robotics Transformer for Real-World Control at Scale
- GR00T N1: An Open Foundation Model for Generalist Humanoid Robots
- Do You Know Where Your Camera Is? View-Invariant Policy Learning with Camera Conditioning
- VLA-LPAF: Lightweight Perspective-Adaptive Fusion for Vision-Language-Action to Enable More Unconstrained Robotic Manipulation
- DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset
- Multi-Camera View Scaling for Data-Efficient Robot Imitation Learning
- NeRF in the Palm of Your Hand: Corrective Augmentation for Robotics via Novel-View Synthesis
- AnyCamVLA: Zero-Shot Camera Adaptation for Viewpoint Robust Vision-Language-Action Models
- 3D Diffusion Policy: Generalizable Visuomotor Policy Learning via Simple 3D Representations
- ActiveVLA: Injecting Active Perception into Vision-Language-Action Models for Precise 3D Robotic Manipulation
- GeoPredict: Leveraging Predictive Kinematics and 3D Gaussian Geometry for Precise VLA Manipulation
- Adapt3R: Adaptive 3D Scene Representation for Domain Transfer in Imitation Learning
- GeoAware-VLA: Implicit Geometry Aware Vision-Language-Action Model
- Spatial Forcing: Implicit Spatial Representation Alignment for Vision-language-action Model
- Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware
- SEM: Enhancing Spatial Understanding for Robust Robot Manipulation
- PEAfowl: Perception-Enhanced Multi-View Vision-Language-Action for Bimanual Manipulation
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models