MAVP: Map-Aware Visuomotor Policies for Mobile Manipulation

arXiv:2609.26378 · cs.RO, cs.AI, cs.CV, cs.LG · Submitted 2026-09-22 · Read on arXiv

cs.RO, cs.AI, cs.CV, cs.LG

Submitted: 2026-09-22

Updated: 2026-09-22

Project page: https://123qwedsa123.github.io/mavp

License: http://creativecommons.org/licenses/by/4.0/

The gist: Successful mobile manipulation requires coordinated base and arm motion while maintaining accurate spatial positioning.

Terminology

Abstract

Successful mobile manipulation requires coordinated base and arm motion while maintaining accurate spatial positioning. However, demonstration-trained policies can struggle to realise the intended base motion reliably, leading to spatial misalignment and subsequent manipulation failures. We present MAVP (Map-Aware Visuomotor Policies), a framework that improves execution reliability by predicting explicit base-pose targets and tracking them using localisation feedback. MAVP reconstructs a static map from teleoperated demonstrations and expresses demonstrated base trajectories in a shared map frame, providing consistent spatial supervision across demonstrations. At execution time, the policy receives RGB observations, joint states, and the robot's current map-frame base pose, and jointly predicts target base poses, arm actions, and gripper actions. A low-level controller tracks the predicted base targets using feedforward motion and pose error feedback, enabling correction of execution deviations. We additionally use pose-noise augmentation during training to improve robustness to errors in the policy's pose input. Across six real-world manipulation tasks and three policy families, MAVP achieves higher task success rates than unanchored velocity control in all tasks. Videos and additional results are available at https://123qwedsa123.github.io/mavp/.

Sources

Related papers