UrbanVLA: A Vision-Language-Action Model for Urban Micromobility
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.
Rosa: Today's paper: "UrbanVLA: A Vision-Language-Action Model for Urban Micromobility".
Dev: UrbanVLA introduces a route-conditioned Vision-Language-Action (VLA) framework designed for scalable urban navigation by explicitly aligning noisy route waypoints with visual observations during execution and subsequently planning trajectories to drive the…
Rosa: First, who's behind it and why it matters.
Title and authors: Dev: Now let's look at who did this research. The authors listed are Anqi Li, Zhiyong Wang, Jiazhao Zhang, Minghan Li, Yunpeng Qi, and Zhibo Chen, along with He Wang. It shows a solid team from different backgrounds contributing to this work.
Rosa: That’s right; it’s a collaborative effort involving people from Peking University and USTC across the project. The authors are clearly tackling this problem with diverse expertise in mind.
Taro: It’s interesting seeing that mix of expertise, especially since they're combining vision, language, and action planning into one unified framework. That level of integration is what I think makes this paper stand out in the autonomy research area.
Dev: From a control engineering viewpoint, having people from different institutions working together on such a complex system suggests they’re looking at the problem from multiple angles—perception, language grounding, and execution control.
Rosa: And that’s what I find compelling; it shows they're not just tweaking one part of the stack but are building something holistic to handle the messy reality of urban navigation.
Taro: I think that holistic approach is what allows them to move beyond just short-term obstacle avoidance and into more complex, rule-based navigation required for city life.
The paper's summary: Rosa: To go deeper into the UrbanVLA: A Vision-Language-Action Model for Urban Micromobility paper, the core idea is that they are proposing a VLA model that takes structured route descriptions as input and directly predicts trajectory waypoints for route following.
Dev: So, in essence, it's not just interpreting a map; it's taking those abstract instructions—the roadbooks—and using visual observations to ground them into actual movement commands.
Taro: That ability to align those noisy navigation tools with real-time visual cues is the central innovation they are highlighting because existing VLAs often fail when the route itself is inaccurate or dynamic.
Rosa: They achieve this by integrating high-level guidance from navigation tools with on-board vision and then learning to plan trajectories that follow those instructions, which allows for reliable, long-horizon navigation over large areas.
Dev: That means the system learns how to interpret those 'turn right in thirty meters' instructions by looking at what’s happening visually right now and adjusting the path accordingly.
Taro: And they address the complexity of real urban environments by specifically mentioning that VLAs need to adhere to a complex set of rules, like traffic signals and sidewalk etiquette, while also adapting to dynamic obstacles in real time.
The paper's improvements: Rosa: When we look at how they improve the system, they introduce a dual-stage training pipeline starting with Supervised Fine-Tuning using simulated environments and web videos, followed by Reinforcement Fine-Tuning on a mixture of simulation and real-world data.
Dev: That two-stage approach is smart because it allows them to first learn the basic navigation skills in a controlled setting before pushing it toward the complexity of real-world scenarios.
Taro: The SFT stage uses Heuristic Trajectory Lifting, or HTL, which is a heuristic algorithm designed to lift high-level route information from raw trajectory data by denoising and removing low-quality paths.
Rosa: That HTL process seems crucial because it helps generate an abstracted route R for training via a Mean Squared Error loss, which cleans up the input data before the model gets trained on it.
Dev: Then they move into the RFT stage using Implicit Q-Learning, or IQL, where the task is formulated as a Partially Observable Markov Decision Process to learn from expert demonstrations in both simulated and real environments.
Conclusion: Rosa: So, to wrap up our discussion on UrbanVLA: A Vision-Language-Action Model for Urban Micromobility, the paper demonstrates how integrating route conditioning with a two-stage training pipeline allows for reliable long-horizon navigation in complex city settings.
Dev: I think the main implication is that we can make robots much more capable of handling the messy, unstructured nature of real urban areas because they can dynamically align abstract instructions with visual reality.
Taro: What really stands out to me is the robustness gained from that refinement process; it shows how essential it is to have both supervised learning and reinforcement learning working together for true adaptability.
Rosa: Exactly. The paper shows that by focusing on route-visual alignment, we can leverage existing navigation tools much more effectively, even when those tools provide noisy data.
Dev: I just hope they keep pushing the loop rate and latency down during deployment; if the planning takes too long, those real-time adjustments won't work as well as they do in simulation.
Taro: And what about the future? I think this work opens up possibilities for systems that can handle much more nuanced social navigation tasks, which is where we need to go next.
Rosa: It’s been a really informative session discussing the UrbanVLA: A Vision-Language-Action Model for Urban Micromobility paper and how it sets a new direction for urban mobility research.
Peking University · Galbot Institute of Technology (USTC) · Beijing Academy of Artificial Intelligence (BAAI)
cs.RO, cs.AI, cs.CV
Submitted: 2025-10-27
Updated: 2026-10-01
Project page: https://pku-epic.github.io/UrbanVLA-Web/End
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 81/100
The gist: UrbanVLA introduces a route-conditioned Vision-Language-Action (VLA) framework designed for scalable urban navigation by explicitly aligning noisy route waypoints with visual observations during
Key concepts
- Route Conditioned VLA
- This framework takes structured route descriptions, including waypoints and turn instructions, as input. It directly predicts the necessary trajectory waypoints required to follow the route while simultaneously processing real-time visual data from the robot's sensors.
- Heuristic Trajectory Lifting (HTL)
- HTL is a technique used during Supervised Fine-Tuning to create an abstracted route representation. It cleans raw trajectory data by denoising paths, removing poor segments, and identifying key turning points to generate a coarser, more reliable route for training.
- Reinforcement Fine-Tuning (RFT)
- RFT refines the model using expert demonstrations from simulation and real-world data. It treats navigation as a Partially Observable Markov Decision Process (POMDP), optimizing actions based on a reward function that balances path completion, collision avoidance, and deviation minimization.
- Visual Token Concatenation
- This involves using two separate pre-trained vision encoders (DINOv2 and SigLIP) to process visual observations. The resulting features from these encoders are combined (concatenated) to create a rich set of visual tokens that are then fed into the Large Language Model alongside language instructions.
Terminology
Summary
UrbanVLA introduces a route-conditioned Vision-Language-Action (VLA) framework designed for scalable urban navigation by explicitly aligning noisy route waypoints with visual observations during execution and subsequently planning trajectories to drive the robot. This method addresses the challenge of long-horizon navigation in dynamic, unstructured city environments by integrating high-level guidance from navigation tools with on-board vision and learning a dual-stage training pipeline involving supervised fine-tuning (SFT) and reinforcement fine-tuning (RFT).
The gist
UrbanVLA proposes a routeconditioned VisionLanguageAction VLA framework that takes structured route descriptions as input and directly predicts trajectory waypoints for route following.
How it works
The framework operates through a two-stage training pipeline to master both low-level capabilities like point-goal reaching and high-level capabilities like route–visual alignment. The process begins with Supervised FineTuning (SFT) using simulated environments and trajectories parsed from web videos, followed by Reinforcement Fine-Tuning (RFT) on a mixture of simulation and real-world data to enhance safety and adaptability.
-
The input to the VLA model is structured route descriptions, which are converted into a structured linguistic representation comprising two components: a set of waypoints sampled from the high-level route, and distance and direction instructions for transitions between blocks (e.g., ‘turn right in 30 meters’).
-
Visual observations are processed using two pre-trained vision encoders (DINOv2 and SigLIP), whose features are concatenated to form visual tokens, which are then projected into the Large Language Model (LLM) backbone alongside the language instruction tokens.
-
The LLM backbone (Qwen2) generates action tokens; for navigation tasks, an action token is captured at the current time step and decoded through an MLP-based action model to obtain the navigation trajectory, denoted as a sequence of waypoints, via Equation (1).
Training Strategy
The training strategy employs two distinct stages:
(i) Supervised Fine-Tuning (SFT):
The SFT stage is applied to the base model NavFoM [19] using urban navigation demonstrations generated by a PPO expert in simulation and web-scale urban travel data. To address the challenge of obtaining direct navigation ‘roadbooks’ from raw trajectories, the authors introduce Heuristic Trajectory Lifting (HTL). HTL is a heuristic algorithm that lifts high-level route information from raw trajectory data by denoising, removing low-quality paths, detecting significant turning points to form coarse waypoints, and perturbing segments with Gaussian positional noise to capture the ambiguity of real-world navigation. This process results in an abstracted route R used for training via a Mean Squared Error (MSE) loss.
(ii) Reinforcement Fine-Tuning (RFT):
The RFT stage refines the model on a hybrid dataset combining expert demonstrations from simulated and real environments using Implicit Q-Learning (IQL) [27]. The task is formulated as a Partially Observable Markov Decision Process (POMDP), where the state s is constructed from the hidden representation of the LLM backbone, H(n)T. The reward function r(s, a) is designed to consider both trajectory efficiency and navigation safety:
(3)
r(s, a) = λcomp lcompletion − λcoll 1collision − λdev 1deviation.
The IQL algorithm learns the value function Vψ(s) and Q-function Qθ(s, a) from this offline dataset to update the policy π(s) via an advantage-weighted regression (AWR) objective (Equation 2). The reward weights are set to λcomp=0.5, λcoll=1, and λdev=1 for real-world data.
UrbanVLA Architecture
The architecture leverages a pre-trained navigation foundation model NavFoM [19] as the base model. The VLA model forward process involves:
(i) High-Level Route Encoding:
Route instructions are converted into a structured linguistic representation, including waypoints and distance/direction cues for transitions. This involves resampling upcoming route segments and applying corner detection algorithms to segment the route into blocks, deriving block-level distance and direction cues.
(ii) VLA Model Forwarding:
Given multi-view RGB observations Ovis, a visual sliding window is applied to retain the nearest k frames. Visual information is encoded using two pre-trained vision encoders (DINOv2 and SigLIP), features are concatenated, downsampled via grid pooling, and projected into the LLM embedding space to obtain visual tokens E1:C1:T. These tokens are fed into the LLM backbone alongside language tokens EL.
Improvements for AI systems
Based on the provided scientific paper UrbanVLA: A Vision-Language-Action Model for Urban Micromobility,
here are specific, actionable improvements for AI systems and a description of what those improved systems can achieve.
) Improved AI System Capabilities
The core improvement lies in developing an end-to-end, route-conditioned VLA framework capable of robust long-horizon navigation in complex, unstructured urban environments by effectively bridging the gap between high-level linguistic instructions and low-level physical actions.
Here are the specific improvements derived from the UrbanVLA methodology:
-
A foundation model that integrates visual perception, language understanding, and action planning (a
NavFoM
). -
A two-stage training pipeline combining Supervised Fine-Tuning (SFT) on simulation/web data and Reinforcement Fine-Tuning (RFT) using a sim-real aggregated dataset with Implicit Q-Learning (IQL).
-
A novel Route Lifting algorithm (HTL) to generate large, diverse, and noisy route conditions from raw trajectories.
-
A unified state representation for the LLM backbone that combines language instructions and visual observations to enable cross-modal reasoning before action decoding.
-
Specific system capabilities enabled by these improvements:
The improved AI system can perform the following high-level tasks with unprecedented reliability in real-world urban settings:
-
Reliable Long-Horizon Route Following (500m+): The system can successfully follow complex, multi-segment route instructions derived from navigation apps, even when the provided route waypoints are noisy or geometrically inaccurate, by aligning them dynamically with real-time visual observations.
-
Robust Generalization to Unseen Environments: By leveraging pre-trained foundation models and training on diverse web videos and simulation data (MetaUrban), the system demonstrates superior performance (e.g., 94% SR in PointNav) in novel urban layouts, lighting conditions, and dynamic obstacle scenarios that it has never encountered during training.
-
High Social Compliance: The model can adhere to complex social navigation norms—such as maintaining appropriate distances from pedestrians and yielding at intersections—even when relying solely on RGB visual inputs, significantly exceeding performance of LiDAR-based baselines (e.g., achieving a high SNS score).
-
Safety-Critical Obstacle Avoidance: The system is enhanced with IQL-based reinforcement fine-tuning, allowing it to learn
safety-aware
decision-making that explicitly prioritizes collision avoidance and pedestrian interaction, making it highly effective in dynamic urban flows. -
Simulated Data Efficiency (Sim-to-Real Transfer): The two-stage training pipeline ensures the model effectively transfers its learned policies from simulation to the real world, minimizing the need for extensive real-world data collection while maintaining high performance.
Sources
- From Seeing to Experiencing: Scaling Navigation Foundation Models with Reinforcement Learning
- StreamVLN: Streaming Vision-and-Language Navigation via SlowFast Context Modeling
- Sekai: A Video Dataset towards World Exploration
- LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding
- LOVON: Legged Open-Vocabulary Object Navigator
- NavDP: Learning Sim-to-Real Navigation Diffusion Policy with Privileged Information Guidance
- Qwen2 Technical Report
- Qwen2.5-VL Technical Report
- Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
- ConRFT: A Reinforced Fine-tuning Method for VLA Models via Consistency Policy
- What Can RL Bring to VLA Generalization? An Empirical Study
- CO-RFT: Efficient Fine-Tuning of Vision-Language-Action Models through Chunked Offline Reinforcement Learning
- Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models
- VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning
- OctoNav: Towards Generalist Embodied Navigation
- LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models
- OpenVLA: An Open-Source Vision-Language-Action Model
- DINOv2: Learning Robust Visual Features without Supervision
- On Evaluation of Embodied Navigation Agents
- Retrospectives on the Embodied AI Workshop
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving