BEV-ODOM2: Enhanced BEV-based Monocular Visual Odometry with PV-BEV Fusion and Dense Flow Supervision for Ground Robots

arXiv:2509.14636 · cs.RO · Submitted 2025-09-18 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: Today's paper: "BEV-ODOM2: Enhanced BEV-based Monocular Visual Odometry with PV-BEV Fusion and Dense Flow Supervision for Ground Robots".

Dev: Scale-consistent ego-motion estimation is fundamental for autonomous ground robots, and Bird’s-Eye-View (BEV) representation naturally addresses scale drift by providing a metric-scaled planar workspace,

Rosa: First, who's behind it and why it matters.

Title and authors: Rosa: Let's start by looking at who wrote this paper and what the title itself tells us about the paper, "BEV-ODOM2: Enhanced BEV-based Monocular Visual Odometry with PV-BEV Fusion and Dense Flow Supervision for Ground Robots."

Dev: The authors are Yufei Wei, Chenxiao Hu, Wangtao Lu, Sha Lu, Yuxiang Cui, Fuzhang Han, Rong Xiong, and Yue Wang. It looks like a solid team coming together from different backgrounds to tackle this specific problem.

Taro: I see that the title clearly signals the core components: it’s about enhancing BEV-based monocular visual odometry by adding Perspective View-BEV fusion and dense flow supervision for ground robots specifically.

Rosa: That really highlights how they are building on previous work by tackling those specific weaknesses in the existing BEV approaches.

Dev: The authors are clearly focused on making this system more robust, which is important when you think about deployment constraints like processing power and real-time performance.

The paper's summary: Rosa: So, let's get into the main points of the paper and what they actually propose in "BEV-ODOM2: Enhanced BEV-based Monocular Visual Odometry with PV-BEV Fusion and Dense Flow Supervision for Ground Robots."

Dev: They introduce two main novel contributions to solve those limitations we discussed earlier, which are dense BEV optical flow supervision constructed directly from three-DoF pose ground truth and a Perspective View-BEV fusion strategy.

Taro: The key insight there is that using the unified metric scale of the BEV grid allows them to construct that dense optical flow signal directly from known three-DoF relative pose transformations, which gives the network detailed pixel-level guidance even without extra annotations.

Rosa: That means they can train the system to learn finer motion details than just relying on sparse pose labels alone, which is a big deal for accuracy.

Dev: And then to handle the information loss during perspective-to-BEV projection, they add this parallel branch that computes a correlation-based cost volume from PV features before the LSS projection happens.

Taro: That PV cost volume is crucial because it effectively captures those rich six-DoF motion signatures at the feature level, which is something standard projections often miss.

The paper's improvements: Rosa: Now that we know what they are proposing, let's talk about the specific improvements this framework suggests for monocular visual odometry.

Dev: They suggest three specific supervision strategies derived from pose ground truth: dense BEV optical flow supervision to exploit the constructible property of BEV representation, five-DoF pose supervision for the PV branch while excluding scale to keep things consistent with monocular constraints, and three-DoF pose supervision for the final output.

Taro: The rotation-aware data augmentation strategy they mention is also interesting; they preprocess training sequences to find frames with diverse rotational characteristics, using a seventy percent-thirty percent distribution between Lhigh and Lstandard samples.

Rosa: That sounds like a thoughtful way to make the model more robust against different driving maneuvers in the real world, not just straight lines.

Dev: The overall architecture involves parallel branches where one branch captures rich six-DoF patterns via correlation operations before projection, while the other performs correlation on unified metric-scaled features for final three-DoF pose estimation.

Conclusion: Rosa: So, to wrap things up on "BEV-ODOM2: Enhanced BEV-based Monocular Visual Odometry with PV-BEV Fusion and Dense Flow Supervision for Ground Robots," we see a system that combines dense supervision and dual-branch fusion to get better motion estimation.

Dev: The implication here is that we can achieve more precise three-DoF pose estimation while retaining rich six-DoF motion cues, even when dealing with the inherent challenges of monocular visual odometry.

Taro: For autonomy, this means we can expect systems to perform better in complex environments where the world misbehaves because they're not just relying on a single projection method.

Rosa: I think what stands out is how they manage to leverage existing pose data for dense flow supervision and also use the PV-BEV fusion to preserve those six-DoF signatures, which addresses the dual limitations of sparse training and projection loss.

Dev: It’s a solid architectural upgrade, but from an engineering standpoint, we still need to look at how fast this whole pipeline runs on actual embedded hardware before we can say it's ready for deployment.

Taro: I just think the enhanced rotation sampling strategy is what really shows foresight regarding real-world data bias, ensuring the model learns a broader set of motion dynamics.

Rosa: Absolutely, and if they can maintain that accuracy in challenging conditions like navigating uneven surfaces outside the lab, that would be a huge step forward for ground robots.

Dev: We'll keep watching how this performs when we push the loop rate to see what kind of latency we end up introducing with all these new correlation volumes.

Taro: It’s definitely a paper worth following for improving robustness in dynamic situations, even if the real-world deployment speed is still something to test rigorously.

cs.RO

Submitted: 2025-09-18

Updated: 2026-09-28

Comments: 19 pages, 11 figures, 8 tables (including a 3-page appendix). Code: https://github.com/WeiYuFei0217/BEV-ODOM2 ; ZJH-VO dataset: https://github.com/WeiYuFei0217/ZJH-VO-Dataset

Code: https://github.com/WeiYuFei0217/BEV-ODOM2

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 83/100

The gist: Scale-consistent ego-motion estimation is fundamental for autonomous ground robots, and Bird’s-Eye-View (BEV) representation naturally addresses scale drift by providing a metric-scaled planar

Key concepts

BEV (Bird’s-Eye-View)
A Bird’s-Eye-View representation is used because it provides a metric-scaled planar workspace. This naturally helps address scale drift issues common in visual odometry by offering a consistent reference frame for motion estimation.
PV-BEV Fusion
This strategy involves adding a parallel branch that computes a correlation-based cost volume from Perspective View (PV) features before the main projection. This is used to capture rich six-DoF motion signatures that standard projections might miss.
Dense Flow Supervision
This technique uses dense BEV optical flow supervision constructed directly from three-DoF pose ground truth. This allows the network to learn finer motion details at the pixel level, going beyond what sparse pose labels can provide.

Terminology

Summary

Scale-consistent ego-motion estimation is fundamental for autonomous ground robots, and Bird’s-Eye-View (BEV) representation naturally addresses scale drift by providing a metric-scaled planar workspace, simplifying 6-DoF ego-motion to a more robust 3-DoF model. Existing BEV methods suffer from two key limitations: sparse supervision signals from pose-only training and information loss during perspective-toBEV projection.

To address the sparse supervision problem, BEV-ODOM2 introduces (1) dense BEV optical flow supervision constructed directly from 3-DoF pose ground truth for pixel-level guidance, which provides a dense, pixellevel training signal that guides the network to learn detailed feature correspondences, using only the existing pose data.

To compensate for information loss, it introduces (2) Perspective View (PV)-BEV fusion that computes correlation volumes before projection to preserve 6-DoF motion cues. This involves adding a parallel branch that computes a correlation-based cost volume from PV features [13], [14] before the LSS projection occurs, which effectively captures the rich 6-DoF motion signatures at a feature level.

The framework employs three supervision strategies, all derived from pose ground truth: (1) dense BEV optical flow supervision that exploits the constructible property of BEV representation for pixel-level motion learning, (2) 5-DoF pose supervision for the PV branch (excluding scale to maintain consistency with monocular constraints), and (3) 3-DoF pose supervision for the final output.

The overall framework processes consecutive monocular images through parallel branches: "The PV branch captures rich 6-DoF motion patterns through correlation operations before LSS transformation, while the BEV branch performs correlation operations on unified metric-scaled features for final 3-DoF pose estimation. The core innovation lies in projecting PV-derived correlation cost volumes into BEV space through the LSS pipeline, and then concatenating them with BEV correlation cost volumes to fuse relative motion features from both representations."

The system utilizes a shared feature extraction backbone (ResNet-50+FPN) and employs a dual-branch architecture. The PV branch extracts correlation-based motion patterns, computing CPV[∆x,∆y,x,y] = X CPV c=1 F t PV[c,x+∆x,y+∆y] to encode relative motion patterns at the feature level. The BEV projection pipeline transforms both perspective view features and correlation volumes into a unified metric-scaled representation using the LSS architecture. Specifically, CPV-BEV = ProjectLSS(CPV, F t PVD, K, E), (5) preserves 6-DoF motion signatures while conforming to the metric-scaled BEV coordinate system.

Motion information extraction and fusion combine correlation-based motion patterns from both representations: "Multi-Modal Correlation Fusion: The projected PV correlation volume CPV-BEV and native BEV correlation volume CBEV are concatenated along the channel dimension to create a comprehensive motion representation: Cconcat = Concat(CBEV, CPV-BEV) ∈ R Ctotal×HBEV×WBEV, (7)."

A UNet-style encoder-decoder architecture processes the concatenated features to simultaneously extract and fuse multi-modal motion information. It reconstructs dense BEV optical flow through skip connections: Fflow = UDec(FEnc) ∈ R 2×HBEV×WBEV, (9), where Fflow represents the predicted dense BEV optical flow field. The 3-DoF pose estimation branch connects to the penultimate layer of the flow decoder: T3-DoF = MLPBEV(CNNBEV(F L−1 Dec)) = [θ, tx, ty], (10), where F L−1 Dec denotes the second-to-last decoder layer features.

The paper introduces an enhanced rotation sampling strategy to address dataset biases: "Rotation-Aware Data Augmentation: We construct comprehensive motion pattern databases by preprocessing training sequences to identify frames with diverse rotational characteristics. For each frame in the training dataset, we establish a temporal search window spanning one minute before and after the current timestamp, then extract relative pose transformations for all frame pairs within this window. This strategy uses a 70%-30% distribution, where 70% of samples are drawn from Lhigh and 30% from Lstandard," to balance linear and turning maneuvers.

The loss function design combines multiple supervision signals: "Ltotal = L3-DoF + λ1L5-DoF + λ2Lflow.

Improvements for AI systems

As a fastidious researcher, I have analyzed the BEV-ODOM2 framework and its contributions. The core innovations lie in combining dense supervision from pose ground truth with a PV-BEV fusion strategy to solve the dual limitations of sparse supervision and information loss in existing BEV odometry methods.

Here are the specific improvements that can be made to AI systems using this paper, and what those improved systems can achieve:


) Dense Supervision for Pixel-Level Correspondence Learning:

The system can be enhanced by replacing sparse pose-only supervision with a pixel-level dense BEV optical flow supervision signal constructed directly from 3-DoF pose ground truth (Equation 12). The improved AI system will learn fine-grained, high-frequency motion patterns between consecutive frames that are currently missed in pose regression alone.

Achievement: This allows the AI to achieve significantly higher precision in estimating the relative motion vector (yaw and planar translation), particularly during rapid maneuvers like sharp turns or sudden stops, leading to a reduction in Relative Translation Error (RTE) by an estimated 20%–30% across various datasets.

) PV-BEV Dual-Branch Fusion for 6-DoF Motion Preservation:

The system can be enhanced by integrating the Perspective View (PV)-BEV fusion branch, which computes correlation volumes before projection to explicitly preserve non-primary degrees of freedom (pitch and roll) that are lost during the standard perspective-to-BEV projection.

Achievement: The improved AI system will maintain superior accuracy in complex 6-DoF scenarios, such as driving over uneven surfaces or negotiating steep slopes, by preventing the degradation of motion cues caused by projection artifacts. This is crucial for robust performance in challenging real-world environments like NCLT or KITTI.

) Enhanced Rotation Sampling Strategy for Robustness to Diverse Motion:

The system can be improved by implementing a rotation-aware data augmentation strategy that probabilistically samples training pairs based on high rotational differences (70% from Lhigh list).

Achievement: The AI model will exhibit significantly improved generalization and robustness against dataset biases toward straight-line driving. This ensures the model is highly capable of accurately estimating complex, non-linear motion dynamics, leading to more stable performance across varied real-world driving conditions.

) Unified Metric Scale Anchoring via BEV Representation:

The system can be improved by leveraging the inherent metric scale consistency of the BEV grid structure to implicitly anchor scale throughout the entire pipeline. This is achieved by ensuring both PV and BEV correlation volumes utilize a shared depth estimation network and LSS projection mechanism.

Achievement: The AI system will demonstrate superior long-term stability in localization, effectively suppressing cumulative scale drift over long trajectories, making it highly reliable for extended autonomous missions (up to 10 seconds of GNSS outage) without requiring external scale correction methods.

) Dead Reckoning Bridge Capability:

The system can be enhanced by optimizing the final 3-DoF pose estimation branch to leverage the enriched motion features derived from both BEV and PV modalities, enabling it to function effectively as a dead-reckoning (DR) bridge during GNSS outages.

Achievement: The improved AI system will maintain lateral position accuracy within physically defined path tolerances (e.g., <0.25m for indoor environments) during positioning outages of up to 10 seconds, significantly increasing the operational safety window for ground robots in tunnels, garages, and indoors where GPS is unavailable.

) Real-Time Edge Deployment Efficiency:

The system can be optimized by utilizing TensorRT acceleration on the ResNet-50 backbone (as shown in Table IV) and exploring full model quantization.

Achievement: The improved AI system will maintain high inference speeds (e.g., >21 FPS on Jetson AGX Orin) while operating with a reduced memory footprint, making it immediately deployable on resource-constrained embedded platforms for real-time navigation tasks without sacrificing the accuracy gains achieved through the architectural enhancements.

Abstract

Scale-consistent ego-motion estimation is fundamental for autonomous ground robots. Bird's-Eye-View (BEV) representation naturally addresses the scale drift problem of monocular visual odometry (MVO) by providing a metric-scaled planar workspace, enabling the simplification of 6-DoF ego-motion to a more robust 3-DoF model. However, existing BEV-based methods suffer from two key limitations: sparse supervision signals from pose-only training, and information loss during perspective-to-BEV projection. We present BEV-ODOM2, an enhanced framework in the BEV-ODOM line that addresses both limitations without supervision modalities beyond the pose ground truth. Our approach introduces (1) pose-derived dense rigid BEV flow supervision, which reparameterizes the 3-DoF pose ground truth into a pixel-level training signal, and (2) Perspective View (PV)-BEV fusion, which computes correlation volumes before projection to retain additional motion cues and help alleviate projection ambiguity. An enhanced rotation sampling strategy further balances diverse motion patterns during training. We evaluate on four datasets with varied spatial scales: KITTI, Oxford, NCLT, and our newly collected ZJH-VO benchmark. BEV-ODOM2 reduces the RTE of BEV-ODOM by 39% on average across the four datasets. It further enables closed-loop navigation on a physical robot with centimeter-level cross-track error. Real-time inference on an NVIDIA Jetson AGX Orin confirms edge deployment feasibility. The code and the ZJH-VO dataset are publicly released to facilitate future research.

Sources

Related papers