Cross-View Sequential Visual Localization with Spatio-Temporal Context Modeling for Autonomous Driving

arXiv:2608.10660 · cs.CV, cs.AI · Submitted 2026-08-11 · Read on arXiv

Jiaping Wang, Shaobo Li, Zhen Wang

School of Information Engineering, Chang'an University

cs.CV, cs.AI

Submitted: 2026-08-11

Updated: 2026-08-12

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 75/100

The gist: This paper proposes a temporal-context-enhanced framework for cross-view sequential visual localization in autonomous driving.

Terminology

Summary

This paper proposes a temporal-context-enhanced framework for cross-view sequential visual localization in autonomous driving. The framework takes one satellite map and a continuous ground-image sequence as input, and introduces directional temporal propagation into the cross-view matching pipeline. Each current-frame feature adaptively aggregates historical scene cues from the previous recurrent state before candidate-region discrimination. The enhanced sequence features are then fed into candidate-region discrimination and local offset estimation modules, enabling fine-grained localization of consecutive frames on the satellite map.

The main contributions are as follows:

(1) This study proposes a cross-view sequence localization framework that integrates hierarchical feature representation, temporal context enhancement, and two-stage localization. The framework uses temporally enhanced coarse features for satellite candidate-region discrimination and fine-grained features containing local structure and texture cues for candidate-conditioned offset estimation, forming a continuous localization pipeline from coarse region localization to precise position prediction.

(2) This study designs a temporal context enhancement module for continuous time windows. At each timestamp, the current-frame feature retrieves historical information from the previous recurrent state through cross-frame attention, while a residual update preserves the current observation as the main representation. The recurrent structure progressively propagates historical context along the sequence, improving the discriminability of ground-view features.

(3) Extensive experiments on CVIS and KITTI-CVL demonstrate the effectiveness and generalization capability of the proposed method. On CVIS, the proposed method reduces mean distance error from 3.80 m to 1.57 m and improves R@1 m by 32.08 percentage points over SOTA. On KITTI-CVL, direct transfer reduces the mean error from 3.57 m to 2.61 m, while target-domain fine-tuning further reduces it to 2.27 m and improves R@1 m to 35.69%.

(4) A real-vehicle field experiment covers nine representative urban driving scenarios. The model uses low-accuracy GPS only to provide a coarse satellite-map prior and requires no additional training or fine-tuning. It achieves a mean error of 2.84 m, a median error of 2.92 m, and an R@5 m of 96.86%, while maintaining effective localization across the evaluated scenarios. In high-rise and occluded areas where RTK measurements drift, the proposed method still produces stable, road-aligned localization results.

The methodology consists of three main components. First, the local feature extraction module uses pretrained DINOv2 to extract coarse- and fine-grained representations from the satellite image and the ground-image sequence within a continuous time window. Deep DINOv2 layers mainly encode abstract global semantics, whereas intermediate layers preserve stronger spatial layout, texture, and structural cues. The last backbone layer produces coarse-grained semantic features, while features from the 5th, 6th, 7th, and 8th intermediate layers are concatenated as fine-grained representations. Second, the proposed temporal context enhancement module models historical context in coarse ground-sequence features, allowing each current frame to preserve its spatial layout while adaptively aggregating informative contextual cues from the historical state. The recurrent state is updated using an asymmetric Query-Key/Value design where the current frame determines what information to retrieve, whereas the previous recurrent state provides retrievable historical context. The current frame is then fused with the historical context through a residual update that preserves the current frame as the main representation and uses historical context as complementary context. Finally, the two-stage localization module performs feature fusion and mask-guided cascaded regression across the entire temporal clip, predicting continuous vehicle trajectory coordinates. Stage 1 uses coarse temporal features to select global candidate regions, identifying satellite grids likely to contain the vehicle. Stage 2 performs candidate-conditioned offset refinement, in which the candidate mask from Stage 1 guides fine-grained feature interactions and the corresponding local offsets refine the selected coarse grids.

The two localization stages are jointly trained using coarse-grid classification supervision and candidate-conditioned offset regression supervision. The total objective aggregates classification and regression losses over all timestamps. The coarse classification loss uses cross-entropy loss to supervise the sequence-level prediction and maximize the probability of the ground-truth grid. The masked regression loss is defined on the ground-truth grid, and if the ground-truth grid is included in the candidate mask, the predicted offset at that grid is used to compute the regression loss; otherwise, the regression loss of that sample is ignored.

Experiments were conducted on the CVIS and KITTI-CVL datasets. On CVIS, the proposed method achieves a mean distance error of 1.57 m and a median distance error of 1.21 m on the test set, reducing the errors of the strongest baseline, TACV, by 2.23 m and 0.71 m, respectively. The proposed method achieves R@1 m of 40.22%, R@2 m of 77.51%, and R@5 m of 98.99% on the test set. On KITTI-CVL, the full model achieves a mean error of 2.61 m and a median error of 1.97 m on the test set, reducing the corresponding errors of TACV by 0.96 m and 0.78 m. After fine-tuning on KITTI-CVL, the model further reduces the test-set mean error to 2.27 m and improves R@1 m to 35.69%.

Ablation studies on CVIS show that using only final-layer DINOv2 features gives a test-set mean error of 12.25 m, indicating that a single semantic feature level cannot simultaneously support coarse candidate-region recognition and fine offset estimation. Introducing hierarchical features reduces the mean error from 12.25 m to 5.92 m and improves R@1 m from 6.37% to 24.44%. Adding temporal context enhancement reduces the mean error to 1.57 m and increases R@1 m and R@5 m to 40.22% and 98.99%, respectively. The full model achieves higher candidate coverage for all K values compared with the model without temporal context. Under the top-64 setting, the full model uses only 17.73% of the 19×19 search space while covering 99.98% of ground-truth grids.

In real-vehicle field experiments across nine typical scenarios, the proposed method achieves a mean error of 2.84 m, a median error of 2.92 m, and R@5 m of 96.86% across all scenarios. The best performance is achieved in the internal-road scenario, with a mean error of 1.38 m, a median error of 1.12 m, R@1 m of 43.20%, and R@5 m of 100.00%. In the urban arterial roads and side roads under viaducts scenarios, the model yields mean errors of 1.78 m and 2.95 m, respectively. In contrast, the model produces relatively greater localization errors in intersections (mean error of 3.02 m) and elevated roads (mean error of 3.67 m). In some severe scenarios where RTK GPS is unable to provide correct localization results due to obstruction by buildings and trees, the visual localization model is still able to provide precise localization results.

The conclusion states that the proposed temporal-context-enhanced cross-view localization framework provides an effective, generalizable, and practical solution for continuous visual localization. Future work includes integrating multimodal sensor cues such as LiDAR, IMU, or HD map priors to improve robustness under heavy occlusion, severe illumination variation, and dynamic traffic; incorporating motion constraints or trajectory-smoothing mechanisms into future temporal modeling to further improve long-horizon trajectory-level consistency; and developing more efficient temporal feature modeling and lightweight deployment strategies to reduce the additional cost introduced by continuous attention.

Improvements for AI systems

Improvements to AI Systems:

  1. Hierarchical Feature Fusion for Vision Models: Integrate multi-scale feature extraction (e.g., combining intermediate-layer spatial/texture cues with final-layer semantic cues) into any visual backbone. This enables simultaneous coarse region recognition and fine-grained offset regression, reducing localization error from 12.25 m to 5.92 m in the paper’s ablation.

  2. Recurrent Temporal Context Enhancement Module: Add a lightweight recurrent state that uses cross-frame attention (query from current frame, key/value from historical state) plus a residual update to preserve current observations. This improves sequential perception tasks (e.g., video object tracking, SLAM) by adaptively aggregating historical cues without losing current-frame fidelity, cutting mean error from 5.92 m to 1.57 m.

  3. Two-Stage Coarse-to-Fine Localization Pipeline: Implement a mask-guided cascaded regression: Stage 1 selects candidate regions via coarse semantic features; Stage 2 uses those masks to condition fine-grained offset prediction. This reduces search space (only 17.73% of grid) while achieving 99.98% ground-truth coverage, enabling efficient and accurate continuous positioning.

  4. Asymmetric Query-Key/Value Attention for Recurrent States: Use a design where the current frame controls what to retrieve from history, while the historical state only provides context. This prevents over-reliance on past observations and improves robustness in dynamic scenes, as shown by stable performance under GPS drift and occlusion.

  5. Joint Training with Masked Regression Loss: Train classification and regression losses jointly, but ignore regression loss for samples where the ground-truth grid is not in the candidate mask. This prevents noisy gradients from incorrect coarse predictions, improving convergence and final accuracy (R@1 m improved to 40.22% on CVIS).

  6. Zero-Shot Transfer with Pretrained Features: Use DINOv2 features (frozen or fine-tuned) to enable direct cross-dataset generalization without retraining. The paper shows direct transfer reduces mean error from 3.57 m to 2.61 m on KITTI-CVL, and fine-tuning further improves it—useful for adapting to new environments with minimal data.

What the Improved AI System Can Do:

  • Autonomous Driving Localization: Continuously localize a vehicle on a satellite map using only a ground camera sequence and low-accuracy GPS, achieving 1.5–2.8 m mean error in urban, highway, and occluded scenarios, even where RTK GPS fails.

  • Robust Visual Odometry/Tracking: Maintain accurate position estimates over time by fusing hierarchical visual features with recurrent temporal context, reducing drift and improving performance under changing lighting, shadows, and dynamic objects.

  • Efficient Search and Retrieval: For any image-to-map or image-to-database matching task, use coarse-to-fine candidate selection to reduce computational cost (e.g., only 17.73% of search space) while maintaining near-perfect recall.

  • Cross-Domain Adaptation: Deploy the system in new cities or conditions without retraining, leveraging pretrained features and temporal modeling to generalize from one dataset to another.

  • Real-Time Deployment: Use the lightweight recurrent module and two-stage pipeline to run on embedded systems for real-time vehicle localization, with minimal additional latency from continuous attention.

  • Failure Recovery: In GPS-denied or multipath environments (e.g., high-rise areas, tunnels), the system provides stable, road-aligned localization, serving as a reliable fallback for navigation and safety-critical decisions.

Sources

Related papers