Road Maps as Free Geometric Priors: Weather-Invariant Drone Geo-Localization with GeoFuse

arXiv:2605.14925 · cs.CV, cs.LG · Submitted 2026-05-14 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "Road Maps as Free Geometric Priors".

Tom: The gist: GeoFuse introduces a cross-modal fusion framework that integrates precisely aligned road map tiles with satellite imagery to yield more discriminative and weather-resilient representations for drone geo-localization.

Jane: First, who's behind it and why it matters.

Title and authors: Tom: So Jane, so this paper is called "Road Maps as Free Geometric Priors: Weather-Invariant Drone Geo-Localization with GeoFuse." I mean, just hearing those words makes me think about how much we rely on perfect conditions for drones to figure out where they are.

Jane: Yeah, it sounds like they're proposing a way to make drone localization work even when the weather is terrible. They're talking about using road maps as these free geometric priors to help the system stay accurate regardless of fog or rain.

Lu: It’s interesting because it moves away from just trying to fix the image noise itself, and instead introduces a completely different kind of information source for matching drone views with satellite images.

Meng: So, like they're not just using more data augmentation to make the drone picture look better in the rain, but they're using something external that is supposed to be stable even when things look bad?

Lalam: It sounds like a way to anchor the location based on something structural rather than purely visual features that get messed up by weather.

The paper's summary: Tom: Right, so they are introducing this framework called GeoFuse. The core idea is integrating road map tiles directly with satellite imagery to create representations that are more discriminative and weather-resilient. It’s about combining the visual data with this external geometric information.

Jane: What I like is how they handle the fusion of those two very different types of data—the image and the map—by looking at them at both token and channel levels. That suggests a really deep interaction between the features, not just simple overlaying.

Lu: They use a flexible fusion module controlled by this lightweight dynamic gating mechanism that adaptively weights how much they trust each modality for every single input instance. That’s smart because you don't want one source to completely dominate the other constantly.

Meng: So, if it’s raining really hard, does the system automatically lean more on the road map cues because it knows those are more stable?

Lalam: Yes, that's exactly what that adaptive weighting is supposed to do; it lets the system dynamically decide which modality contributes most to the final location guess for that specific drone image.

The paper's improvements: Tom: Okay, so what are the actual technical improvements they claim? They focus on this dual-level feature fusion—token-level and channel-level—to get a dual-level fused feature. That means they capture both the fine spatial details and how different parts of the features relate to each other.

Jane: And they pair that up with class-level crossview contrastive learning, which is supposed to encourage drone features to line up consistently with those fused satellite and road map representations, even when things are degraded.

Lu: The paper claims this combination allows them to produce more discriminative and robust representations across varying weather conditions. They report pretty solid numbers on the benchmarks too; they got about a three point four six percent gain in Recall@one on University-one thousand six hundred fifty-two and a twenty-three point one eight percent gain on DenseUAV when testing under different weather conditions <ref:2605.14925#pg3>.

Meng: Those percentage gains sound significant, especially when you look at how much better they perform in those challenging scenarios compared to the existing state-of-the-art methods they are comparing it against.

Lalam: So, the main improvement is that by using road maps as these geometric priors, they're getting a stable localization result that stays high even when the visual input is noisy or partially obscured by weather.

Conclusion: Tom: So, to wrap up on this paper, GeoFuse uses those precisely aligned road map tiles and its adaptive fusion module to create representations that are much better at handling weather noise. It’s about using the road map as a reliable geometric backbone for drone geo-localization.

Jane: It really emphasizes how combining structural knowledge from maps with visual data helps keep the localization accurate across different weather situations, which is something we all face in real-world applications.

Lu: The main contribution here is showing that this approach consistently outperforms current methods on benchmarks like University-one thousand six hundred fifty-two and DenseUAV when those weather conditions are included in the test set <ref:2605.14925#pg3>.

Meng: From an engineering side, it’s interesting because it relies on having those road maps precisely aligned with the satellite data, which can be a tricky setup to get right in really new or complex areas.

Lalam: And they also point out a limitation: this method still depends on the availability and precise spatial alignment of those road maps, which can be an issue in remote or newly developed areas where map data isn't perfect.

Tom: So, the big picture is that this work shows a promising path toward reliable drone geo-localization by using readily available road maps as a weather-invariant geometric prior. We’ll talk about what else is happening in this field next.

University of Macau

cs.CV, cs.LG

Submitted: 2026-05-14

Updated: 2026-10-08

Comments: 18 pages, 4 figures

Code: https://github.com/YsongF/GeoFuse

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 88/100

The gist: The gist: GeoFuse introduces a cross-modal fusion framework that integrates precisely aligned road map tiles with satellite imagery to yield more discriminative and weather-resilient representations

Key concepts

Cross-Modal Fusion Framework
This is the core technique where the system combines information from two different types of data—in this case, satellite images and road maps—to create a richer representation. GeoFuse specifically fuses these modalities using attention mechanisms to ensure that spatial geometric cues from the road map are effectively integrated with visual data.
Token-Level Fusion
This fusion method focuses on matching specific elements (tokens) between the two inputs. In GeoFuse, it uses cross-attention where the road map acts as a key and value to update the tokens derived from the satellite image. This helps model precise spatial correspondences between features in both modalities.
Weather-Invariant Geometric Prior
Road maps serve as a stable, geometric reference that is not easily affected by visual noise or weather changes like fog or rain. By incorporating these road map features into the drone localization process, the framework gains a reliable structural constraint that remains consistent even when the satellite imagery becomes degraded.
Dynamic Gating Mechanism
This lightweight mechanism controls how much influence each modality (satellite vs. road map) has on the final fused representation for every specific instance. It adaptively weights the contributions of each input based on the context, allowing the model to prioritize reliable information when visual conditions are poor.

Terminology

Summary

The gist: GeoFuse introduces a cross-modal fusion framework that integrates precisely aligned road map tiles with satellite imagery to yield more discriminative and weather-resilient representations for drone geo-localization.

Introduction and Motivation

Drone-based geo-localization aims to determine the precise geographic location of a drone-captured image by matching it against reference geo-tagged satellite images [6,36,37,39,53]. However, due to significant viewpoint differences, appearance variations, and environmental changes between drone and satellite images, achieving reliable matching across domains remains highly challenging [11]. Real-world multiple weather conditions introduce visual noise and visibility loss in drone imagery [7]. Existing methods have largely neglected road map data, which provides strong, inherently weather-invariant geometric layout cues at negligible additional cost [2].

Proposed Method: GeoFuse

GeoFuse is a novel cross-modal fusion framework that leverages road map semantics alongside satellite imagery to enrich satellite representations with spatial geometric cues that remain robust under varying weather conditions [3]. Central to GeoFuse is a flexible fusion module that integrates satellite and road map features through interactions at both token and channel levels, controlled by a lightweight dynamic gating mechanism that adaptively weights modality contributions per instance [3]. The framework processes two branches: a drone-text alignment branch for multi-weather cross-modal matching, and a satellite-road fusion branch for geometric-aware representation learning [3].

Feature Fusion Details

The cross-modal learning module consists of two branches: a drone-text alignment branch and a satellite-road fusion branch [3]. The satellite feature Fs and the road map feature Fr are fused via an attention-based fusion mechanism that consists of two complementary components: token-level fusion for modeling spatial correspondences and channel-level fusion for capturing interdependency across feature channels [7]. Token-level fusion involves cross-attention followed by self-attention refinement, where the road map is employed as the key and value in the crossattention to update the satellite image tokens [7]. Channel-level fusion updates features via channel-level cross-attention, where road map feature channels guide this interaction [8]. The dual-level fused feature is obtained by applying global average pooling over the token dimension [9].

Optimization Objectives

The joint training objectives of the multi-modal interaction framework consist of three components:

  1. For the multi-weather drone-text branch, two cross-modal objectives are adopted: a contrastive loss LITC and a matching loss LITM, summed as LIT = LITC + LITM [10].

  2. To enhance cross-view feature discriminability and robustness under challenging conditions, a class-level contrastive loss LCC based on the InfoNCE [25] is introduced [10].

  3. The instance loss LCE, which is a softmax cross-entropy loss with weight-shared classifiers [53], is applied to supervise the drone, satellite, and fused multi-modal embeddings [10].

The total loss is formulated as the sum of these components: Ltotal = LIT + LCE + lambda · LCC [10].

Experimental Validation

Experiments validate that GeoFuse consistently surpasses current state-of-the-art methods, achieving approximately +3.46% and +23.18% Recall@1 gains on the University-1652 and DenseUAV benchmarks against diverse weather conditions, respectively [4]. The method achieves a mean accuracy of R@1 of 80.60% and a mean AP of 83.35% for the Drone → Satellite retrieval task on University-1652 under multiple weathers [11]. Similarly, in the Satellite → Drone retrieval task, our method attains a mean accuracy of R@1 of 90.39% and a mean AP of 80.59% [4]. The results show that GeoFuse achieves greater robustness by fusing satellite imagery with road maps, introducing accurate, stable geographic structural constraints that are largely unaffected by visual noise [12].

Conclusion

In this work, we present GeoFuse, a robust drone-view geo-localization framework that exploits freely available road maps as a weather-invariant geometric prior to counter adverse weather effects [5]. By fusing precisely aligned road-map and satellite features via token- and channel-level interactions, our adaptive fusion module effectively balances modality contributions, while class-level crossview contrastive learning aligns degraded drone views with the fused representations [10]. Experiments on University-1652 and DenseUAV verify consistent gains in localization accuracy across diverse weather conditions and strong cross-dataset generalization [4]. These results highlight the value of road maps as lightweight, accessible auxiliary priors, opening a promising path toward reliable multi-weather drone geo-localization.

Limitations

Our method still has certain limitations that warrant future investigation. It relies on the availability and precise spatial alignment of road maps with satellite imagery, which can be unavailable or misaligned in remote mountainous regions, newly developed urban areas, conflict zones, or scenes affected by recent infrastructure changes and low-resolution data [14]. More critically, in extremely unstructured environments where road networks are sparse or entirely absent, the geometric guidance degrades to effectively blank input [14].

References

  1. Ahn, W.J., Park, S.Y., Pae, D.S., Choi, H.D., Lim, M.T.: Bridging viewpoints in cross-view geo-localization with siamese vision transformer

2605.14925

  1. Sun, J., Huang, J., Jiang, X., Zhou, Y., VONG, C.M.: Cgsi: Context-guided and uav’s status informed multimodal framework for generalizable cross-view geolocalization

  2. Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, Ł., Polosukhin, I.: Attention is all you need

  3. Wen, J., Yu, H., Zheng, Z.: Weatherprompt: Multi-modality representation learning for all-weather drone visual geo-localization

  4. Xia, P., Wan, Y., Zheng, Z., Zhang, Y., Deng, J.: Enhancing cross-view geolocalization with domain alignment and scene consistency

  5. Yang, L., Kang, B., Huang, Z., Xu, X., Feng, J., Zhao, H.: Depth anything: Unleashing the power of large-scale unlabeled data

  6. Wang, T., Zheng, Z., Yan, C., Zhang, J., Sun, Y., Zheng, B., Yang, Y.: Each part matters: Local patterns facilitate cross-view geo-localization

  7. Wang, X., Xu, R., Cui, Z., Wan, Z., Zhang, Y.: Fine-grained cross-view geolocalization using a correlation-aware homography estimator

  8. Wang, T., Zheng, Z., Zhu, Z., Sun, Y., Yan, C., Yang, Y.: Learning cross-view geolocalization embeddings via dynamic weighted decorrelation regularization

  9. Wang, T., Zheng, Z., Sun, Y., Yan, C., Yang, Y., Chua, T.S.: Multiple-environment self-adaptive network for aerial-view geo-localization

  10. Tian, Y., Chen, C., Shah, M.: Cross-view image matching for geo-localization in urban environments

  11. Shi, Y., Yu, X., Liu, L., Campbell, D., Koniusz, P., Li, H.: Accurate 3-dof camera geo-localization via ground-to-satellite image matching

  12. Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J.: Learning transferable visual models from natural language supervision

  13. Vaswani, A.: Attention is all you need

  14. Wang, T., Zheng, Z., Zhu, Z., Sun, Y., Yan, C., Yang, Y.: Learning cross-view geolocalization embeddings via dynamic weighted decorrelation regularization

  15. Wang, X., Xu, R., Cui, Z., Wan, Z.: Fine-grained cross-view geolocalization using a correlation-aware homography estimator

  16. Wen, J., Yu, H., Zheng, Z.: Weatherprompt: Multi-modality representation learning for all-weather drone visual geo-localization

Improvements for AI systems

  1. textbfWeight-Invariant Geometric Prior Integration for Weather Robustness in Geo-localization: The GeoFuse Framework enables models to yield more discriminative and weatherresilient representations by integrating precisely aligned road map tiles with satellite imagery. This allows the AI system to maintain high accuracy when faced with adverse conditions like fog, rain, or snow, which typically cause feature degradation in drone imagery.

  2. textbfAdaptive Multi-Modal Feature Fusion: The framework utilizes a flexible fusion module that combines satellite and road map features via token-level and channel-level interactions controlled by a lightweight dynamic gating mechanism that adaptively weights modality contributions per instance. This specific mechanism allows the system to dynamically decide how much to trust the road map versus the satellite image based on the input instance, leading to superior feature representation.

  3. textbfClass-Level Contrastive Learning for Stable Cross-View Alignment: The implementation of class-level crossview contrastive learning promotes robust alignment between weather-degraded drone features and the fused satellite-roadmap representations. This strategy provides a stable supervision signal independent of the batch composition, which prevents performance instability common in standard contrastive methods, ensuring consistent feature matching across different weather scenarios.

  4. textbf Text-Free Road Map Curation for Generalization: The process of removing textual content from road maps prevents information leakage and makes the maps more agnostic to specific locales, thereby strengthening the model’s generalization ability to unknown or novel environments. This ensures the AI system relies solely on structural cues like road intersections and layouts rather than dataset-specific annotations.

  5. textbf Enhanced Feature Discrimination via Dual-Level Fusion: The combination of token-level fusion (where road maps update satellite tokens) and channel-level fusion (guided by road map feature channels) allows the model to capture both fine-grained spatial correspondences and interdependency across feature channels. This dual approach results in a final representation, described as dual-level fused feature, that is highly discriminative for geographic location prediction.

Sources

Related papers