Cross-View Feature Matching: Survey, Benchmarking, and Foundation-Model Perspectives
Songlin Du, Xiaoyong Lu, Zeyu Wu, Xiaobo Lu, Guobao Xiao, Bin Fan, Jiayi Ma, Takeshi Ikenaga
Southeast University · Tongji University · University of Science and Technology Beijing · Wuhan University · Waseda University
cs.LG, cs.CV
Submitted: 2026-08-11
Updated: 2026-08-12
Comments: This manuscript goes beyond a conventional survey. It proposes a new taxonomy for cross-view feature matching, provides extensive benchmarking under unified datasets and protocols, and offers original analysis from the perspective of vision foundation models. These contributions provide substantive methodological synthesis, empirical findings, and new research insights
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 95/100
The gist: This survey presents a unified review of cross-view feature matching, a fundamental computer vision problem concerned with establishing reliable correspondences between images with large viewpoint
Terminology
Summary
This survey presents a unified review of cross-view feature matching, a fundamental computer vision problem concerned with establishing reliable correspondences between images with large viewpoint variations. The field has evolved from task-specific models toward unified and generalizable correspondence models, with recent progress driven by vision foundation models (VFMs). The authors note that existing studies remain highly diverse in their problem formulations, model architectures, training paradigms, and evaluation protocols, making it difficult to obtain a unified understanding of the field.
The paper introduces a structured taxonomy organizing methods along six key dimensions: feature extraction, single-type feature matcher (including sparse, semi-dense, and dense matching), multi-type feature matcher, VFMs based methods, training strategy, and robust estimation. The main contributions are summarized as: a hierarchical taxonomy, methodological analysis of recent advances, unified benchmarking under consistent protocols, and identification of challenges and trends.
Feature Extraction: The paper categorizes feature detectors and descriptors into handcrafted and learned methods. Handcrafted detectors evolved from Harris corners to SIFT, SURF, MSER, FAST, KAZE, and BRISK. Learned detectors include SuperPoint, D2-Net, R2D2, KeyNet, ALIKE, DISK, DeDoDe, and DaD. Handcrafted descriptors include SIFT, SURF, BRIEF, BRISK, FREAK, and LIOP, while learned descriptors include DeepDesc, L2-Net, HardNet, HyNet, SOSNet, ZippyPoint, and XFeat.
Sparse Matcher: Sparse matching establishes correspondences between keypoint sets. Methods are categorized by motivation:
-
Accuracy-motivated: SuperGlue pioneered Transformer-based matching with self- and cross-attention. Other methods include NCTR, DNC-Net, IMP, SAM, S2H-GNN, HTMatch, DiffGlue, ResMatch, SemaGlue, SemMatcher, CoMatcher, and AAPMatcher.
-
Efficiency-motivated: SGMNet, KeyGNN, ClusterGNN, ClusterGNN+, ParaFormer, LightGlue, AMatFormer, LightSGM, FilterGNN, MaKeGNN, MambaGlue, and Ada-Matcher focus on reducing computational cost.
-
Generalization-motivated: OmniGlue uses DINOv2 features, LFM-3D incorporates 3D priors, and MapGlue develops multimodal matching for remote sensing.
Semi-Dense Matcher: These methods use coarse-to-fine strategies. Categories include:
-
Accuracy-motivated: Patch2Pix, LoFTR, MatchFormer, ASpanFormer, 3DG-STFM, MSFormer, TopicFM, TopicFM+, ASTR, SEM, AffineFormer, FineFormer, PRISM, ContextMatcher, DeepMatcher, FMAP, OAMatcher, DSAP, CorMatcher, MR-Matcher, FMRT, WinMRSI, LGFCTR, CoMatch, and IMD.
-
Efficiency-motivated: EcoMatcher, ELoFTR, MAIM, JamMa, EDM, CasP, VD-Matcher, LiteSAM, and ProMa.
-
Distribution-motivated: PATS, AdaMatcher, RCM, CasMTR, HomoMatcher, IMAmatch, and RCM+.
Dense Matcher: These methods estimate dense correspondence fields. Categories include:
-
Accuracy-motivated: Rocco et al., DGC-Net, GLU-Net, GOCor, PDC-Net, DKM, RGM, RoMa, PanMatch, BEAMER, UFM, and RoMa v2.
-
Efficiency-motivated: GFNet and ArgMatch.
Multi-Type Feature Matcher: Methods incorporating complementary feature types:
-
Point-Line: HDPL, GlueStick, RCM+.
-
Point-Semantic: TopicFM, TopicFM+, SGAM, MESA, OmniGlue, SemaGlue, SemMatcher, SGAD.
-
Point-Depth: Toft et al., 3DG-STFM, LFM-3D, Wang et al., LiftFeat.
Vision Foundation Models Based Matcher: Methods exploiting VFMs:
-
DINO-based: OmniGlue, RoMa, SRMatcher, Cadar et al., DINO-VO, SigMa, SGAT, DistillMatch, TextFM, RIM, RoMa v2, MV-RoMa.
-
SAM-based: MESA, DMESA, SAMFeat, SGAD, SAMatcher.
-
Diffusion-based: DIFT, MATCHA, IMD.
-
Geometry-VFM-based: MASt3R, SGPFeat.
Training Strategy: PMatch introduces paired masked image modeling, GIM uses self-training from internet videos, MINIMA generates pseudo multimodal datasets, MatchAnything proposes universal cross-modality pre-training, and L2M lifts 2D images into 3D space.
Robust Estimation: Methods are classified into:
-
Sampling consensus: RANSAC, MSAC, MLESAC, MAPSAC, LO-RANSAC, ANSAC, BANSAC.
-
Deterministic: Olsson et al., Le et al., Cai et al., Graph-Cut RANSAC, Prasad et al.
-
Learnable: DSAC, MQ-Net, progressive consensus learning, Neural Guided RANSAC, ∇-RANSAC, Probst et al.
Benchmarking: The paper evaluates methods on MegaDepth, ScanNet, HPatches, Aachen Day-Night, and InLoc datasets. Key findings from relative pose estimation on MegaDepth and ScanNet show that the performance of two-view pose estimation methods systematically improves with increasing matching density: dense methods consistently outperform semi-dense approaches, which in turn generally surpass sparse methods.
On MegaDepth, RoMa v2 achieves the highest accuracy with AUC@20° of 86.6%, while on ScanNet it reaches 73.8%. For homography estimation on HPatches, three paradigms exhibit a clear hierarchical pattern in overall performance
with dense methods achieving the highest AUC. For visual localization, semi-dense methods deliver the best performance across indoor and outdoor environments.
Open Problems: The paper identifies seven key challenges:
-
Representation gap between geometric precision and semantic invariance.
-
Ambiguity and uncertainty in correspondence estimation.
-
Lack of explicit correspondence-level reasoning.
-
Limited generalization of correspondence models.
-
Incomplete modeling of non-rigid and dynamic correspondence.
-
Lack of foundation models pretrained for geometric correspondence.
-
Computational bottlenecks in global correspondence modeling.
The survey concludes by stating it "provides a clear and structured foundation for understanding recent advances in cross-view feature matching and serves as a useful reference for future research toward more unified, efficient, and generalizable correspondence systems."
Improvements for AI systems
Improvements to AI Systems:
-
Unified Correspondence Reasoning Module: Build an AI system that integrates sparse, semi-dense, and dense matchers into a single adaptive pipeline. The system dynamically selects the matching density based on task requirements (e.g., real-time vs. high-accuracy) and fuses outputs from all three paradigms to maximize robustness. This improved system can handle arbitrary viewpoint changes in robotics, SLAM, and 3D reconstruction with superior accuracy and efficiency.
-
Foundation Model for Geometric Correspondence: Develop a pretrained VFM specifically for geometric matching (inspired by the gap identified for DINO/SAM/diffusion-based methods). The system would be pretrained on large-scale multi-modal data (RGB, depth, semantic, line) using self-supervised objectives like paired masked image modeling and 3D lifting. This improved system can generalize zero-shot to unseen scenes, non-rigid objects, and cross-domain tasks (e.g., satellite-to-street view, medical imaging) without fine-tuning.
-
Uncertainty-Aware Matching with Explicit Reasoning: Enhance the matcher with an explicit uncertainty estimation module and a correspondence-level reasoning layer (e.g., graph neural network with attention over match hypotheses). The system outputs confidence maps and rejects ambiguous matches, reducing false positives in challenging conditions (occlusions, repetitive textures, dynamic scenes). This improved system can be deployed in safety-critical applications like autonomous driving and augmented reality.
-
Efficient Global Correspondence Transformer: Replace quadratic-complexity attention in global matching with a linear-complexity, Mamba-based or cluster-aware architecture (as suggested by MambaGlue and ClusterGNN). The system maintains global context while scaling to high-resolution images. This improved system can process 4K/8K images in real-time on edge devices, enabling drone navigation, wide-baseline stereo, and large-scale 3D mapping.
-
Cross-Modal and Multi-Type Feature Fusion: Build a system that jointly learns and fuses point, line, semantic, and depth features via a learnable fusion layer (inspired by multi-type matchers like GlueStick and TopicFM). The system uses a unified embedding space to combine complementary cues, improving matching under large illumination changes, seasonal variations, and modality gaps (e.g., RGB-to-infrared). This improved system can perform robust visual localization in changing environments (day/night, weather) for long-term autonomy.
-
Self-Training from Unlabeled Video and Pseudo-Label Generation: Implement a continuous learning system that uses internet videos and pseudo-multimodal datasets (as in GIM and MINIMA) to self-train. The system iteratively refines its correspondence predictions on new data, adapting to novel environments without human annotation. This improved system can be deployed in lifelong learning scenarios, such as warehouse robots or planetary rovers, where environments evolve over time.
-
Hierarchical Robust Estimation with Learnable Consensus: Integrate learnable robust estimators (e.g., Neural Guided RANSAC, DSAC) into the matching pipeline, replacing fixed RANSAC. The system learns to propose inlier sets and optimize model parameters end-to-end. This improved system achieves higher pose accuracy under high outlier ratios (e.g., >90% outliers in wide-baseline matching) and can be used for precise camera calibration, structure-from-motion, and visual odometry.
-
Task-Adaptive Benchmarking and Auto-Configuration: Develop a meta-learning system that analyzes the input image pair (e.g., texture, depth variation, overlap) and automatically configures the optimal matcher (sparse, semi-dense, dense) and hyperparameters (e.g., feature count, threshold). The system uses the unified benchmark results (MegaDepth, ScanNet, HPatches) as a prior. This improved system can be used as a plug-and-play module in existing computer vision pipelines, ensuring near-optimal performance across diverse applications without manual tuning.
Sources
- DaD: Distilled Reinforcement Learning for Diverse Keypoint Detection
- MapGlue: Multimodal Remote Sensing Image Matching
- CoMatch: Dynamic Covisibility-Aware Transformer for Bilateral Subpixel-Level Semi-Dense Image Matching
- HomoMatcher: Dense Feature Matching Results with Semi-Dense Efficiency by Homography Estimation
- RGM: A Robust Generalizable Matching Model
- PanMatch: Unleashing the Potential of Large Vision Models for Unified Matching Models
- UFM: A Simple Path towards Unified Dense Correspondence with Flow
- RoMa v2: Harder Better Faster Denser Feature Matching
- Searching from Area to Point: A Hierarchical Framework for Semantic-Geometric Combined Feature Matching
- DistillMatch: Leveraging Knowledge Distillation from Vision Foundation Model for Multimodal Image Matching
- RIM: A Retrieval-In-Matching Framework for Cross-Domain Global Visual Localization of UAVs
- MV-RoMa: From Pairwise Matching into Multi-View Track Reconstruction
- MESA: Effective Matching Redundancy Reduction by Semantic Area Segmentation
- SAMatcher: Dense Co-Visibility Modeling via Cross-View Fusion for Scale-Imbalance Image Matching
- GIM: Learning Generalizable Image Matcher From Internet Videos
- MatchAnything: Universal Cross-Modality Image Matching with Large-Scale Pre-Training
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks