Research
Cross-View Feature Matching: Survey, Benchmarking, and Foundation-Model Perspectives
arXiv:2608.11093v1 Announce Type: cross Abstract: Cross-view feature matching aims to establish reliable correspondences across images with large viewpoint variations. Over the past decade, the field
arXiv:2608.11093v1 Announce Type: cross Abstract: Cross-view feature matching aims to establish reliable correspondences across images with large viewpoint variations. Over the past decade, the field has evolved from task-specific models toward increasingly unified and generalizable correspondence models, with recent progress further driven by the emergence of vision foundation models (VFMs). Despite these advances, existing studies remain highly diverse in their problem formulations, model architectures, training paradigms, and evaluation protocols, making it difficult to obtain a unified understanding of the field. In this survey, we present a unified review of cross-view feature matching. We first introduce a structured taxonomy covering feature extraction, single-type feature matcher, multi-type feature matcher, VFMs based methods, training strategy and robust estimation, providing a coherent framework for analysis and comparison. We further examine recent advances, distilling key design principles and highlighting the shift toward unified and generalizable correspondence models. We also provide a unified experimental benchmarking of representative state-of-the-art methods under consistent protocols, enabling fair and comprehensive performance comparisons. In addition, we discuss open challenges and future directions, including efficiency, robustness under extreme conditions, and cross-domain generalization. This survey aims to provide a comprehensive and structured reference for understanding the evolution, current landscape, and future development of cross-view feature matching in the era of vision foundation models.
Related
- SceneGlue: Scene-Aware Transformer for Feature Matching without Scene-Level Annotation
- TeaMatch: Teachable Cross-Modal Representation Learning for 2D-3D Matching
- Geometry Reinforced Efficient Attention Tuning Equipped with Normals for Robust Stereo Matching
- Anatomy-Grounded Synthetic Coronary Angiography for Geometry-Informed Multi-View Matching
Source: arXiv cs.CV | 2026-08-12