Pith. sign in

REVIEW 4 major objections 5 minor 255 references

Deep Learning Reforms Image Matching: A Survey and Outlook

T0 review · 4 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read The paper claims that deep learning reforms image matching by replacing and merging pipeline stages, with end-to-end semi-dense/dense matchers proving most accurate and generalizable on hard benchmarks.

desk verdict Useful pipeline-aligned survey with an up-to-date taxonomy, but the benchmark's headline claim that dense matchers 'excel' is undercut by per-method protocol differences that confound architecture with resolution, match budget, and RANSAC threshold. read the letter →

arxiv 2506.04619 v1 pith:I5CUG4Y3 submitted 2025-06-05 cs.CV

classification cs.CV
keywords imagematchingdeeplearningdetector-freesparsedenserelativeposeestimationvisuallocalizationcorrespondence
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper attempts to show that deep learning has not merely improved individual feature detectors or outlier filters; it has progressively dissolved the classical detector-descriptor-matcher-filter-estimator pipeline. It organizes the field into two reform directions—swapping single steps for learnable counterparts and merging multiple steps into end-to-end modules—and benchmarks representative methods on relative pose recovery, homography estimation, and visual localization. The load-bearing result is that fully end-to-end semi-dense/dense matchers, which skip explicit keypoint detection, give the strongest pose accuracy on hard indoor and outdoor scenes and generalize across datasets, while sparse matchers remain limited by keypoint quality and dense matchers by speed.

What carries the argument

The carrying object is a pipeline-aligned taxonomy: two reform directions, alternative learnable steps and merged learnable modules, mapped onto the classical detector-descriptor-to-estimator chain. The survey uses this taxonomy to structure its review and its experiments, and the experiments themselves are carried by standard metrics—pose-error AUC at 5, 10, and 20 degrees, homography corner reprojection accuracy and AUC, PCK for dense matching, and localization recall at distance and orientation thresholds—across MegaDepth, YFCC100M, ScanNet, SUN3D, HPatches, Aachen Day-Night, and InLoc.

What would settle it

Rerun the pose, homography, and localization experiments under a single protocol: one image resolution, one keypoint budget, and one fitting threshold for all methods, plus repeated trials to estimate noise. If semi-dense/dense matchers no longer lead, the paper's central conclusion is wrong.

Watch

Extended reading notes

Core claim

On its own terms, the paper's central discovery is that deep learning reforms image matching structurally: learnable replacements for detector-descriptor, outlier filter, and geometric estimator yield gains, but merging stages into end-to-end units goes further. In its benchmarks, dense matchers such as RoMa and DKM lead on MegaDepth, ScanNet, HPatches, and InLoc, while sparse matchers like SuperGlue and LightGlue remain strong on daytime Aachen localization and easier scenes. The paper concludes that semi-dense/dense frameworks excel in challenging scenarios and generalize well across datasets, with efficiency and multi-view keypoint consistency as open bottlenecks.

Load-bearing premise

That the benchmark is fair across method families: different methods are run at different image sizes, with different numbers of keypoints and different geometric fitting thresholds, so the conclusion that dense matchers are better assumes these settings do not bias the ranking.

Editorial extensions

If this is right

  • If the conclusion holds, applications that need robustness under nighttime lighting, low-texture indoor scenes, or wide baselines should prefer detector-free semi-dense/dense matchers over sparse pipelines.
  • Sparse matchers will remain a default when speed or multi-view 3D consistency matters, because their accuracy ceiling is set by keypoint repeatability and descriptor quality.
  • The remaining barrier for dense matchers is computational cost, so lightweight architectures, pruning, quantization, and knowledge distillation become natural next targets.
  • Learnable outlier filters and geometric estimators are useful upgrades but cannot recover matches that were never proposed, which is why merging stages removes a real ceiling.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A fairer benchmark with uniform image resolutions, equal keypoint budgets, and a single fitting threshold might shrink the dense matchers' lead, since the survey runs top dense models at lower resolutions than sparse pipelines.
  • If dense matching keeps improving, visual SLAM and structure-from-motion systems could replace sparse feature tracking with dense flow, but they would need new machinery to enforce multi-view consistency.
  • The localization results, where a sparse matcher rivals dense ones on daytime Aachen, suggest the advantage of dense methods is scene-dependent rather than universal.
  • The same taxonomy implies that large pretrained geometric models, trained on massive image data, may absorb both step replacement and merging by supplying global priors directly from image pairs.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. This paper surveys deep learning methods for two-view image matching using a taxonomy aligned with the classical pipeline: learnable replacements of individual stages (detector-descriptor, outlier filter, geometric estimator) and merged end-to-end modules (middle-end sparse matcher, semi-dense/dense matcher, pose regressor). It reviews representative methods in each category and reports experiments on relative pose estimation (MegaDepth, YFCC100M, ScanNet, SUN3D), homography estimation (HPatches), matching accuracy (MegaDepth PCK), and visual localization (Aachen Day-Night, InLoc). The paper's central empirical conclusion, stated in Section 5.3.4, is that semi-dense/dense frameworks excel in challenging scenarios and generalize well across datasets, while sparse matchers are limited by keypoint quality.

Significance. The survey has genuine value as a reference: the pipeline-aligned taxonomy is a useful organizing contribution, the coverage includes many 2023-2025 methods, and Section 5.4 is unusually explicit about the evaluation protocols used for different method families. If the empirical comparison were controlled, the conclusion that detector-free dense matchers are the current accuracy leader would be an informative field-level statement. The paper does not ship code or machine-checked proofs, but the tables collate a large body of external results and some new runs; the main weakness is that the protocol heterogeneity described in Section 5.4 makes the headline ranking hard to interpret as a comparison of method architectures.

major comments (4)
  1. [§5.3.4, §5.4.1, Tables 1-2] The headline claim that semi-dense/dense frameworks 'excel' is read from Tables 1, 2, and 4, but those tables vary method architecture together with input resolution, keypoint budget, and RANSAC threshold. In Section 5.4.1, sparse pipelines on MegaDepth use a 1600-pixel longest side and up to 2048 keypoints, while DKM runs at 880x660 and RoMa at 672x672; on ScanNet/SUN3D the dense matchers use a 480-pixel shortest side while sparse matchers are capped at 1024 keypoints, and RANSAC thresholds differ (0.5/f versus 1/f). The homography protocol in Section 5.4.2 similarly assigns 480 shortest side and 2048 keypoints to sparse methods but 640 longest side, 880x660, or 672x672 to dense methods, with a 3/f RANSAC threshold. The comparison therefore does not isolate the 'semi-dense/dense' design choice, and the stated ranking could change under matched protocols; this is load-bearing for the central conclusion.
  2. [§5.3.4, Table 4] On Aachen Day-Night, the paper itself notes that semi-dense/dense matchers are 'not always superior': ALIKED+LightGlue reaches 89.9% daytime and 76.4% nighttime at (0.25m, 2°), at or above DKM (88.1/72.3) and RoMa (88.1/71.7). The global conclusion that dense frameworks 'excel' is then carried mainly by indoor InLoc rows and by the relative-pose/homography tables, where the resolution and RANSAC-threshold confounds from Sections 5.4.1 and 5.4.2 are also present. A conclusion stated as 'Collectively' should separate dataset category from method category, or explicitly qualify the claim to indoor and pose-estimation settings.
  3. [§5.3, Tables 1-4] No error bars, confidence intervals, or repeated-run statistics are reported. Many cross-method gaps in the tables are small (e.g., RoMa 62.76 vs DKM 60.89 vs ELoFTR 56.38 at 5° on MegaDepth; RoMa 59.5 vs ELoFTR 59.5 vs TopicFM+ 59.5 at 1.0m/10° on InLoc DUC2), and it is not possible from the paper to tell whether these differences are reproducible or within run-to-run noise. Because the benchmark code is not released, this uncertainty cannot be resolved by the reader; the authors should either provide the evaluation code and variance estimates or soften the precise ranking claims.
  4. [§5.3.1, Table 1] The selection of 'representative algorithms' is not governed by stated inclusion criteria, and several methods in the tables come from the authors' own group. In addition, the row SIFT+U-Match+* adjusts the inlier prediction threshold from the default 0 to 2.0, an intervention not applied to other outlier filters. This makes it hard to rule out selection or tuning bias in the comparative tables. The authors should state inclusion criteria, release the exact evaluation script, and apply identical post-processing to all methods.
minor comments (5)
  1. [§4.2, Figures 7 and 8] The framework figures contain repeated header text from 'CoMatch: Dynamic Covisibility-Aware Transformer for Bilateral Subpixel-Level Semi-Dense Image Matching' embedded in the image, while the captions credit only 'Image refers to [165]'; this appears to be an editing artifact and should be removed or replaced with the actual figures.
  2. [Table 2 header] The header says 'The default estimator is RANSAC [130]', but reference [130] is NG-RANSAC and the surrounding text refers to RANSAC [33]; please correct the reference.
  3. [§5.4.1] The statement that 'some methods additionally pad images to ensure specific resolution requirements' is too vague; specify per-method padding and resizing choices for reproducibility.
  4. [§5.4.4] The sentence 'For the sake of fairness, we meticulously comply with the pipeline and evaluation settings of the online visual localization benchmark' is at odds with the immediately preceding per-method differences in resolution and keypoint budget; please rephrase or justify those differences.
  5. [Throughout] There are several typos and awkward phrasings, including 'shwon' in §4.1, 'Nignt' in the Table 4 header, 'that interleaves that interleaves' in §4.2.2, and 'inappositeness and unconsistency' in the Introduction.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the survey's conclusions summarize external benchmark experiments rather than deriving results from fitted inputs, self-definitions, or load-bearing self-citations.

full rationale

The paper's central inference, that 'semi-dense/dense frameworks excel in challenging scenarios and generalize well across datasets' (Section 5.3.4), is an empirical summary of Tables 1-4. Those tables report evaluations of externally published methods on standard datasets (MegaDepth, ScanNet, HPatches, Aachen, InLoc) using established metrics (AUC, Acc., PCK). No fitted parameter is renamed as a prediction, and no equation in the paper defines the conclusion in terms of its own inputs. The authors cite several of their own methods (e.g., U-Match, ConvMatch, DeMatch, DiffGlue) and include them as baselines, but these self-citations are not used to justify the load-bearing claim: the claim is supported by the overall table rankings, which include many external methods such as SuperGlue, LightGlue, LoFTR, DKM, and RoMa. No 'uniqueness theorem' or prior author-defined ansatz is invoked to force the chosen taxonomy or conclusion. The identified weaknesses in the benchmark, such as inconsistent resolutions, keypoint budgets, and RANSAC thresholds across method categories, and the absence of error bars, are genuine threats to the validity of the comparison, but they are matters of experimental fairness and significance, not circularity under the definition used here. The survey does not claim to derive a new method from first principles; it organizes existing work and evaluates it. Therefore, no specific circular step can be exhibited, and the appropriate finding is no significant circularity.

Assumptions & free parameters 3 free parameters · 4 assumptions · 0 invented entities

No free parameters are fitted to support a derivation; the ledger instead lists benchmark protocol choices and dataset assumptions that the survey's comparative conclusions depend on.

free parameters (3)
  • RANSAC inlier threshold schedule = 0.5/f, 1/f, 3/f, 12 px, 48 px depending on dataset and method
    Hand-chosen per method and dataset in Section 5.4; affects pose and localization recall and can shift rankings.
  • Image resizing and padding settings = 480 shortest side, 640 longest side, 672x672, 880x660, 1024, 1152x1152, 1600 longest side, etc.
    Varied per method and dataset; dense matcher accuracy is resolution-dependent, so the comparison is not fully controlled.
  • Keypoint extraction budget = 1024, 2048, or 4096 keypoints per image
    Applied differently across detector-based methods and benchmarks; higher budgets generally help pose estimation.
assumptions (4)
  • domain assumption Ground-truth poses, depths, and reconstructions from benchmark datasets are accurate enough to rank methods.
    Section 5.1 relies on COLMAP and bundle adjustment outputs from MegaDepth, YFCC100M, ScanNet, SUN3D, HPatches, Aachen, and InLoc.
  • domain assumption Released model checkpoints and open-source evaluation pipelines reproduce published performance.
    Section 5.4 references external repos such as DenseMatching and HLoc and uses them to produce all tables.
  • domain assumption Testing outdoor-trained models on indoor datasets measures cross-scene generalizability rather than unfairness.
    Section 5.3.1 explicitly uses outdoor models on ScanNet and SUN3D.
  • domain assumption The selected representative methods fairly represent their categories.
    Section 5.3.1 says 'some representative algorithms' are selected; no formal inclusion criteria are given.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Deep Learning Reforms Image Matching: A Survey and Outlook." pith.science (2026). https://pith.science/paper/I5CUG4Y3

@misc{pith2026250604619,
  author       = {Pith},
  title        = {Pith review of: Deep Learning Reforms Image Matching: A Survey and Outlook},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/I5CUG4Y3}},
  note         = {Machine review of arXiv:2506.04619}
}
read the original abstract

Image matching, which establishes correspondences between two-view images to recover 3D structure and camera geometry, serves as a cornerstone in computer vision and underpins a wide range of applications, including visual localization, 3D reconstruction, and simultaneous localization and mapping (SLAM). Traditional pipelines composed of ``detector-descriptor, feature matcher, outlier filter, and geometric estimator'' falter in challenging scenarios. Recent deep-learning advances have significantly boosted both robustness and accuracy. This survey adopts a unique perspective by comprehensively reviewing how deep learning has incrementally transformed the classical image matching pipeline. Our taxonomy highly aligns with the traditional pipeline in two key aspects: i) the replacement of individual steps in the traditional pipeline with learnable alternatives, including learnable detector-descriptor, outlier filter, and geometric estimator; and ii) the merging of multiple steps into end-to-end learnable modules, encompassing middle-end sparse matcher, end-to-end semi-dense/dense matcher, and pose regressor. We first examine the design principles, advantages, and limitations of both aspects, and then benchmark representative methods on relative pose recovery, homography estimation, and visual localization tasks. Finally, we discuss open challenges and outline promising directions for future research. By systematically categorizing and evaluating deep learning-driven strategies, this survey offers a clear overview of the evolving image matching landscape and highlights key avenues for further innovation.

Figures

Figures reproduced from arXiv: 2506.04619 by the authors.

Figure 1
Figure 1. Taxonomy of image feature matching. The orange boxes mark the focus of this paper. [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Timelines of alternative learnable steps (Section 3) and merged learnable modules (Section 4). [PITH_FULL_IMAGE:figures/full_fig_p002_2.png] view at source ↗
Figure 3
Figure 3. Frameworks of different learnable detector-descriptors. [PITH_FULL_IMAGE:figures/full_fig_p003_3.png] view at source ↗
Figures from the paper (7 more)
Figure 4
Figure 4. Figure 4: Frameworks of different learnable outlier filters. [PITH_FULL_IMAGE:figures/full_fig_p006_4.png]
Figure 5
Figure 5. Figure 5: Frameworks of different learnable geometric estimators. [PITH_FULL_IMAGE:figures/full_fig_p007_5.png]
Figure 6
Figure 6. Figure 6: Framework of middle-end sparse matchers. [PITH_FULL_IMAGE:figures/full_fig_p008_6.png]
Figure 7
Figure 7. Figure 7: As the pioneering work in this paradigm, LoFTR [171] uses a ResNet-FPN [172] backbone to extract coarse features at 1/8 resolution and fine features at 1/2 resolution. The coarse features are processed by N layers of interleaved lin￾ear self- and cross-attention [173] …
Figure 7
Figure 7. Figure 7: Frameworks of different end-to-end semi-dense matchers. Image refers to [165]. [PITH_FULL_IMAGE:figures/full_fig_p010_7.png]
Figure 8
Figure 8. Figure 8: Frameworks of different end-to-end dense matchers. Image refers to [165]. [PITH_FULL_IMAGE:figures/full_fig_p011_8.png]
Figure 9
Figure 9. Figure 9: Frameworks of different pose regressors. [PITH_FULL_IMAGE:figures/full_fig_p012_9.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

255 extracted references · 75 canonical work pages

  1. [1]

    Image retrieval: Ideas, influences, and trends of the new age,

    R. Datta, D. Joshi, J. Li, and J. Z. Wang, “Image retrieval: Ideas, influences, and trends of the new age,”ACM CSUR, vol. 40, no. 2, pp. 1–60, 2008

  2. [2]

    Benchmarking 6dof outdoor visual localization in changing conditions,

    T. Sattler, W. Maddern, C. Toft, A. Torii, L. Hammarstrand, E. Stenborg, D. Safari, M. Okutomi, M. Pollefeys, J. Sivicet al., “Benchmarking 6dof outdoor visual localization in changing conditions,” inCVPR, 2018, pp. 8601–8610

  3. [3]

    Hartley and A

    R. Hartley and A. Zisserman,Multiple view geometry in computer vision. Cambridge University Press, 2003

  4. [4]

    Global structure-from-motion revisited,

    L. Pan, D. Bar ´ath, M. Pollefeys, and J. L. Sch ¨onberger, “Global structure-from-motion revisited,” inECCV, 2024, pp. 58–77

  5. [5]

    Lsd-slam: Large-scale direct monocular slam,

    J. Engel, T. Sch ¨ops, and D. Cremers, “Lsd-slam: Large-scale direct monocular slam,” inECCV, 2014, pp. 834–849

  6. [6]

    3d gaussian splatting for real-time radiance field rendering

    B. Kerbl, G. Kopanas, T. Leimk ¨uhler, and G. Drettakis, “3d gaussian splatting for real-time radiance field rendering.”ACM TOG, vol. 42, no. 4, pp. 139:1–139:14, 2023

  7. [7]

    Cooperative computation of stereo disparity: A cooperative algorithm is derived for extracting dis- parity information from stereo image pairs

    D. Marr and T. Poggio, “Cooperative computation of stereo disparity: A cooperative algorithm is derived for extracting dis- parity information from stereo image pairs.”Science, vol. 194, no. 4262, pp. 283–287, 1976

  8. [8]

    Scene analysis using regions,

    C. R. Brice and C. L. Fennema, “Scene analysis using regions,” AI, vol. 1, no. 3-4, pp. 205–226, 1970

Show all 255 references
  1. [9]

    Adaptive least squares correlation: a powerful image matching technique,

    A. Gruen, “Adaptive least squares correlation: a powerful image matching technique,”SAJPRSC, vol. 14, no. 3, pp. 175–187, 1985

  2. [10]

    Rover visual obstacle avoidance

    H. P . Moravec, “Rover visual obstacle avoidance.” inIJCAI, vol. 81, 1981, pp. 785–790

  3. [11]

    Interesting interest points: A comparative study of interest point perfor- mance on a unique data set,

    H. Aanæs, A. L. Dahl, and K. Steenstrup Pedersen, “Interesting interest points: A comparative study of interest point perfor- mance on a unique data set,”IJCV, vol. 97, pp. 18–35, 2012

  4. [12]

    Comparative evaluation of binary features,

    J. Heinly, E. Dunn, and J.-M. Frahm, “Comparative evaluation of binary features,” inECCV, 2012, pp. 759–773

  5. [13]

    Performance compar- isons of contour-based corner detectors,

    M. Awrangjeb, G. Lu, and C. S. Fraser, “Performance compar- isons of contour-based corner detectors,”IEEE TIP, vol. 21, no. 9, pp. 4167–4179, 2012

  6. [14]

    Hpatches: A benchmark and evaluation of handcrafted and learned local descriptors,

    V . Balntas, K. Lenc, A. Vedaldi, and K. Mikolajczyk, “Hpatches: A benchmark and evaluation of handcrafted and learned local descriptors,” inCVPR, 2017, pp. 5173–5182

  7. [15]

    Comparative evaluation of hand-crafted and learned local fea- tures,

    J. L. Schonberger, H. Hardmeier, T. Sattler, and M. Pollefeys, “Comparative evaluation of hand-crafted and learned local fea- tures,” inCVPR, 2017, pp. 1482–1491

  8. [16]

    Image registration methods: a survey,

    B. Zitova and J. Flusser, “Image registration methods: a survey,” IVC, vol. 21, no. 11, pp. 977–1000, 2003

  9. [17]

    Image matching from handcrafted to deep features: A survey,

    J. Ma, X. Jiang, A. Fan, J. Jiang, and J. Yan, “Image matching from handcrafted to deep features: A survey,”IJCV, vol. 129, no. 1, pp. 23–79, 2021

  10. [18]

    Local feature matching using deep learning: A survey,

    S. Xu, S. Chen, R. Xu, C. Wang, P . Lu, and L. Guo, “Local feature matching using deep learning: A survey,”IF, vol. 107, p. 102344, 2024

  11. [19]

    Local feature matching from detector-based to detector- free: a survey,

    Y. Liao, Y. Di, K. Zhu, H. Zhou, M. Lu, Y. Zhang, Q. Duan, and J. Liu, “Local feature matching from detector-based to detector- free: a survey,”AI, vol. 54, no. 5, pp. 3954–3989, 2024

  12. [20]

    Learning to find good correspondences,

    K. M. Yi, E. Trulls, Y. Ono, V . Lepetit, M. Salzmann, and P . Fua, “Learning to find good correspondences,” inCVPR, 2018, pp. 2666–2674

  13. [21]

    Superpoint: Self- supervised interest point detection and description,

    D. DeTone, T. Malisiewicz, and A. Rabinovich, “Superpoint: Self- supervised interest point detection and description,” inCVPRW, 2018, pp. 224–236. 19

  14. [22]

    Pdc-net+: Enhanced probabilistic dense correspondence network,

    P . Truong, M. Danelljan, R. Timofte, and L. Van Gool, “Pdc-net+: Enhanced probabilistic dense correspondence network,”IEEE TP AMI, vol. 45, no. 8, pp. 10 247–10 266, 2023

  15. [23]

    From coarse to fine: Robust hierarchical localization at large scale,

    P .-E. Sarlin, C. Cadena, R. Siegwart, and M. Dymczyk, “From coarse to fine: Robust hierarchical localization at large scale,” in CVPR, 2019, pp. 12 716–12 725

  16. [24]

    Distinctive image features from scale-invariant keypoints,

    D. G. Lowe, “Distinctive image features from scale-invariant keypoints,”IJCV, vol. 60, pp. 91–110, 2004

  17. [25]

    Orb: An efficient alternative to sift or surf,

    E. Rublee, V . Rabaud, K. Konolige, and G. Bradski, “Orb: An efficient alternative to sift or surf,” inICCV, 2011, pp. 2564–2571

  18. [26]

    Robust wide- baseline stereo from maximally stable extremal regions,

    J. Matas, O. Chum, M. Urban, and T. Pajdla, “Robust wide- baseline stereo from maximally stable extremal regions,”IVC, vol. 22, no. 10, pp. 761–767, 2004

  19. [27]

    A performance evaluation of local descriptors,

    K. Mikolajczyk and C. Schmid, “A performance evaluation of local descriptors,”IEEE TP AMI, vol. 27, no. 10, pp. 1615–1630, 2005

  20. [28]

    Object recognition from local scale-invariant fea- tures,

    D. G. Lowe, “Object recognition from local scale-invariant fea- tures,” inICCV, vol. 2, 1999, pp. 1150–1157

  21. [29]

    A combined corner and edge detec- tor,

    C. Harris and M. Stephens, “A combined corner and edge detec- tor,” inAVC, vol. 15, no. 50, 1988, pp. 10–5244

  22. [30]

    Robust point matching via vector field consensus,

    J. Ma, J. Zhao, J. Tian, A. L. Yuille, and Z. Tu, “Robust point matching via vector field consensus,”IEEE TIP, vol. 23, no. 4, pp. 1706–1721, 2014

  23. [31]

    Gms: Grid-based motion statistics for fast, ultra- robust feature correspondence,

    J. Bian, W.-Y. Lin, Y. Matsushita, S.-K. Yeung, T.-D. Nguyen, and M.-M. Cheng, “Gms: Grid-based motion statistics for fast, ultra- robust feature correspondence,” inCVPR, 2017, pp. 4181–4190

  24. [32]

    The development and comparison of robust methods for estimating the fundamental matrix,

    P . H. Torr and D. W. Murray, “The development and comparison of robust methods for estimating the fundamental matrix,”IJCV, vol. 24, pp. 271–300, 1997

  25. [33]

    Random sample consensus: a paradigm for model fitting with applications to image analysis and automated cartography,

    M. A. Fischler and R. C. Bolles, “Random sample consensus: a paradigm for model fitting with applications to image analysis and automated cartography,”CACM, vol. 24, no. 6, pp. 381–395, 1981

  26. [34]

    Usac: A universal framework for random sample consensus,

    R. Raguram, O. Chum, M. Pollefeys, J. Matas, and J.-M. Frahm, “Usac: A universal framework for random sample consensus,” IEEE TP AMI, vol. 35, no. 8, pp. 2022–2038, 2012

  27. [35]

    Magsac++, a fast, reliable and accurate robust estimator,

    D. Barath, J. Noskova, M. Ivashechkin, and J. Matas, “Magsac++, a fast, reliable and accurate robust estimator,” inCVPR, 2020, pp. 1304–1312

  28. [36]

    Fast corner detection,

    M. Trajkovi ´c and M. Hedley, “Fast corner detection,”IVC, vol. 16, no. 2, pp. 75–87, 1998

  29. [37]

    Learning an interest operator from human eye movements,

    W. Kienzle, F. A. Wichmann, B. Scholkopf, and M. O. Franz, “Learning an interest operator from human eye movements,” in CVPRW, 2006, pp. 24–24

  30. [38]

    Faster and better: A machine learning approach to corner detection,

    E. Rosten, R. Porter, and T. Drummond, “Faster and better: A machine learning approach to corner detection,”IEEE TP AMI, vol. 32, no. 1, pp. 105–119, 2008

  31. [39]

    Predicting match- ability,

    W. Hartmann, M. Havlena, and K. Schindler, “Predicting match- ability,” inCVPR, 2014, pp. 9–16

  32. [40]

    A survey of convo- lutional neural networks: analysis, applications, and prospects,

    Z. Li, F. Liu, W. Yang, S. Peng, and J. Zhou, “A survey of convo- lutional neural networks: analysis, applications, and prospects,” IEEE TNNLS, vol. 33, no. 12, pp. 6999–7019, 2021

  33. [41]

    Learning convolutional filters for interest point detection,

    A. Richardson and E. Olson, “Learning convolutional filters for interest point detection,” inICRA, 2013, pp. 631–637

  34. [42]

    Tilde: A temporally invariant learned detector,

    Y. Verdie, K. Yi, P . Fua, and V . Lepetit, “Tilde: A temporally invariant learned detector,” inCVPR, 2015, pp. 5279–5288

  35. [43]

    Toward geomet- ric deep slam,

    D. DeTone, T. Malisiewicz, and A. Rabinovich, “Toward geomet- ric deep slam,”arXiv:1707.07410, pp. 1–14, 2017

  36. [44]

    Learning covariant feature detectors,

    K. Lenc and A. Vedaldi, “Learning covariant feature detectors,” inECCVW, 2016, pp. 100–117

  37. [45]

    Learning dis- criminative and transformation covariant local feature detectors,

    X. Zhang, F. X. Yu, S. Karaman, and S.-F. Chang, “Learning dis- criminative and transformation covariant local feature detectors,” inCVPR, 2017, pp. 6818–6826

  38. [46]

    Key. net: Keypoint detection by handcrafted and learned cnn filters,

    A. Barroso-Laguna, E. Riba, D. Ponsa, and K. Mikolajczyk, “Key. net: Keypoint detection by handcrafted and learned cnn filters,” inICCV, 2019, pp. 5836–5844

  39. [47]

    Key. net: Keypoint detection by handcrafted and learned cnn filters revisited,

    A. Barroso-Laguna and K. Mikolajczyk, “Key. net: Keypoint detection by handcrafted and learned cnn filters revisited,”IEEE TP AMI, vol. 45, no. 1, pp. 698–711, 2022

  40. [48]

    Ness-st: Detecting good and stable keypoints with a neural stability score and the shi- tomasi detector,

    K. Pakulev, A. Vakhitov, and G. Ferrer, “Ness-st: Detecting good and stable keypoints with a neural stability score and the shi- tomasi detector,” inICCV, 2023, pp. 9578–9588

  41. [49]

    Good features to track,

    J. Shiet al., “Good features to track,” inCVPR, 1994, pp. 593–600

  42. [50]

    Self-supervised equivariant learning for oriented keypoint detection,

    J. Lee, B. Kim, and M. Cho, “Self-supervised equivariant learning for oriented keypoint detection,” inCVPR, 2022, pp. 4847–4857

  43. [51]

    Scale-free image keypoints using differentiable persistent homology,

    G. Barbarani, F. Vaccarino, G. Trivigno, M. Guerra, G. Berton, and C. Masone, “Scale-free image keypoints using differentiable persistent homology,” inICML, 2024, pp. 1–13

  44. [52]

    Pca-sift: A more distinctive representa- tion for local image descriptors,

    Y. Ke and R. Sukthankar, “Pca-sift: A more distinctive representa- tion for local image descriptors,” inCVPR, vol. 2, 2004, pp. II–II

  45. [53]

    Learning linear discrimi- nant projections for dimensionality reduction of image descrip- tors,

    H. Cai, K. Mikolajczyk, and J. Matas, “Learning linear discrimi- nant projections for dimensionality reduction of image descrip- tors,”IEEE TP AMI, vol. 33, no. 2, pp. 338–352, 2010

  46. [54]

    Discriminative learning of local image descriptors,

    M. Brown, G. Hua, and S. Winder, “Discriminative learning of local image descriptors,”IEEE TP AMI, vol. 33, no. 1, pp. 43–57, 2010

  47. [55]

    Discriminant embedding for local image descriptors,

    G. Hua, M. Brown, and S. Winder, “Discriminant embedding for local image descriptors,” inICCV, 2007, pp. 1–8

  48. [56]

    Learning local image descriptors,

    S. A. Winder and M. Brown, “Learning local image descriptors,” inCVPR, 2007, pp. 1–8

  49. [57]

    Ldahash: Im- proved matching with smaller descriptors,

    C. Strecha, A. Bronstein, M. Bronstein, and P . Fua, “Ldahash: Im- proved matching with smaller descriptors,”IEEE TP AMI, vol. 34, no. 1, pp. 66–78, 2011

  50. [58]

    Boosting binary keypoint descriptors,

    T. Trzcinski, M. Christoudias, P . Fua, and V . Lepetit, “Boosting binary keypoint descriptors,” inCVPR, 2013, pp. 2874–2881

  51. [59]

    Learning local feature descriptors using convex optimisation,

    K. Simonyan, A. Vedaldi, and A. Zisserman, “Learning local feature descriptors using convex optimisation,”IEEE TP AMI, vol. 36, no. 8, pp. 1573–1585, 2014

  52. [60]

    Signa- ture verification using a

    J. Bromley, I. Guyon, Y. LeCun, E. S ¨ackinger, and R. Shah, “Signa- ture verification using a” siamese” time delay neural network,” inNeurIPS, vol. 6, 1993, pp. 1–8

  53. [61]

    Learning to compare image patches via convolutional neural networks,

    S. Zagoruyko and N. Komodakis, “Learning to compare image patches via convolutional neural networks,” inCVPR, 2015, pp. 4353–4361

  54. [62]

    Match- net: Unifying feature and metric learning for patch-based match- ing,

    X. Han, T. Leung, Y. Jia, R. Sukthankar, and A. C. Berg, “Match- net: Unifying feature and metric learning for patch-based match- ing,” inCVPR, 2015, pp. 3279–3286

  55. [63]

    Discriminative learning of deep convolu- tional feature point descriptors,

    E. Simo-Serra, E. Trulls, L. Ferraz, I. Kokkinos, P . Fua, and F. Moreno-Noguer, “Discriminative learning of deep convolu- tional feature point descriptors,” inICCV, 2015, pp. 118–126

  56. [64]

    Learning spread- out local feature descriptors,

    X. Zhang, F. X. Yu, S. Kumar, and S.-F. Chang, “Learning spread- out local feature descriptors,” inICCV, 2017, pp. 4595–4603

  57. [65]

    Learning local feature descriptors with triplets and shallow convolutional neural networks,

    V . Balntas, E. Riba, D. Ponsa, and K. Mikolajczyk, “Learning local feature descriptors with triplets and shallow convolutional neural networks,” inBMVC, vol. 1, no. 2, 2016, p. 3

  58. [66]

    Learning local image de- scriptors with deep siamese and triplet convolutional networks by minimising global loss functions,

    V . Kumar BG, G. Carneiro, and I. Reid, “Learning local image de- scriptors with deep siamese and triplet convolutional networks by minimising global loss functions,” inCVPR, 2016, pp. 5385– 5394

  59. [67]

    L2-net: Deep learning of discrimi- native patch descriptor in euclidean space,

    Y. Tian, B. Fan, and F. Wu, “L2-net: Deep learning of discrimi- native patch descriptor in euclidean space,” inCVPR, 2017, pp. 661–669

  60. [68]

    Working hard to know your neighbor’s margins: Local descriptor learning loss,

    A. Mishchuk, D. Mishkin, F. Radenovic, and J. Matas, “Working hard to know your neighbor’s margins: Local descriptor learning loss,” inNeurIPS, vol. 30, 2017, pp. 1–12

  61. [69]

    Sosnet: Second order similarity regularization for local descriptor learn- ing,

    Y. Tian, X. Yu, B. Fan, F. Wu, H. Heijnen, and V . Balntas, “Sosnet: Second order similarity regularization for local descriptor learn- ing,” inCVPR, 2019, pp. 11 016–11 025

  62. [70]

    Hynet: Learning local descriptor with hybrid similarity measure and triplet loss,

    Y. Tian, A. Barroso Laguna, T. Ng, V . Balntas, and K. Mikolajczyk, “Hynet: Learning local descriptor with hybrid similarity measure and triplet loss,” inNeurIPS, vol. 33, 2020, pp. 7401–7412

  63. [71]

    Local descriptors optimized for average precision,

    K. He, Y. Lu, and S. Sclaroff, “Local descriptors optimized for average precision,” inCVPR, 2018, pp. 596–605

  64. [72]

    Geodesc: Learning local descriptors by integrating geometry constraints,

    Z. Luo, T. Shen, L. Zhou, S. Zhu, R. Zhang, Y. Yao, T. Fang, and L. Quan, “Geodesc: Learning local descriptors by integrating geometry constraints,” inECCV, 2018, pp. 168–183

  65. [73]

    Learning feature descriptors using camera pose supervision,

    Q. Wang, X. Zhou, B. Hariharan, and N. Snavely, “Learning feature descriptors using camera pose supervision,” inECCV, 2020, pp. 757–774

  66. [74]

    Steerers: A framework for rotation equivariant keypoint descriptors,

    G. B ¨okman, J. Edstedt, M. Felsberg, and F. Kahl, “Steerers: A framework for rotation equivariant keypoint descriptors,” in CVPR, 2024, pp. 4885–4895

  67. [75]

    Affine steerers for structured keypoint description,

    ——, “Affine steerers for structured keypoint description,” in ECCV, 2025, pp. 449–468

  68. [76]

    Contextdesc: Local descriptor augmentation with cross-modality context,

    Z. Luo, T. Shen, L. Zhou, J. Zhang, Y. Yao, S. Li, T. Fang, and L. Quan, “Contextdesc: Local descriptor augmentation with cross-modality context,” inCVPR, 2019, pp. 2527–2536

  69. [77]

    Repeatability is not enough: Learning affine regions via discriminability,

    D. Mishkin, F. Radenovic, and J. Matas, “Repeatability is not enough: Learning affine regions via discriminability,” inECCV, 2018, pp. 284–300. 20

  70. [78]

    Beyond cartesian representations for local descriptors,

    P . Ebel, A. Mishchuk, K. M. Yi, P . Fua, and E. Trulls, “Beyond cartesian representations for local descriptors,” inICCV, 2019, pp. 253–262

  71. [79]

    Gift: Learning transformation-invariant dense visual descriptors via group cnns,

    Y. Liu, Z. Shen, Z. Lin, S. Peng, H. Bao, and X. Zhou, “Gift: Learning transformation-invariant dense visual descriptors via group cnns,” inNeurIPS, vol. 32, 2019, pp. 1–12

  72. [80]

    Group equivariant convolutional networks,

    T. Cohen and M. Welling, “Group equivariant convolutional networks,” inICML, 2016, pp. 2990–2999

  73. [81]

    Learning rotation- equivariant features for visual correspondence,

    J. Lee, B. Kim, S. Kim, and M. Cho, “Learning rotation- equivariant features for visual correspondence,” inCVPR, 2023, pp. 21 887–21 897

  74. [82]

    General e (2)-equivariant steerable cnns,

    M. Weiler and G. Cesa, “General e (2)-equivariant steerable cnns,” inNeurIPS, vol. 32, 2019, pp. 1–12

  75. [83]

    Online invariance selection for local feature descriptors,

    R. Pautrat, V . Larsson, M. R. Oswald, and M. Pollefeys, “Online invariance selection for local feature descriptors,” inECCV, 2020, pp. 707–724

  76. [84]

    Lift: Learned invariant feature transform,

    K. M. Yi, E. Trulls, V . Lepetit, and P . Fua, “Lift: Learned invariant feature transform,” inECCV, 2016, pp. 467–483

  77. [85]

    Lf-net: Learning local features from images,

    Y. Ono, E. Trulls, P . Fua, and K. M. Yi, “Lf-net: Learning local features from images,” inNeurIPS, vol. 31, 2018, pp. 1–13

  78. [86]

    Spatial trans- former networks,

    M. Jaderberg, K. Simonyan, A. Zissermanet al., “Spatial trans- former networks,” inNeurIPS, vol. 28, 2015, pp. 1–9

  79. [87]

    Rf-net: An end-to-end image matching network based on receptive field,

    X. Shen, C. Wang, X. Li, Z. Yu, J. Li, C. Wen, M. Cheng, and Z. He, “Rf-net: An end-to-end image matching network based on receptive field,” inCVPR, 2019, pp. 8132–8140

  80. [88]

    Alike: Accurate and lightweight keypoint detection and descriptor ex- traction,

    X. Zhao, X. Wu, J. Miao, W. Chen, P . C. Chen, and Z. Li, “Alike: Accurate and lightweight keypoint detection and descriptor ex- traction,”IEEE TMM, vol. 25, pp. 3101–3112, 2022

  81. [89]

    Aliked: A lighter keypoint and descriptor extraction network via de- formable transformation,

    X. Zhao, X. Wu, W. Chen, P . C. Chen, Q. Xu, and Z. Li, “Aliked: A lighter keypoint and descriptor extraction network via de- formable transformation,”IEEE TIM, vol. 72, pp. 1–16, 2023

  82. [90]

    D2-net: A trainable cnn for joint description and detection of local features,

    M. Dusmanu, I. Rocco, T. Pajdla, M. Pollefeys, J. Sivic, A. Torii, and T. Sattler, “D2-net: A trainable cnn for joint description and detection of local features,” inCVPR, 2019, pp. 8092–8101

  83. [91]

    Aslfeat: Learning local features of accurate shape and localization,

    Z. Luo, L. Zhou, X. Bai, H. Chen, J. Zhang, Y. Yao, S. Li, T. Fang, and L. Quan, “Aslfeat: Learning local features of accurate shape and localization,” inCVPR, 2020, pp. 6589–6598

  84. [92]

    Deformable convolutional networks,

    J. Dai, H. Qi, Y. Xiong, Y. Li, G. Zhang, H. Hu, and Y. Wei, “Deformable convolutional networks,” inICML, 2017, pp. 764– 773

  85. [93]

    Redfeat: Recoupling detection and descrip- tion for multimodal feature learning,

    Y. Deng and J. Ma, “Redfeat: Recoupling detection and descrip- tion for multimodal feature learning,”IEEE TIP, vol. 32, pp. 591– 602, 2022

  86. [94]

    Disk: Learning local features with policy gradient,

    M. Tyszkiewicz, P . Fua, and E. Trulls, “Disk: Learning local features with policy gradient,” inNeurIPS, vol. 33, 2020, pp. 14 254–14 265

  87. [95]

    R. S. Sutton and A. G. Barto,Reinforcement learning: An introduc- tion. MIT press, 2018

  88. [96]

    R2d2: Reliable and repeatable detector and descriptor,

    J. Revaud, C. De Souza, M. Humenberger, and P . Weinzaepfel, “R2d2: Reliable and repeatable detector and descriptor,” in NeurIPS, vol. 32, 2019, pp. 1–11

  89. [97]

    Sfd2: Semantic-guided fea- ture detection and description,

    F. Xue, I. Budvytis, and R. Cipolla, “Sfd2: Semantic-guided fea- ture detection and description,” inCVPR, 2023, pp. 5206–5216

  90. [98]

    Dedode: Detect, don’t describe—describe, don’t detect for local feature matching,

    J. Edstedt, G. B ¨okman, M. Wadenb¨ack, and M. Felsberg, “Dedode: Detect, don’t describe—describe, don’t detect for local feature matching,” in3DV, 2024, pp. 148–157

  91. [99]

    Dedode v2: Analyzing and improving the dedode keypoint detector,

    J. Edstedt, G. B ¨okman, and Z. Zhao, “Dedode v2: Analyzing and improving the dedode keypoint detector,” inCVPRW, 2024, pp. 4245–4253

  92. [100]

    Xfeat: Accelerated features for lightweight image matching,

    G. Potje, F. Cadar, A. Araujo, R. Martins, and E. R. Nascimento, “Xfeat: Accelerated features for lightweight image matching,” in CVPR, 2024, pp. 2682–2691

  93. [101]

    Learning to make keypoints sub-pixel accurate,

    S. Kim, M. Pollefeys, and D. Barath, “Learning to make keypoints sub-pixel accurate,” inECCV, 2024, pp. 413–431

  94. [102]

    Pointnet: Deep learning on point sets for 3d classification and segmentation,

    C. R. Qi, H. Su, K. Mo, and L. J. Guibas, “Pointnet: Deep learning on point sets for 3d classification and segmentation,” inCVPR, 2017, pp. 652–660

  95. [103]

    Lmr: Learning a two-class classifier for mismatch removal,

    J. Ma, X. Jiang, J. Jiang, J. Zhao, and X. Guo, “Lmr: Learning a two-class classifier for mismatch removal,”IEEE TIP, vol. 28, no. 8, pp. 4045–4059, 2019

  96. [104]

    Learning two-view correspondences and geometry using order-aware network,

    J. Zhang, D. Sun, Z. Luo, A. Yao, L. Zhou, T. Shen, Y. Chen, L. Quan, and H. Liao, “Learning two-view correspondences and geometry using order-aware network,” inICCV, 2019, pp. 5845– 5854

  97. [105]

    Oanet: Learning two-view cor- respondences and geometry using order-aware network,

    J. Zhang, D. Sun, Z. Luo, A. Yao, H. Chen, L. Zhou, T. Shen, Y. Chen, L. Quan, and H. Liao, “Oanet: Learning two-view cor- respondences and geometry using order-aware network,”IEEE TP AMI, vol. 44, no. 6, pp. 3110–3122, 2020

  98. [106]

    Acne: Attentive context normalization for robust permutation- equivariant learning,

    W. Sun, W. Jiang, E. Trulls, A. Tagliasacchi, and K. M. Yi, “Acne: Attentive context normalization for robust permutation- equivariant learning,” inCVPR, 2020, pp. 11 286–11 295

  99. [107]

    T-net++: Effective permutation-equivariance network for two-view corre- spondence pruning,

    G. Xiao, X. Liu, Z. Zhong, X. Zhang, J. Ma, and H. Ling, “T-net++: Effective permutation-equivariance network for two-view corre- spondence pruning,”IEEE TP AMI, vol. 46, no. 12, pp. 10 629– 10 644, 2024

  100. [108]

    Learnable motion coherence for correspondence pruning,

    Y. Liu, L. Liu, C. Lin, Z. Dong, and W. Wang, “Learnable motion coherence for correspondence pruning,” inCVPR, 2021, pp. 3237– 3246

  101. [109]

    Nm-net: Mining reliable neighbors for robust feature correspondences,

    C. Zhao, Z. Cao, C. Li, X. Li, and J. Yang, “Nm-net: Mining reliable neighbors for robust feature correspondences,” inCVPR, 2019, pp. 215–224

  102. [110]

    Progressive correspondence pruning by consensus learning,

    C. Zhao, Y. Ge, F. Zhu, R. Zhao, H. Li, and M. Salzmann, “Progressive correspondence pruning by consensus learning,” in ICCV, 2021, pp. 6464–6473

  103. [111]

    Semi-supervised classification with graph convolutional networks,

    T. N. Kipf and M. Welling, “Semi-supervised classification with graph convolutional networks,” inICLR, 2017, pp. 1–14

  104. [112]

    Ms2dg-net: Progressive correspondence learning via multiple sparse semantics dynamic graph,

    L. Dai, Y. Liu, J. Ma, L. Wei, T. Lai, C. Yang, and R. Chen, “Ms2dg-net: Progressive correspondence learning via multiple sparse semantics dynamic graph,” inCVPR, 2022, pp. 8973–8982

  105. [113]

    Progressive neighbor consistency mining for correspondence pruning,

    X. Liu and J. Yang, “Progressive neighbor consistency mining for correspondence pruning,” inCVPR, 2023, pp. 9527–9537

  106. [114]

    Mgnet: Learning correspondences via multiple graphs,

    D. Luanyuan, X. Du, H. Zhang, and J. Tang, “Mgnet: Learning correspondences via multiple graphs,” inAAAI, vol. 38, no. 4, 2024, pp. 3945–3953

  107. [115]

    Attention is all you need,

    A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” inNeurIPS, vol. 30, 2017, pp. 1–11

  108. [116]

    Learning for mismatch removal via graph attention networks,

    X. Jiang, Y. Wang, A. Fan, and J. Ma, “Learning for mismatch removal via graph attention networks,”ISPRS P&RS, vol. 190, pp. 181–195, 2022

  109. [117]

    Learning second-order attentive context for efficient correspondence pruning,

    X. Ye, W. Zhao, H. Lu, and Z. Cao, “Learning second-order attentive context for efficient correspondence pruning,” inAAAI, vol. 37, no. 3, 2023, pp. 3250–3258

  110. [118]

    U-match: two-view correspondence learning with hierarchy-aware local context aggregation,

    Z. Li, S. Zhang, and J. Ma, “U-match: two-view correspondence learning with hierarchy-aware local context aggregation,” inIJ- CAI, 2023, pp. 1169–1176

  111. [119]

    Graph u-nets,

    H. Gao and S. Ji, “Graph u-nets,” inICML, 2019, pp. 2083–2092

  112. [120]

    U-match: Exploring hierarchy- aware local context for two-view correspondence learning,

    Z. Li, S. Zhang, and J. Ma, “U-match: Exploring hierarchy- aware local context for two-view correspondence learning,”IEEE TP AMI, vol. 46, no. 12, pp. 10 960–10 977, 2024

  113. [121]

    Bclnet: Bilateral consensus learning for two-view correspondence pruning,

    X. Miao, G. Xiao, S. Wang, and J. Yu, “Bclnet: Bilateral consensus learning for two-view correspondence pruning,” inAAAI, vol. 38, no. 5, 2024, pp. 4225–4232

  114. [122]

    Vsformer: Visual-spatial fusion transformer for correspondence pruning,

    T. Liao, X. Zhang, L. Zhao, T. Wang, and G. Xiao, “Vsformer: Visual-spatial fusion transformer for correspondence pruning,” inAAAI, vol. 38, no. 4, 2024, pp. 3369–3377

  115. [123]

    Convmatch: Rethinking network design for two-view correspondence learning,

    S. Zhang and J. Ma, “Convmatch: Rethinking network design for two-view correspondence learning,” inAAAI, 2023, pp. 3472– 3479

  116. [124]

    Convmatch: Rethinking network design for two-view correspondence learning,

    ——, “Convmatch: Rethinking network design for two-view correspondence learning,”IEEE TP AMI, vol. 46, no. 5, pp. 2920– 2935, 2024

  117. [125]

    Demo: Deep motion field consensus with learnable kernels for two-view correspondence learning,

    Y. Lu, J. Le, Z. Li, Y. Yuan, and J. Ma, “Demo: Deep motion field consensus with learnable kernels for two-view correspondence learning,” inAAAI, 2025, pp. 1–9

  118. [126]

    Dematch: Deep decompo- sition of motion field for two-view correspondence learning,

    S. Zhang, Z. Li, Y. Gao, and J. Ma, “Dematch: Deep decompo- sition of motion field for two-view correspondence learning,” in CVPR, 2024, pp. 20 278–20 287

  119. [127]

    Deep fundamental matrix estimation,

    R. Ranftl and V . Koltun, “Deep fundamental matrix estimation,” inECCV, 2018, pp. 284–299

  120. [128]

    Dsac-differentiable ransac for cam- era localization,

    E. Brachmann, A. Krull, S. Nowozin, J. Shotton, F. Michel, S. Gumhold, and C. Rother, “Dsac-differentiable ransac for cam- era localization,” inCVPR, 2017, pp. 6684–6692

  121. [129]

    A survey on monocular re- localization: From the perspective of scene map representation,

    J. Miao, K. Jiang, T. Wen, Y. Wang, P . Jia, B. Wijaya, X. Zhao, Q. Cheng, Z. Xiao, J. Huanget al., “A survey on monocular re- localization: From the perspective of scene map representation,” IEEE TIV, pp. 1–33, 2024

  122. [130]

    Neural-guided ransac: Learning where to sample model hypotheses,

    E. Brachmann and C. Rother, “Neural-guided ransac: Learning where to sample model hypotheses,” inICCV, 2019, pp. 4322– 4331

  123. [131]

    Adaptive reordering sampler with neurally guided magsac,

    T. Wei, J. Matas, and D. Barath, “Adaptive reordering sampler with neurally guided magsac,” inICCV, 2023, pp. 18 163–18 173. 21

  124. [132]

    Bansac: A dynamic bayesian network for adaptive sample consensus,

    V . Piedade and P . Miraldo, “Bansac: A dynamic bayesian network for adaptive sample consensus,” inICCV, 2023, pp. 3738–3747

  125. [133]

    Gener- alized differentiable ransac,

    T. Wei, Y. Patel, A. Shekhovtsov, J. Matas, and D. Barath, “Gener- alized differentiable ransac,” inICCV, 2023, pp. 17 649–17 660

  126. [134]

    Categorical reparametrization with gumble-softmax,

    E. Jang, S. Gu, and B. Poole, “Categorical reparametrization with gumble-softmax,” inICLR, 2017, pp. 1–12

  127. [135]

    Learning to find good models in ransac,

    D. Barath, L. Cavalli, and M. Pollefeys, “Learning to find good models in ransac,” inCVPR, 2022, pp. 15 744–15 753

  128. [136]

    Nefsac: Neurally filtered minimal samples,

    L. Cavalli, M. Pollefeys, and D. Barath, “Nefsac: Neurally filtered minimal samples,” inECCV, 2022, pp. 351–366

  129. [137]

    Two-view geometry scoring without correspondences,

    A. Barroso-Laguna, E. Brachmann, V . A. Prisacariu, G. J. Brostow, and D. Turmukhambetov, “Two-view geometry scoring without correspondences,” inCVPR, 2023, pp. 8979–8989

  130. [138]

    Un- supervised learning of consensus maximization for 3d vision problems,

    T. Probst, D. P . Paudel, A. Chhatkuli, and L. V . Gool, “Un- supervised learning of consensus maximization for 3d vision problems,” inCVPR, 2019, pp. 929–938

  131. [139]

    Fast and accurate matrix completion via truncated nuclear norm regularization,

    Y. Hu, D. Zhang, J. Ye, X. Li, and X. He, “Fast and accurate matrix completion via truncated nuclear norm regularization,” IEEE TP AMI, vol. 35, no. 9, pp. 2117–2130, 2012

  132. [140]

    Un- supervised learning for robust fitting: A reinforcement learning approach,

    G. Truong, H. Le, D. Suter, E. Zhang, and S. Z. Gilani, “Un- supervised learning for robust fitting: A reinforcement learning approach,” inCVPR, 2021, pp. 10 348–10 357

  133. [141]

    Playing atari with deep reinforcement learning,

    V . Mnih, “Playing atari with deep reinforcement learning,” in NeurIPSW, 2013, pp. 1–9

  134. [142]

    Dynamic graph cnn for learning on point clouds,

    Y. Wang, Y. Sun, Z. Liu, S. E. Sarma, M. M. Bronstein, and J. M. Solomon, “Dynamic graph cnn for learning on point clouds,” ACM TOG, vol. 38, no. 5, pp. 1–12, 2019

  135. [143]

    Un- supervised learning for maximum consensus robust fitting: A reinforcement learning approach,

    G. Truong, H. Le, E. Zhang, D. Suter, and S. Z. Gilani, “Un- supervised learning for maximum consensus robust fitting: A reinforcement learning approach,”IEEE TP AMI, vol. 45, no. 3, pp. 3890–3903, 2022

  136. [144]

    Rlsac: Reinforcement learning enhanced sample consensus for end-to-end robust estimation,

    C. Nie, G. Wang, Z. Liu, L. Cavalli, M. Pollefeys, and H. Wang, “Rlsac: Reinforcement learning enhanced sample consensus for end-to-end robust estimation,” inICCV, 2023, pp. 9891–9900

  137. [145]

    Superglue: Learning feature matching with graph neural net- works,

    P .-E. Sarlin, D. DeTone, T. Malisiewicz, and A. Rabinovich, “Superglue: Learning feature matching with graph neural net- works,” inCVPR, 2020, pp. 4938–4947

  138. [146]

    Learning to match features with seeded graph matching network,

    H. Chen, Z. Luo, J. Zhang, L. Zhou, X. Bai, Z. Hu, C.-L. Tai, and L. Quan, “Learning to match features with seeded graph matching network,” inICCV, 2021, pp. 6301–6310

  139. [147]

    Clustergnn: Cluster-based coarse-to-fine graph neural network for efficient feature matching,

    Y. Shi, J.-X. Cai, Y. Shavit, T.-J. Mu, W. Feng, and K. Zhang, “Clustergnn: Cluster-based coarse-to-fine graph neural network for efficient feature matching,” inCVPR, 2022, pp. 12 517–12 526

  140. [148]

    Lightglue: Local feature matching at light speed,

    P . Lindenberger, P .-E. Sarlin, and M. Pollefeys, “Lightglue: Local feature matching at light speed,” inICCV, 2023, pp. 17 627–17 638

  141. [149]

    Om- niglue: Generalizable feature matching with foundation model guidance,

    H. Jiang, A. Karpur, B. Cao, Q. Huang, and A. Araujo, “Om- niglue: Generalizable feature matching with foundation model guidance,” inCVPR, 2024, pp. 19 865–19 875

  142. [150]

    A comprehensive survey on graph neural networks,

    Z. Wu, S. Pan, F. Chen, G. Long, C. Zhang, and S. Y. Philip, “A comprehensive survey on graph neural networks,”IEEE TNNLS, vol. 32, no. 1, pp. 4–24, 2020

  143. [151]

    Sinkhorn distances: Lightspeed computation of opti- mal transport,

    M. Cuturi, “Sinkhorn distances: Lightspeed computation of opti- mal transport,” inNeurIPS, vol. 26, 2013, pp. 1–9

  144. [152]

    Paraformer: Parallel attention transformer for efficient feature matching,

    X. Lu, Y. Yan, B. Kang, and S. Du, “Paraformer: Parallel attention transformer for efficient feature matching,” inAAAI, vol. 37, no. 2, 2023, pp. 1853–1860

  145. [153]

    Imp: Iterative matching and pose estimation with adaptive pooling,

    F. Xue, I. Budvytis, and R. Cipolla, “Imp: Iterative matching and pose estimation with adaptive pooling,” inCVPR, 2023, pp. 21 317–21 326

  146. [154]

    Roformer: Enhanced transformer with rotary position embedding,

    J. Su, M. Ahmed, Y. Lu, S. Pan, W. Bo, and Y. Liu, “Roformer: Enhanced transformer with rotary position embedding,”Neuro- computing, vol. 568, p. 127063, 2024

  147. [155]

    Learning feature matching via matchable keypoint-assisted graph neural network,

    Z. Li and J. Ma, “Learning feature matching via matchable keypoint-assisted graph neural network,”IEEE TIP, vol. 34, pp. 154–169, 2025

  148. [156]

    Mambaglue: Fast and robust local feature matching with mamba,

    K. Ryoo, H. Lim, and H. Myung, “Mambaglue: Fast and robust local feature matching with mamba,” inICRA, 2025, pp. 1–8

  149. [157]

    Mamba: Linear-time sequence modeling with selective state spaces,

    A. Gu and T. Dao, “Mamba: Linear-time sequence modeling with selective state spaces,” inCOLM, 2024, pp. 1–32

  150. [158]

    Scene-aware feature matching,

    X. Lu, Y. Yan, T. Wei, and S. Du, “Scene-aware feature matching,” inICCV, 2023, pp. 3704–3713

  151. [159]

    Resmatch: Resid- ual attention learning for feature matching,

    Y. Deng, K. Zhang, S. Zhang, Y. Li, and J. Ma, “Resmatch: Resid- ual attention learning for feature matching,” inAAAI, vol. 38, no. 2, 2024, pp. 1501–1509

  152. [160]

    Dinov2: Learning robust visual features without supervision,

    M. Oquab, T. Darcet, T. Moutakanni, H. Vo, M. Szafraniec, V . Khalidov, P . Fernandez, D. Haziza, F. Massa, A. El-Noubyet al., “Dinov2: Learning robust visual features without supervision,” TMLR, pp. 1–31, 2024

  153. [161]

    Matching while perceiv- ing: Enhance image feature matching with applicable semantic amalgamation,

    S. Zhang, Z. Zhu, Z. Li, T. Lu, and J. Ma, “Matching while perceiv- ing: Enhance image feature matching with applicable semantic amalgamation,” inAAAI, vol. 39, no. 10, 2025, pp. 10 094–10 102

  154. [162]

    Segnext: Rethinking convolutional attention design for semantic segmentation,

    M.-H. Guo, C.-Z. Lu, Q. Hou, Z. Liu, M.-M. Cheng, and S.-M. Hu, “Segnext: Rethinking convolutional attention design for semantic segmentation,” inNeurIPS, vol. 35, 2022, pp. 1140–1156

  155. [163]

    Diffglue: Diffusion-aided image feature matching,

    S. Zhang and J. Ma, “Diffglue: Diffusion-aided image feature matching,” inACM MM, 2024, pp. 8451–8460

  156. [164]

    Diffusion models in vision: A survey,

    F.-A. Croitoru, V . Hondru, R. T. Ionescu, and M. Shah, “Diffusion models in vision: A survey,”IEEE TP AMI, vol. 45, no. 9, pp. 10 850–10 869, 2023

  157. [165]

    Matchanything: Universal cross-modality image matching with large-scale pre-training,

    X. He, H. Yu, S. Peng, D. Tan, Z. Shen, H. Bao, and X. Zhou, “Matchanything: Universal cross-modality image matching with large-scale pre-training,”arXiv:2501.07556, 2025

  158. [166]

    Neighbourhood consensus networks,

    I. Rocco, M. Cimpoi, R. Arandjelovi ´c, A. Torii, T. Pajdla, and J. Sivic, “Neighbourhood consensus networks,” inNeurIPS, vol. 31, 2018, pp. 1–12

  159. [167]

    Efficient neighbourhood consensus networks via submanifold sparse convolutions,

    I. Rocco, R. Arandjelovi ´c, and J. Sivic, “Efficient neighbourhood consensus networks via submanifold sparse convolutions,” in ECCV, 2020, pp. 605–621

  160. [168]

    Dual-resolution correspon- dence networks,

    X. Li, K. Han, S. Li, and V . Prisacariu, “Dual-resolution correspon- dence networks,” inNeurIPS, vol. 33, 2020, pp. 17 346–17 357

  161. [169]

    Dualrc: A dual-resolution learning framework with neigh- bourhood consensus for visual correspondences,

    ——, “Dualrc: A dual-resolution learning framework with neigh- bourhood consensus for visual correspondences,”IEEE TP AMI, vol. 46, no. 1, pp. 236–249, 2024

  162. [170]

    Efficient dynamic correspondence network,

    J. He, T. Zhang, Z. Zhang, T. Yu, and Y. Zhang, “Efficient dynamic correspondence network,”IEEE TIP, vol. 33, pp. 228–240, 2024

  163. [171]

    Loftr: Detector- free local feature matching with transformers,

    J. Sun, Z. Shen, Y. Wang, H. Bao, and X. Zhou, “Loftr: Detector- free local feature matching with transformers,” inCVPR, 2021, pp. 8922–8931

  164. [172]

    Feature pyramid networks for object detection,

    T.-Y. Lin, P . Doll ´ar, R. Girshick, K. He, B. Hariharan, and S. Belongie, “Feature pyramid networks for object detection,” in CVPR, 2017, pp. 2117–2125

  165. [173]

    Trans- formers are rnns: Fast autoregressive transformers with linear attention,

    A. Katharopoulos, A. Vyas, N. Pappas, and F. Fleuret, “Trans- formers are rnns: Fast autoregressive transformers with linear attention,” inICML, 2020, pp. 5156–5165

  166. [174]

    Image transformer,

    N. Parmar, A. Vaswani, J. Uszkoreit, L. Kaiser, N. Shazeer, A. Ku, and D. Tran, “Image transformer,” inICML, 2018, pp. 4055–4064

  167. [175]

    Matchformer: Interleaving attention in transformers for feature matching,

    Q. Wang, J. Zhang, K. Yang, K. Peng, and R. Stiefelhagen, “Matchformer: Interleaving attention in transformers for feature matching,” inACCV, 2022, pp. 2746–2762

  168. [176]

    Aspanformer: Detector-free image matching with adaptive span transformer,

    H. Chen, Z. Luo, L. Zhou, Y. Tian, M. Zhen, T. Fang, D. Mck- innon, Y. Tsin, and L. Quan, “Aspanformer: Detector-free image matching with adaptive span transformer,” inECCV, 2022, pp. 20–36

  169. [177]

    Affine-based deformable attention and selective fusion for semi-dense matching,

    H. Chen, Z. Luo, Y. Tian, X. Bai, Z. Wang, L. Zhou, M. Zhen, T. Fang, D. Mckinnon, Y. Tsinet al., “Affine-based deformable attention and selective fusion for semi-dense matching,” in CVPRW, 2024, pp. 4254–4263

  170. [178]

    3dg-stfm: 3d geometric guided student-teacher feature matching,

    R. Mao, C. Bai, Y. An, F. Zhu, and C. Lu, “3dg-stfm: 3d geometric guided student-teacher feature matching,” inECCV, 2022, pp. 125–142

  171. [179]

    Guiding local feature matching with surface curvature,

    S. Wang, J. Kannala, M. Pollefeys, and D. Barath, “Guiding local feature matching with surface curvature,” inICCV, 2023, pp. 17 981–17 991

  172. [180]

    Vision transformers for dense prediction,

    R. Ranftl, A. Bochkovskiy, and V . Koltun, “Vision transformers for dense prediction,” inICCV, 2021, pp. 12 179–12 188

  173. [181]

    Topicfm+: Boosting accuracy and efficiency of topic-assisted feature matching,

    K. T. Giang, S. Song, and S. Jo, “Topicfm+: Boosting accuracy and efficiency of topic-assisted feature matching,”IEEE TIP, vol. 33, pp. 6016–6028, 2024

  174. [182]

    Improving transformer-based image matching by cascaded capturing spatially informative keypoints,

    C. Cao and Y. Fu, “Improving transformer-based image matching by cascaded capturing spatially informative keypoints,” inICCV, 2023, pp. 12 129–12 139

  175. [183]

    Adaptive assignment for geometry aware local feature matching,

    D. Huang, Y. Chen, Y. Liu, J. Liu, S. Xu, W. Wu, Y. Ding, F. Tang, and C. Wang, “Adaptive assignment for geometry aware local feature matching,” inCVPR, 2023, pp. 5425–5434

  176. [184]

    Pats: Patch area transportation with subdivision for local feature matching,

    J. Ni, Y. Li, Z. Huang, H. Li, H. Bao, Z. Cui, and G. Zhang, “Pats: Patch area transportation with subdivision for local feature matching,” inCVPR, 2023, pp. 17 776–17 786

  177. [185]

    Adaptive spot-guided transformer for consistent local feature matching,

    J. Yu, J. Chang, J. He, T. Zhang, J. Yu, and F. Wu, “Adaptive spot-guided transformer for consistent local feature matching,” inCVPR, 2023, pp. 21 898–21 908

  178. [186]

    Prism: Progressive dependency maximization for scale- invariant image matching,

    X. Cai, Y. Wang, L. Luo, M. Wang, D. Li, J. Xu, W. Gu, and R. Ai, “Prism: Progressive dependency maximization for scale- invariant image matching,” inACM MM, 2024, pp. 5250–5259. 22

  179. [187]

    Homomatcher: Achieving dense feature matching with semi-dense efficiency by homography estima- tion,

    X. Wang, L. Yu, Y. Zhang, J. Lao, L. Ru, L. Zhong, J. Chen, Y. Zhang, and M. Yang, “Homomatcher: Achieving dense feature matching with semi-dense efficiency by homography estima- tion,” inAAAI, vol. 39, no. 8, 2025, pp. 7952–7960

  180. [188]

    Quadtree attention for vision transformers,

    S. Tang, J. Zhang, S. Zhu, and P . Tan, “Quadtree attention for vision transformers,” inICML, 2022, pp. 1–16

  181. [189]

    Topicfm: Robust and interpretable topic-assisted feature matching,

    K. T. Giang, S. Song, and S. Jo, “Topicfm: Robust and interpretable topic-assisted feature matching,” inAAAI, vol. 37, no. 2, 2023, pp. 2447–2455

  182. [190]

    Ecomatcher: Efficient clustering oriented matcher for detector-free image matching,

    P . Chen, L. Yu, Y. Wan, Y. Zhang, J. Wang, L. Zhong, J. Chen, and M. Yang, “Ecomatcher: Efficient clustering oriented matcher for detector-free image matching,” inECCV, 2024, pp. 344–360

  183. [191]

    Efficient loftr: Semi-dense local feature matching with sparse-like speed,

    Y. Wang, X. He, S. Peng, D. Tan, and X. Zhou, “Efficient loftr: Semi-dense local feature matching with sparse-like speed,” in CVPR, 2024, pp. 21 666–21 675

  184. [192]

    Repvgg: Making vgg-style convnets great again,

    X. Ding, X. Zhang, N. Ma, J. Han, G. Ding, and J. Sun, “Repvgg: Making vgg-style convnets great again,” inCVPR, 2021, pp. 13 733–13 742

  185. [193]

    ETO:efficient transformer-based local feature matching by or- ganizing multiple homography hypotheses,

    J. Ni, G. Zhang, G. Li, Y. Li, X. Liu, Z. Huang, and H. Bao, “ETO:efficient transformer-based local feature matching by or- ganizing multiple homography hypotheses,” inNeurIPS, 2024, pp. 1–13

  186. [194]

    Jamma: Ultra-lightweight local feature match- ing with joint mamba,

    X. Lu and S. Du, “Jamma: Ultra-lightweight local feature match- ing with joint mamba,” inCVPR, 2025, pp. 1–8

  187. [195]

    Flownet: Learning optical flow with convolutional networks,

    A. Dosovitskiy, P . Fischer, E. Ilg, P . Hausser, C. Hazirbas, V . Golkov, P . Van Der Smagt, D. Cremers, and T. Brox, “Flownet: Learning optical flow with convolutional networks,” inICCV, 2015, pp. 2758–2766

  188. [196]

    Dgc-net: Dense geometric correspondence network,

    I. Melekhov, A. Tiulpin, T. Sattler, M. Pollefeys, E. Rahtu, and J. Kannala, “Dgc-net: Dense geometric correspondence network,” inWACV, 2019, pp. 1034–1042

  189. [197]

    Glu-net: Global-local universal network for dense flow and correspondences,

    P . Truong, M. Danelljan, and R. Timofte, “Glu-net: Global-local universal network for dense flow and correspondences,” in CVPR, 2020, pp. 6258–6268

  190. [198]

    Gocor: Bringing globally optimized correspondence volumes into your neural network,

    P . Truong, M. Danelljan, L. V . Gool, and R. Timofte, “Gocor: Bringing globally optimized correspondence volumes into your neural network,” inNeurIPS, vol. 33, 2020, pp. 14 278–14 290

  191. [199]

    Ransac-flow: generic two-stage image alignment,

    X. Shen, F. Darmon, A. A. Efros, and M. Aubry, “Ransac-flow: generic two-stage image alignment,” inECCV, 2020, pp. 618–637

  192. [200]

    Warp consis- tency for unsupervised learning of dense correspondences,

    P . Truong, M. Danelljan, F. Yu, and L. Van Gool, “Warp consis- tency for unsupervised learning of dense correspondences,” in ICCV, 2021, pp. 10 346–10 356

  193. [201]

    Learning accurate dense correspondences and when to trust them,

    P . Truong, M. Danelljan, L. Van Gool, and R. Timofte, “Learning accurate dense correspondences and when to trust them,” in CVPR, 2021, pp. 5714–5724

  194. [202]

    Dkm: Dense kernelized feature matching for geometry estima- tion,

    J. Edstedt, I. Athanasiadis, M. Wadenb ¨ack, and M. Felsberg, “Dkm: Dense kernelized feature matching for geometry estima- tion,” inCVPR, 2023, pp. 17 765–17 775

  195. [203]

    Pmatch: Paired masked image modeling for dense geometric matching,

    S. Zhu and X. Liu, “Pmatch: Paired masked image modeling for dense geometric matching,” inCVPR, 2023, pp. 21 909–21 918

  196. [204]

    Roma: Robust dense feature matching,

    J. Edstedt, Q. Sun, G. B ¨okman, M. Wadenb¨ack, and M. Felsberg, “Roma: Robust dense feature matching,” inCVPR, 2024, pp. 19 790–19 800

  197. [205]

    Learning affine correspondences by integrating geometric constraints,

    P . Sun, B. Guan, Z. Yu, Y. Shang, Q. Yu, and D. Barath, “Learning affine correspondences by integrating geometric constraints,” in CVPR, 2025, pp. 1–8

  198. [206]

    Cotr: Correspondence transformer for matching across images,

    W. Jiang, E. Trulls, J. Hosang, A. Tagliasacchi, and K. M. Yi, “Cotr: Correspondence transformer for matching across images,” inICCV, 2021, pp. 6207–6217

  199. [207]

    Eco-tr: Efficient correspondences finding via coarse-to-fine refinement,

    D. Tan, J.-J. Liu, X. Chen, C. Chen, R. Zhang, Y. Shen, S. Ding, and R. Ji, “Eco-tr: Efficient correspondences finding via coarse-to-fine refinement,” inECCV, 2022, pp. 317–334

  200. [208]

    Iterative deep homography estimation,

    S.-Y. Cao, J. Hu, Z. Sheng, and H.-L. Shen, “Iterative deep homography estimation,” inCVPR, 2022, pp. 1879–1888

  201. [209]

    Unsupervised deep homography: A fast and robust homography estimation model,

    T. Nguyen, S. W. Chen, S. S. Shivakumar, C. J. Taylor, and V . Kumar, “Unsupervised deep homography: A fast and robust homography estimation model,”RA-L, vol. 3, no. 3, pp. 2346– 2353, 2018

  202. [210]

    Content-aware unsupervised deep homography estimation and its extensions,

    S. Liu, N. Ye, C. Wang, J. Zhang, L. Jia, K. Luo, J. Wang, and J. Sun, “Content-aware unsupervised deep homography estimation and its extensions,”IEEE TP AMI, vol. 45, no. 3, pp. 2849–2863, 2022

  203. [211]

    Adapting dense match- ing for homography estimation with grid-based acceleration,

    K. Zhang, Y. Deng, J. Ma, and P . Favaro, “Adapting dense match- ing for homography estimation with grid-based acceleration,” in CVPR, 2025, pp. 1–8

  204. [212]

    Dmhomo: Learning homography with diffusion models,

    H. Li, H. Jiang, A. Luo, P . Tan, H. Fan, B. Zeng, and S. Liu, “Dmhomo: Learning homography with diffusion models,”ACM TOG, vol. 43, no. 3, pp. 1–16, 2024

  205. [213]

    Deep image homography estimation,

    D. DeTone, T. Malisiewicz, and A. Rabinovich, “Deep image homography estimation,”arXiv:1606.03798, pp. 1–6, 2016

  206. [214]

    Deep homography estimation for dynamic scenes,

    H. Le, F. Liu, S. Zhang, and A. Agarwala, “Deep homography estimation for dynamic scenes,” inCVPR, 2020, pp. 7652–7661

  207. [215]

    Recurrent homography estimation using homography-guided image warping and focus transformer,

    S.-Y. Cao, R. Zhang, L. Luo, B. Yu, Z. Sheng, J. Li, and H.-L. Shen, “Recurrent homography estimation using homography-guided image warping and focus transformer,” inCVPR, 2023, pp. 9833– 9842

  208. [216]

    Deep lucas-kanade homogra- phy for multimodal image alignment,

    Y. Zhao, X. Huang, and Z. Zhang, “Deep lucas-kanade homogra- phy for multimodal image alignment,” inCVPR, 2021, pp. 15 950– 15 959

  209. [217]

    Sparse-to-dense multimodal image regis- tration via multi-task learning,

    K. Zhang and J. Ma, “Sparse-to-dense multimodal image regis- tration via multi-task learning,” inICML, 2024, pp. 1–15

  210. [218]

    Lucas-kanade 20 years on: A unifying framework,

    S. Baker and I. Matthews, “Lucas-kanade 20 years on: A unifying framework,”IJCV, vol. 56, pp. 221–255, 2004

  211. [219]

    Unsuper- vised global and local homography estimation with motion basis learning,

    S. Liu, Y. Lu, H. Jiang, N. Ye, C. Wang, and B. Zeng, “Unsuper- vised global and local homography estimation with motion basis learning,”IEEE TP AMI, vol. 45, no. 6, pp. 7885–7899, 2022

  212. [220]

    Megadepth: Learning single-view depth prediction from internet photos,

    Z. Li and N. Snavely, “Megadepth: Learning single-view depth prediction from internet photos,” inCVPR, 2018, pp. 2041–2050

  213. [221]

    Yfcc100m: The new data in multimedia research,

    B. Thomee, D. A. Shamma, G. Friedland, B. Elizalde, K. Ni, D. Poland, D. Borth, and L.-J. Li, “Yfcc100m: The new data in multimedia research,”CACM, vol. 59, no. 2, pp. 64–73, 2016

  214. [222]

    Scannet: Richly-annotated 3d reconstructions of indoor scenes,

    A. Dai, A. X. Chang, M. Savva, M. Halber, T. Funkhouser, and M. Nießner, “Scannet: Richly-annotated 3d reconstructions of indoor scenes,” inCVPR, 2017, pp. 5828–5839

  215. [223]

    Sun3d: A database of big spaces reconstructed using sfm and object labels,

    J. Xiao, A. Owens, and A. Torralba, “Sun3d: A database of big spaces reconstructed using sfm and object labels,” inICCV, 2013, pp. 1625–1632

  216. [224]

    Ncmnet: Neighbor consistency mining network for two-view correspondence pruning,

    X. Liu, R. Qin, J. Yan, and J. Yang, “Ncmnet: Neighbor consistency mining network for two-view correspondence pruning,”IEEE TP AMI, vol. 46, no. 12, pp. 11 254–11 272, 2024

  217. [225]

    Relative camera pose estimation using convolutional neural networks,

    I. Melekhov, J. Ylioinas, J. Kannala, and E. Rahtu, “Relative camera pose estimation using convolutional neural networks,” inACIVS, 2017, pp. 675–687

  218. [226]

    Learn- ing deep features for scene recognition using places database,

    B. Zhou, A. Lapedriza, J. Xiao, A. Torralba, and A. Oliva, “Learn- ing deep features for scene recognition using places database,” in NeurIPS, vol. 27, 2014, pp. 1–9

  219. [227]

    Rpnet: An end-to-end network for relative camera pose estimation,

    S. En, A. Lechervy, and F. Jurie, “Rpnet: An end-to-end network for relative camera pose estimation,” inECCVW, 2018, pp. 1–8

  220. [228]

    Wide-baseline relative camera pose estimation with directional learning,

    K. Chen, N. Snavely, and A. Makadia, “Wide-baseline relative camera pose estimation with directional learning,” inCVPR, 2021, pp. 3258–3268

  221. [229]

    A generalized solution of the orthogonal procrustes problem,

    P . H. Sch ¨onemann, “A generalized solution of the orthogonal procrustes problem,”Psychometrika, vol. 31, no. 1, pp. 1–10, 1966

  222. [230]

    Map-free visual relocalization: Metric pose relative to a single image,

    E. Arnold, J. Wynn, S. Vicente, G. Garcia-Hernando, A. Mon- szpart, V . Prisacariu, D. Turmukhambetov, and E. Brachmann, “Map-free visual relocalization: Metric pose relative to a single image,” inECCV, 2022, pp. 690–708

  223. [231]

    Srpose: Two- view relative pose estimation with sparse keypoints,

    R. Yin, Y. Zhang, Z. Pan, J. Zhu, C. Wang, and B. Jia, “Srpose: Two- view relative pose estimation with sparse keypoints,” inECCV, 2025, pp. 88–107

  224. [232]

    On the continuity of rotation representations in neural networks,

    Y. Zhou, C. Barnes, J. Lu, J. Yang, and H. Li, “On the continuity of rotation representations in neural networks,” inCVPR, 2019, pp. 5745–5753

  225. [233]

    An analysis of svd for deep rotation estimation,

    J. Levinson, C. Esteves, K. Chen, N. Snavely, A. Kanazawa, A. Rostamizadeh, and A. Makadia, “An analysis of svd for deep rotation estimation,” inNeurIPS, vol. 33, 2020, pp. 22 554–22 565

  226. [234]

    Extreme rotation estimation using dense correlation volumes,

    R. Cai, B. Hariharan, N. Snavely, and H. Averbuch-Elor, “Extreme rotation estimation using dense correlation volumes,” inCVPR, 2021, pp. 14 566–14 575

  227. [235]

    Deep fundamental matrix estimation without correspondences,

    O. Poursaeed, G. Yang, A. Prakash, Q. Fang, H. Jiang, B. Har- iharan, and S. Belongie, “Deep fundamental matrix estimation without correspondences,” inECCVW, 2018, pp. 1–13

  228. [236]

    To learn or not to learn: Visual localization from essential matrices,

    Q. Zhou, T. Sattler, M. Pollefeys, and L. Leal-Taixe, “To learn or not to learn: Visual localization from essential matrices,” inICRA, 2020, pp. 3319–3326

  229. [237]

    Structure-from-motion revis- ited,

    J. L. Schonberger and J.-M. Frahm, “Structure-from-motion revis- ited,” inCVPR, 2016, pp. 4104–4113

  230. [238]

    Pixelwise view selection for unstructured multi-view stereo,

    J. L. Sch ¨onberger, E. Zheng, J.-M. Frahm, and M. Pollefeys, “Pixelwise view selection for unstructured multi-view stereo,” inECCV, 2016, pp. 501–518

  231. [239]

    Inloc: Indoor visual localization with dense matching and view synthesis,

    H. Taira, M. Okutomi, T. Sattler, M. Cimpoi, M. Pollefeys, J. Sivic, T. Pajdla, and A. Torii, “Inloc: Indoor visual localization with dense matching and view synthesis,” inCVPR, 2018, pp. 7199– 7209. 23

  232. [240]

    Universal correspondence network,

    C. B. Choy, J. Gwak, S. Savarese, and M. Chandraker, “Universal correspondence network,” inNeurIPS, vol. 29, 2016, pp. 1–9

  233. [241]

    Netvlad: Cnn architecture for weakly supervised place recog- nition,

    R. Arandjelovic, P . Gronat, A. Torii, T. Pajdla, and J. Sivic, “Netvlad: Cnn architecture for weakly supervised place recog- nition,” inCVPR, 2016, pp. 5297–5307

  234. [242]

    Semi- dense feature matching with transformers and its applications in multiple-view geometry,

    Z. Shen, J. Sun, Y. Wang, X. He, H. Bao, and X. Zhou, “Semi- dense feature matching with transformers and its applications in multiple-view geometry,”IEEE TP AMI, vol. 45, no. 6, pp. 7726– 7738, 2022

  235. [243]

    Gim: Learning generalizable image matcher from internet videos,

    X. Shen, Z. Cai, W. Yin, M. M ¨uller, Z. Li, K. Wang, X. Chen, and C. Wang, “Gim: Learning generalizable image matcher from internet videos,” inICLR, 2024, pp. 1–16

  236. [244]

    Aerialmegadepth: Learning aerial-ground reconstruction and view synthesis,

    K. Vuong, A. Ghosh, D. Ramanan, S. Narasimhan, and S. Tulsiani, “Aerialmegadepth: Learning aerial-ground reconstruction and view synthesis,” inCVPR, 2025, pp. 1–8

  237. [245]

    Event-based vision: A survey,

    G. Gallego, T. Delbr ¨uck, G. Orchard, C. Bartolozzi, B. Taba, A. Censi, S. Leutenegger, A. J. Davison, J. Conradt, K. Daniilidis et al., “Event-based vision: A survey,”IEEE TP AMI, vol. 44, no. 1, pp. 154–180, 2020

  238. [246]

    Image fusion meets deep learning: A survey and perspective,

    H. Zhang, H. Xu, X. Tian, J. Jiang, and J. Ma, “Image fusion meets deep learning: A survey and perspective,”IF, vol. 76, pp. 323– 336, 2021

  239. [247]

    Minima: Modality invariant image matching,

    J. Ren, X. Jiang, Z. Li, D. Liang, X. Zhou, and X. Bai, “Minima: Modality invariant image matching,” inCVPR, 2025, pp. 1–8

  240. [248]

    Dgc-gnn: leveraging geom- etry and color cues for visual descriptor-free 2d-3d matching,

    S. Wang, J. Kannala, and D. Barath, “Dgc-gnn: leveraging geom- etry and color cues for visual descriptor-free 2d-3d matching,” in CVPR, 2024, pp. 20 881–20 891

  241. [249]

    Croco v2: Improved cross-view completion pre-training for stereo match- ing and optical flow,

    P . Weinzaepfel, T. Lucas, V . Leroy, Y. Cabon, V . Arora, R. Br´egier, G. Csurka, L. Antsfeld, B. Chidlovskii, and J. Revaud, “Croco v2: Improved cross-view completion pre-training for stereo match- ing and optical flow,” inICCV, 2023, pp. 17 969–17 980

  242. [250]

    Dust3r: Geometric 3d vision made easy,

    S. Wang, V . Leroy, Y. Cabon, B. Chidlovskii, and J. Revaud, “Dust3r: Geometric 3d vision made easy,” inCVPR, 2024, pp. 20 697–20 709

  243. [251]

    Grounding image matching in 3d with mast3r,

    V . Leroy, Y. Cabon, and J. Revaud, “Grounding image matching in 3d with mast3r,” inECCV, 2024, pp. 71–91

  244. [252]

    Vggt: Visual geometry grounded transformer,

    J. Wang, M. Chen, N. Karaev, A. Vedaldi, C. Rupprecht, and D. Novotny, “Vggt: Visual geometry grounded transformer,” CVPR, pp. 1–8, 2025

  245. [253]

    Deep learning in remote sensing image matching: A survey,

    L. Li, L. Han, Y. Ye, Y. Xiang, and T. Zhang, “Deep learning in remote sensing image matching: A survey,”ISPRS P&RS, vol. 225, pp. 88–112, 2025

  246. [254]

    A review of multimodal image matching: Methods and applications,

    X. Jiang, J. Ma, G. Xiao, Z. Shao, and X. Guo, “A review of multimodal image matching: Methods and applications,”IF, vol. 73, pp. 22–71, 2021

  247. [255]

    Spatiotemporal modeling of molecular holograms,

    X. Qiu, D. Y. Zhu, Y. Lu, J. Yao, Z. Jing, K. H. Min, M. Cheng, H. Pan, L. Zuo, S. Kinget al., “Spatiotemporal modeling of molecular holograms,”Cell, vol. 187, no. 26, pp. 7351–7373, 2024

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.