Pith. sign in

REVIEW 3 major objections 44 references

EpiMask raises satellite image matching accuracy by up to 30% by restricting cross-attention to epipolar bands taken from camera metadata.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-13 20:51 UTC pith:3TVWWKQU

load-bearing objection Solid, geometry-aware LoFTR adaptation that delivers real gains on SatDepth; the epipolar mask is the real contribution and the experiments back it. the 3 major comments →

arxiv 2603.21463 v2 pith:3TVWWKQU submitted 2026-03-23 cs.CV

EpiMask: Leveraging Epipolar Distance Based Masks in Cross-Attention for Satellite Image Matching

classification cs.CV
keywords satellite image matchingepipolar geometrycross-attention maskpushbroom camerasemi-dense matchingaffine fundamental matrixSatDepthLoRA fine-tuning
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Ground-based deep matchers are trained on pinhole images and therefore bake in straight-line epipolar geometry. Satellite pushbroom cameras synthesize each image line by line, so their epipolar curves are nonlinear and those same networks underperform even after re-training. EpiMask adapts a LoFTR-style semi-dense matcher by replacing unconstrained cross-attention with an epipolar-distance mask derived from patch-wise affine approximations of the RPC camera model, and by LoRA-fine-tuning a satellite foundation encoder. On the SatDepth benchmark the resulting matches are denser and more precise, lifting pose-estimation accuracy by as much as thirty percent relative to the best re-trained baselines. Anyone who needs reliable multi-view geometry from satellite pairs—for alignment, sparse reconstruction, or change detection—gains a concrete, geometry-aware tool rather than a generic network that merely saw more satellite pixels.

Core claim

Simply re-training a ground-based matcher on satellite data is suboptimal. Explicitly injecting the nonlinear epipolar geometry of pushbroom cameras—through patch-wise affine fundamental matrices and an epipolar-distance attention mask—together with lightweight fine-tuning of a satellite-pretrained encoder, yields up to 30% higher matching accuracy (pose AUC and 1-pixel precision) than the strongest re-trained baselines on SatDepth.

What carries the argument

Masked cross-attention (MXA): for each query pixel an initial affine fundamental matrix F0 computed from RPC metadata defines a band of width b that shrinks across transformer layers; only keys inside that band receive attention, so the network is forced to search only geometrically plausible locations while still learning the correspondence.

Load-bearing premise

The initial affine fundamental matrix taken from satellite metadata must be accurate enough that the true match always falls inside the chosen epipolar band; if metadata noise or terrain relief push the correspondence outside the band, the mask becomes a hard error.

What would settle it

Perturb the RPC-derived F0 on held-out pairs so that a non-trivial fraction of ground-truth matches lie outside the γp band, then retrain and re-evaluate; a clear drop in precision and pose AUC relative to the unperturbed run would falsify the claim that the mask is helpful under realistic metadata error.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • High-resolution EpiMask variants reach roughly 90% pose-estimation AUC, making them practical for satellite image alignment.
  • The same pairs produce far more true-positive matches, yielding denser sparse point clouds without extra post-processing.
  • Narrower bands suit modern satellites with accurate pose; wider bands tolerate noisier older sensors.
  • The masked-attention idea extends immediately to any multi-view system that supplies reliable camera metadata (UAVs with onboard pose, calibrated X-ray angiography).
  • LoRA fine-tuning of a frozen satellite foundation encoder is sufficient to adapt appearance features for matching without full backbone retraining.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • An adaptive band-width schedule conditioned on estimated pose uncertainty could reduce sensitivity to metadata quality across different satellite constellations.
  • The progressive narrowing of the mask across layers is a general curriculum that may transfer to other geometry-constrained matching tasks beyond satellites.
  • Performance still degrades under extreme view-angle differences on unseen AOIs, indicating that data diversity remains a co-equal bottleneck with architecture.
  • Porting the same epipolar mask into a fully dense matcher could improve satellite stereo depth without classical post-filtering.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 0 minor

Summary. EpiMask adapts a LoFTR-style semi-dense matcher to satellite (pushbroom) imagery by (1) using patch-wise affine approximations of the RPC camera model to obtain an initial affine fundamental matrix F0, (2) inserting an epipolar-distance attention mask into coarse cross-attention (and into dual-softmax matching) with a warm-up and linearly annealed band width, and (3) LoRA-fine-tuning a Satlas-Pretrain Swin encoder inside an FPN feature extractor. Trained and evaluated on the official SatDepth splits against re-trained ground-based baselines (satLoFTR, satMatchFormer, satDualRC-Net, SIFT+satCAPS), the method reports higher pose-estimation AUC, 1-px precision, and denser true-positive matches, with the abstract’s “up to 30%” figure reflecting the largest relative gains on selected metrics/AOIs. Ablations cover resolution, mask width γ, positional encoding, LoRA rank, skip-fusion, and two-stage training; attention and confidence visualizations are provided in the supplement.

Significance. The work targets a genuine and under-addressed mismatch between pinhole-trained matchers and pushbroom epipolar geometry, while exploiting metadata that is routinely available with satellite products. The empirical package is solid for an architecture paper: same-protocol re-training of strong baselines, multiple AOIs, simulated-rotation stress tests, and component ablations that isolate the mask, encoder adaptation, and resolution choices. If the reported lifts hold under broader geographic and sensor conditions, EpiMask is immediately useful for satellite image alignment and sparse 3D reconstruction. Strengths include the geometry-aware mask design with warm-up/annealing, LoRA adaptation of a satellite foundation encoder, and the promise of code release. The main caveats are geographic coverage of SatDepth and dependence on RPC-derived F0 quality—both already partly acknowledged.

major comments (3)
  1. Abstract and §1 claim “up to 30% improvement in matching accuracy” versus re-trained ground-based models. Table 1 and Fig. 18 show large but metric- and AOI-dependent gains (e.g., Jacksonville AUC@5° ~81→93; San Fernando AUC@5° ~55→88; 1-px precision Jacksonville ~62→83). Please state explicitly which metric, baseline, and AOI produce the 30% figure (relative vs absolute), and prefer reporting a small set of primary metrics with relative/absolute deltas so the claim cannot be read as uniform across all settings.
  2. §3.1–3.2 and §3.4: the load-bearing assumption is that the RPC-derived affine F0 yields an epipolar band of width γp that contains the true correspondence after warm-up. Ablations (Fig. 9, §4.2) show γ=0.4 vs 0.6 are similar, and warm-up softens the hard mask, but there is no controlled experiment with noisy or biased F0 (e.g., perturbed RPC / affine parameters, or older sensors with poorer pose). A short sensitivity study—or at least quantitative failure analysis when the true match falls outside the final band—would make the central geometric claim much more robust.
  3. §3.4 / §4 evaluation protocol: the paper switches from SatDepth’s K=200 top matches to K=2000 because EpiMask produces more correspondences. Pose AUC and precision@1px can depend on how many matches enter RANSAC/pose estimation. Please confirm that baseline numbers in Table 1 / Fig. 18 were recomputed under the same K (or report both K=200 and K=2000 for all methods) so the comparison remains protocol-fair.

Circularity Check

0 steps flagged

No significant circularity: empirical architecture paper whose gains are measured on held-out data against re-trained baselines; epipolar mask is constructed from external RPC metadata.

full rationale

EpiMask is a standard empirical deep-learning architecture paper (LoFTR-style transformer with two satellite-specific modifications). The epipolar-distance mask M_epi is computed from the RPC-derived affine fundamental matrix F0 that ships with the imagery (Sec. 3.1–3.2, Eq. 1); it is not fitted to matching labels and is only softened by a warm-up schedule and linear band shrinkage. Coarse and fine losses (Eq. 4) are ordinary supervised cross-entropy / weighted MSE on ground-truth matches obtained by warping SatDepth maps through the same affine cameras. All quantitative claims (pose AUC, 1-px precision, #TP matches, up to 30 % lift) are obtained by training and evaluating on the official SatDepth splits against re-trained baselines (satLoFTR, satMatchFormer, …) and are supported by multiple ablations (resolution, γ, LoRA rank, skip-fusion, training stages). Self-citation of the authors’ prior SatDepth dataset is expected for evaluation and does not close any logical loop: the architectural novelty and the measured gains remain independent of that citation. No equation, uniqueness claim, or “prediction” reduces to its own inputs by construction.

Axiom & Free-Parameter Ledger

6 free parameters · 4 axioms · 1 invented entities

As an empirical deep-learning paper the central claim rests on a small set of standard geometric approximations, a handful of hand-chosen architectural hyperparameters, and the usual supervised matching losses. No new physical entities are postulated; the free parameters are the usual training knobs plus the mask-width schedule that is unique to this work.

free parameters (6)
  • epipolar band fraction γ = 0.4 or 0.6
    Final mask width is set to γp with γ ∈ {0.4, 0.6}; both values are chosen by ablation rather than derived.
  • mask warm-up epochs Nm = 5
    No masking for the first 5 epochs; chosen by hand.
  • coarse confidence threshold δc = 0.3
    Matches kept only if Pc ≥ 0.3; standard but free.
  • LoRA rank / alpha = 16/8 or 32/16
    Rank 16 (α=8) or 32 (α=16) selected by ablation.
  • learning-rate schedule milestones and base lr = lr_true=5e-4, γ_lr=0.5
    Multi-step decay at epochs {8,12,16,20,24}, lr scaled from 8e-3 reference; conventional but free.
  • 3-D distance threshold δ3D for GT matches
    Used to label coarse ground-truth correspondences from SatDepth maps; value inherited from SatDepth but still a free threshold.
axioms (4)
  • domain assumption For sufficiently small image patches the nonlinear RPC camera model can be replaced by an affine camera, yielding a usable affine fundamental matrix F0.
    Stated in Sec. 3.1 and used to construct every epipolar mask; classic result from photogrammetry literature but load-bearing for the mask to be valid.
  • domain assumption Dual-softmax + mutual nearest neighbor with a fixed confidence threshold produces a reliable set of coarse matches.
    Inherited from LoFTR / SuperGlue and used without re-derivation (Eq. 3).
  • domain assumption The Satlas-Pretrain Swin-B encoder already contains features that are useful for matching after light LoRA adaptation.
    Justifies freezing the backbone and only training LoRA + decoder + transformers.
  • standard math Linear attention is an adequate approximation to full attention for the self-attention layers.
    Taken from Katharopoulos et al. and LoFTR; used for efficiency.
invented entities (1)
  • EpiMask / masked cross-attention (MXA) with linearly annealed epipolar band no independent evidence
    purpose: Restricts each query pixel’s attention to a geometrically plausible band whose width shrinks across layers.
    The specific combination of epipolar-distance mask + linear annealing schedule is introduced by this paper; no independent physical existence outside the architecture.

pith-pipeline@v1.1.0-grok45 · 30231 in / 3249 out tokens · 35186 ms · 2026-07-13T20:51:07.405209+00:00 · methodology

0 comments
read the original abstract

The deep-learning based image matching networks can now handle significantly larger variations in viewpoints and illuminations while providing matched pairs of pixels with sub-pixel precision. These networks have been trained with ground-based image datasets and, implicitly, their performance is optimized for the pinhole camera geometry. Consequently, you get suboptimal performance when such networks are used to match satellite images since those images are synthesized as a moving satellite camera records one line at a time of the points on the ground. In this paper, we present EpiMask, a semi-dense image matching network for satellite images that (1) Incorporates patch-wise affine approximations to the camera modeling geometry; (2) Uses an epipolar distance-based attention mask to restrict cross-attention to geometrically plausible regions; and (3) That fine-tunes a foundational pretrained image encoder for robust feature extraction. Experiments on the SatDepth dataset demonstrate up to 30% improvement in matching accuracy compared to re-trained ground-based models.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

44 extracted references · 2 linked inside Pith

  1. [1]

    HPatches: A benchmark and evaluation of handcrafted and learned local descriptors

    Vassileios Balntas, Karel Lenc, Andrea Vedaldi, and Krystian Mikolajczyk. HPatches: A benchmark and evaluation of handcrafted and learned local descriptors. InProceedings of IEEE Intl. Conf. on Computer Vision and Pattern Recognition (CVPR), 2017. 3

  2. [2]

    Axel Barroso-Laguna, Edgar Riba, Daniel Ponsa, and Krystian Mikolajczyk. Key. Net: Keypoint detection by handcrafted and learned CNN filters. InProceedings of IEEE Intl. Conf. on Computer Vision and Pattern Recognition (CVPR), 2019. 1, 3

  3. [3]

    Satlaspretrain: A large-scale dataset for remote sensing image understanding

    Favyen Bastani, Piper Wolters, Ritwik Gupta, Joe Ferdinando, and Aniruddha Kembhavi. Satlaspretrain: A large-scale dataset for remote sensing image understanding. InProceedings of Intl. Conf. on Computer Vision (ICCV), 2023. 3, 4, 9

  4. [4]

    SURF: Speeded Up Robust Features

    Herbert Bay, Tinne Tuytelaars, and Luc Van Gool. SURF: Speeded Up Robust Features. InProceedings of the European Conference on Computer Vision (ECCV), 2006. 3

  5. [5]

    End-to-end object detection with transformers

    Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. End-to-end object detection with transformers. InProceedings of the European Conference on Computer Vision (ECCV), 2020. 4

  6. [6]

    Aspanformer: Detector-free image matching with adaptive span transformer

    Hongkai Chen, Zixin Luo, Lei Zhou, Yurun Tian, Mingmin Zhen, Tian Fang, David McKinnon, Yanghai Tsin, and Long Quan. Aspanformer: Detector-free image matching with adaptive span transformer. InProceedings of the European Conference on Computer Vision (ECCV), 2022. 1, 2

  7. [7]

    ScanNet: Richly-annotated 3D Reconstructions of Indoor Scenes

    Angela Dai, Angel X Chang, Manolis Savva, Maciej Halber, Thomas Funkhouser, and Matthias Nießner. ScanNet: Richly-annotated 3D Reconstructions of Indoor Scenes. InProceedings of IEEE Intl. Conf. on Computer Vision and Pattern Recognition (CVPR), 2017. 3

  8. [8]

    An automatic and modular stereo pipeline for pushbroom images

    Carlo De Franchis, Enric Meinhardt-Llopis, Julien Michel, Jean-Michel Morel, and Gabriele Facciolo. An automatic and modular stereo pipeline for pushbroom images. InISPRS Annals of the Photogrammetry, Remote Sensing and Spatial Information Sciences,

  9. [9]

    Rahul Deshmukh and Avinash C. Kak. SatDepth: A Novel Dataset for Satellite Image Matching.IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 19:894–903, 2026. 1, 2, 3, 5, 6, 10, 15

  10. [10]

    SuperPoint: Self-Supervised Interest Point Detection and Description

    Daniel DeTone, Tomasz Malisiewicz, and Andrew Rabinovich. SuperPoint: Self-Supervised Interest Point Detection and Description. InProceedings of IEEE Intl. Conf. on Computer Vision and Pattern Recognition (CVPR), 2018. 3

  11. [11]

    D2-Net: A Trainable CNN for Joint Detection and Description of Local Features

    Mihai Dusmanu, Ignacio Rocco, Tomas Pajdla, Marc Pollefeys, Josef Sivic, Akihiko Torii, and Torsten Sattler. D2-Net: A Trainable CNN for Joint Detection and Description of Local Features. InProceedings of IEEE Intl. Conf. on Computer Vision and Pattern Recognition (CVPR), 2019. 1, 3

  12. [12]

    DKM: Dense kernelized feature matching for geometry estimation

    Johan Edstedt, Ioannis Athanasiadis, Mårten Wadenbäck, and Michael Felsberg. DKM: Dense kernelized feature matching for geometry estimation. InProceedings of IEEE Intl. Conf. on Computer Vision and Pattern Recognition (CVPR), 2023. 1, 2, 3

  13. [13]

    Roma: Robust dense feature matching

    Johan Edstedt, Qiyu Sun, Georg Bökman, Mårten Wadenbäck, and Michael Felsberg. Roma: Robust dense feature matching. In Proceedings of IEEE Intl. Conf. on Computer Vision and Pattern Recognition (CVPR), 2024. 1, 2, 3

  14. [14]

    A pipeline for automated processing of declassified corona kh-4 (1962–1972) stereo imagery.IEEE Transactions on Geoscience and Remote Sensing, 60:1–14, 2022

    Sajid Ghuffar, Tobias Bolch, Ewelina Rupnik, and Atanu Bhattacharya. A pipeline for automated processing of declassified corona kh-4 (1962–1972) stereo imagery.IEEE Transactions on Geoscience and Remote Sensing, 60:1–14, 2022. 3

  15. [15]

    Cambridge university press, 2003

    Richard Hartley and Andrew Zisserman.Multiple View Geometry in Computer Vision. Cambridge university press, 2003. 3

  16. [16]

    Lora: Low-rank adaptation of large language models

    Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, Weizhu Chen, et al. Lora: Low-rank adaptation of large language models. InProceedings of Intl. Conf. on Learning Representations (ICLR), 2022. 4, 9

  17. [17]

    Transformers are rnns: Fast autoregressive transformers with linear attention

    Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, and François Fleuret. Transformers are rnns: Fast autoregressive transformers with linear attention. 2020. 4, 10

  18. [18]

    Whu-stereo: A challenging benchmark for stereo matching of high-resolution satellite images.IEEE Transactions on Geoscience and Remote Sensing, 61:1–14, 2023

    Shenhong Li, Sheng He, San Jiang, Wanshou Jiang, and Lin Zhang. Whu-stereo: A challenging benchmark for stereo matching of high-resolution satellite images.IEEE Transactions on Geoscience and Remote Sensing, 61:1–14, 2023. 1

  19. [19]

    Dual-Resolution Correspondence Networks

    Xinghui Li, Kai Han, Shuda Li, and Victor Prisacariu. Dual-Resolution Correspondence Networks. InProceedings of Conf. on Neural Information Processing Systems (NeurIPS), 2020. 1, 2, 3, 6, 15

  20. [20]

    MegaDepth: Learning Single-View Depth Prediction from Internet Photos

    Zhengqi Li and Noah Snavely. MegaDepth: Learning Single-View Depth Prediction from Internet Photos. InProceedings of IEEE Intl. Conf. on Computer Vision and Pattern Recognition (CVPR), 2018. 2, 3

  21. [21]

    Feature pyramid networks for object detection

    Tsung-Yi Lin, Piotr Dollár, Ross Girshick, Kaiming He, Bharath Hariharan, and Serge Belongie. Feature pyramid networks for object detection. InProceedings of IEEE Intl. Conf. on Computer Vision and Pattern Recognition (CVPR), 2017. 4, 9

  22. [22]

    LightGlue: Local Feature Matching at Light Speed

    Philipp Lindenberger, Paul-Edouard Sarlin, and Marc Pollefeys. LightGlue: Local Feature Matching at Light Speed. InProceedings of Intl. Conf. on Computer Vision (ICCV), 2023. 1, 2, 3

  23. [23]

    Swin transformer: Hierarchical vision transformer using shifted windows

    Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin transformer: Hierarchical vision transformer using shifted windows. InProceedings of Intl. Conf. on Computer Vision (ICCV), 2021. 3

  24. [24]

    Distinctive Image Features from Scale-Invariant Keypoints

    David G Lowe. Distinctive Image Features from Scale-Invariant Keypoints. InIntl. Journal of Computer Vision (IJCV), 2004. 3

  25. [25]

    GeoDesc: Learning Local Descriptors by Integrating Geometry Constraints

    Zixin Luo, Tianwei Shen, Lei Zhou, Siyu Zhu, Runze Zhang, Yao Yao, Tian Fang, and Long Quan. GeoDesc: Learning Local Descriptors by Integrating Geometry Constraints. InProceedings of the European Conference on Computer Vision (ECCV), 2018. 3

  26. [26]

    Orientation theory for satellite CCD line-scanner imageries of hilly terrains

    Atsushi Okamoto, Si Akamatu, and Hiroyuki Hasegawa. Orientation theory for satellite CCD line-scanner imageries of hilly terrains. International Archives of Photogrammetry and Remote Sensing, 1993. 3

  27. [27]

    LF-Net: Learning local features from images

    Yuki Ono, Eduard Trulls, Pascal Fua, and Kwang Moo Yi. LF-Net: Learning local features from images. InProceedings of Conf. on Neural Information Processing Systems (NeurIPS), 2018. 3

  28. [28]

    Dinov2: Learning robust visual features without supervision.arXiv preprint arXiv:2304.07193,

    Maxime Oquab, Timothée Darcet, Théo Moutakanni, Huy V o, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Fran- cisco Massa, Alaaeldin El-Nouby, et al. Dinov2: Learning robust visual features without supervision.arXiv preprint arXiv:2304.07193,

  29. [29]

    Sonali Patil, Bharath Comandur, Tanmay Prakash, and Avinash C. Kak. A New Stereo Benchmarking Dataset for Satellite Images. arXiv preprint arXiv:1907.04404, 2019. 1

  30. [30]

    Neighbourhood Consensus Networks

    Ignacio Rocco, Mircea Cimpoi, Relja Arandjelovi´c, Akihiko Torii, Tomas Pajdla, and Josef Sivic. Neighbourhood Consensus Networks. InProceedings of Conf. on Neural Information Processing Systems (NeurIPS), 2018. 2, 4

  31. [31]

    Efficient Neighbourhood Consensus Networks via Submanifold Sparse Convolu- tions

    Ignacio Rocco, Relja Arandjelovic, and Josef Sivic. Efficient Neighbourhood Consensus Networks via Submanifold Sparse Convolu- tions. InProceedings of the European Conference on Computer Vision (ECCV), 2020. 2

  32. [32]

    ORB: An efficient alternative to SIFT or SURF

    Ethan Rublee, Vincent Rabaud, Kurt Konolige, and Gary Bradski. ORB: An efficient alternative to SIFT or SURF. InProceedings of Intl. Conf. on Computer Vision (ICCV), 2011. 3

  33. [33]

    SuperGlue: Learning Feature Matching with Graph Neural Networks

    Paul-Edouard Sarlin, Daniel DeTone, Tomasz Malisiewicz, and Andrew Rabinovich. SuperGlue: Learning Feature Matching with Graph Neural Networks. InProceedings of IEEE Intl. Conf. on Computer Vision and Pattern Recognition (CVPR), 2020. 1, 2, 3, 4

  34. [34]

    Image Retrieval for Image-Based Localization Revisited

    Torsten Sattler, Tobias Weyand, Bastian Leibe, and Leif Kobbelt. Image Retrieval for Image-Based Localization Revisited. In Proceedings of British Machine Vision Conference (BMVC), 2012. 3

  35. [35]

    Gim: Learning generalizable image matcher from internet videos

    Xuelun Shen, Zhipeng Cai, Wei Yin, Matthias Müller, Zijun Li, Kaixuan Wang, Xiaozhi Chen, and Cheng Wang. Gim: Learning generalizable image matcher from internet videos. InProceedings of Intl. Conf. on Learning Representations (ICLR), 2024. 1, 3

  36. [36]

    Deep Learning Meets Satellite Images – An Evaluation on Handcrafted and Learning-based Features for Multi-date Satellite Stereo Images, 2024

    Shuang Song, Luca Morelli, Xinyi Wu, Rongjun Qin, Hessah Albanwan, and Fabio Remondino. Deep Learning Meets Satellite Images – An Evaluation on Handcrafted and Learning-based Features for Multi-date Satellite Stereo Images, 2024. 1, 3 25

  37. [37]

    LoFTR: Detector-free local feature matching with transformers

    Jiaming Sun, Zehong Shen, Yuang Wang, Hujun Bao, and Xiaowei Zhou. LoFTR: Detector-free local feature matching with transformers. InProceedings of IEEE Intl. Conf. on Computer Vision and Pattern Recognition (CVPR), 2021. 1, 2, 3, 4, 5, 6, 10, 15

  38. [38]

    InLoc: Indoor Visual Localization with Dense Matching and View Synthesis

    Hajime Taira, Masatoshi Okutomi, Torsten Sattler, Mircea Cimpoi, Marc Pollefeys, Josef Sivic, Tomas Pajdla, and Akihiko Torii. InLoc: Indoor Visual Localization with Dense Matching and View Synthesis. InProceedings of IEEE Intl. Conf. on Computer Vision and Pattern Recognition (CVPR), 2018. 3

  39. [39]

    DISK: Learning local features with policy gradient

    Michał Tyszkiewicz, Pascal Fua, and Eduard Trulls. DISK: Learning local features with policy gradient. InProceedings of Conf. on Neural Information Processing Systems (NeurIPS), 2020. 1

  40. [40]

    Attention is all you need

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. InProceedings of Conf. on Neural Information Processing Systems (NeurIPS), 2017. 3, 4

  41. [41]

    Learning feature descriptors using camera pose supervision

    Qianqian Wang, Xiaowei Zhou, Bharath Hariharan, and Noah Snavely. Learning feature descriptors using camera pose supervision. In Proceedings of the European Conference on Computer Vision (ECCV), 2020. 3, 5, 6, 15

  42. [42]

    Matchformer: Interleaving attention in transformers for feature matching

    Qing Wang, Jiaming Zhang, Kailun Yang, Kunyu Peng, and Rainer Stiefelhagen. Matchformer: Interleaving attention in transformers for feature matching. InProceedings of the Asian Conference on Computer Vision, 2022. 1, 2, 3, 4, 6, 15

  43. [43]

    LIFT: Learned invariant feature transform

    Kwang Moo Yi, Eduard Trulls, Vincent Lepetit, and Pascal Fua. LIFT: Learned invariant feature transform. InProceedings of the European Conference on Computer Vision (ECCV), 2016. 1, 3

  44. [44]

    Patch2Pix: Epipolar-Guided Pixel-Level Correspondences

    Qunjie Zhou, Torsten Sattler, and Laura Leal-Taixé. Patch2Pix: Epipolar-Guided Pixel-Level Correspondences. InProceedings of IEEE Intl. Conf. on Computer Vision and Pattern Recognition (CVPR), 2020. 2, 3 26