Pith. sign in

REVIEW 2 major objections 5 minor 1 cited by

A Guide to Structureless Visual Localization

T0 review · 2 major / 5 minor · reviewed 2026-08-16 · deepseek-v4-flash

Pith's one-line read This paper's benchmark shows that structureless visual localization is most accurate when query poses come from classical multi-view geometry rather than from neural pose regression.

desk verdict Useful first systematic benchmark of structureless localization, but the abstract's 'wide margin' regression claim rests on a single dataset and a self-admittedly suboptimal baseline. read the letter →

arxiv 2504.17636 v1 pith:XFZ3LPCI submitted 2025-04-24 cs.CV

classification cs.CV
keywords visuallocalizationstructurelessrelativeposeestimationregressionimageretrievalessentialmatrixlocaltriangulationbenchmark
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper maps out a family of visual localization methods that avoid storing a 3D scene model, representing the scene only as database images with known poses. It compares four families—pose triangulation from pairwise relative poses, semi-generalized relative pose estimation, locally triangulating 3D points on the fly, and relative pose regression—across outdoor and indoor benchmarks. The central finding is that the more classical geometry a method uses, the more accurate it tends to be, and that recent regression-based approaches trail the classical families by a wide margin on the dataset where they were tested. The practical stakes are that structureless methods, which are easy to update as scenes change, can come close to the accuracy of model-based pipelines, making them a viable alternative for self-driving and augmented-reality systems.

What carries the argument

The comparison is organized around how much explicit geometric reasoning each structureless pipeline applies to 2D-2D matches between the query and retrieved database images. The key mechanisms are: pairwise essential-matrix estimation with rotation averaging and translation triangulation; the E5+1 semi-generalized solver, which computes the query pose with respect to two database images at once (five matches to one image fix the relative pose, one match to the other fixes scale); on-the-fly triangulation of 3D points followed by a P3P absolute pose solver inside RANSAC; and regression models that predict depths or relative poses directly. The measured quantities are localization recall at standard pose-error thresholds and average runtime per query.

What would settle it

Re-run the comparison with a regression pipeline that is allowed to use known intrinsics and reference poses during alignment, or that is fine-tuned to be scale-aware, and evaluate it on the other outdoor and indoor datasets used in the paper; if its localization recall matches or exceeds the E5+1 and local-triangulation results reported here, the paper's main accuracy ordering fails.

Watch

Extended reading notes

Core claim

The paper's central claim is that among structureless localization approaches, accuracy follows the amount of geometric reasoning in the pipeline. Pose triangulation, which fuses pairwise relative poses, is consistently less accurate than semi-generalized relative pose estimation, where the query pose is computed jointly with respect to two database images and the translation scale is recovered. Locally triangulating a small 3D model per query and then applying absolute pose estimation yields the best accuracy among structureless methods, effectively catching up with a mesh-based structure-based pipeline on the Aachen day split and remaining only a few points behind at night. Relative pose regression, represented by MASt3R-based pipelines and Reloc3r, performs worst on Aachen Day-Night, with even the best regression variant notably behind the classical five-point essential-matrix method. The paper also reports that the accuracy-versus-runtime trade-off favors the E5+1 solver, since local triangulation is most accurate but slower.

Load-bearing premise

The ranking of pose regression as the least accurate family assumes that the released MASt3R pipeline, which cannot exploit the known intrinsics and poses of database images, is a fair representative of that family.

Editorial extensions

If this is right

  • Practitioners who value accuracy can use local on-the-fly triangulation as a structureless method that rivals a structure-based mesh pipeline on the Aachen day split.
  • Practitioners who value runtime can use the E5+1 semi-generalized solver, which keeps most of the accuracy of local triangulation at a fraction of the runtime.
  • Released relative pose regression models are not yet a drop-in replacement for geometric solvers in this setting: on Aachen Day-Night they trail even the pairwise essential-matrix baseline.
  • The best feature type is method- and scene-dependent, so a robust structureless system should expose matcher choice rather than hard-code one feature.
  • Structureless localization's flexibility—adding or removing database images—can be obtained with only a modest accuracy loss relative to structure-based state of the art.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the regression gap is mostly an artifact of unmodified inference, a version of MASt3R that consumes known intrinsics and reference poses could close much of the reported margin; the paper itself flags this as a suboptimal aspect of the released implementation.
  • Because the regression families were only evaluated on Aachen Day-Night, the accuracy ordering may not transfer to indoor or seasonal-change datasets, where day-night appearance shifts are less central.
  • A natural next experiment is a hybrid: use regression models to propose dense matches, then feed those matches to E5+1 or local triangulation; the ablation data suggest matches, not pose predictions, are where learned models already help.
  • The scene-dependent failure of the depth-aware E3+1 variant hints that adaptive selection of a depth-aware solver—trusting monocular depth only when scene structure is close to the camera—could beat either fixed choice.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. This paper presents a review and experimental comparison of structureless visual localization methods, which estimate a query camera pose from a database of images with known poses and no stored 3D model. The authors compare pose triangulation (Ess. mat., LazyLoc), semi-generalized relative pose estimation (E5+1, E3+1), local SfM-on-the-fly triangulation, and relative pose regression (MASt3R-based variants, Reloc3r) on Aachen Day-Night v1.1, Extended CMU Seasons, and NAVER indoor datasets. The evaluation uses multiple sparse and dense feature matchers and two depth estimators, and reports localization recall at standard thresholds, runtimes, and a comparison with structure-based baselines. The main claims are that more classical geometric reasoning tends to improve pose accuracy, that local triangulation and E5+1 are the most accurate structureless options, that regression-based methods underperform on the tested Aachen setup, and that structureless methods are competitive with, but slightly less accurate than, structure-based methods.

Significance. If the comparative findings hold, the paper would be a useful reference for practitioners and a benchmark for structureless localization research. Its strengths are the breadth of the evaluation across three datasets and multiple feature matchers, the consistent evaluation protocol, and the explicit acknowledgment of limitations, including the suboptimal out-of-the-box MASt3R pose alignment and the single-dataset regression evaluation. The paper does not introduce a new method, and its headline claim about regression-based methods is broader than the current evidence; the contribution is therefore primarily a well-organized empirical study whose conclusions need to be scoped precisely.

major comments (2)
  1. [Abstract; Sec. 4.1, Tab. 1] The central claim that classical geometry-based structureless methods outperform pose-regression methods 'by a wide margin' is supported only on Aachen Day-Night v1.1. Tab. 1 is the only regression evaluation; Sec. 4.1 states that the regression approaches were not evaluated on other datasets because of inaccurate pose estimates and long run-times. Since Sec. 2 itself notes that recent regression methods 'can significantly outperform classical approaches' under conditions of little visual overlap, and the Aachen top-10-retrieval setup is not such a regime, the abstract's unqualified margin claim overstates the evidence. The MASt3R pose align representative is further acknowledged in Sec. 3 to be suboptimal because the released implementation cannot use known intrinsics and database poses, so the family-level ranking is at least partly a statement about out-of-the-box pipelines. I recommend either adding regression evaluations on the other benchmarks or restricting the claim to Aachen and clearly attributing it to the evaluated variants.
  2. [Sec. 4.2, Tabs. 2-6] The accuracy-runtime trade-off conclusion relies on runtimes measured with a different configuration than the accuracy tables. Tab. 2 reports runtimes with pre-computed SuperPoint features matched by LightGlue, while the per-method best setups in Tabs. 3-6 use RoMa or MASt3R for most methods (e.g., E5+1 and Local triangulation use RoMa on Aachen; LazyLoc uses MASt3R on Aachen). Because feature matching is a dominant cost in these pipelines, the claims that 'E5+1 offers a better trade-off between pose accuracy and run-time' and that 'LazyLoc ... provides the fastest run-times' are not justified by consistent measurements. Runtimes should be reported for the exact configurations used in the accuracy tables, or the trade-off claims should be explicitly scoped to the SP+LG configuration.
minor comments (5)
  1. [Sec. 4.2, Tabs. 3-4] The sentence 'E5+1 and E3+1 directly compute the query pose w.r.t. multiple database images, which further improves performance' is contradicted by E3+1's results on Aachen Day-Night (Tab. 3) and on the Extended CMU park scene (Tab. 4), where E3+1 is worse than several pose-triangulation baselines. Since the paragraph immediately discusses this scene dependence, please name E5+1 in the opening sentence or move the caveat before the claim.
  2. [Sec. 4.1] The sentence 'Given the inaccurate pose estimates observed for the Aachen dataset' should specify 'inaccurate pose estimates of the regression-based approaches', because as written it could be misread as a statement about the Aachen dataset itself.
  3. [Sec. 3] The claim that the MASt3R limitation 'applies to all other 3D reconstruction approaches based on the relative pose regression [132, 129, 40]' cites Reloc3r [40], which is a relative pose regression method but not a 3D reconstruction approach; the citation should be adjusted to avoid implying Reloc3r builds a 3D model.
  4. [Throughout] There are several typos and formatting inconsistencies: 'NA VER' should be 'NAVER', 'RoMA' should be 'RoMa', 'ocal depth errors' is missing 'l', 'MAST3R depth' in the Fig. 3 caption should be 'MASt3R depth', and 'theE5+1' in Sec. 4.2 needs a space.
  5. [Sec. 1, contribution (c)] The phrase 'methods using less geometric reasoning can offer a better performance' is ambiguous; it should say 'better runtime performance' or 'a better accuracy-runtime trade-off'.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the paper is an empirical benchmark whose conclusions are measured against external datasets and standard baselines.

full rationale

This is a survey and benchmarking paper, not a derivation. The central claims (classical geometric reasoning beats pose regression by a wide margin; structureless methods are competitive with structure-based ones) are supported by measured localization recalls on Aachen Day-Night v1.1, Extended CMU Seasons, and the NAVER indoor datasets, reported in Tables 1-7. No equation is shown to reduce to its own inputs, and no fitted parameter is renamed as a prediction. The authors' own prior work appears as baselines (Ess. mat. [147], MeshLoc [85]), but it is not privileged: Ess. mat. is among the weaker structureless methods and MeshLoc trails Hloc, while the ranking among structureless families is driven by independent implementations (5Pt, E5+1, Local triangulation, LazyLoc, MASt3R pose align). The paper explicitly limits the regression evaluation to Aachen ('Given the inaccurate pose estimates observed for the Aachen dataset, coupled with long run-times, we did not evaluate the pose regression-based approaches on other datasets') and openly acknowledges that MASt3R pose align is suboptimal because the released implementation cannot use known intrinsics and database poses. These are evidence-strength and external-validity concerns, not circularity: they weaken the generality of the 'wide margin' claim, but the claim is not constructed from the method definitions or from self-citations. No self-citation is invoked as an unverified load-bearing premise, no uniqueness theorem is imported from the authors, and no ansatz is smuggled in via citation. The comparison is self-contained against external benchmarks, so the appropriate finding is no significant circularity.

Assumptions & free parameters 7 free parameters · 6 assumptions · 0 invented entities

Benchmark paper: no derivations or fitted models. The free parameters are evaluation thresholds and pipeline hyperparameters, several of which are dataset-dependent or chosen by grid search. Axioms are standard benchmark assumptions: reliable ground truth, calibrated cameras, correct external implementations, and the choice of MASt3R pose align as a representative of the regression family, which the paper itself flags as suboptimal.

free parameters (7)
  • Dense matcher track distance threshold = 5 px
    Threshold for merging dense matcher keypoints into tracks for Local triangulation; explicitly selected by grid search (Sec 3).
  • Local triangulation reprojection thresholds = 2 px (SuperPoint/ALIKED), 8 px (RoMa/MASt3R)
    Per-feature reprojection error thresholds for triangulation in Local triangulation variants (Sec 4, Implementation details).
  • Inlier counting threshold = 12 px
    Reprojection and epipolar error threshold for inlier counting across all methods (Sec 4, Implementation details).
  • RANSAC iteration bounds = 1,000 to 100,000
    Iteration limits for locally optimized RANSAC in E5+1, E3+1 and Local triangulation (Sec 4, Implementation details).
  • MASt3R pose align sampling = 3 retrieved images, 10 iterations, min 50 correspondences
    Heuristics to avoid wrong reference images skewing the alignment (Sec 3 and Sec 4, Implementation details).
  • MASt3R pose align optimization hyperparameters = 300 iterations per stage, learning rates 0.2 and 0.02
    Two-stage camera/point-map optimization settings (Sec 4, Implementation details).
  • Retrieval set size = top-10 (main), top-10/top-20 (structure-based comparison)
    Fixed retrieval count for all methods; results may shift with a different retrieval pool (Sec 4).
assumptions (6)
  • domain assumption Benchmark ground-truth poses are accurate enough to support centimeter/degree-level recall reporting
    The evaluation assumes reference poses of Aachen v1.1, CMU Seasons, and NAVER are reliable ground truth (Sec 4, Evaluation protocol).
  • domain assumption Cameras are calibrated
    All pose solvers used (5Pt, 3Pt, P3P, E5+1) assume known intrinsics; the paper states reference images have known camera intrinsics (Sec 3 and 4).
  • domain assumption External implementations are correct
    The paper relies on PoseLib, OpenCV, LightGlue, RoMa, MASt3R, LazyLoc (author-provided), and Reloc3r (author-provided) implementations without independent verification (Sec 3).
  • domain assumption Retrieval via EigenPlaces top-10 is sufficient for all methods
    All methods share the same retrieval step; the comparison assumes the retrieved set is adequate for each method family (Sec 4, Implementation details).
  • domain assumption Depth maps can be used in 3Pt solvers ignoring absolute scale
    Ess. mat. (3Pt+depth) and E3+1 use relative depth directions only; the scale of Metric3D vs MASt3R depth is noted as irrelevant (Sec 4.1).
  • ad hoc to paper The suboptimal out-of-the-box MASt3R pipeline represents the regression family
    The paper assumes MASt3R pose align, which cannot use known intrinsics/poses, is a fair representative of relative-pose regression; the paper itself flags this as suboptimal (Sec 3).

how reviews work

0 comments
Cite this review

Pith. "Pith review of A Guide to Structureless Visual Localization." pith.science (2026). https://pith.science/paper/XFZ3LPCI

@misc{pith2026250417636,
  author       = {Pith},
  title        = {Pith review of: A Guide to Structureless Visual Localization},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/XFZ3LPCI}},
  note         = {Machine review of arXiv:2504.17636}
}
read the original abstract

Visual localization algorithms, i.e., methods that estimate the camera pose of a query image in a known scene, are core components of many applications, including self-driving cars and augmented / mixed reality systems. State-of-the-art visual localization algorithms are structure-based, i.e., they store a 3D model of the scene and use 2D-3D correspondences between the query image and 3D points in the model for camera pose estimation. While such approaches are highly accurate, they are also rather inflexible when it comes to adjusting the underlying 3D model after changes in the scene. Structureless localization approaches represent the scene as a database of images with known poses and thus offer a much more flexible representation that can be easily updated by adding or removing images. Although there is a large amount of literature on structure-based approaches, there is significantly less work on structureless methods. Hence, this paper is dedicated to providing the, to the best of our knowledge, first comprehensive discussion and comparison of structureless methods. Extensive experiments show that approaches that use a higher degree of classical geometric reasoning generally achieve higher pose accuracy. In particular, approaches based on classical absolute or semi-generalized relative pose estimation outperform very recent methods based on pose regression by a wide margin. Compared with state-of-the-art structure-based approaches, the flexibility of structureless methods comes at the cost of (slightly) lower pose accuracy, indicating an interesting direction for future work.

Figures

Figures reproduced from arXiv: 2504.17636 by the authors.

Figure 1
Figure 1. Comparison of depth maps from different sources - on the top left is the source image. The corresponding source [PITH_FULL_IMAGE:figures/full_fig_p006_1.png] view at source ↗
Figure 2
Figure 2. Localization results for the Ess. mat. (5Pt) approach for different features. We report localization recalls (higher is better) on the Y-axis at multiple pose thresholds (X-axis). For the outdoor scenes, the best results are obtained with the RoMa matcher. For the indoor scenes, the MASt3R matcher performs best for the coarser thresholds. Pose triangulation. We present the results for Ess. mat. us￾ing the 5Pt solver… view at source ↗
Figure 3
Figure 3. Localization results for the Ess. mat. (3Pt + depth) approach for different features and monocular depth predictors. We report localization recalls (higher is better) on the Y-axis at multiple pose thresholds (X-axis). For most scenes, the choice of the depth predictor is not critical. For outdoor scenes, RoMa yields the best results. For indoor scenes, MASt3R leads to the highest pose accuracy in most cases. Aachen… view at source ↗
Figures from the paper (6 more)
Figure 4
Figure 4. Figure 4: LazyLoc localization results for different features. We report localization recalls (higher is better) on the Y-axis at multiple pose thresholds (X-axis). There is no type of feature that performs best in all scenes. However, the MASt3R matcher performs well in general…
Figure 5
Figure 5. Figure 5: E5+1 localization results for different features. We report localization recalls (higher is better) on the Y-axis at multiple pose thresholds (X-axis). For the outdoor scenes, the best results are typically obtained with the RoMa matcher. For the indoor scenes, the MAS…
Figure 6
Figure 6. Figure 6: E3+1 localization results for different features and depth predictors. We report localization recalls (higher is better) on the Y-axis at multiple pose thresholds (X-axis). The choice of the depth predictor is not critical. estingly, both SuperPoint and ALIKED outperfo…
Figure 7
Figure 7. Figure 7: Localization results for the local 3D point triangulation from all retrieved images ( [PITH_FULL_IMAGE:figures/full_fig_p009_7.png]
Figure 8
Figure 8. Figure 8: Localization results for the local 3D point triangulation from reference image pairs ( [PITH_FULL_IMAGE:figures/full_fig_p009_8.png]
Figure 9
Figure 9. Figure 9: Comparison of the query images with high camera position error (top row) and the images with low camera position [PITH_FULL_IMAGE:figures/full_fig_p012_9.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. RoMa v2: Harder Better Faster Denser Feature Matching

    cs.CV 2025-11 conditional novelty 5.0 of 10

    RoMa v2 reports state-of-the-art dense matching on six benchmarks and top relative-pose accuracy on MegaDepth-1500 and ScanNet-1500, runs 1.7× faster than RoMa, and adds pixel-wise covariance predictions.

Reference graph

Works this paper leans on

148 extracted references · 69 canonical work pages · cited by 1 Pith paper

  1. [1]

    Objects can move: 3d change detection by geometric transformation consistency

    Aikaterini Adam, Torsten Sattler, Konstantinos Karantzalos, and Tomas Pajdla. Objects can move: 3d change detection by geometric transformation consistency. In European Confer- ence on Computer Vision (ECCV) , pages 108–124. Springer,

  2. [2]

    NetVLAD: CNN Architecture for Weakly Supervised Place Recognition

    Relja Arandjelovi ´c, Petr Gronát, Akihiko Torii, Tomás Pa- jdla, and Josef Sivic. NetVLAD: CNN Architecture for Weakly Supervised Place Recognition. 2016 IEEE Confer- ence on Computer Vision and Pattern Recognition (CVPR) , pages 5297–5307, 2015. 1, 2

  3. [3]

    All about VLAD

    Relja Arandjelovic and Andrew Zisserman. All about VLAD. In Proceedings of the IEEE conference on Computer Vision and Pattern Recognition, pages 1578–1585, 2013. 1

  4. [4]

    Arandjelovi ´c and A

    R. Arandjelovi ´c and A. Zisserman. Visual vocabulary with a semantic twist. In Asian Conference on Computer Vision (ACCV), 2014. 2

  5. [5]

    GIS-Assisted Object Detection and Geospatial Localization

    Shervin Ardeshir, Amir Roshan Zamir, Alejandro Torroella, and Mubarak Shah. GIS-Assisted Object Detection and Geospatial Localization. In ECCV, 2014. 2

  6. [6]

    Map-free Vi- sual Relocalization: Metric Pose Relative to a Single Image

    Eduardo Arnold, Jamie Wynn, Sara Vicente, Guillermo Garcia-Hernando, Aron Monszpart, Victor Prisacariu, Dani- yar Turmukhambetov, and Eric Brachmann. Map-free Vi- sual Relocalization: Metric Pose Relative to a Single Image. In European Conference on Computer Vision (ECCV), pages 690–708. Springer, 2022. 1, 10

  7. [7]

    C. Arth, D. Wagner, M. Klopschitz, A. Irschara, and D. Schmalstieg. Wide area localization on mobile phones. In ISMAR, 2009. 2

  8. [8]

    Baatz, K

    G. Baatz, K. Köser, D. Chen, R. Grzeszczuk, and M. Polle- feys. Handling Urban Location Recognition as a 2D Homo- thetic Problem. In ECCV, 2010. 2

Show all 148 references
  1. [9]

    Chen, Radek Grzeszczuk, and Marc Pollefeys

    Georges Baatz, Kevin Köser, David M. Chen, Radek Grzeszczuk, and Marc Pollefeys. Leveraging 3D City Models for Rotation Invariant Place-of-Interest Recognition. IJCV, 96:315–334, 2011. 2

  2. [10]

    The CMU Visual Localization Data Set

    Hernan Badino, Daniel Huber, and Takeo Kanade. The CMU Visual Localization Data Set. http://3dvis.ri.cmu. edu/data-sets/localization, 2011. 5, 10, 11

  3. [11]

    Reloc- Net: Continuous Metric Learning Relocalisation using Neu- ral Nets

    Vassileios Balntas, Shuda Li, and Victor Prisacariu. Reloc- Net: Continuous Metric Learning Relocalisation using Neu- ral Nets. In The European Conference on Computer Vision (ECCV), September 2018. 1, 3

  4. [12]

    MAGSAC++, a Fast, Reliable and Accurate Robust Estimator

    Dániel Baráth, Jana Noskova, Maksym Ivashechkin, and Jiri Matas. MAGSAC++, a Fast, Reliable and Accurate Robust Estimator. 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 1301–1309, 2019. 1, 2

  5. [13]

    Re- thinking Visual Geo-Localization for Large-Scale Applica- tions

    Gabriele Berton, Carlo Masone, and Barbara Caputo. Re- thinking Visual Geo-Localization for Large-Scale Applica- tions. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022. 2

  6. [14]

    EigenPlaces: Training Viewpoint Robust Models for Visual Place Recognition

    Gabriele Berton, Gabriele Trivigno, Barbara Caputo, and Carlo Masone. EigenPlaces: Training Viewpoint Robust Models for Visual Place Recognition. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 11080–11090, October 2023. 1, 2, 5, 9, 11, 12

  7. [15]

    Heikkila, and Zuzana Kukelova

    Snehal Bhayani, Torsten Sattler, Dániel Baráth, Patrik Be- liansky, J. Heikkila, and Zuzana Kukelova. Calibrated and Partially Calibrated Semi-Generalized Homographies. 2021 IEEE/CVF International Conference on Computer Vision (ICCV), pages 5916–5925, 2021. 1, 3

  8. [16]

    Gsloc: Visual localization with 3d gaussian splatting

    Kazii Botashev, Vladislav Pyatov, Gonzalo Ferrer, and Sta- matios Lefkimmiatis. Gsloc: Visual localization with 3d gaussian splatting. In 2024 IEEE/RSJ International Confer- ence on Intelligent Robots and Systems (IROS) , pages 5664–

  9. [17]

    Accelerated Coordinate Encoding: Learning to Relocalize in Minutes Using RGB and Poses

    Eric Brachmann, Tommaso Cavallari, and Victor Adrian Prisacariu. Accelerated Coordinate Encoding: Learning to Relocalize in Minutes Using RGB and Poses. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 5044–5053, 2023. 1, 2

  10. [18]

    Humenberger, Carsten Rother, and Torsten Sattler

    Eric Brachmann, M. Humenberger, Carsten Rother, and Torsten Sattler. On the Limits of Pseudo Ground Truth in Vi- sual Camera Re-localisation. 2021 IEEE/CVF International Conference on Computer Vision (ICCV) , pages 6198–6208,

  11. [19]

    DSAC - Differentiable RANSAC for Camera Local- ization

    Eric Brachmann, Alexander Krull, Sebastian Nowozin, Jamie Shotton, Frank Michel, Stefan Gumhold, and Carsten Rother. DSAC - Differentiable RANSAC for Camera Local- ization. In CVPR, 2017. 2

  12. [20]

    Learning Less is More - 6D Camera Localization via 3D Surface Regression

    Eric Brachmann and Carsten Rother. Learning Less is More - 6D Camera Localization via 3D Surface Regression. In CVPR, 2018. 2

  13. [21]

    Visual Camera Re- Localization From RGB and RGB-D Images Using DSAC

    Eric Brachmann and Carsten Rother. Visual Camera Re- Localization From RGB and RGB-D Images Using DSAC. IEEE Transactions on Pattern Analysis and Machine Intelli- gence, 44:5847–5865, 2020. 1, 2

  14. [22]

    G. Bradski. The OpenCV Library. Dr . Dobb’s Journal of Software Tools, 2000. 4

  15. [23]

    Geometry-aware learning of maps for camera localization

    Samarth Brahmbhatt, Jinwei Gu, Kihwan Kim, James Hays, and Jan Kautz. Geometry-aware learning of maps for camera localization. In CVPR, 2018. 3

  16. [24]

    Large Scale Joint Semantic Re-Localisation and Scene Understanding via Globally Unique Instance Co- ordinate Regression

    Ignas Budvytis, Marvin Teichmann, Tomas V ojir, and Roberto Cipolla. Large Scale Joint Semantic Re-Localisation and Scene Understanding via Globally Unique Instance Co- ordinate Regression. In BMVC, 2019. 2

  17. [25]

    Cao and N

    S. Cao and N. Snavely. Graph-Based Discriminative Learn- ing for Location Recognition. In CVPR, 2013. 2

  18. [26]

    Let’s take this online: Adapting scene coordinate regression network predictions for online RGB-D camera relocalisation

    Tommaso Cavallari, Luca Bertinetto, Jishnu Mukhoti, Philip Torr, and Stuart Golodetz. Let’s take this online: Adapting scene coordinate regression network predictions for online RGB-D camera relocalisation. In 3DV, 2019. 2

  19. [27]

    Lord, Julien Valentin, Victor A

    Tommaso Cavallari, Stuart Golodetz, Nicholas A. Lord, Julien Valentin, Victor A. Prisacariu, Luigi Di Stefano, and Philip H. S. Torr. Real-time RGB-D camera pose estimation in novel scenes using a relocalisation cascade. TPAMI, 2019. 1, 2

  20. [28]

    Chen, Georges Baatz, Kevin Köser, Sam S

    David M. Chen, Georges Baatz, Kevin Köser, Sam S. Tsai, Ramakrishna Vedantham, Timo Pylvänäinen, Kimmo Roimela, Xin Chen, Jeff Bach, Marc Pollefeys, Bernd Girod, and Radek Grzeszczuk. City-Scale Landmark Identification on Mobile Devices. In CVPR, 2011. 2

  21. [29]

    Refinement for absolute pose regression with neural feature synthesis

    Shuai Chen, Yash Bhalgat, Xinghui Li, Jiawang Bian, Kejie Li, Zirui Wang, and Victor Adrian Prisacariu. Refinement for absolute pose regression with neural feature synthesis. arXiv preprint arXiv:2303.10087, 2023. 1, 2, 3

  22. [30]

    Dfnet: Enhance absolute pose regression with di- rect feature matching

    Shuai Chen, Xinghui Li, Zirui Wang, and Victor A Prisacariu. Dfnet: Enhance absolute pose regression with di- rect feature matching. In European Conference on Computer 13 Vision, pages 1–17. Springer Nature Switzerland Cham,

  23. [31]

    Reid, and Michael Milford

    Zetao Chen, Adam Jacobson, Niko Sünderhauf, Ben Up- croft, Lingqiao Liu, Chunhua Shen, Ian D. Reid, and Michael Milford. Deep Learning Features at Scale for Visual Place Recognition. ICRA, 2017. 2

  24. [32]

    Choudhary and P

    S. Choudhary and P. J. Narayanan. Visibility probability structure from sfm datasets and applications. InECCV, 2012. 2

  25. [33]

    Optimal randomized ransac

    Ond ˇrej Chum and Ji ˇrí Matas. Optimal randomized ransac. IEEE Transactions on Pattern Analysis and Machine Intelli- gence, 30(8):1472–1482, 2008. 1, 2

  26. [34]

    Global structure-from-motion by similarity averaging

    Zhaopeng Cui and Ping Tan. Global structure-from-motion by similarity averaging. In Proceedings of the IEEE interna- tional conference on computer vision , pages 864–872, 2015. 3

  27. [35]

    SuperPoint: Self-Supervised Interest Point Detection and Description

    Daniel DeTone, Tomasz Malisiewicz, and Andrew Rabi- novich. SuperPoint: Self-Supervised Interest Point Detection and Description. 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), pages 337–33712, 2017. 4, 5, 10, 11

  28. [36]

    CamNet: Coarse-to-fine retrieval for camera re- localization

    Mingyu Ding, Zhe Wang, Jiankai Sun, Jianping Shi, and Ping Luo. CamNet: Coarse-to-fine retrieval for camera re- localization. In ICCV, 2019. 3

  29. [37]

    RePoseD: Efficient Relative Pose Estimation With Known Depth Information

    Yaqing Ding, Viktor Kocur, Václav Vávra, Zuzana Berger Haladová, Jian Yang, Torsten Sattler, and Zuzana Kukelova. RePoseD: Efficient Relative Pose Estimation With Known Depth Information. arXiv preprint arXiv:2501.07742, 2025. 4, 10

  30. [38]

    Revisiting the P3P problem

    Yaqing Ding, Jian Yang, Viktor Larsson, Carl Olsson, and Kalle Åström. Revisiting the P3P problem. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4872–4880, 2023. 1, 2, 4

  31. [39]

    Lazy Visual Localization via Motion Aver- aging

    Siyan Dong, Shaohui Liu, Hengkai Guo, Baoquan Chen, and Marc Pollefeys. Lazy Visual Localization via Motion Aver- aging. arXiv:2307.09981, 2023. 1, 3, 4, 10, 11, 12

  32. [40]

    Reloc3r: Large-Scale Training of Relative Camera Pose Regression for Generaliz- able, Fast, and Accurate Visual Localization

    Siyan Dong, Shuzhe Wang, Shaohui Liu, Lulu Cai, Qingnan Fan, Juho Kannala, and Yanchao Yang. Reloc3r: Large-Scale Training of Relative Camera Pose Regression for Generaliz- able, Fast, and Accurate Visual Localization. arXiv preprint arXiv:2412.08376, 2024. 2, 3, 4, 5, 9

  33. [41]

    Visual localization via few-shot scene region classification

    Siyan Dong, Shuzhe Wang, Yixin Zhuang, Juho Kannala, Marc Pollefeys, and Baoquan Chen. Visual localization via few-shot scene region classification. In 2022 International Conference on 3D Vision (3DV), pages 393–402. IEEE, 2022. 2

  34. [42]

    RoMa: Robust Dense Feature Matching

    Johan Edstedt, Qiyu Sun, Georg Bökman, Mårten Waden- bäck, and Michael Felsberg. RoMa: Robust Dense Feature Matching. IEEE Conference on Computer Vision and Pattern Recognition, 2024. 4, 5, 9

  35. [43]

    Fischler and Robert C

    Martin A. Fischler and Robert C. Bolles. Random sample consensus: a paradigm for model fitting with applications to image analysis and automated cartography. Commun. ACM, 24:381–395, 1981. 1, 2, 3

  36. [44]

    Gordo, J

    A. Gordo, J. Almazan, J. Revaud, and D. Larlus. End-to- end Learning of Deep Visual Representations for Image Re- trieval. IJCV, 2017. 1

  37. [45]

    Das pothenotische problem in er- weiterter gestalt nebst bber seine anwendungen in der geo- dasie

    Johann August Grunert. Das pothenotische problem in er- weiterter gestalt nebst bber seine anwendungen in der geo- dasie. Grunerts Archiv fur Mathematik und Physik , pages 238–248, 1841. 1

  38. [46]

    Multi-output learning for camera relocalization

    Abner Guzman-Rivera, Pushmeet Kohli, Ben Glocker, Jamie Shotton, Toby Sharp, Andrew Fitzgibbon, and Shahram Izadi. Multi-output learning for camera relocalization. In CVPR, 2014. 2

  39. [47]

    Review and analysis of solutions of the three point perspective pose estimation problem

    Bert M Haralick, Chung-Nan Lee, Karsten Ottenberg, and Michael Nölle. Review and analysis of solutions of the three point perspective pose estimation problem. International journal of computer vision (IJCV) , 13:331–356, 1994. 1, 2

  40. [48]

    Multiple View Ge- ometry in Computer Vision

    Richard Hartley and Andrew Zisserman. Multiple View Ge- ometry in Computer Vision. Cambridge University Press, 2nd edition, 2001. 4

  41. [49]

    Patch-netvlad: Multi-scale fusion of locally-global descriptors for place recognition

    Stephen Hausler, Sourav Garg, Ming Xu, Michael Milford, and Tobias Fischer. Patch-netvlad: Multi-scale fusion of locally-global descriptors for place recognition. In Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pages 14141–14152, June

  42. [50]

    Project AutoVision: Localiza- tion and 3D Scene Perception for an Autonomous Vehicle with a Multi-Camera System

    Lionel Heng, Benjamin Choi, Zhaopeng Cui, Marcel Gep- pert, Sixing Hu, Benson Kuan, Peidong Liu, Rang Ho Man Nguyen, Ye Chuan Yeo, Andreas Geiger, Gim Hee Lee, Marc Pollefeys, and Torsten Sattler. Project AutoVision: Localiza- tion and 3D Scene Perception for an Autonomous Veh...

  43. [51]

    Xiaoyan Zhang, Zhipeng Cai, Xi- aoxiao Long, Hao Chen, Kaixuan Wang, Gang Yu, Chunhua Shen, and Shaojie Shen

    Mu Hu, Wei Yin, China. Xiaoyan Zhang, Zhipeng Cai, Xi- aoxiao Long, Hao Chen, Kaixuan Wang, Gang Yu, Chunhua Shen, and Shaojie Shen. Metric3d v2: A versatile monoc- ular geometric foundation model for zero-shot metric depth and surface normal estimation. IEEE transactions on p...

  44. [52]

    Robust Image Retrieval-based Visual Localization using Kapture

    Martin Humenberger, Yohann Cabon, Nicolas Guerin, Julien Morat, Jérôme Revaud, Philippe Rerole, Noé Pion, Cesar de Souza, Vincent Leroy, and Gabriela Csurka. Robust Image Retrieval-based Visual Localization using Kapture. arXiv:2007.13867, 2022. 2

  45. [53]

    Investigating the Role of Im- age Retrieval for Visual Localization: An Exhaustive Bench- mark

    Martin Humenberger, Yohann Cabon, Noé Pion, Philippe Weinzaepfel, Donghwan Lee, Nicolas Guérin, Torsten Sat- tler, and Gabriela Csurka. Investigating the Role of Im- age Retrieval for Visual Localization: An Exhaustive Bench- mark. IJCV, 130(7):1811–1836, Jul 2022. 3

  46. [54]

    Irschara, C

    A. Irschara, C. Zach, J.-M. Frahm, and H. Bischof. From Structure-from-Motion Point Clouds to Fast Location Recog- nition. In CVPR, 2009. 2

  47. [55]

    Aggregating local descriptors into a compact image representation

    Hervé Jégou, Matthijs Douze, Cordelia Schmid, and Patrick Pérez. Aggregating local descriptors into a compact image representation. 2010 IEEE Computer Society Conference on Computer Vision and Pattern Recognition, pages 3304–3311,

  48. [56]

    A solution for the best rotation to relate two sets of vectors

    Wolfgang Kabsch. A solution for the best rotation to relate two sets of vectors. Acta Crystallographica Section A: Crys- tal Physics, Diffraction, Theoretical and General Crystallog- raphy, 1976. 4

  49. [57]

    Screened Poisson Sur- face Reconstruction

    Michael Kazhdan and Hugues Hoppe. Screened Poisson Sur- face Reconstruction. ACM Trans. Graph., 32(3), July 2013. 1

  50. [58]

    Geometric loss functions for camera pose regression with deep learning

    Alex Kendall and Roberto Cipolla. Geometric loss functions for camera pose regression with deep learning. In CVPR,

  51. [59]

    PoseNet: A Convolutional Network for Real-Time 6-DOF Camera Relocalization

    Alex Kendall, Matthew Grimes, and Roberto Cipolla. PoseNet: A Convolutional Network for Real-Time 6-DOF Camera Relocalization. In ICCV, 2015. 3

  52. [60]

    3D Gaussian Splatting for Real-Time Radiance Field Rendering

    Bernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, and George Drettakis. 3D Gaussian Splatting for Real-Time Radiance Field Rendering. ACM Trans. Graph., 42(4):139– 1, 2023. 1

  53. [61]

    PoseLib - Minimal Solvers for Camera Pose Estimation, 2020

    Viktor Larsson and contributors. PoseLib - Minimal Solvers for Camera Pose Estimation, 2020. 4

  54. [62]

    Camera Relocalization by Computing Pairwise 14 Relative Poses Using Convolutional Neural Network

    Zakaria Laskar, Iaroslav Melekhov, Surya Kalia, and Juho Kannala. Camera Relocalization by Computing Pairwise 14 Relative Poses Using Convolutional Neural Network. In ICCV Workshops, 2017. 1, 3

  55. [63]

    Sala Matas, and Ond ˇrej Chum

    Karel Lebeda, Juan E. Sala Matas, and Ond ˇrej Chum. Fixing the Locally Optimized RANSAC. In BMVC, 2012. 1, 2, 3, 4

  56. [64]

    Hu- menberger

    Donghwan Lee, Soohyun Ryu, Suyong Yeon, Yonghan Lee, Deok-Won Kim, Cheolho Han, Yohann Cabon, Philippe Weinzaepfel, Nicolas Gu’erin, Gabriela Csurka, and M. Hu- menberger. Large-scale Localization Datasets in Crowded Indoor Spaces. 2021 IEEE/CVF Conference on Computer Vision a...

  57. [65]

    Ground- ing Image Matching in 3D with MASt3R, 2024

    Vincent Leroy, Yohann Cabon, and Jerome Revaud. Ground- ing Image Matching in 3D with MASt3R, 2024. 2, 3, 4, 5, 6, 9, 12

  58. [66]

    Hierarchical scene coordinate classification and regression for visual localization

    Xiaotian Li, Shuzhe Wang, Yi Zhao, Jakob Verbeek, and Juho Kannala. Hierarchical scene coordinate classification and regression for visual localization. In CVPR, 2020. 2

  59. [67]

    Huttenlocher

    Yunpeng Li, Noah Snavely, and Dan P. Huttenlocher. Lo- cation Recognition using Prioritized Feature Matching. In ECCV, 2010. 2

  60. [68]

    Huttenlocher, and Pas- cal V

    Yunpeng Li, Noah Snavely, Daniel P. Huttenlocher, and Pas- cal V . Fua. Worldwide Pose Estimation Using 3D Point Clouds. In European Conference on Computer Vision, 2012. 1, 2

  61. [69]

    Sinha, Michael F

    Hyon Lim, Sudipta N. Sinha, Michael F. Cohen, and Matthew Uyttendaele. Real-time image-based 6-DOF local- ization in large-scale environments. 2012 IEEE Conference on Computer Vision and Pattern Recognition , pages 1043– 1050, 2012. 1

  62. [70]

    Florence, Jonathan T

    Yen-Chen Lin, Peter R. Florence, Jonathan T. Barron, Al- berto Rodriguez, Phillip Isola, and Tsung-Yi Lin. iNeRF: Inverting Neural Radiance Fields for Pose Estimation. 2021 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pages 1323–1330, 2020. 1, 3

  63. [71]

    LightGlue: Local Feature Matching at Light Speed

    Philipp Lindenberger, Paul-Edouard Sarlin, and Marc Polle- feys. LightGlue: Local Feature Matching at Light Speed. In ICCV, 2023. 5, 10, 11

  64. [72]

    Gs-cpr: Efficient camera pose refinement via 3d gaussian splatting

    Changkun Liu, Shuai Chen, Yash Sanjay Bhalgat, Siyan Hu, Ming Cheng, Zirui Wang, Victor Adrian Prisacariu, and Tris- tan Braud. Gs-cpr: Efficient camera pose refinement via 3d gaussian splatting. In The Thirteenth International Confer- ence on Learning Representations, 2024. 1, 3

  65. [73]

    Nerf- loc: Visual localization with conditional neural radiance field

    Jianlin Liu, Qiang Nie, Yong Liu, and Chengjie Wang. Nerf- loc: Visual localization with conditional neural radiance field. 2023 IEEE International Conference on Robotics and Automation (ICRA), 2023. 1, 2

  66. [74]

    Visual place recognition: A survey

    Stephanie Lowry, Niko Sünderhauf, Paul Newman, John J Leonard, David Cox, Peter Corke, and Michael J Milford. Visual place recognition: A survey. IEEE Transactions on Robotics, 32(1):1–19, 2016. 2

  67. [75]

    Hesch, Marc Pollefeys, and Roland Y

    Simon Lynen, Torsten Sattler, Michael Bosse, Joel A. Hesch, Marc Pollefeys, and Roland Y . Siegwart. Get Out of My Lab: Large-scale, Real-Time Visual-Inertial Localization. In Robotics: Science and Systems , 2015. 1, 2

  68. [76]

    Massiceti, A

    D. Massiceti, A. Krull, E. Brachmann, C. Rother, and P. H.S. Torr. Random Forests versus Neural Networks - What’s Best for Camera Relocalization? In ICRA, 2017. 2

  69. [77]

    6dgs: 6d pose estimation from a single image and a 3d gaussian splatting model

    Bortolon Matteo, Theodore Tsesmelis, Stuart James, Fabio Poiesi, and Alessio Del Bue. 6dgs: 6d pose estimation from a single image and a 3d gaussian splatting model. In Aleš Leonardis, Elisa Ricci, Stefan Roth, Olga Russakovsky, Torsten Sattler, and Gül Varol, editors, Compute...

  70. [78]

    Sven Middelberg, Torsten Sattler, Ole Untzelmann, and Leif P. Kobbelt. Scalable 6-DOF Localization on Mobile De- vices. In European Conference on Computer Vision, 2014. 1

  71. [79]

    Srinivasan, Matthew Tancik, Jonathan T

    Ben Mildenhall, Pratul P. Srinivasan, Matthew Tancik, Jonathan T. Barron, Ravi Ramamoorthi, and Ren Ng. NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis. Commun. ACM, 65:99–106, 2020. 1

  72. [80]

    LENS: Localiza- tion enhanced by neRF synthesis

    Arthur Moreau, Nathan Piasco, Dzmitry Tsishkou, Bogdan Stanciulescu, and Arnaud de La Fortelle. LENS: Localiza- tion enhanced by neRF synthesis. In CoRL, 2021. 3

  73. [81]

    OoD-Pose: Camera Pose Regression From Out-of-Distribution Synthetic Views

    Tony Ng, Adrian Lopez-Rodriguez, Vassileios Balntas, and Krystian Mikolajczyk. OoD-Pose: Camera Pose Regression From Out-of-Distribution Synthetic Views. In 2022 Interna- tional Conference on 3D Vision (3DV), 2022. 3

  74. [82]

    D. Nistér. An Efficient Solution to the Five-Point Relative Pose Problem. IEEE Transactions on Pattern Analysis and Machine Intelligence, 26(6):756–770, June 2004. 3, 4

  75. [83]

    An efficient solution to the five-point relative pose problem

    David Nistér. An efficient solution to the five-point relative pose problem. IEEE Transactions on Pattern Analysis and Machine Intelligence, 26:756–770, 2004. 4

  76. [84]

    Global structure-from-motion revisited

    Linfei Pan, Dániel Baráth, Marc Pollefeys, and Johannes L Schönberger. Global structure-from-motion revisited. In European Conference on Computer Vision , pages 58–77. Springer, 2024. 3

  77. [85]

    Meshloc: Mesh-based visual localization

    V ojtech Panek, Zuzana Kukelova, and Torsten Sattler. Meshloc: Mesh-based visual localization. In ECCV, 2022. 1, 2, 5, 6, 11, 12

  78. [86]

    Visual Localization using Imperfect 3D Models from the Internet

    V ojtech Panek, Zuzana Kukelova, and Torsten Sattler. Visual Localization using Imperfect 3D Models from the Internet. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 13175–13186, 2023. 1

  79. [87]

    Vi- sual Localization Using Imperfect 3D Models From the In- ternet

    V ojtech Panek, Zuzana Kukelova, and Torsten Sattler. Vi- sual Localization Using Imperfect 3D Models From the In- ternet. In Proceedings of the IEEE/CVF Conference on Com- puter Vision and Pattern Recognition (CVPR), pages 13175– 13186, June 2023. 2

  80. [88]

    Lambda twist: An accu- rate fast robust perspective three point (p3p) solver

    Mikael Persson and Klas Nordberg. Lambda twist: An accu- rate fast robust perspective three point (p3p) solver. In Euro- pean Conference on Computer Vision (ECCV), 2018. 1, 2, 4, 5, 9

  81. [89]

    Philbin, O

    J. Philbin, O. Chum, M. Isard, J. Sivic, and A. Zisserman. Object Retrieval with Large V ocabularies and Fast Spatial Matching. In CVPR, 2007. 2

  82. [90]

    Philbin, M

    J. Philbin, M. Isard, J. Sivic, and A. Zisserman. Descriptor learning for efficient retrieval. In ECCV, 2010. 2

  83. [91]

    Self-Supervised Learning of Neural Im- plicit Feature Fields for Camera Pose Refinement

    Maxime Pietrantoni, Gabriela Csurka, Martin Humenberger, and Torsten Sattler. Self-Supervised Learning of Neural Im- plicit Feature Fields for Camera Pose Refinement. In 2024 International Conference on 3D Vision (3DV) , pages 484–

  84. [92]

    Segloc: Learning segmentation-based representations for privacy-preserving visual localization

    Maxime Pietrantoni, Martin Humenberger, Torsten Sattler, and Gabriela Csurka. Segloc: Learning segmentation-based representations for privacy-preserving visual localization. In Proceedings of the IEEE/CVF Conference on Computer Vi- sion and Pattern Recognition (CVPR) , pages 1...

  85. [93]

    Benchmarking Image Retrieval for Visual Localization

    Noé Pion, Martin Humenberger, Gabriela Csurka, Yohann Cabon, and Torsten Sattler. Benchmarking Image Retrieval for Visual Localization. In 3DV, 2020. 3

  86. [94]

    Fine- Tuning CNN Image Retrieval with No Human Annotation

    Filip Radenovi ´c, Giorgos Tolias, and Ond ˇrej Chum. Fine- Tuning CNN Image Retrieval with No Human Annotation. TPAMI, 2019. 2

  87. [95]

    Revaud, J

    J. Revaud, J. Almazan, R.S. Rezende, and C.R. de Souza. Learning with Average Precision: Training Image Retrieval with a Listwise Loss. In ICCV, 2019. 1

  88. [96]

    R2D2: repeatable and reli- able detector and descriptor

    Jerome Revaud, Philippe Weinzaepfel, César Roberto de Souza, and Martin Humenberger. R2D2: repeatable and reli- able detector and descriptor. In NeurIPS, 2019. 12 15

  89. [97]

    From Coarse to Fine: Robust Hierarchical Localization at Large Scale

    Paul-Edouard Sarlin, Cesar Cadena, Roland Siegwart, and Marcin Dymczyk. From Coarse to Fine: Robust Hierarchical Localization at Large Scale. In CVPR, 2019. 1, 2, 11, 12

  90. [98]

    Leveraging Deep Vi- sual Descriptors for Hierarchical Efficient Localization

    Paul-Edouard Sarlin, Frédéric Debraine, Marcin Dymczyk, Roland Siegwart, and Cesar Cadena. Leveraging Deep Vi- sual Descriptors for Hierarchical Efficient Localization. In Conference on Robot Learning (CoRL) , 2018. 2

  91. [99]

    SuperGlue: Learning Feature Matching with Graph Neural Networks

    Paul-Edouard Sarlin, Daniel DeTone, Tomasz Malisiewicz, and Andrew Rabinovich. SuperGlue: Learning Feature Matching with Graph Neural Networks. In CVPR, 2020. 1, 11, 12

  92. [100]

    Back to the Feature: Learning Robust Cam- era Localization from Pixels to Pose

    Paul-Edouard Sarlin, Ajaykumar Unagar, Måns Larsson, Hugo Germain, Carl Toft, Victor Larsson, Marc Pollefeys, Vincent Lepetit, Lars Hammarstrand, Fredrik Kahl, and Torsten Sattler. Back to the Feature: Learning Robust Cam- era Localization from Pixels to Pose. In CVPR, 2021. 1, 3

  93. [101]

    Hyperpoints and fine vocab- ularies for large-scale location recognition

    Torsten Sattler, Michal Havlena, Filip Radenovic, Konrad Schindler, and Marc Pollefeys. Hyperpoints and fine vocab- ularies for large-scale location recognition. In ICCV, 2015. 2

  94. [102]

    Sattler, B

    T. Sattler, B. Leibe, and L. Kobbelt. Fast Image-Based Lo- calization using Direct 2D-to-3D Matching. In ICCV, 2011. 2

  95. [103]

    Improv- ing Image-Based Localization by Active Correspondence Search

    Torsten Sattler, Bastian Leibe, and Leif Kobbelt. Improv- ing Image-Based Localization by Active Correspondence Search. In ECCV, 2012. 2

  96. [104]

    Leibe, and Leif P

    Torsten Sattler, B. Leibe, and Leif P. Kobbelt. Efficient & Effective Prioritized Matching for Large-Scale Image-Based Localization. IEEE Transactions on Pattern Analysis and Machine Intelligence, 39:1744–1756, 2017. 1

  97. [105]

    Benchmarking 6DOF Urban Visual Local- ization in Changing Conditions

    Torsten Sattler, Will Maddern, Carl Toft, Akihiko Torii, Lars Hammarstrand, Erik Stenborg, Daniel Safari, Masatoshi Okutomi, Marc Pollefeys, Josef Sivic, Fredrik Kahl, and Tomas Pajdla. Benchmarking 6DOF Urban Visual Local- ization in Changing Conditions. In CVPR, 2018. 5, 8, ...

  98. [106]

    Image Retrieval for Image-Based Localization Re- visited

    Torsten Sattler, Tobias Weyand, Bastian Leibe, and Leif Kobbelt. Image Retrieval for Image-Based Localization Re- visited. In BMVC, 2012. 2, 5, 8, 9, 10, 11, 12

  99. [107]

    Understanding the limitations of cnn-based ab- solute camera pose regression

    Torsten Sattler, Qunjie Zhou, Marc Pollefeys, and Laura Leal-Taixé. Understanding the limitations of cnn-based ab- solute camera pose regression. In CVPR, 2019. 3

  100. [108]

    Learning Multi- Scene Absolute Pose Regression With Transformers

    Yoli Shavit, Ron Ferens, and Yosi Keller. Learning Multi- Scene Absolute Pose Regression With Transformers. In ICCV, 2021. 3

  101. [109]

    Scene Coordinate Regression Forests for Camera Relocaliza- tion in RGB-D Images

    Jamie Shotton, Ben Glocker, Christopher Zach, Shahram Izadi, Antonio Criminisi, and Andrew William Fitzgibbon. Scene Coordinate Regression Forests for Camera Relocaliza- tion in RGB-D Images. 2013 IEEE Conference on Computer Vision and Pattern Recognition, pages 2930–2937, 2013. 1, 2

  102. [110]

    Semantically Guided Geo- location and Modeling in Urban Environments

    Gautam Singh and Jana Košecká. Semantically Guided Geo- location and Modeling in Urban Environments. In Large- Scale Visual Geo-Localization, 2016. 2

  103. [111]

    Sivic and A

    J. Sivic and A. Zisserman. Video Google: A Text Retrieval Approach to Object Matching in Videos. In ICCV, 2003. 2

  104. [112]

    Lowe, and J

    Stephen Se, D. Lowe, and J. Little. Global localization using distinctive visual features. In IEEE/RSJ International Con- ference on Intelligent Robots and Systems , 2002. 2

  105. [113]

    LoFTR: Detector-Free Local Feature Match- ing with Transformers

    Jiaming Sun, Zehong Shen, Yuang Wang, Hujun Bao, and Xiaowei Zhou. LoFTR: Detector-Free Local Feature Match- ing with Transformers. 2021 IEEE/CVF Conference on Com- puter Vision and Pattern Recognition (CVPR) , pages 8918– 8927, 2021. 4

  106. [114]

    icomma: Inverting 3d gaus- sian splatting for camera pose estimation via comparing and matching

    Yuan Sun, Xuan Wang, Yunfan Zhang, Jie Zhang, Caigui Jiang, Yu Guo, and Fei Wang. icomma: Inverting 3d gaus- sian splatting for camera pose estimation via comparing and matching. arXiv preprint arXiv:2312.09031, 2023. 1, 3

  107. [115]

    Svärm, O

    L. Svärm, O. Enqvist, F. Kahl, and M. Oskarsson. City- Scale Localization for Cameras with Known Vertical Direc- tion. PAMI, 39(7):1455–1461, 2017. 2

  108. [116]

    Okutomi, Torsten Sattler, Mircea Cimpoi, Marc Pollefeys, Josef Sivic, Tomás Pajdla, and Akihiko Torii

    Hajime Taira, M. Okutomi, Torsten Sattler, Mircea Cimpoi, Marc Pollefeys, Josef Sivic, Tomás Pajdla, and Akihiko Torii. InLoc: Indoor Visual Localization with Dense Matching and View Synthesis. IEEE Transactions on Pattern Analysis and Machine Intelligence, 43:1293–1307, 2019. 2

  109. [117]

    InLoc: Indoor Visual Localization with Dense Matching and View Synthesis

    Hajime Taira, Masatoshi Okutomi, Torsten Sattler, Mircea Cimpoi, Marc Pollefeys, Josef Sivic, Tomas Pajdla, and Ak- ihiko Torii. InLoc: Indoor Visual Localization with Dense Matching and View Synthesis. TPAMI, 2021. 1

  110. [118]

    Learning camera localization via dense scene matching

    Shitao Tang, Chengzhou Tang, Rui Huang, Siyu Zhu, and Ping Tan. Learning camera localization via dense scene matching. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1831–1841,

  111. [119]

    Geometrically mappable image features

    Janine Thoma, Danda Pani Paudel, Ajad Chhatkuli, and Luc Van Gool. Geometrically mappable image features. IEEE Robotics and Automation Letters , 5(2):2062–2069,

  112. [120]

    Avrithis, and Hervé Jégou

    Giorgos Tolias, Yannis S. Avrithis, and Hervé Jégou. Image Search with Selective Match Kernels: Aggregation Across Single and Multiple Images. IJCV, 116(3):247–261, 2016. 2

  113. [121]

    24/7 place recognition by view synthesis

    Akihiko Torii, Relja Arandjelovi ´c, Josef Sivic, Masatoshi Okutomi, and Tomas Pajdla. 24/7 place recognition by view synthesis. In CVPR, 2015. 2, 3

  114. [122]

    Are Large-Scale 3D Models Really Necessary for Accurate Vi- sual Localization? TPAMI, 2021

    Akihiko Torii, Hajime Taira, Josef Sivic, Marc Pollefeys, Masatoshi Okutomi, Tomas Pajdla, and Torsten Sattler. Are Large-Scale 3D Models Really Necessary for Accurate Vi- sual Localization? TPAMI, 2021. 1, 3

  115. [123]

    The Unreasonable Effectiveness of Pre- Trained Features for Camera Pose Refinement

    Gabriele Trivigno, Carlo Masone, Barbara Caputo, and Torsten Sattler. The Unreasonable Effectiveness of Pre- Trained Features for Camera Pose Refinement. In Proceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 12786–12798, 2024. 1, 3

  116. [124]

    Least-Squares Estimation of Transfor- mation Parameters Between Two Point Patterns

    Shinji Umeyama. Least-Squares Estimation of Transfor- mation Parameters Between Two Point Patterns. IEEE Transactions on Pattern Analysis & Machine Intelligence , 13(04):376–380, 1991. 4

  117. [125]

    Exploiting Un- certainty in Regression Forests for Accurate Camera Relo- calization

    Julien Valentin, Matthias Nießner, Jamie Shotton, Andrew Fitzgibbon, Shahram Izadi, and Philip Torr. Exploiting Un- certainty in Regression Forests for Accurate Camera Relo- calization. In CVPR, 2015. 2

  118. [126]

    GN-Net: The Gauss-Newton Loss for Multi-Weather Relocalization

    Lukas V on Stumberg, Patrick Wenzel, Qadeer Khan, and Daniel Cremers. GN-Net: The Gauss-Newton Loss for Multi-Weather Relocalization. IEEE Robotics and Automa- tion Letters, 5(2):890–897, 2020. 3

  119. [127]

    LM-Reloc: Levenberg-Marquardt Based Direct Vi- sual Relocalization

    Lukas V on Stumberg, Patrick Wenzel, Nan Yang, and Daniel Cremers. LM-Reloc: Levenberg-Marquardt Based Direct Vi- sual Relocalization. In 2020 International Conference on 3D Vision (3DV), pages 968–977. IEEE, 2020. 3

  120. [128]

    Image- Based Localization Using LSTMs for Structured Feature Correlation

    Florian Walch, Caner Hazirbas, Laura Leal-Taixé, Torsten Sattler, Sebastian Hilsenbeck, and Daniel Cremers. Image- Based Localization Using LSTMs for Structured Feature Correlation. In ICCV, 2017. 3

  121. [129]

    VGGT: Visual Geometry Grounded Transformer

    Jianyuan Wang, Minghao Chen, Nikita Karaev, Andrea Vedaldi, Christian Rupprecht, and David Novotny. VGGT: Visual Geometry Grounded Transformer. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2025. 3, 5 16

  122. [130]

    Dgc-gnn: leveraging geometry and color cues for visual descriptor-free 2d-3d matching

    Shuzhe Wang, Juho Kannala, and Daniel Barath. Dgc-gnn: leveraging geometry and color cues for visual descriptor-free 2d-3d matching. In Proceedings of the IEEE/CVF Confer- ence on Computer Vision and Pattern Recognition , pages 20881–20891, 2024. 1

  123. [131]

    HSCNet++: Hierarchical Scene Coordinate Classification and Regression for Visual Localization with Transformer.arXiv:2305.03595,

    Shuzhe Wang, Zakaria Laskar, Iaroslav Melekhov, Xiaotian Li, Yi Zhao, Giorgos Tolias, and Juho Kannala. HSCNet++: Hierarchical Scene Coordinate Classification and Regression for Visual Localization with Transformer.arXiv:2305.03595,

  124. [132]

    DUSt3R: Geometric 3D Vision Made Easy

    Shuzhe Wang, Vincent Leroy, Yohann Cabon, Boris Chidlovskii, and Jerome Revaud. DUSt3R: Geometric 3D Vision Made Easy. In CVPR, 2024. 2, 3, 4, 5, 6, 9

  125. [133]

    City-scale scene change de- tection using point clouds

    Zi Jian Yew and Gim Hee Lee. City-scale scene change de- tection using point clouds. In 2021 IEEE International Con- ference on Robotics and Automation (ICRA) , pages 13362– 13369. IEEE, 2021. 1

  126. [134]

    Metric3d: Towards zero-shot metric 3d prediction from a single image

    Wei Yin, Chi Zhang, Hao Chen, Zhipeng Cai, Gang Yu, Kaixuan Wang, Xiaozhi Chen, and Chunhua Shen. Metric3d: Towards zero-shot metric 3d prediction from a single image. In ICCV, 2023. 5, 6, 10

  127. [135]

    A. R. Zamir and M. Shah. Accurate Image Localization Based on Google Maps Street View. In ECCV, 2010. 2

  128. [136]

    Image Geo- Localization Based on MultipleNearest Neighbor Feature Matching Using Generalized Graphs

    Amir Roshan Zamir and Mubarak Shah. Image Geo- Localization Based on MultipleNearest Neighbor Feature Matching Using Generalized Graphs. PAMI, 36(8):1546– 1558, 2014. 2

  129. [137]

    Cam- era pose voting for large-scale image-based localization

    Bernhard Zeisl, Torsten Sattler, and Marc Pollefeys. Cam- era pose voting for large-scale image-based localization. In ICCV, 2015. 2

  130. [138]

    Gsplatloc: Ultra-precise camera localization via 3d gaussian splatting

    Atticus J Zeller. Gsplatloc: Ultra-precise camera localization via 3d gaussian splatting. arXiv preprint arXiv:2412.20056,

  131. [139]

    Image Based Localization in Urban Environments

    Wei Zhang and Jana Kosecka. Image Based Localization in Urban Environments. Third International Symposium on 3D Data Processing, Visualization, and Transmission (3DPVT’06), pages 33–40, 2006. 1, 3

  132. [140]

    Ref- erence Pose Generation for Visual Localization via Learned Features and View Synthesis

    Zichao Zhang, Torsten Sattler, and Davide Scaramuzza. Ref- erence Pose Generation for Visual Localization via Learned Features and View Synthesis. arXiv, 2005.05179, 2020. 5, 8, 9, 10, 11, 12

  133. [141]

    Xiaoming Zhao, Xingming Wu, Weihai Chen, Peter C. Y . Chen, Qingsong Xu, and Zhengguo Li. Aliked: A lighter keypoint and descriptor extraction network via deformable transformation. IEEE Transactions on Instrumentation & Measurement, 72:1–16, 2023. 4, 5

  134. [142]

    Xiaoming Zhao, Xingming Wu, Jinyu Miao, Weihai Chen, Peter C. Y . Chen, and Zhengguo Li. Alike: Accurate and lightweight keypoint detection and descriptor extraction. IEEE Transactions on Multimedia, 3 2022. 4, 5

  135. [143]

    Structure From Motion Using Structure-Less Resection

    Enliang Zheng and Changchang Wu. Structure From Motion Using Structure-Less Resection. In ICCV, 2015. 1, 3, 4

  136. [144]

    Is geometry enough for matching in visual lo- calization? In European Conference on Computer Vision , pages 407–425

    Qunjie Zhou, Sérgio Agostinho, Aljoša Ošep, and Laura Leal-Taixé. Is geometry enough for matching in visual lo- calization? In European Conference on Computer Vision , pages 407–425. Springer, 2022. 1

  137. [145]

    The nerfect match: Exploring nerf features for visual localization

    Qunjie Zhou, Maxim Maximov, Or Litany, and Laura Leal- Taixé. The nerfect match: Exploring nerf features for visual localization. In European Conference on Computer Vision , pages 108–127. Springer, 2024. 1, 2

  138. [146]

    Patch2Pix: Epipolar-Guided Pixel-Level Correspondences

    Qunjie Zhou, Torsten Sattler, and Laura Leal-Taixé. Patch2Pix: Epipolar-Guided Pixel-Level Correspondences. 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 4667–4676, 2020. 4

  139. [147]

    To Learn or Not to Learn: Visual Localization from Essential Matrices

    Qunjie Zhou, Torsten Sattler, Marc Pollefeys, and Laura Leal-Taixé. To Learn or Not to Learn: Visual Localization from Essential Matrices. In ICRA, 2019. 1, 3, 4, 9, 11

  140. [148]

    Very large-scale global sfm by dis- tributed motion averaging

    Siyu Zhu, Runze Zhang, Lei Zhou, Tianwei Shen, Tian Fang, Ping Tan, and Long Quan. Very large-scale global sfm by dis- tributed motion averaging. In Proceedings of the IEEE con- ference on computer vision and pattern recognition , pages 4568–4577, 2018. 3 17

Pith tools

Reviewed August 16, 2026 · model on record in the stance chip above.