Pith. sign in

REVIEW 3 major objections 45 references

GeoGS-SLAM builds online monocular maps by seeding 3D Gaussians from feed-forward geometry priors, then refining them with both photometric and geometric losses so RGB-only SLAM finally keeps high-fidelity rendering.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

An online monocular SLAM system that samples 3D Gaussians from RGB plus VGGT geometric priors and jointly optimizes poses and map with photometric and geometric losses plus loop closure, beating prior monocular 3DGS and prior-based SLAM on rendering and tracking.

T0 review reviewed 2026-07-14 challenge →

load-bearing objection Solid integrative monocular 3DGS+VGGT SLAM with real PSNR/ATE gains; novelty is the closed photometric loop, not new primitives, and the uncalibrated comparison needs a careful read. the 3 major comments →

arxiv 2607.11184 v1 pith:CB6YBHNM submitted 2026-07-13 cs.RO

GeoGS-SLAM: Online Monocular Reconstruction Using Gaussian Splatting with Geometric Priors

classification cs.RO
keywords monocular SLAM3D Gaussian Splattinggeometric priorsfeed-forward reconstructiononline dense mappingphotometric-geometric optimizationloop closure
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Most Gaussian-splatting SLAM systems need depth sensors, while pure feed-forward SLAM systems throw away the original RGB once they have predicted geometry. GeoGS-SLAM closes that loop: a pre-trained visual-geometry model first supplies camera and scene priors from uncalibrated RGB; new Gaussian primitives are sampled directly from both the image and those priors; poses and the map are then jointly optimized with a coarse-to-fine mix of photometric and geometric losses, plus online loop closure. On indoor and outdoor benchmarks the system reports higher rendering quality and lower tracking error than prior monocular methods while still running in real time. A reader who cares about dense mapping for robotics or AR should care because the method claims to remove the depth-sensor requirement without sacrificing photorealism.

Core claim

By seeding a 3D Gaussian map from feed-forward camera and scene priors and then refining both the map and the poses against the original RGB images, monocular SLAM can simultaneously obtain accurate trajectories and high-fidelity novel-view renderings without external depth sensors or calibrated intrinsics.

What carries the argument

Closed-loop pipeline that samples Gaussians from RGB-plus-priors, then jointly minimizes photometric (L1+SSIM) and geometric (depth) losses inside a sliding keyframe window under coarse-to-fine pyramid optimization, finished by online loop-closure pose-graph correction.

Load-bearing premise

The feed-forward model’s camera and depth priors, after simple scale alignment between overlapping windows, are accurate and consistent enough across indoor and outdoor scenes to bootstrap monocular 3D Gaussian optimization.

What would settle it

Run the identical pipeline on a new outdoor sequence whose domain differs sharply from the feed-forward model’s training data; if the resulting absolute trajectory error jumps by more than a factor of two relative to the paper’s Waymo numbers while a calibrated RGB-D baseline remains stable, the central claim fails.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 0 minor

Summary. GeoGS-SLAM is an online monocular dense SLAM system that fuses a 3D Gaussian Splatting map with frozen feed-forward geometric priors (VGGT). From uncalibrated RGB, it predicts camera and scene priors over overlapping keyframe windows, aligns scale via high-confidence point-map norms (Eq. 5), expands the Gaussian map by DoG-guided direct primitive sampling from RGB and priors, jointly optimizes poses and Gaussians with photometric plus geometric losses under a coarse-to-fine pyramid (Eqs. 8–10), and applies MegaLoc-based loop closure with pose-graph optimization and Gaussian pose updates. On Replica, TUM RGB-D, and Waymo the paper reports state-of-the-art monocular rendering (Table I; e.g. PSNR +4.94 on Replica) and strong uncalibrated tracking (Table II vs MASt3R-SLAM/VGGT-SLAM/SLAM3R), with ablations (Table IV) and a per-keyframe runtime breakdown (Table III).

Significance. If the empirical claims hold under fair monocular uncalibrated conditions, the paper offers a practical closed-loop paradigm that keeps photometric evidence in the loop while using modern feed-forward geometry for bootstrapping—addressing a clear gap between depth-dependent 3DGS-SLAM and prior-only systems that discard RGB. Strengths include a clean modular design, transparent Calib./Uncalib. partitioning in Table II, component ablations that move metrics in the expected direction (camera priors are load-bearing), multi-domain evaluation including outdoor Waymo, and an explicit runtime table. The contribution is primarily systems/empirical rather than theoretical; novelty lies in the integration (direct sampling + joint photo-geo optimization + online loop closure on top of VGGT priors), which is a useful and timely step for monocular dense reconstruction in robotics.

major comments (3)
  1. Tables I–II and §IV.B: the headline claim of superiority over “SOTA monocular SLAM methods” needs a stricter protocol statement. Table II correctly separates Calib. vs Uncalib., and several strong GS baselines (Photo-SLAM, Splat-SLAM, S3PO-GS) sit under Calib. with better or comparable ATE on some sets (e.g. Replica Photo-SLAM 0.022 / Splat-SLAM 0.018 vs Ours 0.024; Waymo Photo-SLAM 0.366 / Splat-SLAM 0.495 vs Ours 0.861). Please state explicitly which baselines were re-run by the authors under identical uncalibrated monocular settings (no GT intrinsics/depth, same sequences/splits) versus numbers taken from prior papers, and align abstract/conclusion wording with the Uncalib. subset when claiming tracking SOTA.
  2. §III-B, Eq. (5) and Table IV (w/o camera priors): the system’s metric scale and outdoor results rest on VGGT priors after mutual-keyframe scale alignment. The ablation (ATE 0.922 without camera priors) shows necessity but not residual quality. Please report a direct prior-quality diagnostic—e.g. residual scale error and relative-pose error of aligned pT/pK/pD versus GT on Replica vs Waymo before joint optimization—so readers can judge domain-shift risk and how much photometric refinement (Eq. 10) actually corrects. Without this, the outdoor Waymo gains remain hard to attribute.
  3. Table III and abstract “online real-time” claim: 582.1 ms per keyframe (joint optimization 480.6 ms) is only real-time if keyframe rate is low. Report end-to-end throughput on continuous streams (average FPS, keyframe fraction under τ_disp, latency to first map update) for each dataset, and clarify whether non-keyframes are tracked with a cheaper pose-only step. As written, “real-time performance” is underspecified relative to the robotics claim.

Circularity Check

0 steps flagged

No circularity: empirical monocular SLAM system evaluated on external benchmarks with frozen external priors; reported PSNR/ATE gains do not reduce to fitted inputs by construction.

full rationale

GeoGS-SLAM is a systems paper whose central claims are empirical superiority of rendering (PSNR/SSIM/LPIPS) and tracking (ATE RMSE) on Replica, TUM RGB-D and Waymo versus third-party monocular baselines (MonoGS, Photo-SLAM, Splat-SLAM, MASt3R-SLAM, VGGT-SLAM, etc.). Geometric priors come from a frozen pre-trained VGGT model (Sec. III-B); scale alignment (Eq. 5) is a simple mutual-keyframe ratio of high-confidence points; map expansion samples primitives from RGB + those priors (Sec. III-C); joint optimization minimizes photometric L1+SSIM plus geometric L1 depth losses (Eqs. 8–10) via coarse-to-fine rendering; loop closure uses MegaLoc descriptors + pose-graph LM. None of these steps defines the reported metrics in terms of themselves, fits a free parameter on the evaluation sets and then “predicts” a related quantity, or rests on a load-bearing self-citation uniqueness theorem. Ablations (Table IV) simply remove components and re-measure the same external metrics; the catastrophic drop without camera priors confirms dependence on VGGT but does not create a circular derivation. The paper is therefore self-contained against external data and baselines; score 0 is the correct, expected outcome for this class of work.

Axiom & Free-Parameter Ledger

7 free parameters · 5 axioms · 0 invented entities

The central empirical claim rests on standard 3DGS rendering math, a frozen external geometry network (VGGT), place recognition (MegaLoc), hand-chosen thresholds and loss weights, and the modeling assumption that photometric+geometric joint optimization on keyframe windows plus pose-graph loop closure yields globally consistent monocular maps. No new physical entities are postulated; free parameters are engineering thresholds and weights that affect map density and optimization balance.

free parameters (7)
  • τ_disp (keyframe optical-flow displacement threshold)
    Controls which frames enter the sliding window; not derived, chosen to balance coverage and compute.
  • τ_conf (prior confidence threshold for scale alignment)
    Selects pixels used in ρ scale factor (Eq. 5); hand-set, affects scale consistency.
  • τ_prim (primitive placement probability gap)
    Gates new Gaussian spawning from DoG difference; controls map density and redundancy.
  • τ_sim (loop descriptor cosine similarity threshold)
    Accepts/rejects loop candidates from MegaLoc; false positives/negatives change global consistency.
  • λ_SSIM and λ_geo (loss weights)
    Balance L1/SSIM photometric terms and geometric depth loss in joint optimization; fitted/chosen for reported quality.
  • window size w and pyramid levels n
    Structural hyperparameters of multi-view prior prediction and coarse-to-fine training; not theoretically fixed.
  • DoG σ1=0.5, σ2=1.5 and scale ε
    Fixed sampling and Gaussian size constants taken from design choice / prior practice.
axioms (5)
  • domain assumption Differentiable 3D Gaussian rasterization (Kerbl et al.) correctly models appearance and depth via α-blending for SLAM map representation.
    Sec. III-A adopts standard 3DGS projection and blending as the scene model without re-deriving validity conditions.
  • domain assumption VGGT multi-view predictions of extrinsics, intrinsics, depths, and points are sufficiently accurate after scale alignment to bootstrap monocular uncalibrated SLAM.
    Sec. III-B treats frozen VGGT outputs as geometric priors; ablation shows collapse without camera priors.
  • ad hoc to paper Overlapping-window scale factor ρ from high-confidence point norms (Eq. 5) restores metric consistency across windows.
    Heuristic alignment procedure specific to this pipeline; not proven to eliminate scale drift in general.
  • domain assumption MegaLoc global descriptors plus cosine threshold yield true loop closures usable as pose-graph constraints.
    Sec. III-E relies on pre-trained place recognition for global consistency.
  • domain assumption Joint minimization of photometric (L1+SSIM) and geometric (depth L1) losses with coarse-to-fine pyramids improves poses and Gaussians without destructive local minima.
    Sec. III-D optimization strategy assumed effective for online keyframe windows.

reviewed 2026-07-14 · how reviews work

0 comments
Cite this review

Pith. "Pith review of GeoGS-SLAM: Online Monocular Reconstruction Using Gaussian Splatting with Geometric Priors." pith.science (2026). https://pith.science/paper/CB6YBHNM

@misc{pith2026260711184,
  author       = {Pith},
  title        = {Pith review of: GeoGS-SLAM: Online Monocular Reconstruction Using Gaussian Splatting with Geometric Priors},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/CB6YBHNM}},
  note         = {Machine review of arXiv:2607.11184}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

SLAM methods based on 3D Gaussian Splatting (3DGS) have demonstrated impressive tracking and mapping performance, but typically require additional geometric information from external depth sensors. Meanwhile, recent SLAM systems that leverage geometric priors from pre-trained feed-forward models enable real-time dense reconstruction, yet often discard original RGB information during optimization, thus degrading overall reconstruction quality. We present GeoGS-SLAM, an online monocular dense reconstruction system that combines the 3DGS-based map representation with learned geometric priors. Given uncalibrated RGB input, we first employ a feed-forward visual geometry model to predict camera and scene priors. The Gaussian scene map is then expanded by directly sampling Gaussian primitives from both RGB input and geometric priors. Camera poses and the scene map are jointly optimized through a coarse-to-fine strategy that minimizes both photometric and geometric losses. To ensure global consistency, we further incorporate online loop closure detection and pose graph optimization. Extensive experiments across indoor and outdoor benchmarks demonstrate that GeoGS-SLAM achieves superior rendering quality and tracking accuracy compared to state-of-the-art methods while maintaining online real-time performance. Project page: https://rlgao.github.io/geogs_slam.

Figures

Figures reproduced from arXiv: 2607.11184 by Letian Jin, Ruilan Gao, Yu Zhang.

Figure 1
Figure 1. Figure 1: Rendering results on three datasets. Our method produces high-fidelity reconstructions on both indoor and outdoor benchmarks, outperforming state-of-the-art monocular 3DGS-based SLAM methods. Abstract— SLAM methods based on 3D Gaussian Splatting (3DGS) have demonstrated impressive tracking and mapping performance, but typically require additional geometric infor￾mation from external depth sensors. Meanwhil… view at source ↗
Figure 2
Figure 2. Figure 2 [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Comparison of rendering results. Our method produces photorealistic reconstructions for both indoor and outdoor scenes. For indoor environments, it captures fine-grained textures and geometric details with minimal artifacts. For outdoor scenarios, it successfully handles complex driving scenes, preserving architectural structures and vehicle details. These substantial tracking improvements across diverse e… view at source ↗
Figure 4
Figure 4. Figure 4: Comparison of estimated trajectories. Each tracking result is projected onto the x-y plane, with ground truth shown as a dashed line. Our method achieves superior accuracy and robustness across indoor and outdoor environments. TABLE II: Tracking accuracy comparison across three datasets. ATE RMSE [m] (Ó) is reported. Our method achieves superior tracking performance compared to methods operating on uncalib… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

45 extracted references · 6 linked inside Pith

  1. [1]

    How nerfs and 3d gaussian splatting are reshaping slam: a survey,

    F. Tosi, Y . Zhang, Z. Gong, E. Sandstr¨om, S. Mattoccia, M. R. Oswald, and M. Poggi, “How nerfs and 3d gaussian splatting are reshaping slam: a survey,”arXiv preprint arXiv:2402.13255, vol. 4, p. 1, 2024. TABLE III:Runtime evaluation results.Average processing time per keyframe across datasets is reported, demonstrating the computational efficiency of ou...

  2. [2]

    Advances in feed-forward 3d reconstruction and view synthesis: A survey,

    J. Zhang, Y . Li, A. Chen, M. Xu, K. Liu, J. Wang, X.-X. Long, H. Liang, Z. Xu, H. Suet al., “Advances in feed-forward 3d reconstruction and view synthesis: A survey,”arXiv preprint arXiv:2507.14501, 2025

  3. [3]

    Nerf: Representing scenes as neural radiance fields for view synthesis,

    B. Mildenhall, P. P. Srinivasan, M. Tancik, J. T. Barron, R. Ramamoor- thi, and R. Ng, “Nerf: Representing scenes as neural radiance fields for view synthesis,”Communications of the ACM, vol. 65, no. 1, pp. 99–106, 2021

  4. [4]

    3d gaussian splatting for real-time radiance field rendering

    B. Kerbl, G. Kopanas, T. Leimk ¨uhler, and G. Drettakis, “3d gaussian splatting for real-time radiance field rendering.”ACM Trans. Graph., vol. 42, no. 4, pp. 139–1, 2023

  5. [5]

    Gaussian splatting slam,

    H. Matsuki, R. Murai, P. H. Kelly, and A. J. Davison, “Gaussian splatting slam,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 18 039–18 048

  6. [6]

    Splatam: Splat track & map 3d gaussians for dense rgb-d slam,

    N. Keetha, J. Karhade, K. M. Jatavallabhula, G. Yang, S. Scherer, D. Ramanan, and J. Luiten, “Splatam: Splat track & map 3d gaussians for dense rgb-d slam,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 21 357–21 366

  7. [7]

    Rgb-only gaussian splatting slam for unbounded outdoor scenes,

    S. Yu, C. Cheng, Y . Zhou, X. Yang, and H. Wang, “Rgb-only gaussian splatting slam for unbounded outdoor scenes,” in2025 IEEE International Conference on Robotics and Automation (ICRA), 2025, pp. 11 068–11 074

  8. [8]

    Dust3r: Geometric 3d vision made easy,

    S. Wang, V . Leroy, Y . Cabon, B. Chidlovskii, and J. Revaud, “Dust3r: Geometric 3d vision made easy,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 20 697–20 709

  9. [9]

    Vggt: Visual geometry grounded transformer,

    J. Wang, M. Chen, N. Karaev, A. Vedaldi, C. Rupprecht, and D. Novotny, “Vggt: Visual geometry grounded transformer,” inPro- ceedings of the Computer Vision and Pattern Recognition Conference, 2025, pp. 5294–5306

  10. [10]

    Mast3r-slam: Real- time dense slam with 3d reconstruction priors,

    R. Murai, E. Dexheimer, and A. J. Davison, “Mast3r-slam: Real- time dense slam with 3d reconstruction priors,” inProceedings of the Computer Vision and Pattern Recognition Conference, 2025, pp. 16 695–16 705

  11. [11]

    Slam3r: Real-time dense scene reconstruction from monocular rgb videos,

    Y . Liu, S. Dong, S. Wang, Y . Yin, Y . Yang, Q. Fan, and B. Chen, “Slam3r: Real-time dense scene reconstruction from monocular rgb videos,” inProceedings of the Computer Vision and Pattern Recogni- tion Conference, 2025, pp. 16 651–16 662

  12. [12]

    Vggt-slam: Dense rgb slam optimized on the sl (4) manifold,

    D. Maggio, H. Lim, and L. Carlone, “Vggt-slam: Dense rgb slam optimized on the sl (4) manifold,”arXiv preprint arXiv:2505.12549, 2025

  13. [13]

    The replica dataset: A digital replica of indoor spaces,

    J. Straub, T. Whelan, L. Ma, Y . Chen, E. Wijmans, S. Green, J. J. Engel, R. Mur-Artal, C. Ren, S. Vermaet al., “The replica dataset: A digital replica of indoor spaces,”arXiv preprint arXiv:1906.05797, 2019

  14. [14]

    A benchmark for the evaluation of rgb-d slam systems,

    J. Sturm, N. Engelhard, F. Endres, W. Burgard, and D. Cremers, “A benchmark for the evaluation of rgb-d slam systems,” in2012 IEEE/RSJ international conference on intelligent robots and systems. IEEE, 2012, pp. 573–580

  15. [15]

    Scalability in perception for autonomous driving: Waymo open dataset,

    P. Sun, H. Kretzschmar, X. Dotiwalla, A. Chouard, V . Patnaik, P. Tsui, J. Guo, Y . Zhou, Y . Chai, B. Caineet al., “Scalability in perception for autonomous driving: Waymo open dataset,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020, pp. 2446–2454

  16. [16]

    Parallel tracking and mapping for small ar workspaces,

    G. Klein and D. Murray, “Parallel tracking and mapping for small ar workspaces,” in2007 6th IEEE and ACM international symposium on mixed and augmented reality. IEEE, 2007, pp. 225–234

  17. [17]

    Orb-slam: A versatile and accurate monocular slam system,

    R. Mur-Artal, J. M. M. Montiel, and J. D. Tardos, “Orb-slam: A versatile and accurate monocular slam system,”IEEE transactions on robotics, vol. 31, no. 5, pp. 1147–1163, 2015

  18. [18]

    Orb-slam2: An open-source slam system for monocular, stereo, and rgb-d cameras,

    R. Mur-Artal and J. D. Tard ´os, “Orb-slam2: An open-source slam system for monocular, stereo, and rgb-d cameras,”IEEE transactions on robotics, vol. 33, no. 5, pp. 1255–1262, 2017

  19. [19]

    Orb-slam3: An accurate open-source library for visual, visual–inertial, and multimap slam,

    C. Campos, R. Elvira, J. J. G. Rodr ´ıguez, J. M. Montiel, and J. D. Tard ´os, “Orb-slam3: An accurate open-source library for visual, visual–inertial, and multimap slam,”IEEE transactions on robotics, vol. 37, no. 6, pp. 1874–1890, 2021

  20. [20]

    Lsd-slam: Large-scale di- rect monocular slam,

    J. Engel, T. Sch ¨ops, and D. Cremers, “Lsd-slam: Large-scale di- rect monocular slam,” inEuropean conference on computer vision. Springer, 2014, pp. 834–849

  21. [21]

    Direct sparse odometry,

    J. Engel, V . Koltun, and D. Cremers, “Direct sparse odometry,”IEEE transactions on pattern analysis and machine intelligence, vol. 40, no. 3, pp. 611–625, 2017

  22. [22]

    Kinectfusion: Real-time dense surface mapping and tracking,

    R. A. Newcombe, S. Izadi, O. Hilliges, D. Molyneaux, D. Kim, A. J. Davison, P. Kohi, J. Shotton, S. Hodges, and A. Fitzgibbon, “Kinectfusion: Real-time dense surface mapping and tracking,” in 2011 10th IEEE international symposium on mixed and augmented reality. Ieee, 2011, pp. 127–136

  23. [23]

    Bundle- fusion: Real-time globally consistent 3d reconstruction using on- the-fly surface reintegration,

    A. Dai, M. Nießner, M. Zollh ¨ofer, S. Izadi, and C. Theobalt, “Bundle- fusion: Real-time globally consistent 3d reconstruction using on- the-fly surface reintegration,”ACM Transactions on Graphics (ToG), vol. 36, no. 4, p. 1, 2017

  24. [24]

    Orbeez-slam: A real-time monocular visual slam with orb features and nerf-realized mapping,

    C.-M. Chung, Y .-C. Tseng, Y .-C. Hsu, X.-Q. Shi, Y .-H. Hua, J.-F. Yeh, W.-C. Chen, Y .-T. Chen, and W. H. Hsu, “Orbeez-slam: A real-time monocular visual slam with orb features and nerf-realized mapping,” in2023 IEEE International Conference on Robotics and Automation (ICRA), 2023, pp. 9400–9406

  25. [25]

    Hero-slam: Hybrid enhanced robust optimization of neural slam,

    Z. Xin, Y . Yue, L. Zhang, and C. Wu, “Hero-slam: Hybrid enhanced robust optimization of neural slam,” in2024 IEEE International Conference on Robotics and Automation (ICRA), 2024, pp. 8610– 8616

  26. [26]

    imap: Implicit map- ping and positioning in real-time,

    E. Sucar, S. Liu, J. Ortiz, and A. J. Davison, “imap: Implicit map- ping and positioning in real-time,” inProceedings of the IEEE/CVF international conference on computer vision, 2021, pp. 6229–6238

  27. [27]

    Nice-slam: Neural implicit scalable encoding for slam,

    Z. Zhu, S. Peng, V . Larsson, W. Xu, H. Bao, Z. Cui, M. R. Oswald, and M. Pollefeys, “Nice-slam: Neural implicit scalable encoding for slam,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022, pp. 12 786–12 796

  28. [28]

    V ox- fusion: Dense tracking and mapping with voxel-based neural implicit representation,

    X. Yang, H. Li, H. Zhai, Y . Ming, Y . Liu, and G. Zhang, “V ox- fusion: Dense tracking and mapping with voxel-based neural implicit representation,” in2022 IEEE International Symposium on Mixed and Augmented Reality (ISMAR). IEEE, 2022, pp. 499–507

  29. [29]

    Hier-slam: Scaling-up semantics in slam with a hierarchically categorical gaussian splatting,

    B. Li, Z. Cai, Y .-F. Li, I. Reid, and H. Rezatofighi, “Hier-slam: Scaling-up semantics in slam with a hierarchically categorical gaussian splatting,” in2025 IEEE International Conference on Robotics and Automation (ICRA), 2025, pp. 9748–9754

  30. [30]

    Opengs-slam: Open-set dense semantic slam with 3d gaussian splatting for object- level scene understanding,

    D. Yang, Y . Gao, X. Wang, Y . Yue, Y . Yang, and M. Fu, “Opengs-slam: Open-set dense semantic slam with 3d gaussian splatting for object- level scene understanding,” in2025 IEEE International Conference on Robotics and Automation (ICRA), 2025, pp. 8486–8492

  31. [31]

    Gs- slam: Dense visual slam with 3d gaussian splatting,

    C. Yan, D. Qu, D. Xu, B. Zhao, Z. Wang, D. Wang, and X. Li, “Gs- slam: Dense visual slam with 3d gaussian splatting,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recog- nition, 2024, pp. 19 595–19 604

  32. [32]

    Photo-slam: Real-time simultaneous localization and photorealistic mapping for monocular stereo and rgb-d cameras,

    H. Huang, L. Li, H. Cheng, and S.-K. Yeung, “Photo-slam: Real-time simultaneous localization and photorealistic mapping for monocular stereo and rgb-d cameras,” inProceedings of the IEEE/CVF Confer- ence on Computer Vision and Pattern Recognition, 2024, pp. 21 584– 21 593

  33. [33]

    Large- scale gaussian splatting slam,

    Z. Xin, C. Wu, P. Huang, Y . Zhang, Y . Mao, and G. Huang, “Large- scale gaussian splatting slam,” in2025 IEEE International Conference on Robotics and Automation (ICRA), 2025, pp. 8478–8485

  34. [34]

    Outdoor monocular slam with global scale-consistent 3d gaussian pointmaps,

    C. Cheng, S. Yu, Z. Wang, Y . Zhou, and H. Wang, “Outdoor monocular slam with global scale-consistent 3d gaussian pointmaps,”arXiv preprint arXiv:2507.03737, 2025

  35. [35]

    Grounding image matching in 3d with mast3r,

    V . Leroy, Y . Cabon, and J. Revaud, “Grounding image matching in 3d with mast3r,” inEuropean Conference on Computer Vision. Springer, 2024, pp. 71–91

  36. [36]

    Fast3r: Towards 3d reconstruction of 1000+ images in one forward pass,

    J. Yang, A. Sax, K. J. Liang, M. Henaff, H. Tang, A. Cao, J. Chai, F. Meier, and M. Feiszli, “Fast3r: Towards 3d reconstruction of 1000+ images in one forward pass,” inProceedings of the Computer Vision and Pattern Recognition Conference, 2025, pp. 21 924–21 935

  37. [37]

    On-the- fly reconstruction for large-scale novel view synthesis from unposed images,

    A. Meuleman, I. Shah, A. Lanvin, B. Kerbl, and G. Drettakis, “On-the- fly reconstruction for large-scale novel view synthesis from unposed images,”ACM Transactions on Graphics (TOG), vol. 44, no. 4, pp. 1–14, 2025

  38. [38]

    Theory of edge detection,

    D. Marr and E. Hildreth, “Theory of edge detection,”Proceedings of the Royal Society of London. Series B. Biological Sciences, vol. 207, no. 1167, pp. 187–217, 1980

  39. [39]

    Image quality assessment: from error visibility to structural similarity,

    Z. Wang, A. C. Bovik, H. R. Sheikh, and E. P. Simoncelli, “Image quality assessment: from error visibility to structural similarity,”IEEE transactions on image processing, vol. 13, no. 4, pp. 600–612, 2004

  40. [40]

    Megaloc: One retrieval to place them all,

    G. Berton and C. Masone, “Megaloc: One retrieval to place them all,” inProceedings of the Computer Vision and Pattern Recognition Conference, 2025, pp. 2861–2867

  41. [41]

    The unreasonable effectiveness of deep features as a perceptual metric,

    R. Zhang, P. Isola, A. A. Efros, E. Shechtman, and O. Wang, “The unreasonable effectiveness of deep features as a perceptual metric,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 586–595

  42. [42]

    Go-slam: Global optimization for consistent 3d instant reconstruction,

    Y . Zhang, F. Tosi, S. Mattoccia, and M. Poggi, “Go-slam: Global optimization for consistent 3d instant reconstruction,” inProceedings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 3727–3737

  43. [43]

    Glorie-slam: Globally optimized rgb-only implicit encoding point cloud slam,

    G. Zhang, E. Sandstr ¨om, Y . Zhang, M. Patel, L. Van Gool, and M. R. Oswald, “Glorie-slam: Globally optimized rgb-only implicit encoding point cloud slam,”arXiv preprint arXiv:2403.19549, 2024

  44. [44]

    Splat-slam: Globally optimized rgb-only slam with 3d gaussians,

    E. Sandstr ¨om, G. Zhang, K. Tateno, M. Oechsle, M. Niemeyer, Y . Zhang, M. Patel, L. Van Gool, M. Oswald, and F. Tombari, “Splat-slam: Globally optimized rgb-only slam with 3d gaussians,” inProceedings of the Computer Vision and Pattern Recognition Conference, 2025, pp. 1680–1691

  45. [45]

    Droid-splat: Com- bining end-to-end slam with 3d gaussian splatting,

    C. Homeyer, L. Begiristain, and C. Schn ¨orr, “Droid-splat: Com- bining end-to-end slam with 3d gaussian splatting,”arXiv preprint arXiv:2411.17660, 2024

This paper was first reviewed by grok-4.5 on July 14, 2026.