REVIEW 5 major objections 5 minor 1 cited by
6DOPE-GS: Online 6D Object Pose Estimation using Gaussian Splatting
T0 review · 5 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read 6DOPE-GS claims that 2D Gaussian Splatting can replace slow neural-field training for model-free 6D pose tracking, matching baseline accuracy at five times the speed and enabling 4-5 Hz live operation.
desk verdict A real speedup for model-free 6D pose tracking via 2D Gaussian Splatting, but the 'matches SOTA' claim rests on ADD-S and the initialization sensitivity is under-tested. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the Gaussian Object Field, an incremental 2D Gaussian Splatting model in which each particle is a flattened, oriented Gaussian disk (a surfel) with position, rotation, zero-thickness scale, opacity, and color. Differentiable rasterization of these disks lets gradients flow from photometric and depth reconstruction losses back into both the Gaussian parameters and the keyframe poses, so pose refinement and model building happen in one loop. Two mechanisms keep that loop stable: dynamic keyframe selection, which clusters candidate views around an icosahedron to maximize spatial coverage and removes outlier views whose reconstruction loss deviates by more than three times the median absolute deviation, and opacity-percentile-based adaptive density control, which prunes low-opacity Gaussians below the 5th percentile until the 95th percentile opacity exceeds a threshold. The corrected keyframe poses then drive an online pose graph optimization for per-frame object pose output.
What would settle it
On an HO3D sequence, add controlled noise to the coarse initialization, for example 5 degrees of rotation or 2 cm of translation, and measure the final ADD-S error; if these small perturbations make the Gaussian field diverge, the claimed speed advantage holds only under near-perfect initialization.
Extended reading notes
Core claim
The central claim is that a 2D Gaussian Splatting representation can carry both 3D reconstruction and 6D pose refinement in a model-free tracking loop, and that this removes the computational bottleneck that makes prior neural-field trackers slow. Concretely, the paper reports ADD-S AUC of 95.07 percent on HO3D versus BundleSDF's 94.86 percent, at 0.24 seconds per frame versus 2.10 seconds, and 93.79 percent versus 92.82 percent on YCBInEOAT, at 0.22 seconds versus 0.82 seconds per frame. It also reports that on the non-symmetric ADD metric on HO3D it trails BundleSDF (84.33 versus 89.56), attributing this to occlusion in hand-object interactions. The method's pose estimates come from jointly optimizing a Gaussian Object Field and a set of keyframe poses, then feeding corrected keyframe poses into an online pose-graph optimization.
Load-bearing premise
The load-bearing assumption is that the rough starting pose produced by matching image features and then fitting geometry robustly is close enough to the true pose that the joint Gaussian-splatting optimization converges; the paper states that a bad initialization makes the optimization diverge and reports no measurement of how much initial error is tolerable.
Editorial extensions
If this is right
- Live model-free tracking becomes practical: object poses can update at 4-5 Hz from one RGB-D camera, enough for many robotic and augmented-reality loops.
- On the tested datasets, reconstruction reaches sub-centimeter Chamfer distances (0.15-0.41 cm), so the online Gaussian field is a usable dense object model while tracking.
- The fivefold speedup over BundleSDF comes without sacrificing symmetric pose accuracy in the reported experiments.
- Because only selected keyframes are rendered for joint optimization, the per-frame cost stays roughly constant even as the object model grows, rather than scaling with every video frame.
- With a single first-frame mask, the tracker handles novel objects with no CAD model, category-level training, or reference images.
Reading between the lines
- Beyond the paper: the coverage-based keyframe selector could be lifted into other Gaussian Splatting SLAM systems as a general cost-reduction strategy, but the paper does not test that transfer.
- Beyond the paper: because the method never measures how tracking accuracy degrades with initial pose error, a natural next experiment is a controlled noise study to map the basin of convergence of the joint optimization.
- Beyond the paper: the trained Gaussian field is not directly coupled to the online pose graph; connecting them could reduce drift, a direction the paper itself leaves open.
- Beyond the paper: the live rate of 4-5 Hz without GUI leaves a thin margin over object motion, so deployment on embedded hardware would likely need further pruning or frame skipping.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes 6DOPE-GS, a model-free 6D object pose estimation and tracking method that uses 2D Gaussian Splatting to jointly optimize a Gaussian object field and keyframe poses, with a dynamic keyframe selection procedure and an opacity percentile-based pruning mechanism. The method is evaluated on HO3D and YCBInEOAT against SLAM-based baselines and BundleSDF/BundleTrack, reporting competitive pose accuracy, sub-centimeter Chamfer distances, and substantially lower average processing time per frame. A live demonstration on a ZED 2 camera shows 3-5 Hz tracking. The central claim is that the method matches state-of-the-art baselines while being about 5x faster.
Significance. The paper addresses a practical bottleneck in model-free 6D pose tracking: the high computational cost of neural object-field training. By replacing the SDF representation with 2D Gaussian Splatting and adding keyframe selection and pruning, the authors demonstrate a plausible route to live tracking at 3-5 Hz, which is a meaningful step over BundleSDF's roughly 0.4 Hz. The evaluation covers two standard benchmarks and includes an ablation of each proposed component. However, the claim that accuracy 'matches' state-of-the-art is only partially supported: on HO3D the ADD AUC is 5.23 points below BundleSDF, and no error bars or multiple runs are reported. The robustness of the initialization-dependent pipeline to challenging conditions (low texture, fast motion, occlusion) is not experimentally characterized. These gaps are significant for a systems paper whose headline is a speed/accuracy tradeoff.
major comments (5)
- [Sec. 4.3, Table 2] The claim that 6DOPE-GS 'matches the performance of state-of-the-art baselines' is not fully supported by Table 2: on HO3D, the ADD AUC is 84.33 versus BundleSDF's 89.56 (a 5.23-point gap), while ADD-S is 95.07 versus 94.86. Since ADD-S is less sensitive for symmetric objects, the headline claim overstates the result. Please report per-sequence/per-object results and confidence intervals, or soften the claim.
- [Sec. 3.3] The method's reliability hinges on the coarse initial poses from LoFTR+RANSAC lying inside the convergence basin of the 2DGS optimization, as the paper itself states in Sec. 3.3 ('errors in the pose initialization can cause a divergence'). The 3x MAD outlier filter only works when bad poses are a minority; if initialization errors are systematic, both the median and MAD are corrupted. No experiments vary initialization noise, mask noise, or occlusion ratio, so the robustness and speed advantages are not established for such failure-prone scenarios. Please add a sensitivity analysis.
- [Sec. 4.4, Tables 1-2] The speedup claim is inconsistent: the abstract and Sec. 4.4 say '5× speedup' over BundleSDF, but Tables 1-2 show 0.22s vs 0.82s (3.7x) on YCBInEOAT and 0.24s vs 2.10s (8.75x) on HO3D. The intermediate variants BundleSDF-async and BundleSDF-lite are introduced in Sec. 4.4 but their numbers are not in the tables. Please clarify which comparison yields 5x and include a per-component runtime breakdown.
- [Sec. 3.1, 3.3, 3.5] Several parameters are left unspecified: the icosahedron subdivision level, the opacity percentile thresholds, the MAD factor, and the visibility threshold for keyframe selection; the coarse pose initialization and pose graph optimization are delegated to [71] and [69]. Without these values or released code, the method cannot be reproduced or compared fairly. Please report all hyperparameters and release code.
- [Sec. 4.5, Table 3] The ablation in Table 3 reports accuracy only. Since the proposed pruning (Sec. 3.4) and keyframe selection (Sec. 3.3) are claimed to improve efficiency and stability, the ablation should include average time per frame and Gaussian counts for each variant. Otherwise the efficiency benefit of these components is not demonstrated.
minor comments (5)
- [Sec. 4.5] The text uses 'YCBInEoat' where 'YCBInEOAT' is meant; please correct this typo.
- [Sec. 4.4] Figure 4 is mentioned but not described in detail; please add a sentence explaining what each curve or point represents and how the tradeoff should be read.
- [Fig. 3] Only successful qualitative results are shown; consider including representative failure cases or a discussion of failure modes, especially under occlusion and fast motion.
- [Sec. 4.6] The statement that SAM2 runs at 28 FPS is not connected to the system's 3-5 Hz bottleneck; please clarify which components limit the frame rate.
- [Sec. 3.3] Reference [53] is cited for distributing points on a sphere, but the specific construction from icosahedron vertices and face centers is more concrete; please clarify the resolution levels used.
Circularity Check
No circularity: an empirical systems paper whose central claims are evaluated against external benchmarks, with components built from external prior work and hyperparameters treated as ablated engineering choices.
full rationale
This is an empirical systems paper, not a derivation, and I found no step in which a claimed prediction reduces to an input by construction or via self-citation. The central claims, matching state-of-the-art model-free 6D pose tracking while providing a roughly 5x speedup, are supported by comparisons on the HO3D v3 official test set and the YCBInEOAT dataset using standard metrics (ADD/ADD-S AUC, Chamfer distance, and average processing time per frame). The method's components, LoFTR feature matching with RANSAC coarse pose initialization, 2D Gaussian Splatting rendering, MAD-based keyframe filtering, opacity-percentile pruning, and pose graph optimization, are optimization and engineering modules with parameters such as the bottom 5th percentile opacity threshold and 3x MAD cutoff that are set heuristically and examined through ablations in Table 3, not fitted to the reported pose numbers in a way that makes the output equal to an input. The paper explicitly concedes the HO3D ADD gap (84.33 versus BundleSDF's 89.56 in Table 2), so the headline is a qualified comparative empirical finding rather than a forced result. No load-bearing self-citation is present: the cited works for coarse pose details, differentiable pose-gradient propagation, and 2D Gaussian rendering (BundleTrack, BundleSDF, 2DGS, SplaTAM, MonoGS) are external prior work. The limitations stated in Sec. 5, namely that Gaussian rasterization 'may be less effective in gradient computations compared to differentiable ray casting' and that the optimized 2D Gaussians are not directly integrated into the online pose graph optimization, are honest robustness and integration caveats, not admissions of circularity. The initialization-sensitivity concern raised by the reader is a correctness and generalization risk, not a circularity, because the coarse pose is an input to the optimization rather than a quantity derived from the claimed output. Accordingly, no circular step can be quoted, and the appropriate score is 0.
Assumptions & free parameters
free parameters (4)
- Opacity pruning percentile (bottom 5th / 95th threshold) =
bottom 5th percentile; threshold on 95th percentile not specified
- MAD outlier rejection factor =
3
- Keyframe selection anchor resolution (icosahedron subdivisions) =
not specified
- Visibility threshold for pose graph keyframe selection =
not specified
assumptions (5)
- standard math 2D Gaussian Splatting rendering is differentiable with respect to keyframe poses and Gaussian parameters
- domain assumption The tracked object is rigid and its appearance is static during the optimization window
- domain assumption SAM2 provides consistent, accurate object masks on every frame
- domain assumption LoFTR produces sufficient correct correspondences for RANSAC-based coarse pose initialization
- domain assumption Depth observations are noisy but unbiased enough for joint photometric-geometric optimization
Cite this review
Pith. "Pith review of 6DOPE-GS: Online 6D Object Pose Estimation using Gaussian Splatting." pith.science (2026). https://pith.science/paper/JT2LSHIY
@misc{pith2026241201543,
author = {Pith},
title = {Pith review of: 6DOPE-GS: Online 6D Object Pose Estimation using Gaussian Splatting},
year = {2026},
howpublished = {\url{https://pith.science/paper/JT2LSHIY}},
note = {Machine review of arXiv:2412.01543}
}
abstract
Efficient and accurate object pose estimation is an essential component for modern vision systems in many applications such as Augmented Reality, autonomous driving, and robotics. While research in model-based 6D object pose estimation has delivered promising results, model-free methods are hindered by the high computational load in rendering and inferring consistent poses of arbitrary objects in a live RGB-D video stream. To address this issue, we present 6DOPE-GS, a novel method for online 6D object pose estimation \& tracking with a single RGB-D camera by effectively leveraging advances in Gaussian Splatting. Thanks to the fast differentiable rendering capabilities of Gaussian Splatting, 6DOPE-GS can simultaneously optimize for 6D object poses and 3D object reconstruction. To achieve the necessary efficiency and accuracy for live tracking, our method uses incremental 2D Gaussian Splatting with an intelligent dynamic keyframe selection procedure to achieve high spatial object coverage and prevent erroneous pose updates. We also propose an opacity statistic-based pruning mechanism for adaptive Gaussian density control, to ensure training stability and efficiency. We evaluate our method on the HO3D and YCBInEOAT datasets and show that 6DOPE-GS matches the performance of state-of-the-art baselines for model-free simultaneous 6D pose tracking and reconstruction while providing a 5$\times$ speedup. We also demonstrate the method's suitability for live, dynamic object tracking and reconstruction in a real-world setting.
Figures
Figures from the paper (2 more)
Forward citations
Cited by 1 Pith paper
-
One View, Many Worlds: Single-Image to 3D Object Meets Generative Domain Randomization for One-Shot 6D Pose Estimation
Given one RGB-D photo of an unseen object, an AI-generated 3D mesh, aligned jointly in metric scale and pose, yields state-of-the-art one-shot 6D pose estimation on YCBInEOAT, TOYL, and LM-O.
Reference graph
Works this paper leans on
-
[71]
BundleSDF: Neural 6-DoF Tracking and 3D Re- construction of Unknown Objects,
Bowen Wen, Jonathan Tremblay, Valts Blukis, Stephen Tyree, Thomas Muller, Alex Evans, Dieter Fox, Jan Kautz, and Stan Birchfield. BundleSDF: Neural 6-DoF Tracking and 3D Re- construction of Unknown Objects, . 2, 3, 4, 5, 6, 7
-
[69]
BundleTrack: 6D Pose Track- ing for Novel Objects without Instance or Category-Level 3D Models
Bowen Wen and Kostas Bekris. BundleTrack: 6D Pose Track- ing for Novel Objects without Instance or Category-Level 3D Models. In 2021 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pages 8067–8074. 2, 3, 4, 5, 6, 7
work page 2021
-
[1]
Least-squares fitting of two 3-d point sets
K Somani Arun, Thomas S Huang, and Steven D Blostein. Least-squares fitting of two 3-d point sets. IEEE Transactions on pattern analysis and machine intelligence, (5):698–700,
-
[2]
Revising Densification in Gaussian Splatting
Samuel Rota Bul `o, Lorenzo Porzi, and Peter Kontschieder. Revising Densification in Gaussian Splatting. 5
-
[3]
Zeropose: Cad-prompted zero-shot object 6d pose estimation in cluttered scenes
Jianqiu Chen, Zikun Zhou, Mingshan Sun, Rui Zhao, Liwei Wu, Tianpeng Bao, and Zhenyu He. Zeropose: Cad-prompted zero-shot object 6d pose estimation in cluttered scenes. IEEE Transactions on Circuits and Systems for Video Technology,
-
[4]
Sgpa: Structure-guided prior adapta- tion for category-level 6d object pose estimation
Kai Chen and Qi Dou. Sgpa: Structure-guided prior adapta- tion for category-level 6d object pose estimation. In Proceed- ings of the IEEE/CVF International Conference on Computer Vision, pages 2773–2782, 2021. 2
work page 2021
-
[5]
An overview on visual slam: From tradition to semantic
Weifeng Chen, Guangtao Shang, Aihong Ji, Chengjun Zhou, Xiyang Wang, Chonghui Xu, Zhenxiong Li, and Kai Hu. An overview on visual slam: From tradition to semantic. Remote Sensing, 14(13):3010, 2022. 3
work page 2022
-
[6]
Neural unsigned distance fields for implicit function learning
Julian Chibane, Gerard Pons-Moll, et al. Neural unsigned distance fields for implicit function learning. Advances in Neural Information Processing Systems , 33:21638–21652,
Show all 79 references
-
[7]
Compressing ex- plicit voxel grid representations: fast nerfs become also small
Chenxi Lola Deng and Enzo Tartaglione. Compressing ex- plicit voxel grid representations: fast nerfs become also small. In Proceedings of the IEEE/CVF Winter Conference on Ap- plications of Computer Vision, pages 1236–1245, 2023. 5
2023
-
[8]
PoseRBPF: A Rao–Blackwellized Particle Filter for 6-D Object Pose Tracking
Xinke Deng, Arsalan Mousavian, Yu Xiang, Fei Xia, Timothy Bretl, and Dieter Fox. PoseRBPF: A Rao–Blackwellized Particle Filter for 6-D Object Pose Tracking. 37(5):1328–
-
[9]
Self-supervised 6d object pose estimation for robot manipulation
Xinke Deng, Yu Xiang, Arsalan Mousavian, Clemens Eppner, Timothy Bretl, and Dieter Fox. Self-supervised 6d object pose estimation for robot manipulation. In 2020 IEEE In- ternational Conference on Robotics and Automation (ICRA), pages 3665–3671. IEEE, 2020. 1
2020
-
[10]
Lightgaussian: Unbounded 3d gaussian compression with 15x reduction and 200+ fps.arXiv preprint arXiv:2311.17245, 2023
Zhiwen Fan, Kevin Wang, Kairun Wen, Zehao Zhu, Dejia Xu, and Zhangyang Wang. Lightgaussian: Unbounded 3d gaussian compression with 15x reduction and 200+ fps.arXiv preprint arXiv:2311.17245, 2023. 5
2023 arXiv
-
[11]
Random sample consensus: a paradigm for model fitting with applications to image analysis and automated cartography
Martin A Fischler and Robert C Bolles. Random sample consensus: a paradigm for model fitting with applications to image analysis and automated cartography. Communications of the ACM, 24(6):381–395, 1981. 4
1981
-
[12]
Multi-view stereo: A tutorial
Yasutaka Furukawa, Carlos Hern ´andez, et al. Multi-view stereo: A tutorial. Foundations and Trends® in Computer Graphics and Vision, 9(1-2):1–148, 2015. 2
2015
-
[13]
6d object pose regression via supervised learning on point clouds
Ge Gao, Mikko Lauri, Yulong Wang, Xiaolin Hu, Jianwei Zhang, and Simone Frintrop. 6d object pose regression via supervised learning on point clouds. In 2020 IEEE Interna- tional Conference on Robotics and Automation (ICRA), pages 3643–3649. IEEE, 2020. 2
2020
-
[14]
A survey of 6d object detection based on 3d models for indus- trial applications
Felix Gorschl¨uter, Pavel Rojtberg, and Thomas P¨ollabauer. A survey of 6d object detection based on 3d models for indus- trial applications. Journal of Imaging, 8(3):53, 2022. 1
2022
-
[15]
IRGS: Inter-Reflective Gaussian Splatting with 2D Gaussian Ray Tracing
Chun Gu, Xiaofei Wei, Zixuan Zeng, Yuxuan Yao, and Li Zhang. IRGS: Inter-Reflective Gaussian Splatting with 2D Gaussian Ray Tracing. 8
-
[16]
Honnotate: A method for 3d annotation of hand and object poses
Shreyas Hampali, Mahdi Rad, Markus Oberweger, and Vin- cent Lepetit. Honnotate: A method for 3d annotation of hand and object poses. In Proceedings of the IEEE/CVF con- ference on computer vision and pattern recognition , pages 3196–3206, 2020. 6
2020
-
[17]
OnePose++: Keypoint-Free One-Shot Object Pose Estimation without CAD Models
Xingyi He, Jiaming Sun, Yuang Wang, Di Huang, and Xi- aowei Zhou. OnePose++: Keypoint-Free One-Shot Object Pose Estimation without CAD Models. . 2
-
[18]
FFB6D: A Full Flow Bidirectional Fusion Network for 6D Pose Estimation
Yisheng He, Haibin Huang, Haoqiang Fan, Qifeng Chen, and Jian Sun. FFB6D: A Full Flow Bidirectional Fusion Network for 6D Pose Estimation. In 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pages 3002–3012. IEEE, . 2
2021
-
[19]
PVN3D: A Deep Point-Wise 3D Keypoints V oting Network for 6DoF Pose Estimation
Yisheng He, Wei Sun, Haibin Huang, Jianran Liu, Hao- qiang Fan, and Jian Sun. PVN3D: A Deep Point-Wise 3D Keypoints V oting Network for 6DoF Pose Estimation. In 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 11629–11638. IEEE, . 2
2020
-
[20]
FS6D: Few-Shot 6D Pose Estimation of Novel Ob- jects
Yisheng He, Yao Wang, Haoqiang Fan, Jian Sun, and Qifeng Chen. FS6D: Few-Shot 6D Pose Estimation of Novel Ob- jects. In 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 6804–6814. IEEE, . 2, 6
2022
-
[21]
Single-stage 6d object pose estimation
Yinlin Hu, Pascal Fua, Wei Wang, and Mathieu Salzmann. Single-stage 6d object pose estimation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 2930–2939, 2020. 2
2020
-
[22]
2D Gaussian Splatting for Geometrically Accurate Radiance Fields
Binbin Huang, Zehao Yu, Anpei Chen, Andreas Geiger, and Shenghua Gao. 2D Gaussian Splatting for Geometrically Accurate Radiance Fields. 2, 3, 4
-
[23]
A survey of state-of-the-art on visual slam
Iman Abaspur Kazerouni, Luke Fitzgerald, Gerard Dooly, and Daniel Toal. A survey of state-of-the-art on visual slam. Expert Systems with Applications, 205:117734, 2022. 3
2022
-
[24]
Splatam: Splat track & map 3d gaussians for dense rgb-d slam
Nikhil Keetha, Jay Karhade, Krishna Murthy Jatavallabhula, Gengshan Yang, Sebastian Scherer, Deva Ramanan, and Jonathon Luiten. Splatam: Splat track & map 3d gaussians for dense rgb-d slam. In Proceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition,...
2024
-
[25]
3D Gaussian Splatting for Real-Time Radi- ance Field Rendering
Bernhard Kerbl, Georgios Kopanas, Thomas Leimk¨uhler, and George Drettakis. 3D Gaussian Splatting for Real-Time Radi- ance Field Rendering. 2, 3, 4, 5, 8
-
[26]
Large-scale 6d object pose estimation dataset for industrial bin-picking
Kilian Kleeberger, Christian Landgraf, and Marco F Huber. Large-scale 6d object pose estimation dataset for industrial bin-picking. In 2019 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pages 2573–2578. IEEE, 2019. 1
2019
-
[27]
Cosypose: Consistent multi-view multi-object 6d pose estimation
Yann Labb´e, Justin Carpentier, Mathieu Aubry, and Josef Sivic. Cosypose: Consistent multi-view multi-object 6d pose estimation. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XVII 16, pages 574–591. Springer, 2020. 2 9
2020
-
[28]
Megapose: 6d pose estimation of novel objects via render & compare
Yann Labb´e, Lucas Manuelli, Arsalan Mousavian, Stephen Tyree, Stan Birchfield, Jonathan Tremblay, Justin Carpentier, Mathieu Aubry, Dieter Fox, and Josef Sivic. Megapose: 6d pose estimation of novel objects via render & compare. arXiv preprint arXiv:2212.06870, 2022. 2
2022 arXiv
-
[29]
Geogaussian: Geometry-aware gaussian splatting for scene rendering
Yanyan Li, Chenyu Lyu, Yan Di, Guangyao Zhai, Gim Hee Lee, and Federico Tombari. Geogaussian: Geometry-aware gaussian splatting for scene rendering. In European Confer- ence on Computer Vision , pages 441–457. Springer, 2025. 5
2025
-
[30]
SAM-6D: Segment Anything Model Meets Zero-Shot 6D Object Pose Estimation
Jiehong Lin, Lihua Liu, Dekun Lu, and Kui Jia. SAM-6D: Segment Anything Model Meets Zero-Shot 6D Object Pose Estimation. In 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 27906–27916. IEEE. 2
2024
-
[31]
Kdfnet: Learning keypoint distance field for 6d object pose estimation
Xingyu Liu, Shun Iwase, and Kris M Kitani. Kdfnet: Learning keypoint distance field for 6d object pose estimation. In 2021 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pages 4631–4638. IEEE, 2021. 2
2021
-
[32]
Gen6D: Generalizable Model-Free 6-DoF Object Pose Estimation from RGB Images
Yuan Liu, Yilin Wen, Sida Peng, Cheng Lin, Xiaoxiao Long, Taku Komura, and Wenping Wang. Gen6D: Generalizable Model-Free 6-DoF Object Pose Estimation from RGB Images. 2
-
[33]
Dynamic 3d gaussians: Tracking by persistent dynamic view synthesis
Jonathon Luiten, Georgios Kopanas, Bastian Leibe, and Deva Ramanan. Dynamic 3d gaussians: Tracking by persistent dynamic view synthesis. In 2024 International Conference on 3D Vision (3DV), pages 800–809. IEEE, 2024. 3
2024
-
[34]
A comprehensive survey of visual slam algorithms
Andr´ea Macario Barros, Maugan Michel, Yoann Moline, Gwenol´e Corre, and Fr ´ed´erick Carrel. A comprehensive survey of visual slam algorithms. Robotics, 11(1):24, 2022. 3
2022
-
[35]
Hidenobu Matsuki, Riku Murai, Paul H. J. Kelly, and An- drew J. Davison. Gaussian Splatting SLAM. 6, 7
-
[36]
Gaussian splatting slam
Hidenobu Matsuki, Riku Murai, Paul HJ Kelly, and An- drew J Davison. Gaussian splatting slam. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 18039–18048, 2024. 3, 4
2024
-
[37]
Occupancy networks: Learning 3d reconstruction in function space
Lars Mescheder, Michael Oechsle, Michael Niemeyer, Se- bastian Nowozin, and Andreas Geiger. Occupancy networks: Learning 3d reconstruction in function space. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 4460–4470, 2019. 3
2019
-
[38]
Nerf: Representing scenes as neural radiance fields for view syn- thesis
Ben Mildenhall, Pratul P Srinivasan, Matthew Tancik, Jonathan T Barron, Ravi Ramamoorthi, and Ren Ng. Nerf: Representing scenes as neural radiance fields for view syn- thesis. Communications of the ACM , 65(1):99–106, 2021. 3
2021
-
[39]
3d gaussian ray tracing: Fast tracing of particle scenes
Nicolas Moenne-Loccoz, Ashkan Mirzaei, Or Perel, Riccardo de Lutio, Janick Martinez Esturo, Gavriel State, Sanja Fidler, Nicholas Sharp, and Zan Gojcic. 3d gaussian ray tracing: Fast tracing of particle scenes. arXiv preprint arXiv:2407.07090,
-
[40]
Kinect- fusion: Real-time dense surface mapping and tracking
Richard A Newcombe, Shahram Izadi, Otmar Hilliges, David Molyneaux, David Kim, Andrew J Davison, Pushmeet Kohi, Jamie Shotton, Steve Hodges, and Andrew Fitzgibbon. Kinect- fusion: Real-time dense surface mapping and tracking. In 2011 10th IEEE international symposium on mixed ...
2011
-
[41]
GigaPose: Fast and Robust Novel Object Pose Estimation via One Correspondence
Van Nguyen Nguyen, Thibault Groueix, Mathieu Salzmann, and Vincent Lepetit. GigaPose: Fast and Robust Novel Object Pose Estimation via One Correspondence. In 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 9903–9913. IEEE. 2
2024
-
[42]
isdf: Real-time neural signed distance fields for robot percep- tion
Joseph Ortiz, Alexander Clegg, Jing Dong, Edgar Sucar, David Novotny, Michael Zollhoefer, and Mustafa Mukadam. isdf: Real-time neural signed distance fields for robot percep- tion. In Robotics: Science and Systems, 2022. 3
2022
-
[43]
A survey of structure from motion*
Onur ¨Ozyes ¸il, Vladislav V oroninski, Ronen Basri, and Amit Singer. A survey of structure from motion*. Acta Numerica, 26:305–364, 2017. 2
2017
-
[44]
Deepsdf: Learning continuous signed distance functions for shape representation
Jeong Joon Park, Peter Florence, Julian Straub, Richard New- combe, and Steven Lovegrove. Deepsdf: Learning continuous signed distance functions for shape representation. In Pro- ceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 165–174, 2019. 3
2019
-
[45]
Latentfusion: End-to-end differentiable reconstruction and rendering for unseen object pose estimation
Keunhong Park, Arsalan Mousavian, Yu Xiang, and Dieter Fox. Latentfusion: End-to-end differentiable reconstruction and rendering for unseen object pose estimation. In Proceed- ings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10710–10719, 2020. 2
2020
-
[46]
Automatic differentiation in pytorch
Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Al- ban Desmaison, Luca Antiga, and Adam Lerer. Automatic differentiation in pytorch. 2017. 4
2017
-
[47]
Derpanis, and Kostas Daniilidis
Georgios Pavlakos, Xiaowei Zhou, Aaron Chan, Konstanti- nos G. Derpanis, and Kostas Daniilidis. 6-DoF object pose from semantic keypoints. In 2017 IEEE International Confer- ence on Robotics and Automation (ICRA), pages 2011–2018. 2
2017
-
[48]
PVNet: Pixel-Wise V oting Network for 6DoF Pose Estimation
Sida Peng, Yuan Liu, Qixing Huang, Xiaowei Zhou, and Hujun Bao. PVNet: Pixel-Wise V oting Network for 6DoF Pose Estimation. pages 4561–4570. 2
-
[49]
Bb8: A scalable, accurate, robust to partial occlusion method for predicting the 3d poses of challenging objects without using depth
Mahdi Rad and Vincent Lepetit. Bb8: A scalable, accurate, robust to partial occlusion method for predicting the 3d poses of challenging objects without using depth. In Proceedings of the IEEE international conference on computer vision, pages 3828–3836, 2017. 2
2017
-
[50]
SAM 2: Segment Anything in Images and Videos
Nikhila Ravi, Valentin Gabeur, Yuan-Ting Hu, Ronghang Hu, Chaitanya Ryali, Tengyu Ma, Haitham Khedr, Roman R¨adle, Chloe Rolland, Laura Gustafson, Eric Mintun, Junting Pan, Vasudev Alwala, Nicolas Carion, Chao-Yuan Wu, Ross Girshick, Piotr Doll´ar, and Christoph Feichtenhofer....
-
[51]
Deepsfm: Robust deep iterative refinement for structure from motion
Xinlin Ren, Xingkui Wei, Zhuwen Li, Yanwei Fu, Yinda Zhang, and Xiangyang Xue. Deepsfm: Robust deep iterative refinement for structure from motion. IEEE Transactions on Pattern Analysis and Machine Intelligence, 46(6):4058–4074,
-
[52]
Maskfu- sion: Real-time recognition, tracking and reconstruction of multiple moving objects
Martin Runz, Maud Buffier, and Lourdes Agapito. Maskfu- sion: Real-time recognition, tracking and reconstruction of multiple moving objects. In 2018 IEEE International Sympo- sium on Mixed and Augmented Reality (ISMAR), pages 10–20. IEEE, 2018. 7
2018
-
[53]
Distributing many points on a sphere
Edward B Saff and Amo BJ Kuijlaars. Distributing many points on a sphere. The mathematical intelligencer, 19:5–11,
-
[54]
Osop: A multi-stage one shot object pose estimation frame- work
Ivan Shugurov, Fu Li, Benjamin Busam, and Slobodan Ilic. Osop: A multi-stage one shot object pose estimation frame- work. In Proceedings of the IEEE/CVF Conference on Com- puter Vision and Pattern Recognition, pages 6835–6844, 2022. 2
2022
-
[55]
Sdf-2-sdf: Highly accurate 3d object reconstruction
Miroslava Slavcheva, Wadim Kehl, Nassir Navab, and Slobo- dan Ilic. Sdf-2-sdf: Highly accurate 3d object reconstruction. In Computer Vision–ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11–14, 2016, Proceed- ings, Part I 14, pages 680–696. Springer, ...
2016
-
[56]
Learn- ing to assemble: Estimating 6d poses for robotic object-object manipulation
Stefan Stev ˇsi´c, Sammy Christen, and Otmar Hilliges. Learn- ing to assemble: Estimating 6d poses for robotic object-object manipulation. IEEE Robotics and Automation Letters, 5(2): 1159–1166, 2020. 1
2020
-
[57]
Deep multi-state object pose estimation for augmented reality assembly
Yongzhi Su, Jason Rambach, Nareg Minaskan, Paul Lesur, Alain Pagani, and Didier Stricker. Deep multi-state object pose estimation for augmented reality assembly. In2019 IEEE International Symposium on Mixed and Augmented Reality Adjunct (ISMAR-Adjunct), pages 222–227. IEEE, 2019. 1
2019
-
[58]
LoFTR: Detector-Free Local Feature Match- ing with Transformers
Jiaming Sun, Zehong Shen, Yuang Wang, Hujun Bao, and Xiaowei Zhou. LoFTR: Detector-Free Local Feature Match- ing with Transformers. In 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pages 8918–8927. IEEE. 3, 4
2021
-
[59]
Onepose: One-shot object pose estimation without cad mod- els
Jiaming Sun, Zihao Wang, Siyu Zhang, Xingyi He, Hongcheng Zhao, Guofeng Zhang, and Xiaowei Zhou. Onepose: One-shot object pose estimation without cad mod- els. In Proceedings of the IEEE/CVF Conference on Com- puter Vision and Pattern Recognition, pages 6825–6834, 2022. 2
2022
-
[60]
Implicit 3d orientation learning for 6d object detection from rgb images
Martin Sundermeyer, Zoltan-Csaba Marton, Maximilian Durner, Manuel Brucker, and Rudolph Triebel. Implicit 3d orientation learning for 6d object detection from rgb images. In Proceedings of the european conference on computer vi- sion (ECCV), pages 699–715, 2018. 2
2018
-
[61]
Droid-slam: Deep visual slam for monocular, stereo, and rgb-d cameras
Zachary Teed and Jia Deng. Droid-slam: Deep visual slam for monocular, stereo, and rgb-d cameras. Advances in neural information processing systems, 34:16558–16569, 2021. 6, 7
2021
-
[62]
Sinha, and Pascal Fua
Bugra Tekin, Sudipta N. Sinha, and Pascal Fua. Real-Time Seamless Single Shot 6D Object Pose Prediction. pages 292–
-
[63]
6- PACK: Category-level 6D Pose Tracker with Anchor-Based Keypoints,
Chen Wang, Roberto Mart ´ın-Mart´ın, Danfei Xu, Jun Lv, Cewu Lu, Li Fei-Fei, Silvio Savarese, and Yuke Zhu. 6- PACK: Category-level 6D Pose Tracker with Anchor-Based Keypoints, . 2
-
[64]
DenseFusion: 6D Object Pose Estimation by Iterative Dense Fusion
Chen Wang, Danfei Xu, Yuke Zhu, Roberto Martin-Martin, Cewu Lu, Li Fei-Fei, and Silvio Savarese. DenseFusion: 6D Object Pose Estimation by Iterative Dense Fusion. pages 3343–3352, . 2
-
[65]
3d reconstruction with spatial memory
Hengyi Wang and Lourdes Agapito. 3d reconstruction with spatial memory. arXiv preprint arXiv:2408.16061, 2024. 3
2024 arXiv
-
[66]
He Wang, Srinath Sridhar, Jingwei Huang, Julien Valentin, Shuran Song, and Leonidas J. Guibas. Normalized Object Coordinate Space for Category-Level 6D Object Pose and Size Estimation, . 2
-
[67]
Dust3r: Geometric 3d vi- sion made easy
Shuzhe Wang, Vincent Leroy, Yohann Cabon, Boris Chidlovskii, and Jerome Revaud. Dust3r: Geometric 3d vi- sion made easy. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 20697– 20709, 2024. 3
2024
-
[68]
Multi-view stereo in the deep learning era: A comprehensive review
Xiang Wang, Chen Wang, Bing Liu, Xiaoqing Zhou, Liang Zhang, Jin Zheng, and Xiao Bai. Multi-view stereo in the deep learning era: A comprehensive review. Displays, 70: 102102, 2021. 2
2021
-
[70]
Bowen Wen, Chaitanya Mitash, Baozhang Ren, and Kostas E. Bekris. Se(3)-TrackNet: Data-driven 6D Pose Tracking by Calibrating Image Residuals in Synthetic Domains. In 2020 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pages 10367–10373, . 2
2020
-
[72]
Foun- dationPose: Unified 6D Pose Estimation and Tracking of Novel Objects,
Bowen Wen, Wei Yang, Jan Kautz, and Stan Birchfield. Foun- dationPose: Unified 6D Pose Estimation and Tracking of Novel Objects, . 2
-
[73]
se (3)-tracknet: Data-driven 6d pose tracking by calibrating image residuals in synthetic domains
Bowen Wen, Chaitanya Mitash, Baozhang Ren, and Kostas E Bekris. se (3)-tracknet: Data-driven 6d pose tracking by calibrating image residuals in synthetic domains. In 2020 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pages 10367–10373. IEEE, 2020. 5
2020
-
[74]
Posecnn: A convolutional neural network for 6d object pose estimation in cluttered scenes
Yu Xiang, Tanner Schmidt, Venkatraman Narayanan, and Dieter Fox. Posecnn: A convolutional neural network for 6d object pose estimation in cluttered scenes. arXiv preprint arXiv:1711.00199, 2017. 6
2017 arXiv
-
[75]
Gs-slam: Dense visual slam with 3d gaussian splatting
Chi Yan, Delin Qu, Dan Xu, Bin Zhao, Zhigang Wang, Dong Wang, and Xuelong Li. Gs-slam: Dense visual slam with 3d gaussian splatting. In Proceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition, pages 19595–19604, 2024. 3, 4
2024
-
[76]
Monst3r: A simple approach for estimating geometry in the presence of motion
Junyi Zhang, Charles Herrmann, Junhwa Hur, Varun Jampani, Trevor Darrell, Forrester Cole, Deqing Sun, and Ming-Hsuan Yang. Monst3r: A simple approach for estimating geometry in the presence of motion. arXiv preprint arXiv:2410.03825,
-
[77]
Augmented reality system based on real-time object 6d pose estimation
Yan Zhao, Shaobo Zhang, Wanqing Zhao, Ying Wei, and Jinye Peng. Augmented reality system based on real-time object 6d pose estimation. In 2023 2nd International Confer- ence on Image Processing and Media Computing (ICIPMC), pages 27–34. IEEE, 2023. 1
2023
-
[78]
Nice-slam: Neural implicit scalable encoding for slam
Zihan Zhu, Songyou Peng, Viktor Larsson, Weiwei Xu, Hujun Bao, Zhaopeng Cui, Martin R Oswald, and Marc Pollefeys. Nice-slam: Neural implicit scalable encoding for slam. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 12786–12796, 2022. 6, 7
2022
-
[79]
FoundPose: Unseen Object Pose Estimation with Foundation Features
Evin Pınar ¨Ornek, Yann Labb´e, Bugra Tekin, Lingni Ma, Cem Keskin, Christian Forster, and Tomas Hodan. FoundPose: Unseen Object Pose Estimation with Foundation Features. 2 11
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.