REVIEW 3 major objections 6 minor 63 references
Diffusion Features for Zero-Shot 6DoF Object Pose Estimation
T0 review · 3 major / 6 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read Latent diffusion model features, extracted from Stable Diffusion's decoder, improve zero-shot 6DoF object pose estimation over Vision Transformer features, raising Average Recall by up to 27% on standard benchmarks.
desk verdict Diffusion features for zero-shot pose estimation: a useful pipeline paper, but the claim that LDM features beat ViT features is not actually isolated from the pipeline changes that come with it. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central mechanism is the PCA co-projection of aggregated diffusion hyperfeatures. Hyperfeatures are pixel-wise feature vectors formed by aggregating multiple feature maps from different layers of the diffusion decoder. Query and template feature maps from Stable Diffusion's decoder layers 2, 5, 8, and 11 are concatenated and projected with a shared PCA basis; the co-projected features are split back, upsampled to the resolution of the finest layer, and concatenated along the feature dimension. K-means clustering then creates surface clusters in both embeddings, and correspondences are estimated only within clusters that match by cosine similarity, with RANSAC-based fundamental matrix filtering to discard mismatched clusters and a sub-pixel refinement step for small or distant objects. This co-projection is the load-bearing element that creates a common space in which cluster-wise cosine similarity can match query pixels to template pixels.
What would settle it
Take the DZOP pipeline and swap the Stable Diffusion backbone for DINOv2 features, keeping the co-projection, clustering, and pose solving unchanged. If Average Recall on LMO, YCBV, and TLESS matches or exceeds the reported diffusion results, then the main claim that LDM features are more effective than ViT features for this task is refuted. This is directly testable because the infrastructure is already built for the DINO baseline.
Extended reading notes
Core claim
The central claim is that Latent Diffusion Model features, taken from the U-Net decoder of Stable Diffusion, are more effective than Vision Transformer features for both template matching and correspondence estimation in zero-shot 6DoF object pose estimation. Under a controlled comparison that keeps segmentation, templates, geometric correspondence retrieval, and pose solving identical to ZS6D, swapping DINO for Stable Diffusion improves Average Recall by 10.08% on LMO, 12.65% on YCBV, and 27.14% on TLESS, while also improving the template accuracy metric Acc15 on all three datasets. The paper further claims that the co-projection of hyperfeatures from decoder layers 2, 5, 8, and 11 into a shared low-dimensional space, followed by cluster-wise matching, is what makes the diffusion features usable for object-level correspondence; ablations on LMO show that removing co-projection or clustering degrades performance. The result is a template-based zero-shot pose estimator that outperforms the ViT baseline and, on LMO and YCBV, also outperforms zero-shot methods that fine-tune their feature extractor.
Load-bearing premise
The premise that the PCA co-projection of Stable Diffusion decoder features from layers 2, 5, 8, and 11 creates a shared space where cluster-wise cosine similarity reliably matches pixel correspondences; this is validated only on LMO and not shown to transfer to the other datasets.
Editorial extensions
If this is right
- Zero-shot pose estimation can be built on generative-model features without object-specific fine-tuning, improving accuracy on occluded, textureless, and illumination-varying objects over the ViT baseline.
- The PCA co-projection and cluster-wise correspondence matching constitute a reusable strategy for adapting diffusion features to object-level matching tasks, not just scene-level correspondence as previously shown.
- The gains at low error tolerances mean diffusion features are useful for applications requiring precise poses, such as robotic grasping, where tight alignment matters.
- The framework provides a controlled comparison point for evaluating future vision foundation models for pose estimation, as evidenced by the paper's own comparison with other ViT-based methods.
- The method inherits a computational cost of 50 diffusion steps per template and query, so runtime efficiency remains a bottleneck for real-time use.
Reading between the lines
- A likely extension not tested here is whether the co-projection step could be replaced by a learned projection; if so, the pipeline could adapt to new domains with fewer hand-set hyperparameters.
- The 27% gain on TLESS hints that diffusion features may encode shape and geometry priors that are especially valuable when texture is absent; a focused study varying texture and symmetry would test that conjecture.
- The paper's comparison with FoundPose suggests that newer ViT features (DINOv2) applied through a similar template-matching strategy can also outperform DINO; this implies the feature extractor choice, rather than the architecture family, may be the dominant factor. A broader benchmark across more backbones would clarify this.
- If SD features are distilled into a faster extractor, the 50-step denoising cost could be amortized, making the zero-shot pose pipeline practical for online robotics; the paper does not address this but the architecture permits it.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes DZOP, a zero-shot 6DoF object pose estimation pipeline that replaces the DINO ViT feature extractor used in ZS6D with Stable Diffusion (SD) features. The method uses SD decoder layer 2 features for template matching, aggregates hyperfeatures from layers 2, 5, 8, and 11 with PCA co-projection for correspondence estimation, applies k-means cluster-wise matching with RANSAC, and includes sub-pixel refinement. Experiments on LMO, YCBV, and TLESS report Average Recall improvements over ZS6D of 10.08%, 12.65%, and 27.14%, respectively. The authors claim that LDM features are more effective than ViT features for template matching and zero-shot pose estimation, and they release the source code.
Significance. If the central claim were established, the result would be significant: it would demonstrate that latent diffusion backbones can outperform self-supervised ViTs in a zero-shot pose estimation setting, a task currently dominated by ViT-based feature extractors. The paper has several strengths: it uses the same templates, CNOS segmentations, and PnP solver as ZS6D, making the method-level comparison fair; it provides ablations of the main correspondence stages; and it releases code. However, the headline comparison is not a controlled feature comparison because the pipeline changes simultaneously with the feature extractor. The significance of the paper is therefore conditional on additional experiments that isolate the contribution of the LDM features.
major comments (3)
- [Sec. 1, Tables 1-2] The reported AR gains of DZOP over ZS6D (10.08% on LMO, 12.65% on YCBV, 27.14% on TLESS) cannot be attributed to the LDM feature family because the comparison changes the feature extractor and the downstream correspondence pipeline simultaneously. The paper states in Sec. 1 that 'to ensure optimal performance with diffusion features, we also modify the downstream pipeline stages,' and Sec. 3.3 introduces PCA co-projection (Eq. 5), k-means cluster-wise matching, RANSAC-based cluster filtering, and sub-pixel refinement. Table 3 shows that the co-projection step alone changes AR from 0.149 to 0.427 on LMO, demonstrating that these components are load-bearing. To support Contribution 2, the authors need a controlled comparison, e.g., DINO features inside the DZOP correspondence pipeline, or SD features inside the original ZS6D pipeline without the additional components.
- [Sec. 4.4, Fig. 6] The hyperparameters used in the main experiments (200 clusters, PCA dimension 64, 10 correspondences per cluster, 50 SD timesteps) are selected using LMO ground-truth detections and then applied to the LMO result reported in Tables 1-2. This selection-on-test-data inflates the LMO improvement and weakens the cross-dataset evidence for the central claim. The authors should fix these hyperparameters a priori or use a held-out validation split, e.g., tune on one dataset and report the other two as untuned.
- [Sec. 4.4, Table 3] The ablations in Table 3 show that co-projection, clustering, and sub-pixel refinement improve DZOP on LMO, but they do not compare DINO features in the same correspondence pipeline. Without a cross-feature ablation, the paper does not establish that the benefits of these components are specific to SD features, nor that LDM features are 'more effective' for pose estimation. Similarly, the template-matching comparison in Table 2 compares SD layer 2 features against DINO features used by ZS6D, but the matching procedures may differ; the paper should clarify whether the template matching mechanism is otherwise identical.
minor comments (6)
- [Sec. 4.2, after Table 1] The text 'The underlined values in Figure 1 indicate the highest AR compared to the ZS6D baseline' should refer to Table 1, not Figure 1.
- [Sec. 4.4, first paragraph] The sentence 'Figure 3 shows the influence of the correspondence matching functions on the AR and runtime' is a cross-reference error; the ablation results are presented in Table 3, while Figure 3 shows per-object AR.
- [Eq. 5] The notation in Eq. (5) is unclear: it writes the projected features as the product of the union of features with V, but does not define how the union is formed or how the result is partitioned into query and template parts. Please define the concatenation and splitting explicitly.
- [Sec. 4.1, metrics] The sentence 'Where is of the metrics are an average recall over different error thresholds of θ' is grammatically garbled and should be rewritten.
- [Table 1] The table heading describes the methods as 'single-shot monocular,' but HiPose (Ref. [5]) is an RGB-D method; please clarify or correct the categorization.
- [Sec. 3.3, Eq. 6] In the sub-pixel refinement formula, the indices i and j are used both as summation variables and as coordinate indices, which is confusing; please define the local neighborhood coordinates explicitly.
Circularity Check
No significant circularity: the paper is an empirical comparison against external ground-truth pose benchmarks, and no target quantity is defined in terms of the method's own outputs.
full rationale
The derivation chain in DZOP is an empirical pipeline, not a formal derivation from assumptions. Template matching (Eq. 4), correspondence estimation (Eqs. 5-6), and PnP pose retrieval all feed into metrics (AR, Acc15) that are computed against ground-truth poses from LMO, YCBV, and TLESS, so the reported quantities are externally falsifiable rather than equivalent to the inputs by construction. The only self-citation is the ZS6D baseline (Ref. [6]), but that is a published method with its own external evaluation; using it as the comparison baseline is standard practice and is not load-bearing circularity. The PCA co-projection in Eq. 5 is fit to the query and template features, but the correspondence quality is evaluated by final pose accuracy against ground truth, not by the projection itself. Hyperparameters (200 clusters, PCA dimension 64, 10 correspondences, 50 timesteps) are chosen via ablations on LMO (Fig. 6) and then applied to the LMO headline result; this is a selection-bias/statistical concern about the strength of the evidence, not a circular reduction, because the reported AR is still a measurement of an independent ground-truth quantity. The skeptic's point that the DZOP-vs-ZS6D comparison changes both the feature extractor and downstream stages is a valid confound for causal attribution of the gains to LDM features, but it does not make the 'prediction' equal to the input; it is an experimental-control issue rather than circularity. No circular step can be exhibited from the text.
Assumptions & free parameters
free parameters (4)
- number of correspondence clusters =
200
- PCA co-projection dimension =
64
- correspondences per cluster =
10
- Stable Diffusion timesteps for feature extraction =
50
assumptions (4)
- domain assumption SD U-Net decoder feature maps from layers 2, 5, 8, and 11 provide a semantically aligned multiscale representation for object-level correspondence
- domain assumption The rendered template set of 300 views per object covers the viewpoint sphere sufficiently for nearest-template retrieval to recover coarse pose
- domain assumption CNOS masks are accurate enough to serve as object location priors for both methods
- standard math BOP metrics (VSD, MSSD, MSPD) in Eq. 7 are computed correctly and match the BOP challenge definitions
Cite this review
Pith. "Pith review of Diffusion Features for Zero-Shot 6DoF Object Pose Estimation." pith.science (2026). https://pith.science/paper/K6ET7QN5
@misc{pith2026241116668,
author = {Pith},
title = {Pith review of: Diffusion Features for Zero-Shot 6DoF Object Pose Estimation},
year = {2026},
howpublished = {\url{https://pith.science/paper/K6ET7QN5}},
note = {Machine review of arXiv:2411.16668}
}
read the original abstract
Zero-shot object pose estimation enables the retrieval of object poses from images without necessitating object-specific training. In recent approaches this is facilitated by vision foundation models (VFM), which are pre-trained models that are effectively general-purpose feature extractors. The characteristics exhibited by these VFMs vary depending on the training data, network architecture, and training paradigm. The prevailing choice in this field are self-supervised Vision Transformers (ViT). This study assesses the influence of Latent Diffusion Model (LDM) backbones on zero-shot pose estimation. In order to facilitate a comparison between the two families of models on a common ground we adopt and modify a recent approach. Therefore, a template-based multi-staged method for estimating poses in a zero-shot fashion using LDMs is presented. The efficacy of the proposed approach is empirically evaluated on three standard datasets for object-specific 6DoF pose estimation. The experiments demonstrate an Average Recall improvement of up to 27% over the ViT baseline. The source code is available at: https://github.com/BvG1993/DZOP.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[1]
Jian Liu, Wei Sun, Chongpei Liu, Xing Zhang, and Qiang Fu. Robotic continuous grasping system by shape transformer-guided multiobject category-level 6-d pose estimation. IEEE Transactions on Industrial Informatics, 19(11):11171–11181, 2023
work page 2023
-
[2]
Challenges for monocular 6-d object pose estimation in robotics
Stefan Thalhammer, Dominik Bauer, Peter H¨ onig, Jean-Baptiste Weibel, Jos´ e Garc ´ ıa-Rodr ´ ıguez, and Markus Vincze. Challenges for monocular 6-d object pose estimation in robotics. IEEE Transactions on Robotics, 40:4065–4084, 2024
work page 2024
-
[3]
Blenderproc: Reducing the reality gap with photorealistic rendering
Maximilian Denninger, Martin Sundermeyer, Dominik Winkelbauer, Dmitry Olefir, Tomas Hodan, Youssef Zidan, Mohamad Elbadrawy, Markus Knauer, Harinandan Katam, and Ahsan Lodhi. Blenderproc: Reducing the reality gap with photorealistic rendering. 2020
work page 2020
-
[4]
Bop challenge 2023 on detection segmentation and pose estimation of seen and unseen rigid objects
Tomas Hodan, Martin Sundermeyer, Yann Labbe, Van Nguyen Nguyen, Gu Wang, Eric Brach- mann, Bertram Drost, Vincent Lepetit, Carsten Rother, and Jiri Matas. Bop challenge 2023 on detection segmentation and pose estimation of seen and unseen rigid objects. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5610–5619, 2024
work page 2023
-
[5]
Yongliang Lin, Yongzhi Su, Praveen Nathan, Sandeep Inuganti, Yan Di, Martin Sundermeyer, Fabian Manhardt, Didier Stricker, Jason Rambach, and Yu Zhang. Hipose: Hierarchical binary surface encoding and correspondence pruning for rgb-d 6dof object pose estimation. In 2024 Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pag...
work page 2024
-
[6]
Zs6d: Zero-shot 6d object pose estimation using vision transformers
Philipp Ausserlechner, David Haberger, Stefan Thalhammer, Jean-Baptiste Weibel, and Markus Vincze. Zs6d: Zero-shot 6d object pose estimation using vision transformers. In 2024 IEEE International Conference on Robotics and Automation (ICRA), pages 463–469. IEEE, 2024
work page 2024
-
[7]
Megapose: 6d pose estimation of novel objects via render & compare
Yann Labb´ e, Lucas Manuelli, Arsalan Mousavian, Stephen Tyree, Stan Birchfield, Jonathan Trem- blay, Justin Carpentier, Mathieu Aubry, Dieter Fox, and Josef Sivic. Megapose: 6d pose estimation of novel objects via render & compare. In CoRL 2022-Conference on Robot Learning, 2022
work page 2022
-
[8]
Foundpose: Unseen object pose estimation with foundation features
Evin Pınar ¨Ornek, Yann Labb´ e, Bugra Tekin, Lingni Ma, Cem Keskin, Christian Forster, and Tomas Hodan. Foundpose: Unseen object pose estimation with foundation features. pages 163– 182, 2024
work page 2024
Show all 63 references
-
[9]
Osop: A multi-stage one shot object pose estimation framework
Ivan Shugurov, Fu Li, Benjamin Busam, and Slobodan Ilic. Osop: A multi-stage one shot object pose estimation framework. In 2022 Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6835–6844, 2022
2022
-
[10]
Foundationpose: Unified 6d pose estima- tion and tracking of novel objects
Bowen Wen, Wei Yang, Jan Kautz, and Stan Birchfield. Foundationpose: Unified 6d pose estima- tion and tracking of novel objects. In 2024 Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 17868–17879, 2024
2024
-
[11]
Zero123-6d: Zero-shot novel view syn- thesis for rgb category-level 6d pose estimation
Francesco Di Felice, Alberto Remus, Stefano Gasperini, Benjamin Busam, Lionel Ott, Federico Tombari, Roland Siegwart, and Carlo Alberto Avizzano. Zero123-6d: Zero-shot novel view syn- thesis for rgb category-level 6d pose estimation. arXiv preprint arXiv:2403.14279, 2024
2024 arXiv
-
[12]
Emerging properties in self-supervised vision transformers
Mathilde Caron, Hugo Touvron, Ishan Misra, Herv´ e J´ egou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. Emerging properties in self-supervised vision transformers. In 2021Proceedings of the IEEE/CVF international conference on computer vision, pages 9650–9660. IEEE, 2021
2021
-
[13]
Epnp: An accurate o(n) solution to the pnp problem
Vincent Lepetit, Francesc Moreno-Noguer, and Pascal Fua. Epnp: An accurate o(n) solution to the pnp problem. International Journal of Computer Vision, 81(2):155–166, 2009
2009
-
[14]
High- resolution image synthesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bj¨ orn Ommer. High- resolution image synthesis with latent diffusion models. In 2022 Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 10684–10695, June 2022
2022
-
[15]
A tale of two features: Stable diffusion complements dino for zero- shot semantic correspondence
Junyi Zhang, Charles Herrmann, Junhwa Hur, Luisa Polania Cabrera, Varun Jampani, Deqing Sun, and Ming-Hsuan Yang. A tale of two features: Stable diffusion complements dino for zero- shot semantic correspondence. Advances in Neural Information Processing Systems, 36, 2024
2024
-
[16]
Learning 6d object pose estimation using 3d object coordinates
Eric Brachmann, Alexander Krull, Frank Michel, Stefan Gumhold, Jamie Shotton, and Carsten Rother. Learning 6d object pose estimation using 3d object coordinates. In Computer Vision– ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Pa...
2014
-
[17]
Xiang, T
Y. Xiang, T. Schmidt, V. Narayanan, and D. Fox. Posecnn: A convolutional neural network for 6d object pose estimation in cluttered scenes. In 2018 Proceedings of Robotics: Science and Systems (RSS), 2018
2018
-
[18]
T-less: An rgb-d dataset for 6d pose estimation of texture-less objects
Tomas Hodan, Frank Michel, Eric Brachmann, Wadim Kehl, Anders Buch, Dirk Kraft, Bertram Drost, Joao Vidal, Stephan Ihrke, Xenophon Zabulis, et al. T-less: An rgb-d dataset for 6d pose estimation of texture-less objects. In 2017 IEEE Winter Conference on Applications of Compute...
2017
-
[19]
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. In 2021 Internationa...
2021
-
[20]
Cad-model recognition and 6dof pose estimation using 3d cues
Aitor Aldoma, Markus Vincze, Nico Blodow, David Gossow, Suat Gedikli, Radu Bogdan Rusu, and Gary Bradski. Cad-model recognition and 6dof pose estimation using 3d cues. In 2011 IEEE international conference on computer vision workshops (ICCV workshops), pages 585–592. IEEE, 2011
2011
-
[21]
Pose estimation using local structure-specific shape and appearance context
Anders Glent Buch, Dirk Kraft, Joni-Kristian Kamarainen, Henrik Gordon Petersen, and Norbert Kr¨ uger. Pose estimation using local structure-specific shape and appearance context. In 2013 IEEE international conference on robotics and automation, pages 2080–2087. IEEE, 2013
2013
-
[22]
Model globally, match locally: Ef- ficient and robust 3d object recognition
Bertram Drost, Markus Ulrich, Nassir Navab, and Slobodan Ilic. Model globally, match locally: Ef- ficient and robust 3d object recognition. In 2010 IEEE Computer Society Conference on Computer Vision and Pattern Recognition, pages 998–1005. IEEE, 2010
2010
-
[23]
Model based training, detection and pose estimation of texture-less 3d objects in heavily cluttered scenes
Stefan Hinterstoisser, Vincent Lepetit, Slobodan Ilic, Stefan Holzer, Gary Bradski, Kurt Konolige, and Nassir Navab. Model based training, detection and pose estimation of texture-less 3d objects in heavily cluttered scenes. In Computer Vision–ACCV 2012: 11th Asian Conference ...
2012
-
[24]
Faster and finer pose estimation for multiple instance objects in a single rgb image
Lee Aing, Wen-Nung Lie, and Guo-Shiang Lin. Faster and finer pose estimation for multiple instance objects in a single rgb image. Image and Vision Computing, 130:104618, 2023
2023
-
[25]
Irpe: Instance-level reconstruction-based 6d pose estimator
Le Jin, Guoshun Zhou, Zherong Liu, Yuanchao Yu, Teng Zhang, Minghui Yang, and Jun Zhou. Irpe: Instance-level reconstruction-based 6d pose estimator. Image and Vision Computing, page 105340, 2024
2024
-
[26]
Ssd-6d: Making rgb-based 3d detection and 6d pose estimation great again
Wadim Kehl, Fabian Manhardt, Federico Tombari, Slobodan Ilic, and Nassir Navab. Ssd-6d: Making rgb-based 3d detection and 6d pose estimation great again. In 2017 Proceedings of the IEEE international conference on computer vision, pages 1521–1529, 2017
2017
-
[27]
Cosypose: Consistent multi-view multi-object 6d pose estimation
Yann Labb´ e, Justin Carpentier, Mathieu Aubry, and Josef Sivic. Cosypose: Consistent multi-view multi-object 6d pose estimation. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XVII 16, pages 574–591. Springer, 2020
2020
-
[28]
A dynamic keypoint selection network for 6dof pose estimation
Haowen Sun, Taiyong Wang, and Enlin Yu. A dynamic keypoint selection network for 6dof pose estimation. Image and Vision Computing, 118:104372, 2022
2022
-
[29]
Real-time seamless single shot 6d object pose prediction
Bugra Tekin, Sudipta N Sinha, and Pascal Fua. Real-time seamless single shot 6d object pose prediction. In 2018 Proceedings of the IEEE conference on computer vision and pattern recognition, pages 292–301, 2018
2018
-
[30]
Gdr-net: Geometry-guided direct regression network for monocular 6d object pose estimation
Gu Wang, Fabian Manhardt, Federico Tombari, and Xiangyang Ji. Gdr-net: Geometry-guided direct regression network for monocular 6d object pose estimation. In 2021 Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 16611–16621, 2021
2021
-
[31]
Oa-pose: Occlusion-aware monoc- ular 6-dof object pose estimation under geometry alignment for robot manipulation
Jikun Wang, Luqing Luo, Weixiang Liang, and Zhi-Xin Yang. Oa-pose: Occlusion-aware monoc- ular 6-dof object pose estimation under geometry alignment for robot manipulation. Pattern Recognition, 154:110576, 2024
2024
-
[32]
Multiple geometry representations for 6d object pose estimation in occluded or truncated scenes
Jichun Wang, Lemiao Qiu, Guodong Yi, Shuyou Zhang, and Yang Wang. Multiple geometry representations for 6d object pose estimation in occluded or truncated scenes. Pattern Recognition, 132:108903, 2022
2022
-
[33]
Geometric-aware dense matching network for 6d pose estimation of objects from rgb-d images
Chenrui Wu, Long Chen, Shenglong Wang, Han Yang, and Junjie Jiang. Geometric-aware dense matching network for 6d pose estimation of objects from rgb-d images. Pattern Recognition, 137:109293, 2023. 14
2023
-
[34]
Gpv-pose: Category-level object pose estimation via geometry-guided point-wise voting
Yan Di, Ruida Zhang, Zhiqiang Lou, Fabian Manhardt, Xiangyang Ji, Nassir Navab, and Federico Tombari. Gpv-pose: Category-level object pose estimation via geometry-guided point-wise voting. In 2022 Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognitio...
2022
-
[35]
i2c-net: using instance-level neural networks for monocular category-level 6d pose estimation
Alberto Remus, Salvatore D’Avella, Francesco Di Felice, Paolo Tripicchio, and Carlo Alberto Avizzano. i2c-net: using instance-level neural networks for monocular category-level 6d pose estimation. IEEE Robotics and Automation Letters, 8(3):1515–1522, 2023
2023
-
[36]
Normalized object coordinate space for category-level 6d object pose and size estimation
He Wang, Srinath Sridhar, Jingwei Huang, Julien Valentin, Shuran Song, and Leonidas J Guibas. Normalized object coordinate space for category-level 6d object pose and size estimation. In 2019 Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pa...
2019
-
[37]
Templates for 3d object pose estimation revisited: Generalization to new objects and robustness to occlusions
Van Nguyen Nguyen, Yinlin Hu, Yang Xiao, Mathieu Salzmann, and Vincent Lepetit. Templates for 3d object pose estimation revisited: Generalization to new objects and robustness to occlusions. In 2022 Proceedings of the IEEE/CVF conference on computer vision and pattern recognit...
2022
-
[38]
Self- supervised vision transformers for 3d pose estimation of novel objects
Stefan Thalhammer, Jean-Baptiste Weibel, Markus Vincze, and Jose Garcia-Rodriguez. Self- supervised vision transformers for 3d pose estimation of novel objects. Image and Vision Comput- ing, 139:104816, 2023
2023
-
[39]
Zero-shot category-level object pose estimation
Walter Goodwin, Sagar Vaze, Ioannis Havoutis, and Ingmar Posner. Zero-shot category-level object pose estimation. In 2022 Proceedings of the European Conference on Computer Vision (ECCV), 2022
2022
-
[40]
Gigapose: Fast and robust novel object pose estimation via one correspondence
Van Nguyen Nguyen, Thibault Groueix, Mathieu Salzmann, and Vincent Lepetit. Gigapose: Fast and robust novel object pose estimation via one correspondence. In 2024 Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 9903–9913, 2024
2024
-
[41]
Sam-6d: Segment anything model meets zero-shot 6d object pose estimation
Jiehong Lin, Lihua Liu, Dekun Lu, and Kui Jia. Sam-6d: Segment anything model meets zero-shot 6d object pose estimation. In 2024 Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 27906–27916, June 2024
2024
-
[42]
6d-diff: A keypoint diffusion framework for 6d object pose estimation
Li Xu, Haoxuan Qu, Yujun Cai, and Jun Liu. 6d-diff: A keypoint diffusion framework for 6d object pose estimation. In 2024 Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 9676–9686, June 2024
2024
-
[43]
Secondpose: Se (3)-consistent dual-stream feature fusion for category-level pose estimation
Yamei Chen, Yan Di, Guangyao Zhai, Fabian Manhardt, Chenyangguang Zhang, Ruida Zhang, Federico Tombari, Nassir Navab, and Benjamin Busam. Secondpose: Se (3)-consistent dual-stream feature fusion for category-level pose estimation. In 2024 Proceedings of the IEEE/CVF Conference...
2024
-
[44]
Segment anything
Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C Berg, Wan-Yen Lo, et al. Segment anything. In 2023 Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 4015–4026, 2023
2023
-
[45]
Dinov2: Learning robust visual features without supervision
Maxime Oquab, Timoth´ ee Darcet, Th´ eo Moutakanni, Huy Vo, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, et al. Dinov2: Learning robust visual features without supervision. Transactions on Machine Learning Research Jour...
2024
-
[46]
A general protocol to probe large vision models for 3d physical understanding
Guanqi Zhan, Chuanxia Zheng, Weidi Xie, and Andrew Zisserman. A general protocol to probe large vision models for 3d physical understanding. In The Thirty-eighth Annual Conference on Neural Information Processing Systems (2023), 2023. 15
2023
-
[47]
Gomez, Lukasz Kaiser, and Illia Polosukhin
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. Advances in Neural Information Processing Systems, 30:5998–6008, 2017
2017
-
[48]
Glipv2: Unifying localization and vision- language understanding
Haotian Zhang, Pengchuan Zhang, Xiaowei Hu, Yen-Chun Chen, Liunian Li, Xiyang Dai, Lijuan Wang, Lu Yuan, Jenq-Neng Hwang, and Jianfeng Gao. Glipv2: Unifying localization and vision- language understanding. Advances in Neural Information Processing Systems, 35:36067–36080, 2022
2022
-
[49]
Open x-embodiment: Robotic learning datasets and rt-x models
Quan Vuong, Sergey Levine, Homer Rich Walke, Karl Pertsch, Anikait Singh, Ria Doshi, Charles Xu, Jianlan Luo, Liam Tan, Dhruv Shah, et al. Open x-embodiment: Robotic learning datasets and rt-x models. In 2nd Workshop on Language and Robot Learning: Language as Grounding (2023), 2023
2023
-
[50]
Denoising diffusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. Advances in neural information processing systems, 33:6840–6851, 2020
2020
-
[51]
Repaint: Inpainting using denoising diffusion probabilistic models
Andreas Lugmayr, Martin Danelljan, Andres Romero, Fisher Yu, Radu Timofte, and Luc Van Gool. Repaint: Inpainting using denoising diffusion probabilistic models. In 2022 Proceed- ings of the IEEE/CVF conference on computer vision and pattern recognition, pages 11461–11471, 2022
2022
-
[52]
Implicit diffusion models for continuous super-resolution
Sicheng Gao, Xuhui Liu, Bohan Zeng, Sheng Xu, Yanjing Li, Xiaoyan Luo, Jianzhuang Liu, Xiantong Zhen, and Baochang Zhang. Implicit diffusion models for continuous super-resolution. In 2023 Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages...
2023
-
[53]
Inversion-based style transfer with diffusion models
Yuxin Zhang, Nisha Huang, Fan Tang, Haibin Huang, Chongyang Ma, Weiming Dong, and Changsheng Xu. Inversion-based style transfer with diffusion models. In 2023 Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10146–10156, 2023
2023
-
[54]
Diffusion hyperfeatures: Searching through time and space for semantic correspondence
Grace Luo, Lisa Dunlap, Dong Huk Park, Aleksander Holynski, and Trevor Darrell. Diffusion hyperfeatures: Searching through time and space for semantic correspondence. In Advances in Neural Information Processing Systems 36 (2023), 2023
2023
-
[55]
Click to grasp: Zero-shot precise manipulation via visual diffusion descriptors, 2024
Nikolaos Tsagkas, Jack Rome, Subramanian Ramamoorthy, Oisin Mac Aodha, and Chris Xiaoxuan Lu. Click to grasp: Zero-shot precise manipulation via visual diffusion descriptors, 2024
2024
-
[56]
U-net: Convolutional networks for biomed- ical image segmentation
Olaf Ronneberger, Philipp Fischer, and Thomas Brox. U-net: Convolutional networks for biomed- ical image segmentation. In 2015 Proceedings of the International Conference on Medical Image Computing and Computer-Assisted Intervention (MICCAI), pages 234–241, 2015
2015
-
[57]
Multiple View Geometry in Computer Vision
Richard Hartley and Andrew Zisserman. Multiple View Geometry in Computer Vision. Cambridge University Press, 2nd edition, 2003
2003
-
[58]
Cnos: A strong baseline for cad-based novel object segmentation
Van Nguyen Nguyen, Thibault Groueix, Georgy Ponimatkin, Vincent Lepetit, and Tomas Hodan. Cnos: A strong baseline for cad-based novel object segmentation. In 2023 Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 2134–2140, 2023
2023
-
[59]
Pix2pose: Pixel-wise coordinate regression of objects for 6d pose estimation
Kiru Park, Timothy Patten, and Markus Vincze. Pix2pose: Pixel-wise coordinate regression of objects for 6d pose estimation. In 2019 Proceedings of the IEEE/CVF international conference on computer vision, pages 7668–7677, 2019
2019
-
[60]
Laion-5b: An open large-scale dataset for training next generation image-text models
Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Wortsman, et al. Laion-5b: An open large-scale dataset for training next generation image-text models. Advances in Neural In- for...
2022
-
[61]
On the importance of noise scheduling in diffusion models
Arash Vahdat, Karsten Kreis, and Jan Kautz. On the importance of noise scheduling in diffusion models. arXiv preprint arXiv:2112.10759, 2021. 16
2021 arXiv
-
[62]
Random sample consensus: A paradigm for model fitting with applications to image analysis and automated cartography
Martin A Fischler and Robert C Bolles. Random sample consensus: A paradigm for model fitting with applications to image analysis and automated cartography. Communications of the ACM, 24(6):381–395, 1981
1981
-
[63]
Telling left from right: Identifying geometry-aware semantic correspondence
Junyi Zhang, Charles Herrmann, Junhwa Hur, Eric Chen, Varun Jampani, Deqing Sun, and Ming-Hsuan Yang. Telling left from right: Identifying geometry-aware semantic correspondence. In 2024 Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3...
2024
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.