REVIEW 3 major objections 2 minor 2 cited by
A diffusion-based framework reconstructs 3D hand-held object geometry from a single monocular image by using hand-object interaction as geometric guidance.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
A diffusion model guided by hand-object interaction and geometric cues reconstructs 3D hand-held object geometry from monocular RGB images.
T0 review reviewed 2026-08-05 challenge →
load-bearing objection Plausible diffusion-plus-optimization idea for hand-held object reconstruction, but the abstract alone can't support the empirical claims—read the full paper before trusting it. the 3 major comments →
Follow My Hold: Hand-Object Interaction Reconstruction through Geometric Guidance
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
Core claim
The central claim is that hand-object interaction itself can serve as a strong geometric prior for 3D reconstruction. The paper proposes a latent diffusion model conditioned on an inpainted object appearance, and at inference time it supervises the velocity field of the diffusion process while jointly optimizing the transformations of both the hand and the object. This optimization is driven by multi-modal geometric cues—normal and depth alignment, silhouette consistency, and 2D keypoint reprojection—plus signed distance field supervision and explicit contact and non-intersection constraints. The result, the paper states, is accurate, robust, and coherent reconstruction under occlusion, with
What carries the argument
The central mechanism is an optimization-in-the-loop guidance applied to a latent diffusion model. During the diffusion generation process, the hand and object poses are optimized together by applying supervision directly to the velocity field, using geometric cues (normals, depth, silhouette, keypoints) and physical constraints (contact, non-intersection, signed distance field). This joint optimization is what guides the diffusion model to output geometrically plausible and interaction-aware object shapes.
Load-bearing premise
The method assumes that a single photo of a hand-held object contains enough reliable geometric signal—normals, depth, silhouettes, keypoints—and that the optimization-in-the-loop converges to the true object shape rather than a local optimum, especially under severe occlusion.
What would settle it
On a benchmark with known ground-truth shapes and heavy hand occlusion, if reconstructions that match the visible pixels still differ substantially from the ground-truth shape, or if the optimized hand-object poses violate contact/non-intersection constraints, the central claim of accurate and physically coherent reconstruction would fail.
If this is right
- Monocular RGB images become sufficient for reconstructing hand-held object geometry without multi-view or depth sensors.
- Occlusion robustness improves because the hand-object interaction and geometric cues constrain the shape even where the object is hidden.
- The method could generalize to in-the-wild images, making 3D reconstruction useful for AR/VR, robotics, and human-object interaction understanding.
- Physical plausibility is enforced during generation, reducing the common need for post-hoc contact refinement or intersection removal.
Where Pith is reading between the lines
- The optimization-in-the-loop velocity guidance could be adapted to other interaction reconstruction problems, such as two-handed objects or tool use, where spatial constraints between actors and objects carry geometric information.
- One testable extension is to evaluate how sensitive the method is to each geometric cue; ablating normals or keypoints would show which cue carries the reconstruction under occlusion.
- If the method is reliable, it might enable online, interactive reconstruction from a single video frame, which could feed real-time AR applications where heavy post-processing is infeasible.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript proposes a diffusion-based framework for reconstructing 3D geometry of hand-held objects from a single monocular RGB image. The method conditions a latent diffusion model on an inpainted object appearance and applies inference-time guidance with an optimization-in-the-loop design that supervises the velocity field while optimizing hand and object transformations. The guidance uses multi-modal geometric cues (normal and depth alignment, silhouette consistency, 2D keypoint reprojection) plus signed distance field supervision and contact/non-intersection constraints. The abstract claims the method yields accurate, robust, and coherent reconstructions under occlusion and generalizes to in-the-wild scenarios.
Significance. If the claims are supported, the work could advance single-image hand-object reconstruction by replacing heavy post-processing with direct generation during the diffusion process, leveraging hand-object interaction as geometric guidance. The explicit use of optimization-in-the-loop and multi-modal cue alignment is a plausible direction. However, the manuscript as submitted contains only the abstract and provides no technical derivation, implementation details, or quantitative validation. Consequently, the significance cannot be assessed beyond the plausibility of the stated idea.
major comments (3)
- [Abstract] The central empirical claim—'accurate, robust and coherent reconstructions under occlusion while generalizing well to in-the-wild scenarios'—is not supported by any experiments. There are no datasets, metrics, baselines, ablations, or failure analyses. Given that monocular reconstruction is ill-posed and the method relies on predicted cues that are known to be biased in occluded hand-object regions, quantitative evaluation is essential to substantiate this claim.
- [Abstract] The method description is only a list of design components. No equations are given for the velocity-field supervision, the optimization-in-the-loop, the signed distance field supervision, or the contact/non-intersection constraints. Without formal definitions or a reference to derivations, the correctness of the approach cannot be checked, and the 'optimization-in-the-loop' mechanism remains underspecified.
- [Abstract] The robustness-under-occlusion claim relies on the reliability of inpainted appearance and predicted geometric cues in occluded regions. The abstract provides no evidence—such as oracle ablations or analysis of failure cases—that the guidance procedure corrects biased cues rather than reinforcing them. This is a load-bearing point because errors in the predicted cues can be self-reinforcing through the diffusion guidance.
minor comments (2)
- [Abstract] The title 'Follow My Hold' is evocative but does not convey the paper's content; a more descriptive title would aid readability.
- [Abstract] The abstract contains no references to prior work, making it difficult to position the contribution relative to existing hand-object reconstruction methods.
Circularity Check
No circularity identified; abstract-only text provides no derivation or fitted-vs-predicted conflation.
full rationale
The manuscript excerpt consists solely of the abstract and does not contain equations, experimental sections, or citations. The central claim—that a latent diffusion model with inference-time geometric guidance reconstructs hand-held object geometry—does not, on its face, define any output in terms of its own input. The 'optimization-in-the-loop' over hand and object transformations is described as a method component, not as a quantity fitted to data and then re-reported as a prediction. No self-citation, uniqueness theorem, or ansatz-smuggling citation appears in the provided text. The skeptic's concern that predicted geometric cues and inpainted appearance might be biased under occlusion, causing self-reinforcing errors, is a plausible robustness/correctness risk but is not a circularity: it does not reduce the reconstruction result to an input by construction, and the abstract offers no specific equation or reduction that could be exhibited. Per the hard rules, circularity must be demonstrated with exact quotes and a specific reduction; none is available. Hence a score of 0 is appropriate.
Axiom & Free-Parameter Ledger
axioms (3)
- domain assumption Monocular RGB images contain sufficient geometric information (normals, depth, silhouettes, keypoints) to guide 3D reconstruction.
- domain assumption Hand-object contact and non-intersection constraints can be reliably enforced from estimated hand parameters without ground-truth object geometry.
- ad hoc to paper Inference-time guidance with velocity supervision converges to high-quality object geometry.
Cite this review
Pith. "Pith review of Follow My Hold: Hand-Object Interaction Reconstruction through Geometric Guidance." pith.science (2026). https://pith.science/paper/LQGJG2CM
@misc{pith2026250818213,
author = {Pith},
title = {Pith review of: Follow My Hold: Hand-Object Interaction Reconstruction through Geometric Guidance},
year = {2026},
howpublished = {\url{https://pith.science/paper/LQGJG2CM}},
note = {Machine review of arXiv:2508.18213}
}
read the original abstract
We propose a novel diffusion-based framework for reconstructing 3D geometry of hand-held objects from monocular RGB images by leveraging hand-object interaction as geometric guidance. Our method conditions a latent diffusion model on an inpainted object appearance and uses inference-time guidance to optimize the object reconstruction, while simultaneously ensuring plausible hand-object interactions. Unlike prior methods that rely on extensive post-processing or produce low-quality reconstructions, our approach directly generates high-quality object geometry during the diffusion process by introducing guidance with an optimization-in-the-loop design. Specifically, we guide the diffusion model by applying supervision to the velocity field while simultaneously optimizing the transformations of both the hand and the object being reconstructed. This optimization is driven by multi-modal geometric cues, including normal and depth alignment, silhouette consistency, and 2D keypoint reprojection. We further incorporate signed distance field supervision and enforce contact and non-intersection constraints to ensure physical plausibility of hand-object interaction. Our method yields accurate, robust and coherent reconstructions under occlusion while generalizing well to in-the-wild scenarios.
Forward citations
Cited by 2 Pith papers
-
VHOI: Controllable Video Generation of Human-Object Interactions from Sparse Trajectories via Motion Densification
VHOI densifies sparse trajectories into color-encoded HOI mask sequences and conditions a fine-tuned video diffusion model on them to produce controllable human-object interaction videos, including full navigation sequences.
-
Reconstructing Objects along Hand Interaction Timelines in Egocentric Video
A constrained optimisation-and-propagation method that enforces stable hand contact while an object is held improves 3D pose reconstruction of rigid objects across egocentric hand-interaction timelines.
Reference graph
Works this paper leans on
-
[1]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION format.date year duplicate empty "emp...
-
[2]
Pushing the envelope for rgb-based dense 3d hand pose estimation via neural rendering
Seungryul Baek, Kwang In Kim, and Tae-Kyun Kim. Pushing the envelope for rgb-based dense 3d hand pose estimation via neural rendering. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 1067--1076, 2019
work page 2019
-
[3]
3d hand shape and pose from images in the wild
Adnane Boukhayma, Rodrigo de Bem, and Philip HS Torr. 3d hand shape and pose from images in the wild. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10843--10852, 2019
work page 2019
-
[4]
Contactpose: A dataset of grasps with object contact and hand pose
Samarth Brahmbhatt, Chengcheng Tang, Christopher D Twigg, Charles C Kemp, and James Hays. Contactpose: A dataset of grasps with object contact and hand pose. In Computer Vision--ECCV 2020: 16th European Conference, Glasgow, UK, August 23--28, 2020, Proceedings, Part XIII 16, pages 361--378. Springer, 2020
work page 2020
-
[5]
Reconstructing hand-object interactions in the wild
Zhe Cao, Ilija Radosavovic, Angjoo Kanazawa, and Jitendra Malik. Reconstructing hand-object interactions in the wild. In Proceedings of the IEEE/CVF international conference on computer vision, pages 12417--12426, 2021
work page 2021
-
[6]
Dexycb: A benchmark for capturing hand grasping of objects
Yu-Wei Chao, Wei Yang, Yu Xiang, Pavlo Molchanov, Ankur Handa, Jonathan Tremblay, Yashraj S Narang, Karl Van Wyk, Umar Iqbal, Stan Birchfield, et al. Dexycb: A benchmark for capturing hand grasping of objects. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 9044--9053, 2021
work page 2021
-
[7]
Tracking and Reconstructing Hand Object Interactions from Point Cloud Sequences in the Wild
Jiayi Chen, Mi Yan, Jiazhao Zhang, Yinzhen Xu, Xiaolong Li, Yijia Weng, Li Yi, Shuran Song, and He Wang. Tracking and reconstructing hand object interactions from point cloud sequences in the wild. arXiv preprint arXiv:2209.12009, 2022 a
work page internal anchor Pith review Pith/arXiv arXiv 2022
-
[8]
Joint hand-object 3d reconstruction from a single image with cross-branch feature fusion
Yujin Chen, Zhigang Tu, Di Kang, Ruizhi Chen, Linchao Bao, Zhengyou Zhang, and Junsong Yuan. Joint hand-object 3d reconstruction from a single image with cross-branch feature fusion. IEEE Transactions on Image Processing, 30: 0 4008--4021, 2021
work page 2021
-
[9]
Alignsdf: Pose-aligned signed distance fields for hand-object reconstruction
Zerui Chen, Yana Hasson, Cordelia Schmid, and Ivan Laptev. Alignsdf: Pose-aligned signed distance fields for hand-object reconstruction. In Computer Vision -- ECCV 2022, pages 231--248, Cham, 2022 b . Springer Nature Switzerland
work page 2022
-
[10]
Alignsdf: Pose-aligned signed distance fields for hand-object reconstruction
Zerui Chen, Yana Hasson, Cordelia Schmid, and Ivan Laptev. Alignsdf: Pose-aligned signed distance fields for hand-object reconstruction. In European conference on computer vision, pages 231--248. Springer, 2022 c
work page 2022
-
[11]
gsdf: Geometry-driven signed distance functions for 3d hand-object reconstruction
Zerui Chen, Shizhe Chen, Cordelia Schmid, and Ivan Laptev. gsdf: Geometry-driven signed distance functions for 3d hand-object reconstruction. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 12890--12900, 2023
work page 2023
-
[12]
HORT : Monocular hand-held objects reconstruction with transformers
Zerui Chen, Rolandos Alexandros Potamias, Shizhe Chen, and Cordelia Schmid. HORT : Monocular hand-held objects reconstruction with transformers. In ICCV, 2025
work page 2025
-
[13]
Ganhand: Predicting human grasp affordances in multi-object scenes
Enric Corona, Albert Pumarola, Guillem Alenya, Francesc Moreno-Noguer, and Gr \'e gory Rogez. Ganhand: Predicting human grasp affordances in multi-object scenes. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 5031--5041, 2020
work page 2020
-
[14]
Generalizable 3d scene reconstruction via divide and conquer from a single view
Andreea Dogaru, Mert \"O zer, and Bernhard Egger. Generalizable 3d scene reconstruction via divide and conquer from a single view. arXiv preprint arXiv:2404.03421, 2024
Pith/arXiv arXiv 2024
-
[15]
Hamba: Single-view 3d hand reconstruction with graph-guided bi-scanning mamba
Haoye Dong, Aviral Chharia, Wenbo Gou, Francisco Vicente Carrasco, and Fernando D De la Torre. Hamba: Single-view 3d hand reconstruction with graph-guided bi-scanning mamba. Advances in Neural Information Processing Systems, 37: 0 2127--2160, 2024
work page 2024
-
[16]
H. Edelsbrunner, D. Kirkpatrick, and R. Seidel. On the shape of a set of points in the plane. IEEE Transactions on Information Theory, 29 0 (4): 0 551--559, 1983
work page 1983
-
[17]
Scaling rectified flow transformers for high-resolution image synthesis
Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas M \"u ller, Harry Saini, Yam Levi, Dominik Lorenz, Axel Sauer, Frederic Boesel, et al. Scaling rectified flow transformers for high-resolution image synthesis. In Forty-first international conference on machine learning, 2024
2024
-
[18]
Arctic: A dataset for dexterous bimanual hand-object manipulation
Zicong Fan, Omid Taheri, Dimitrios Tzionas, Muhammed Kocabas, Manuel Kaufmann, Michael J Black, and Otmar Hilliges. Arctic: A dataset for dexterous bimanual hand-object manipulation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12943--12954, 2023
2023
-
[19]
Hold: Category-agnostic 3d reconstruction of interacting hands and objects from video
Zicong Fan, Maria Parelli, Maria Eleni Kadoglou, Xu Chen, Muhammed Kocabas, Michael J Black, and Otmar Hilliges. Hold: Category-agnostic 3d reconstruction of interacting hands and objects from video. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 494--504, 2024
work page 2024
-
[20]
First-person hand action benchmark with rgb-d videos and 3d hand pose annotations
Guillermo Garcia-Hernando, Shanxin Yuan, Seungryul Baek, and Tae-Kyun Kim. First-person hand action benchmark with rgb-d videos and 3d hand pose annotations. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 409--419, 2018
work page 2018
-
[21]
Detecting and recognizing human-object interactions
Georgia Gkioxari, Ross Girshick, Piotr Doll \'a r, and Kaiming He. Detecting and recognizing human-object interactions. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 8359--8367, 2018
work page 2018
-
[22]
Contactopt: Optimizing contact to improve grasps
Patrick Grady, Chengcheng Tang, Christopher D Twigg, Minh Vo, Samarth Brahmbhatt, and Charles C Kemp. Contactopt: Optimizing contact to improve grasps. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1471--1481, 2021
work page 2021
-
[23]
An object-dependent hand pose prior from sparse training data
Henning Hamer, Juergen Gall, Thibaut Weise, and Luc Van Gool. An object-dependent hand pose prior from sparse training data. In 2010 IEEE Computer Society Conference on Computer Vision and Pattern Recognition, pages 671--678. IEEE, 2010
work page 2010
-
[24]
Honnotate: A method for 3d annotation of hand and object poses
Shreyas Hampali, Mahdi Rad, Markus Oberweger, and Vincent Lepetit. Honnotate: A method for 3d annotation of hand and object poses. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 3196--3206, 2020
2020
-
[25]
Learning joint reconstruction of hands and manipulated objects
Yana Hasson, Gul Varol, Dimitrios Tzionas, Igor Kalevatykh, Michael J Black, Ivan Laptev, and Cordelia Schmid. Learning joint reconstruction of hands and manipulated objects. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 11807--11816, 2019
work page 2019
-
[26]
Towards unconstrained joint hand-object reconstruction from rgb videos
Yana Hasson, G \"u l Varol, Cordelia Schmid, and Ivan Laptev. Towards unconstrained joint hand-object reconstruction from rgb videos. In 2021 International Conference on 3D Vision (3DV), pages 659--668. IEEE, 2021
work page 2021
-
[27]
Lrm: Large reconstruction model for single image to 3d
Yicong Hong, Kai Zhang, Jiuxiang Gu, Sai Bi, Yang Zhou, Difan Liu, Feng Liu, Kalyan Sunkavalli, Trung Bui, and Hao Tan. Lrm: Large reconstruction model for single image to 3d. arXiv preprint arXiv:2311.04400, 2023
Pith/arXiv arXiv 2023
-
[28]
Grasping field: Learning implicit representations for human grasps
Korrawe Karunratanakul, Jinlong Yang, Yan Zhang, Michael J Black, Krikamol Muandet, and Siyu Tang. Grasping field: Learning implicit representations for human grasps. In 2020 International Conference on 3D Vision (3DV), pages 333--344. IEEE, 2020
work page 2020
-
[29]
Flux.1 kontext: Flow matching for in-context image generation and editing in latent space, 2025
Black Forest Labs, Stephen Batifol, Andreas Blattmann, Frederic Boesel, Saksham Consul, Cyril Diagne, Tim Dockhorn, Jack English, Zion English, Patrick Esser, Sumith Kulal, Kyle Lacey, Yam Levi, Cheng Li, Dominik Lorenz, Jonas Müller, Dustin Podell, Robin Rombach, Harry Saini, Axel Sauer, and Luke Smith. Flux.1 kontext: Flow matching for in-context image ...
2025
-
[30]
Self-supervised single-view 3d reconstruction via semantic consistency
Xueting Li, Sifei Liu, Kihwan Kim, Shalini De Mello, Varun Jampani, Ming-Hsuan Yang, and Jan Kautz. Self-supervised single-view 3d reconstruction via semantic consistency. In ECCV, 2020
2020
-
[31]
Yaron Lipman, Marton Havasi, Peter Holderrieth, Neta Shaul, Matt Le, Brian Karrer, Ricky TQ Chen, David Lopez-Paz, Heli Ben-Hamu, and Itai Gat. Flow matching guide and code. arXiv preprint arXiv:2412.06264, 2024
Pith/arXiv arXiv 2024
-
[32]
Semi-supervised 3d hand-object poses estimation with interactions in time
Shaowei Liu, Hanwen Jiang, Jiarui Xu, Sifei Liu, and Xiaolong Wang. Semi-supervised 3d hand-object poses estimation with interactions in time. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 14687--14697, 2021
work page 2021
-
[33]
Yumeng Liu, Xiaoxiao Long, Zemin Yang, Yuan Liu, Marc Habermann, Christian Theobalt, Yuexin Ma, and Wenping Wang. Easyhoi: Unleashing the power of large models for reconstructing hand-object interactions in the wild. arXiv preprint arXiv:2411.14280, 2024
arXiv 2024
-
[34]
Luca Medeiros. Language segment anything. https://github.com/luca-medeiros/lang-segment-anything, 2023. Accessed: 2025-05-23
work page 2023
-
[35]
Giraffe: Representing scenes as compositional generative neural feature fields
Michael Niemeyer and Andreas Geiger. Giraffe: Representing scenes as compositional generative neural feature fields. In Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), 2021
work page 2021
-
[36]
Dinov2: Learning robust visual features without supervision
Maxime Oquab, Timoth \'e e Darcet, Th \'e o Moutakanni, Huy Vo, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, et al. Dinov2: Learning robust visual features without supervision. arXiv preprint arXiv:2304.07193, 2023
Pith/arXiv arXiv 2023
-
[37]
Learning to imitate object interactions from internet videos, 2022
Austin Patel, Andrew Wang, Ilija Radosavovic, and Jitendra Malik. Learning to imitate object interactions from internet videos, 2022
work page 2022
-
[38]
Reconstructing hands in 3d with transformers
Georgios Pavlakos, Dandan Shan, Ilija Radosavovic, Angjoo Kanazawa, David Fouhey, and Jitendra Malik. Reconstructing hands in 3d with transformers. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 9826--9836, 2024
work page 2024
-
[39]
Wilor: End-to-end 3d hand localization and reconstruction in-the-wild, 2024
Rolandos Alexandros Potamias, Jinglei Zhang, Jiankang Deng, and Stefanos Zafeiriou. Wilor: End-to-end 3d hand localization and reconstruction in-the-wild, 2024
work page 2024
-
[40]
Novel-view synthesis and pose estimation for hand-object interaction from sparse views
Wentian Qu, Zhaopeng Cui, Yinda Zhang, Chenyu Meng, Cuixia Ma, Xiaoming Deng, and Hongan Wang. Novel-view synthesis and pose estimation for hand-object interaction from sparse views. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 15100--15111, 2023
work page 2023
-
[41]
Graf: Generative radiance fields for 3d-aware image synthesis
Katja Schwarz, Yiyi Liao, Michael Niemeyer, and Andreas Geiger. Graf: Generative radiance fields for 3d-aware image synthesis. In Advances in Neural Information Processing Systems, pages 20154--20166. Curran Associates, Inc., 2020
work page 2020
-
[42]
Flexible isosurface extraction for gradient-based mesh optimization
Tianchang Shen, Jacob Munkberg, Jon Hasselgren, Kangxue Yin, Zian Wang, Wenzheng Chen, Zan Gojcic, Sanja Fidler, Nicholas Sharp, and Jun Gao. Flexible isosurface extraction for gradient-based mesh optimization. ACM Transactions on Graphics (TOG), 42 0 (4): 0 1--16, 2023
work page 2023
-
[43]
Lgm: Large multi-view gaussian model for high-resolution 3d content creation
Jiaxiang Tang, Zhaoxi Chen, Xiaokang Chen, Tengfei Wang, Gang Zeng, and Ziwei Liu. Lgm: Large multi-view gaussian model for high-resolution 3d content creation. In European Conference on Computer Vision, pages 1--18. Springer, 2024
work page 2024
-
[44]
Maxim Tatarchenko, Stephan R Richter, Ren \'e Ranftl, Zhuwen Li, Vladlen Koltun, and Thomas Brox. What do single-view 3d reconstruction networks learn? In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 3405--3414, 2019
work page 2019
-
[45]
Gemini: A family of highly capable multimodal models, 2024
Gemini Team. Gemini: A family of highly capable multimodal models, 2024
2024
-
[46]
H+ o: Unified egocentric recognition of 3d hand-object poses and interactions
Bugra Tekin, Federica Bogo, and Marc Pollefeys. H+ o: Unified egocentric recognition of 3d hand-object poses and interactions. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 4511--4520, 2019
work page 2019
-
[47]
Collaborative learning for hand and object reconstruction with attention-guided graph convolution
Tze Ho Elden Tse, Kwang In Kim, Ales Leonardis, and Hyung Jin Chang. Collaborative learning for hand and object reconstruction with attention-guided graph convolution. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1664--1674, 2022
work page 2022
-
[48]
Pf-lrm: Pose-free large reconstruction model for joint pose and shape prediction
Peng Wang, Hao Tan, Sai Bi, Yinghao Xu, Fujun Luan, Kalyan Sunkavalli, Wenping Wang, Zexiang Xu, and Kai Zhang. Pf-lrm: Pose-free large reconstruction model for joint pose and shape prediction. arXiv preprint arXiv:2311.12024, 2023
Pith/arXiv arXiv 2023
-
[49]
Ruicheng Wang, Sicheng Xu, Cassie Dai, Jianfeng Xiang, Yu Deng, Xin Tong, and Jiaolong Yang. Moge: Unlocking accurate monocular geometry estimation for open-domain images with optimal training supervision. arXiv preprint arXiv:2410.19115, 2024 a
Pith/arXiv arXiv 2024
-
[50]
Crm: Single image to 3d textured mesh with convolutional reconstruction model
Zhengyi Wang, Yikai Wang, Yifei Chen, Chendong Xiang, Shuo Chen, Dajiang Yu, Chongxuan Li, Hang Su, and Jun Zhu. Crm: Single image to 3d textured mesh with convolutional reconstruction model. In European Conference on Computer Vision, pages 57--74. Springer, 2024 b
work page 2024
-
[51]
Structured 3d latents for scalable and versatile 3d generation
Jianfeng Xiang, Zelong Lv, Sicheng Xu, Yu Deng, Ruicheng Wang, Bowen Zhang, Dong Chen, Xin Tong, and Jiaolong Yang. Structured 3d latents for scalable and versatile 3d generation. arXiv preprint arXiv:2412.01506, 2024
Pith/arXiv arXiv 2024
-
[52]
Jiale Xu, Weihao Cheng, Yiming Gao, Xintao Wang, Shenghua Gao, and Ying Shan. Instantmesh: Efficient 3d mesh generation from a single image with sparse-view large reconstruction models. arXiv preprint arXiv:2404.07191, 2024
Pith/arXiv arXiv 2024
-
[53]
Oakink: A large-scale knowledge repository for understanding hand-object interaction
Lixin Yang, Kailin Li, Xinyu Zhan, Fei Wu, Anran Xu, Liu Liu, and Cewu Lu. Oakink: A large-scale knowledge repository for understanding hand-object interaction. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 20953--20962, 2022
work page 2022
-
[54]
What's in your hands? 3d reconstruction of generic objects in hands
Yufei Ye, Abhinav Gupta, and Shubham Tulsiani. What's in your hands? 3d reconstruction of generic objects in hands. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 3895--3905, 2022
work page 2022
-
[55]
Perceiving 3d human-object spatial arrangements from a single image in the wild
Jason Y Zhang, Sam Pepose, Hanbyul Joo, Deva Ramanan, Jitendra Malik, and Angjoo Kanazawa. Perceiving 3d human-object spatial arrangements from a single image in the wild. In Computer Vision--ECCV 2020: 16th European Conference, Glasgow, UK, August 23--28, 2020, Proceedings, Part XII 16, pages 34--51. Springer, 2020
work page 2020
-
[56]
End-to-end hand mesh recovery from a monocular rgb image
Xiong Zhang, Qiang Li, Hong Mo, Wenbo Zhang, and Wen Zheng. End-to-end hand mesh recovery from a monocular rgb image. In Proceedings of the IEEE/CVF international conference on computer vision, pages 2354--2364, 2019
work page 2019
-
[57]
Hunyuan3d 2.0: Scaling diffusion models for high resolution textured 3d assets generation
Zibo Zhao, Zeqiang Lai, Qingxiang Lin, Yunfei Zhao, Haolin Liu, Shuhui Yang, Yifei Feng, Mingxin Yang, Sheng Zhang, Xianghui Yang, et al. Hunyuan3d 2.0: Scaling diffusion models for high resolution textured 3d assets generation. arXiv preprint arXiv:2501.12202, 2025
Pith/arXiv arXiv 2025
-
[58]
Monocular real-time hand shape and motion capture using multi-modal data
Yuxiao Zhou, Marc Habermann, Weipeng Xu, Ikhsanul Habibie, Christian Theobalt, and Feng Xu. Monocular real-time hand shape and motion capture using multi-modal data. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 5346--5355, 2020
work page 2020
This paper was first reviewed by deepseek-v4-flash on August 5, 2026.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.