Pith. sign in

REVIEW 3 major objections 2 minor 2 cited by

A diffusion-based framework reconstructs 3D hand-held object geometry from a single monocular image by using hand-object interaction as geometric guidance.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

A diffusion model guided by hand-object interaction and geometric cues reconstructs 3D hand-held object geometry from monocular RGB images.

T0 review reviewed 2026-08-05 challenge →

load-bearing objection Plausible diffusion-plus-optimization idea for hand-held object reconstruction, but the abstract alone can't support the empirical claims—read the full paper before trusting it. the 3 major comments →

arxiv 2508.18213 v1 pith:LQGJG2CM submitted 2025-08-25 cs.CV

Follow My Hold: Hand-Object Interaction Reconstruction through Geometric Guidance

classification cs.CV
keywords hand-object interactiondiffusion model3D reconstructionmonocular RGBgeometric guidanceinference-time optimizationsigned distance fieldocclusion robustness
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper claims that a single RGB image of a hand holding an object is enough to reconstruct the object's full 3D geometry, even when the object is partially hidden. The proposed method uses a diffusion model that generates the object's shape while simultaneously keeping the hand and object physically consistent. It does this by optimizing both hand and object poses during the generation process, guided by geometric cues like surface normals, depth, silhouettes, and keypoints. If the method works as described, it would remove the need for heavy post-processing and make single-image hand-held object reconstruction practical for real-world images.

Core claim

The central claim is that hand-object interaction itself can serve as a strong geometric prior for 3D reconstruction. The paper proposes a latent diffusion model conditioned on an inpainted object appearance, and at inference time it supervises the velocity field of the diffusion process while jointly optimizing the transformations of both the hand and the object. This optimization is driven by multi-modal geometric cues—normal and depth alignment, silhouette consistency, and 2D keypoint reprojection—plus signed distance field supervision and explicit contact and non-intersection constraints. The result, the paper states, is accurate, robust, and coherent reconstruction under occlusion, with

What carries the argument

The central mechanism is an optimization-in-the-loop guidance applied to a latent diffusion model. During the diffusion generation process, the hand and object poses are optimized together by applying supervision directly to the velocity field, using geometric cues (normals, depth, silhouette, keypoints) and physical constraints (contact, non-intersection, signed distance field). This joint optimization is what guides the diffusion model to output geometrically plausible and interaction-aware object shapes.

Load-bearing premise

The method assumes that a single photo of a hand-held object contains enough reliable geometric signal—normals, depth, silhouettes, keypoints—and that the optimization-in-the-loop converges to the true object shape rather than a local optimum, especially under severe occlusion.

What would settle it

On a benchmark with known ground-truth shapes and heavy hand occlusion, if reconstructions that match the visible pixels still differ substantially from the ground-truth shape, or if the optimized hand-object poses violate contact/non-intersection constraints, the central claim of accurate and physically coherent reconstruction would fail.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • Monocular RGB images become sufficient for reconstructing hand-held object geometry without multi-view or depth sensors.
  • Occlusion robustness improves because the hand-object interaction and geometric cues constrain the shape even where the object is hidden.
  • The method could generalize to in-the-wild images, making 3D reconstruction useful for AR/VR, robotics, and human-object interaction understanding.
  • Physical plausibility is enforced during generation, reducing the common need for post-hoc contact refinement or intersection removal.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The optimization-in-the-loop velocity guidance could be adapted to other interaction reconstruction problems, such as two-handed objects or tool use, where spatial constraints between actors and objects carry geometric information.
  • One testable extension is to evaluate how sensitive the method is to each geometric cue; ablating normals or keypoints would show which cue carries the reconstruction under occlusion.
  • If the method is reliable, it might enable online, interactive reconstruction from a single video frame, which could feed real-time AR applications where heavy post-processing is infeasible.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 2 minor

Summary. The manuscript proposes a diffusion-based framework for reconstructing 3D geometry of hand-held objects from a single monocular RGB image. The method conditions a latent diffusion model on an inpainted object appearance and applies inference-time guidance with an optimization-in-the-loop design that supervises the velocity field while optimizing hand and object transformations. The guidance uses multi-modal geometric cues (normal and depth alignment, silhouette consistency, 2D keypoint reprojection) plus signed distance field supervision and contact/non-intersection constraints. The abstract claims the method yields accurate, robust, and coherent reconstructions under occlusion and generalizes to in-the-wild scenarios.

Significance. If the claims are supported, the work could advance single-image hand-object reconstruction by replacing heavy post-processing with direct generation during the diffusion process, leveraging hand-object interaction as geometric guidance. The explicit use of optimization-in-the-loop and multi-modal cue alignment is a plausible direction. However, the manuscript as submitted contains only the abstract and provides no technical derivation, implementation details, or quantitative validation. Consequently, the significance cannot be assessed beyond the plausibility of the stated idea.

major comments (3)
  1. [Abstract] The central empirical claim—'accurate, robust and coherent reconstructions under occlusion while generalizing well to in-the-wild scenarios'—is not supported by any experiments. There are no datasets, metrics, baselines, ablations, or failure analyses. Given that monocular reconstruction is ill-posed and the method relies on predicted cues that are known to be biased in occluded hand-object regions, quantitative evaluation is essential to substantiate this claim.
  2. [Abstract] The method description is only a list of design components. No equations are given for the velocity-field supervision, the optimization-in-the-loop, the signed distance field supervision, or the contact/non-intersection constraints. Without formal definitions or a reference to derivations, the correctness of the approach cannot be checked, and the 'optimization-in-the-loop' mechanism remains underspecified.
  3. [Abstract] The robustness-under-occlusion claim relies on the reliability of inpainted appearance and predicted geometric cues in occluded regions. The abstract provides no evidence—such as oracle ablations or analysis of failure cases—that the guidance procedure corrects biased cues rather than reinforcing them. This is a load-bearing point because errors in the predicted cues can be self-reinforcing through the diffusion guidance.
minor comments (2)
  1. [Abstract] The title 'Follow My Hold' is evocative but does not convey the paper's content; a more descriptive title would aid readability.
  2. [Abstract] The abstract contains no references to prior work, making it difficult to position the contribution relative to existing hand-object reconstruction methods.

Circularity Check

0 steps flagged

No circularity identified; abstract-only text provides no derivation or fitted-vs-predicted conflation.

full rationale

The manuscript excerpt consists solely of the abstract and does not contain equations, experimental sections, or citations. The central claim—that a latent diffusion model with inference-time geometric guidance reconstructs hand-held object geometry—does not, on its face, define any output in terms of its own input. The 'optimization-in-the-loop' over hand and object transformations is described as a method component, not as a quantity fitted to data and then re-reported as a prediction. No self-citation, uniqueness theorem, or ansatz-smuggling citation appears in the provided text. The skeptic's concern that predicted geometric cues and inpainted appearance might be biased under occlusion, causing self-reinforcing errors, is a plausible robustness/correctness risk but is not a circularity: it does not reduce the reconstruction result to an input by construction, and the abstract offers no specific equation or reduction that could be exhibited. Per the hard rules, circularity must be demonstrated with exact quotes and a specific reduction; none is available. Hence a score of 0 is appropriate.

Axiom & Free-Parameter Ledger

0 free parameters · 3 axioms · 0 invented entities

No explicit free parameters or invented entities appear in the abstract. The paper relies on standard assumptions for monocular reconstruction and on the stability of its proposed optimization loop, but these are not verified in the provided text.

axioms (3)
  • domain assumption Monocular RGB images contain sufficient geometric information (normals, depth, silhouettes, keypoints) to guide 3D reconstruction.
    The method relies on multi-modal cues extracted from a single image; if these cues are unreliable, the optimization-in-the-loop guidance fails.
  • domain assumption Hand-object contact and non-intersection constraints can be reliably enforced from estimated hand parameters without ground-truth object geometry.
    The abstract claims these constraints ensure physical plausibility; this presumes the estimated hand state is accurate enough to define contact.
  • ad hoc to paper Inference-time guidance with velocity supervision converges to high-quality object geometry.
    The abstract proposes an optimization-in-the-loop design but provides no convergence analysis or study of local minima.

reviewed 2026-08-05 · how reviews work

0 comments
Cite this review

Pith. "Pith review of Follow My Hold: Hand-Object Interaction Reconstruction through Geometric Guidance." pith.science (2026). https://pith.science/paper/LQGJG2CM

@misc{pith2026250818213,
  author       = {Pith},
  title        = {Pith review of: Follow My Hold: Hand-Object Interaction Reconstruction through Geometric Guidance},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/LQGJG2CM}},
  note         = {Machine review of arXiv:2508.18213}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

We propose a novel diffusion-based framework for reconstructing 3D geometry of hand-held objects from monocular RGB images by leveraging hand-object interaction as geometric guidance. Our method conditions a latent diffusion model on an inpainted object appearance and uses inference-time guidance to optimize the object reconstruction, while simultaneously ensuring plausible hand-object interactions. Unlike prior methods that rely on extensive post-processing or produce low-quality reconstructions, our approach directly generates high-quality object geometry during the diffusion process by introducing guidance with an optimization-in-the-loop design. Specifically, we guide the diffusion model by applying supervision to the velocity field while simultaneously optimizing the transformations of both the hand and the object being reconstructed. This optimization is driven by multi-modal geometric cues, including normal and depth alignment, silhouette consistency, and 2D keypoint reprojection. We further incorporate signed distance field supervision and enforce contact and non-intersection constraints to ensure physical plausibility of hand-object interaction. Our method yields accurate, robust and coherent reconstructions under occlusion while generalizing well to in-the-wild scenarios.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. VHOI: Controllable Video Generation of Human-Object Interactions from Sparse Trajectories via Motion Densification

    cs.CV 2025-12 unverdicted novelty 6.0

    VHOI densifies sparse trajectories into color-encoded HOI mask sequences and conditions a fine-tuned video diffusion model on them to produce controllable human-object interaction videos, including full navigation sequences.

  2. Reconstructing Objects along Hand Interaction Timelines in Egocentric Video

    cs.CV 2025-12 conditional novelty 6.0

    A constrained optimisation-and-propagation method that enforces stable hand contact while an object is held improves 3D pose reconstruction of rigid objects across egocentric hand-interaction timelines.

Reference graph

Works this paper leans on

58 extracted references · 41 canonical work pages · cited by 2 Pith papers · 1 internal anchor

  1. [1]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION format.date year duplicate empty "emp...

  2. [2]

    Pushing the envelope for rgb-based dense 3d hand pose estimation via neural rendering

    Seungryul Baek, Kwang In Kim, and Tae-Kyun Kim. Pushing the envelope for rgb-based dense 3d hand pose estimation via neural rendering. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 1067--1076, 2019

  3. [3]

    3d hand shape and pose from images in the wild

    Adnane Boukhayma, Rodrigo de Bem, and Philip HS Torr. 3d hand shape and pose from images in the wild. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10843--10852, 2019

  4. [4]

    Contactpose: A dataset of grasps with object contact and hand pose

    Samarth Brahmbhatt, Chengcheng Tang, Christopher D Twigg, Charles C Kemp, and James Hays. Contactpose: A dataset of grasps with object contact and hand pose. In Computer Vision--ECCV 2020: 16th European Conference, Glasgow, UK, August 23--28, 2020, Proceedings, Part XIII 16, pages 361--378. Springer, 2020

  5. [5]

    Reconstructing hand-object interactions in the wild

    Zhe Cao, Ilija Radosavovic, Angjoo Kanazawa, and Jitendra Malik. Reconstructing hand-object interactions in the wild. In Proceedings of the IEEE/CVF international conference on computer vision, pages 12417--12426, 2021

  6. [6]

    Dexycb: A benchmark for capturing hand grasping of objects

    Yu-Wei Chao, Wei Yang, Yu Xiang, Pavlo Molchanov, Ankur Handa, Jonathan Tremblay, Yashraj S Narang, Karl Van Wyk, Umar Iqbal, Stan Birchfield, et al. Dexycb: A benchmark for capturing hand grasping of objects. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 9044--9053, 2021

  7. [7]

    Tracking and Reconstructing Hand Object Interactions from Point Cloud Sequences in the Wild

    Jiayi Chen, Mi Yan, Jiazhao Zhang, Yinzhen Xu, Xiaolong Li, Yijia Weng, Li Yi, Shuran Song, and He Wang. Tracking and reconstructing hand object interactions from point cloud sequences in the wild. arXiv preprint arXiv:2209.12009, 2022 a

  8. [8]

    Joint hand-object 3d reconstruction from a single image with cross-branch feature fusion

    Yujin Chen, Zhigang Tu, Di Kang, Ruizhi Chen, Linchao Bao, Zhengyou Zhang, and Junsong Yuan. Joint hand-object 3d reconstruction from a single image with cross-branch feature fusion. IEEE Transactions on Image Processing, 30: 0 4008--4021, 2021

  9. [9]

    Alignsdf: Pose-aligned signed distance fields for hand-object reconstruction

    Zerui Chen, Yana Hasson, Cordelia Schmid, and Ivan Laptev. Alignsdf: Pose-aligned signed distance fields for hand-object reconstruction. In Computer Vision -- ECCV 2022, pages 231--248, Cham, 2022 b . Springer Nature Switzerland

  10. [10]

    Alignsdf: Pose-aligned signed distance fields for hand-object reconstruction

    Zerui Chen, Yana Hasson, Cordelia Schmid, and Ivan Laptev. Alignsdf: Pose-aligned signed distance fields for hand-object reconstruction. In European conference on computer vision, pages 231--248. Springer, 2022 c

  11. [11]

    gsdf: Geometry-driven signed distance functions for 3d hand-object reconstruction

    Zerui Chen, Shizhe Chen, Cordelia Schmid, and Ivan Laptev. gsdf: Geometry-driven signed distance functions for 3d hand-object reconstruction. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 12890--12900, 2023

  12. [12]

    HORT : Monocular hand-held objects reconstruction with transformers

    Zerui Chen, Rolandos Alexandros Potamias, Shizhe Chen, and Cordelia Schmid. HORT : Monocular hand-held objects reconstruction with transformers. In ICCV, 2025

  13. [13]

    Ganhand: Predicting human grasp affordances in multi-object scenes

    Enric Corona, Albert Pumarola, Guillem Alenya, Francesc Moreno-Noguer, and Gr \'e gory Rogez. Ganhand: Predicting human grasp affordances in multi-object scenes. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 5031--5041, 2020

  14. [14]

    Generalizable 3d scene reconstruction via divide and conquer from a single view

    Andreea Dogaru, Mert \"O zer, and Bernhard Egger. Generalizable 3d scene reconstruction via divide and conquer from a single view. arXiv preprint arXiv:2404.03421, 2024

  15. [15]

    Hamba: Single-view 3d hand reconstruction with graph-guided bi-scanning mamba

    Haoye Dong, Aviral Chharia, Wenbo Gou, Francisco Vicente Carrasco, and Fernando D De la Torre. Hamba: Single-view 3d hand reconstruction with graph-guided bi-scanning mamba. Advances in Neural Information Processing Systems, 37: 0 2127--2160, 2024

  16. [16]

    Edelsbrunner, D

    H. Edelsbrunner, D. Kirkpatrick, and R. Seidel. On the shape of a set of points in the plane. IEEE Transactions on Information Theory, 29 0 (4): 0 551--559, 1983

  17. [17]

    Scaling rectified flow transformers for high-resolution image synthesis

    Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas M \"u ller, Harry Saini, Yam Levi, Dominik Lorenz, Axel Sauer, Frederic Boesel, et al. Scaling rectified flow transformers for high-resolution image synthesis. In Forty-first international conference on machine learning, 2024

  18. [18]

    Arctic: A dataset for dexterous bimanual hand-object manipulation

    Zicong Fan, Omid Taheri, Dimitrios Tzionas, Muhammed Kocabas, Manuel Kaufmann, Michael J Black, and Otmar Hilliges. Arctic: A dataset for dexterous bimanual hand-object manipulation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12943--12954, 2023

  19. [19]

    Hold: Category-agnostic 3d reconstruction of interacting hands and objects from video

    Zicong Fan, Maria Parelli, Maria Eleni Kadoglou, Xu Chen, Muhammed Kocabas, Michael J Black, and Otmar Hilliges. Hold: Category-agnostic 3d reconstruction of interacting hands and objects from video. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 494--504, 2024

  20. [20]

    First-person hand action benchmark with rgb-d videos and 3d hand pose annotations

    Guillermo Garcia-Hernando, Shanxin Yuan, Seungryul Baek, and Tae-Kyun Kim. First-person hand action benchmark with rgb-d videos and 3d hand pose annotations. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 409--419, 2018

  21. [21]

    Detecting and recognizing human-object interactions

    Georgia Gkioxari, Ross Girshick, Piotr Doll \'a r, and Kaiming He. Detecting and recognizing human-object interactions. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 8359--8367, 2018

  22. [22]

    Contactopt: Optimizing contact to improve grasps

    Patrick Grady, Chengcheng Tang, Christopher D Twigg, Minh Vo, Samarth Brahmbhatt, and Charles C Kemp. Contactopt: Optimizing contact to improve grasps. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1471--1481, 2021

  23. [23]

    An object-dependent hand pose prior from sparse training data

    Henning Hamer, Juergen Gall, Thibaut Weise, and Luc Van Gool. An object-dependent hand pose prior from sparse training data. In 2010 IEEE Computer Society Conference on Computer Vision and Pattern Recognition, pages 671--678. IEEE, 2010

  24. [24]

    Honnotate: A method for 3d annotation of hand and object poses

    Shreyas Hampali, Mahdi Rad, Markus Oberweger, and Vincent Lepetit. Honnotate: A method for 3d annotation of hand and object poses. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 3196--3206, 2020

  25. [25]

    Learning joint reconstruction of hands and manipulated objects

    Yana Hasson, Gul Varol, Dimitrios Tzionas, Igor Kalevatykh, Michael J Black, Ivan Laptev, and Cordelia Schmid. Learning joint reconstruction of hands and manipulated objects. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 11807--11816, 2019

  26. [26]

    Towards unconstrained joint hand-object reconstruction from rgb videos

    Yana Hasson, G \"u l Varol, Cordelia Schmid, and Ivan Laptev. Towards unconstrained joint hand-object reconstruction from rgb videos. In 2021 International Conference on 3D Vision (3DV), pages 659--668. IEEE, 2021

  27. [27]

    Lrm: Large reconstruction model for single image to 3d

    Yicong Hong, Kai Zhang, Jiuxiang Gu, Sai Bi, Yang Zhou, Difan Liu, Feng Liu, Kalyan Sunkavalli, Trung Bui, and Hao Tan. Lrm: Large reconstruction model for single image to 3d. arXiv preprint arXiv:2311.04400, 2023

  28. [28]

    Grasping field: Learning implicit representations for human grasps

    Korrawe Karunratanakul, Jinlong Yang, Yan Zhang, Michael J Black, Krikamol Muandet, and Siyu Tang. Grasping field: Learning implicit representations for human grasps. In 2020 International Conference on 3D Vision (3DV), pages 333--344. IEEE, 2020

  29. [29]

    Flux.1 kontext: Flow matching for in-context image generation and editing in latent space, 2025

    Black Forest Labs, Stephen Batifol, Andreas Blattmann, Frederic Boesel, Saksham Consul, Cyril Diagne, Tim Dockhorn, Jack English, Zion English, Patrick Esser, Sumith Kulal, Kyle Lacey, Yam Levi, Cheng Li, Dominik Lorenz, Jonas Müller, Dustin Podell, Robin Rombach, Harry Saini, Axel Sauer, and Luke Smith. Flux.1 kontext: Flow matching for in-context image ...

  30. [30]

    Self-supervised single-view 3d reconstruction via semantic consistency

    Xueting Li, Sifei Liu, Kihwan Kim, Shalini De Mello, Varun Jampani, Ming-Hsuan Yang, and Jan Kautz. Self-supervised single-view 3d reconstruction via semantic consistency. In ECCV, 2020

  31. [31]

    Flow matching guide and code

    Yaron Lipman, Marton Havasi, Peter Holderrieth, Neta Shaul, Matt Le, Brian Karrer, Ricky TQ Chen, David Lopez-Paz, Heli Ben-Hamu, and Itai Gat. Flow matching guide and code. arXiv preprint arXiv:2412.06264, 2024

  32. [32]

    Semi-supervised 3d hand-object poses estimation with interactions in time

    Shaowei Liu, Hanwen Jiang, Jiarui Xu, Sifei Liu, and Xiaolong Wang. Semi-supervised 3d hand-object poses estimation with interactions in time. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 14687--14697, 2021

  33. [33]

    Easyhoi: Unleashing the power of large models for reconstructing hand-object interactions in the wild

    Yumeng Liu, Xiaoxiao Long, Zemin Yang, Yuan Liu, Marc Habermann, Christian Theobalt, Yuexin Ma, and Wenping Wang. Easyhoi: Unleashing the power of large models for reconstructing hand-object interactions in the wild. arXiv preprint arXiv:2411.14280, 2024

  34. [34]

    Language segment anything

    Luca Medeiros. Language segment anything. https://github.com/luca-medeiros/lang-segment-anything, 2023. Accessed: 2025-05-23

  35. [35]

    Giraffe: Representing scenes as compositional generative neural feature fields

    Michael Niemeyer and Andreas Geiger. Giraffe: Representing scenes as compositional generative neural feature fields. In Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), 2021

  36. [36]

    Dinov2: Learning robust visual features without supervision

    Maxime Oquab, Timoth \'e e Darcet, Th \'e o Moutakanni, Huy Vo, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, et al. Dinov2: Learning robust visual features without supervision. arXiv preprint arXiv:2304.07193, 2023

  37. [37]

    Learning to imitate object interactions from internet videos, 2022

    Austin Patel, Andrew Wang, Ilija Radosavovic, and Jitendra Malik. Learning to imitate object interactions from internet videos, 2022

  38. [38]

    Reconstructing hands in 3d with transformers

    Georgios Pavlakos, Dandan Shan, Ilija Radosavovic, Angjoo Kanazawa, David Fouhey, and Jitendra Malik. Reconstructing hands in 3d with transformers. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 9826--9836, 2024

  39. [39]

    Wilor: End-to-end 3d hand localization and reconstruction in-the-wild, 2024

    Rolandos Alexandros Potamias, Jinglei Zhang, Jiankang Deng, and Stefanos Zafeiriou. Wilor: End-to-end 3d hand localization and reconstruction in-the-wild, 2024

  40. [40]

    Novel-view synthesis and pose estimation for hand-object interaction from sparse views

    Wentian Qu, Zhaopeng Cui, Yinda Zhang, Chenyu Meng, Cuixia Ma, Xiaoming Deng, and Hongan Wang. Novel-view synthesis and pose estimation for hand-object interaction from sparse views. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 15100--15111, 2023

  41. [41]

    Graf: Generative radiance fields for 3d-aware image synthesis

    Katja Schwarz, Yiyi Liao, Michael Niemeyer, and Andreas Geiger. Graf: Generative radiance fields for 3d-aware image synthesis. In Advances in Neural Information Processing Systems, pages 20154--20166. Curran Associates, Inc., 2020

  42. [42]

    Flexible isosurface extraction for gradient-based mesh optimization

    Tianchang Shen, Jacob Munkberg, Jon Hasselgren, Kangxue Yin, Zian Wang, Wenzheng Chen, Zan Gojcic, Sanja Fidler, Nicholas Sharp, and Jun Gao. Flexible isosurface extraction for gradient-based mesh optimization. ACM Transactions on Graphics (TOG), 42 0 (4): 0 1--16, 2023

  43. [43]

    Lgm: Large multi-view gaussian model for high-resolution 3d content creation

    Jiaxiang Tang, Zhaoxi Chen, Xiaokang Chen, Tengfei Wang, Gang Zeng, and Ziwei Liu. Lgm: Large multi-view gaussian model for high-resolution 3d content creation. In European Conference on Computer Vision, pages 1--18. Springer, 2024

  44. [44]

    What do single-view 3d reconstruction networks learn? In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 3405--3414, 2019

    Maxim Tatarchenko, Stephan R Richter, Ren \'e Ranftl, Zhuwen Li, Vladlen Koltun, and Thomas Brox. What do single-view 3d reconstruction networks learn? In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 3405--3414, 2019

  45. [45]

    Gemini: A family of highly capable multimodal models, 2024

    Gemini Team. Gemini: A family of highly capable multimodal models, 2024

  46. [46]

    H+ o: Unified egocentric recognition of 3d hand-object poses and interactions

    Bugra Tekin, Federica Bogo, and Marc Pollefeys. H+ o: Unified egocentric recognition of 3d hand-object poses and interactions. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 4511--4520, 2019

  47. [47]

    Collaborative learning for hand and object reconstruction with attention-guided graph convolution

    Tze Ho Elden Tse, Kwang In Kim, Ales Leonardis, and Hyung Jin Chang. Collaborative learning for hand and object reconstruction with attention-guided graph convolution. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1664--1674, 2022

  48. [48]

    Pf-lrm: Pose-free large reconstruction model for joint pose and shape prediction

    Peng Wang, Hao Tan, Sai Bi, Yinghao Xu, Fujun Luan, Kalyan Sunkavalli, Wenping Wang, Zexiang Xu, and Kai Zhang. Pf-lrm: Pose-free large reconstruction model for joint pose and shape prediction. arXiv preprint arXiv:2311.12024, 2023

  49. [49]

    Moge: Unlocking accurate monocular geometry estimation for open-domain images with optimal training supervision

    Ruicheng Wang, Sicheng Xu, Cassie Dai, Jianfeng Xiang, Yu Deng, Xin Tong, and Jiaolong Yang. Moge: Unlocking accurate monocular geometry estimation for open-domain images with optimal training supervision. arXiv preprint arXiv:2410.19115, 2024 a

  50. [50]

    Crm: Single image to 3d textured mesh with convolutional reconstruction model

    Zhengyi Wang, Yikai Wang, Yifei Chen, Chendong Xiang, Shuo Chen, Dajiang Yu, Chongxuan Li, Hang Su, and Jun Zhu. Crm: Single image to 3d textured mesh with convolutional reconstruction model. In European Conference on Computer Vision, pages 57--74. Springer, 2024 b

  51. [51]

    Structured 3d latents for scalable and versatile 3d generation

    Jianfeng Xiang, Zelong Lv, Sicheng Xu, Yu Deng, Ruicheng Wang, Bowen Zhang, Dong Chen, Xin Tong, and Jiaolong Yang. Structured 3d latents for scalable and versatile 3d generation. arXiv preprint arXiv:2412.01506, 2024

  52. [52]

    Instantmesh: Efficient 3d mesh generation from a single image with sparse-view large reconstruction models

    Jiale Xu, Weihao Cheng, Yiming Gao, Xintao Wang, Shenghua Gao, and Ying Shan. Instantmesh: Efficient 3d mesh generation from a single image with sparse-view large reconstruction models. arXiv preprint arXiv:2404.07191, 2024

  53. [53]

    Oakink: A large-scale knowledge repository for understanding hand-object interaction

    Lixin Yang, Kailin Li, Xinyu Zhan, Fei Wu, Anran Xu, Liu Liu, and Cewu Lu. Oakink: A large-scale knowledge repository for understanding hand-object interaction. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 20953--20962, 2022

  54. [54]

    What's in your hands? 3d reconstruction of generic objects in hands

    Yufei Ye, Abhinav Gupta, and Shubham Tulsiani. What's in your hands? 3d reconstruction of generic objects in hands. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 3895--3905, 2022

  55. [55]

    Perceiving 3d human-object spatial arrangements from a single image in the wild

    Jason Y Zhang, Sam Pepose, Hanbyul Joo, Deva Ramanan, Jitendra Malik, and Angjoo Kanazawa. Perceiving 3d human-object spatial arrangements from a single image in the wild. In Computer Vision--ECCV 2020: 16th European Conference, Glasgow, UK, August 23--28, 2020, Proceedings, Part XII 16, pages 34--51. Springer, 2020

  56. [56]

    End-to-end hand mesh recovery from a monocular rgb image

    Xiong Zhang, Qiang Li, Hong Mo, Wenbo Zhang, and Wen Zheng. End-to-end hand mesh recovery from a monocular rgb image. In Proceedings of the IEEE/CVF international conference on computer vision, pages 2354--2364, 2019

  57. [57]

    Hunyuan3d 2.0: Scaling diffusion models for high resolution textured 3d assets generation

    Zibo Zhao, Zeqiang Lai, Qingxiang Lin, Yunfei Zhao, Haolin Liu, Shuhui Yang, Yifei Feng, Mingxin Yang, Sheng Zhang, Xianghui Yang, et al. Hunyuan3d 2.0: Scaling diffusion models for high resolution textured 3d assets generation. arXiv preprint arXiv:2501.12202, 2025

  58. [58]

    Monocular real-time hand shape and motion capture using multi-modal data

    Yuxiao Zhou, Marc Habermann, Weipeng Xu, Ikhsanul Habibie, Christian Theobalt, and Feng Xu. Monocular real-time hand shape and motion capture using multi-modal data. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 5346--5355, 2020

This paper was first reviewed by deepseek-v4-flash on August 5, 2026.