REVIEW 3 major objections 3 minor 292 references
This survey argues that the scattered field of foundation-model-assisted hand-object interaction is unified by a taxonomy of eight types of prior knowledge in three families, and that every method can be understood by which prior it injects
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-04 01:18 UTC pith:FNALSTCS
load-bearing objection A useful organizing taxonomy for HOI+foundation models, but the uncertainty-mitigation claims are inferred rather than evidenced. the 3 major comments →
Hand-Object Interaction in the Age of Large Foundation Models:Reconstruction, Generation, and Embodied Transfer
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The paper's central claim is that the apparent fragmentation of HOI-plus-foundation-model work reflects a missing organizing dimension: what knowledge is introduced, where it enters, and which uncertainty it reduces. The authors define a deliberately narrow boundary—a method counts as foundation-model-prior only when an explicitly identified, large-scale general-purpose pretrained model contributes cross-domain knowledge through predictions, representations, transferred parameters, adaptation, or distillation—and on that basis catalog eight sub-priors in three families. They argue this taxonomy covers six HOI tasks (pose estimation, hand-held object reconstruction, dynamic reconstruction, gr
What carries the argument
The central organizing device is the eight-sub-prior taxonomy combined with a pipeline abstraction: each method is characterized by its prior source, the representation it injects, and the injection operator it uses. This vocabulary—initialization, regularization, conditioning, token fusion, score-guided regularization, retargeting—describes how foundation-model knowledge enters the HOI backbone and task head. The taxonomy is linked to a five-uncertainty model (shape, spatial, physical, semantic, dynamic), so each prior family is mapped to the specific ambiguities it mitigates. For embodied transfer, the machinery is a five-route diagram tracing HOI evidence through a transferred signal and
Load-bearing premise
The load-bearing premise is the paper's stipulative boundary for what counts as a foundation-model prior: a method is included only when an explicitly identified, large-scale general-purpose pretrained model contributes cross-domain knowledge through predictions, representations, transferred parameters, adaptation, or distillation, and task-specialized pretraining does not count.
What would settle it
A concrete test would be to re-run the survey's taxonomic tables under a broader inclusion rule that also counts task-specialized pretrained initializations as foundation priors; if the qualitative mapping between prior families and the five uncertainties still holds, the boundary is not load-bearing, but if the taxonomy's cells become crowded and the uncertainty mapping blurs, the paper's organizational claim is weakened. Alternatively, a single benchmark that jointly reports geometry, contact agreement, physical plausibility, and task success across the six tasks could settle whether current
If this is right
- HOI methods should be characterized and compared by the foundation-model priors they exploit, not only by architecture or training data.
- Geometric, semantic, and visual priors are complementary, so multi-prior systems are the natural next step for reducing shape, spatial, physical, semantic, and dynamic uncertainty together.
- Evaluation that reports only geometry or image fidelity is insufficient; contact agreement, physical plausibility, and functional task success must be reported alongside.
- Embodied transfer is best viewed as a downstream consumer of HOI reconstruction and generation outputs, with concrete transfer routes from human evidence to robot policy.
- The taxonomy provides a shared vocabulary that can make future HOI papers comparable and can guide the design of integrated, verifiable HOI systems.
Where Pith is reading between the lines
- The taxonomy's gatekeeping boundary is an author choice rather than an empirical result; re-running the survey's tables with a broader definition of foundation priors (for example, including task-specialized pretrained initializations) would shift coverage and primary/auxiliary assignments.
- The eight sub-priors are not cleanly orthogonal in practice: the same vision-language model counts as a semantic grounding prior when used for localization and as a language reasoning prior when used for intent inference, suggesting the taxonomy is a lens for reading methods rather than a unique partition.
- A concrete extension the paper leaves implicit is a joint benchmark that measures geometry, contact, physical plausibility, and task success together; such a benchmark would directly test the survey's claim that current metrics have blind spots.
- The emerging line of action-conditioned video generation (world models) is identified as an inference pattern rather than an injection operator; editorially, this could mature into a fourth visual sub-prior or a new task category as interactive rollouts grow.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This survey organizes the hand-object interaction (HOI) literature around eight foundation-model sub-priors grouped into geometric, semantic, and visual families, and maps how these priors are represented, injected, and adapted across six HOI reconstruction/generation tasks and embodied transfer to robot learning. The paper claims to be the first systematic review of foundation-model priors for HOI, and supports the taxonomy with representative methods in Tables 3–5, a dataset/evaluation summary, and a live repository.
Significance. If the taxonomy and the uncertainty-mitigation mapping are accepted, the survey would provide a useful organizing principle for a rapidly growing but fragmented field. The manuscript is internally consistent, the gatekeeping definition of 'foundation-model prior' is explicit and applied transparently (e.g., excluding ViTPose initialization in HaMeR), and the paper repeatedly identifies evaluation blind spots. These are strengths. However, the central analytical output — the mapping from each prior to the HOI uncertainties it 'helps reduce' — is largely inferential and not backed by the surveyed papers' evidence, which weakens the strongest claim.
major comments (3)
- [Abstract, Table 3, Secs. 3.3/5.3/7.2.3] The abstract claims the survey reveals 'which HOI uncertainty it helps reduce,' and Table 3's Unc.↓ column operationalizes this. Yet the cited papers rarely ablate the foundation-model component against the listed uncertainty. The paper itself concedes: Sec. 3.3 states a reconstructed shape 'should not be interpreted as interaction evidence by itself'; Sec. 5.3 says 'image realism is not interaction correctness'; Sec. 7.2.3 notes physical metrics 'depend strongly on mesh quality, friction, contact modeling, and simulator settings.' The Unc.↓ assignments thus appear to be authors' inference, not literature evidence. Please either (a) reclassify the Unc.↓ column as a hypothesized mechanism with a clear caveat, or (b) for each Table 3 row, cite an ablation/experiment from the original paper that supports the assignment. This is load-bearing because the abstract frames the entire survey arou
- [Sec. 1] The manuscript claims 'the first systematic review' of foundation-model priors for HOI, but no search/selection methodology is provided: no databases, query terms, inclusion/exclusion criteria, screening process, or date cutoff are documented. Without this, the 'systematic' claim cannot be audited and the survey cannot be distinguished from an author-selected narrative review. Please add a methodology paragraph (or appendix) documenting the protocol, or soften the claim to 'first literature survey' / 'comprehensive review'.
- [Table 3 and Table 4] The representative method list and dataset tables rely on a large number of arXiv preprints and very recent 2026 venue entries (e.g., GeoHand arXiv:2605.17354, ScaleHP arXiv:2606.25619, several CVPR 2026 entries). Given the survey's cutoff is not stated, the 'first systematic' claim and the balanced coverage of the field are difficult to assess. State the literature cutoff date, and mark entries that are preprints or not yet peer-reviewed at that date. This is especially relevant because several Table 3 exemplars are from the authors' own group, and the selection criteria for 'representative' methods are not specified.
minor comments (3)
- [Fig. 7] The figure caption and labels refer to 'Sec. 4.2: Human-Data Pretraining', 'Sec. 4.3: Human-to-Robot Skill Transfer', and 'Sec. 4.4: HOI-to-Robot Data Engines', but the corresponding sections in the text are 6.2, 6.3, and 6.4. Update the figure numbering.
- [Sec. 6.2.1] The phrase 'frame-aligned action chunks' and '6DoF object trajectories' appear without prior definition in the HOI taxonomy; consider adding these to the interaction-representation list in Sec. 2.2.2 for terminological consistency.
- [Throughout] The text frequently uses 'systematically analyze' and 'systematic coverage' without a clear definition of systematicity. Align these terms with the (proposed) methodology, or use more neutral phrasing such as 'structured analysis.'
Circularity Check
No significant circularity: the taxonomy is stipulative and the uncertainty mapping is interpretive, but no result reduces to its inputs by construction.
full rationale
This is a survey, not a derivation. Its central deliverable is a taxonomy of foundation-model priors and a qualitative mapping from priors to five uncertainties. The boundary in Sec. 1 is explicitly stipulative: "we use a deliberately narrow boundary: a method is considered a foundation-model-prior method only when an explicitly identified, large-scale general-purpose pretrained model contributes cross-domain knowledge to the HOI pipeline." Classification decisions are therefore authorial choices, not circular deductions. The Unc.↓ column in Table 3 is interpretive attribution rather than a fitted/predicted quantity; the paper itself qualifies the evidence base in several places: Sec. 3.3 says "the reconstructed shape should not be interpreted as interaction evidence by itself," Sec. 5.3 says "image realism is not interaction correctness," Sec. 7.2.3 says penetration and force-closure measures "depend strongly on mesh quality, friction, contact modeling, and simulator settings," and Sec. 8.4 lists prior reliability as an open problem. These caveats weaken the empirical support for the mapping, but they do not make it circular: no quantity is fitted and then presented as a prediction, and no equation or definition forces the survey's conclusions. The self-citations (GeoHand, HandOS, and ScaleHP in Table 3, with author overlap; MoGe-2 also has overlapping authors) are visible, but these are externally falsifiable method papers used as exemplars, the taxonomy does not depend on them, and per rule 4 self-citation alone is not circularity. The central claim—that the fragmented HOI literature can be organized by eight foundation-model sub-priors in three families—is a transparent analytic frame with independent content.
Axiom & Free-Parameter Ledger
axioms (3)
- ad hoc to paper A method counts as foundation-model-prior only if an explicitly identified, large-scale general-purpose pretrained model contributes cross-domain knowledge through predictions, representations, transferred parameters, adaptation, or distillation.
- domain assumption HOI failure modes decompose into five uncertainties: shape, spatial, physical, semantic, dynamic.
- domain assumption Each method can be assigned a unique primary sub-prior and optional auxiliary sub-priors with a single injection operator.
read the original abstract
Hand-object interaction (HOI) modeling remains challenging because it requires joint reasoning about hand articulation, object geometry, contact, semantics, and dynamics under severe visual uncertainty. Foundation models introduce transferable prior knowledge learned from large-scale cross-domain data, offering new ways to address these challenges beyond task-specific data and models. However, the rapidly growing literature remains fragmented, and existing studies typically describe these methods simply as ``using large models'' without systematically characterizing what knowledge is introduced, where it enters the HOI pipeline, or which HOI uncertainty it helps reduce. This survey presents the first systematic review of foundation-model priors for HOI. We organize the literature into six HOI tasks spanning reconstruction and generation. More importantly, we establish a taxonomy of eight foundation-model sub-priors grouped into geometric, semantic, and visual families. Geometric priors encompass shape retrieval, shape reconstruction, and spatial reconstruction; semantic priors include semantic grounding and language reasoning; and visual priors cover visual representation, image generation, and video generation. Based on this taxonomy, we systematically analyze how different priors are represented, injected, and adapted across HOI pipelines and tasks. Beyond how foundation models empower HOI, we further examine how HOI-derived knowledge is used in robot learning, including human-data pretraining, human-to-robot skill transfer, and HOI-to-robot data generation. Finally, we summarize datasets and evaluation protocols, and discuss limitations and future directions toward more generalizable HOI systems. To support long-term progress, we curate a live repository that continuously aggregates emerging methods and benchmarks.
Figures
Reference graph
Works this paper leans on
-
[1]
H+O: unified egocentric recognition of 3D hand- object poses and interactions
Bugra Tekin, Federica Bogo, and Marc Pollefeys. H+O: unified egocentric recognition of 3D hand- object poses and interactions. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2019. doi: 10.1109/CVPR.2019.00464
arXiv 2019
-
[2]
Learning joint reconstruction of hands and manipulated objects
Yana Hasson, Gul Varol, Dimitrios Tzionas, Igor Kalevatykh, Michael J Black, Ivan Laptev, and Cordelia Schmid. Learning joint reconstruction of hands and manipulated objects. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2019
2019
-
[3]
CPF: Learning a contact potential field to model the hand-object interaction
Lixin Yang, Xinyu Zhan, Kailin Li, Wenqiang Xu, Jiefeng Li, and Cewu Lu. CPF: Learning a contact potential field to model the hand-object interaction. InProc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), 2021
2021
-
[4]
HOLD: Category-agnostic 3D re- construction of interacting hands and objects from video
Zicong Fan, Maria Parelli, Maria Eleni Kadoglou, Xu Chen, Muhammed Kocabas, Michael J Black, and Otmar Hilliges. HOLD: Category-agnostic 3D re- construction of interacting hands and objects from video. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2024
2024
-
[5]
What’s in your hands? 3D reconstruction of generic objects in hands
Yufei Ye, Abhinav Gupta, and Shubham Tulsiani. What’s in your hands? 3D reconstruction of generic objects in hands. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2022
2022
-
[6]
Grasping field: Learning implicit representations for human grasps
Korrawe Karunratanakul, Jinlong Yang, Yan Zhang, Michael J Black, Krikamol Muandet, and Siyu Tang. Grasping field: Learning implicit representations for human grasps. InProc. Int. Conf. 3D Vis. (3DV), 2020
2020
-
[7]
DUSt3R: Geomet- ric 3D vision made easy
Shuzhe Wang, Vincent Leroy, Yohann Cabon, Boris Chidlovskii, and Jerome Revaud. DUSt3R: Geomet- ric 3D vision made easy. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2024
2024
-
[8]
Samarth Brahmbhatt, Chengcheng Tang, Christo- pher D Twigg, Charles C Kemp, and James Hays. 23 Hand-Object Interaction in the Age of Large Foundation Models: Reconstruction, Generation, and Embodied Transfer Contactpose: A dataset of grasps with object con- tact and hand pose. InProc. Eur. Conf. Comput. Vis. (ECCV), 2020
2020
-
[9]
S2contact: Graph-based network for 3D hand-object contact estimation with semi-supervised learning
Tze Ho Elden Tse, Zhongqun Zhang, Kwang In Kim, Ales Leonardis, Feng Zheng, and Hyung Jin Chang. S2contact: Graph-based network for 3D hand-object contact estimation with semi-supervised learning. In Proc. Eur. Conf. Comput. Vis. (ECCV), 2022. doi: 10.1007/978-3-031-19769-7\ 33
-
[10]
Contactopt: Optimizing contact to improve grasps
Patrick Grady, Chengcheng Tang, Christopher D Twigg, Minh V o, Samarth Brahmbhatt, and Charles C Kemp. Contactopt: Optimizing contact to improve grasps. InProc. IEEE/CVF Conf. Comput. Vis. Pat- tern Recognit. (CVPR), 2021
2021
-
[11]
Deep- SimHO: Stable pose estimation for hand-object in- teraction via physics simulation
Rong Wang, Wei Mao, and Hongdong Li. Deep- SimHO: Stable pose estimation for hand-object in- teraction via physics simulation. InAdv. Neural Inf. Process. Syst. (NeurIPS), volume 36, 2023
2023
-
[12]
D- grasp: Physically plausible dynamic grasp synthesis for hand-object interactions
Sammy Christen, Muhammed Kocabas, Emre Aksan, Jemin Hwangbo, Jie Song, and Otmar Hilliges. D- grasp: Physically plausible dynamic grasp synthesis for hand-object interactions. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2022
2022
-
[13]
Grounding DINO: Marrying DINO with grounded pre-training for open-set object detection
Shilong Liu, Zhaoyang Zeng, Tianhe Ren, Feng Li, Hao Zhang, Jie Yang, Qing Jiang, Chunyuan Li, Jian- wei Yang, Hang Su, et al. Grounding DINO: Marrying DINO with grounded pre-training for open-set object detection. InProc. Eur. Conf. Comput. Vis. (ECCV), 2024
2024
-
[14]
Segment anything
Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C Berg, Wan-Yen Lo, et al. Segment anything. InProc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), 2023
2023
-
[15]
SemGrasp: Semantic grasp generation via language aligned discretization
Kailin Li, Jingbo Wang, Lixin Yang, Cewu Lu, and Bo Dai. SemGrasp: Semantic grasp generation via language aligned discretization. InProc. Eur. Conf. Comput. Vis. (ECCV), 2024
2024
-
[16]
Text2grasp: Synthesis of grasps by text prompts for object grasping parts
Xiaoyun Chang and Yi Sun. Text2grasp: Synthesis of grasps by text prompts for object grasping parts. In Proc. Int. Symp. Neural Netw. (ISNN), 2025
2025
-
[17]
Towards unconstrained joint hand-object re- construction from RGB videos
Yana Hasson, G¨ul Varol, Cordelia Schmid, and Ivan Laptev. Towards unconstrained joint hand-object re- construction from RGB videos. InProc. Int. Conf. 3D Vis. (3DV), 2021
2021
-
[18]
Ayce Idil Aytekin, Xu Chen, Zhengyang Shen, Thabo Beeler, Helge Rhodin, Rishabh Dabral, and Christian Theobalt. Grasp in gaussians: Fast monocular recon- struction of dynamic hand-object interactions.arXiv preprint arXiv:2604.12929, 2026
Pith/arXiv arXiv 2026
-
[19]
EasyHOI: Unleashing the power of large models for reconstructing hand-object inter- actions in the wild
Yumeng Liu, Xiaoxiao Long, Zemin Yang, Yuan Liu, Marc Habermann, Christian Theobalt, Yuexin Ma, and Wenping Wang. EasyHOI: Unleashing the power of large models for reconstructing hand-object inter- actions in the wild. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2025
2025
-
[20]
MagicHOI: Leveraging 3D priors for accurate hand-object recon- struction from short monocular video clips
Shibo Wang, Haonan He, Maria Parelli, Christoph Gebhardt, Zicong Fan, and Jie Song. MagicHOI: Leveraging 3D priors for accurate hand-object recon- struction from short monocular video clips. InProc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), 2025
2025
-
[21]
Diffusion-guided reconstruction of everyday hand-object interaction clips
Yufei Ye, Poorvi Hebbar, Abhinav Gupta, and Shub- ham Tulsiani. Diffusion-guided reconstruction of everyday hand-object interaction clips. InProc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), 2023
2023
-
[22]
Hand-object interaction image gen- eration
Hezhen Hu, Weilun Wang, Wengang Zhou, and Houqiang Li. Hand-object interaction image gen- eration. InAdv. Neural Inf. Process. Syst. (NeurIPS), volume 35, 2022
2022
-
[23]
HOIDiffusion: Generating realistic 3D hand-object interaction data
Mengqi Zhang, Yang Fu, Zheng Ding, Sifei Liu, Zhuowen Tu, and Xiaolong Wang. HOIDiffusion: Generating realistic 3D hand-object interaction data. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recog- nit. (CVPR), 2024
2024
-
[24]
Svimo: Synchro- nized diffusion for video and motion generation in hand-object interaction scenarios
Lingwei Dang, Ruizhi Shao, Hongwen Zhang, Wei Min, Yebin Liu, and Qingyao Wu. Svimo: Synchro- nized diffusion for video and motion generation in hand-object interaction scenarios. InAdv. Neural Inf. Process. Syst. (NeurIPS), volume 38, 2025
2025
-
[25]
Jiale Xu, Weihao Cheng, Yiming Gao, Xintao Wang, Shenghua Gao, and Ying Shan. Instantmesh: Efficient 3D mesh generation from a single image with sparse- view large reconstruction models.arXiv preprint arXiv:2404.07191, 2024
Pith/arXiv arXiv 2024
-
[26]
OpenShape: Scaling up 3D shape representation towards open-world understanding
Minghua Liu, Ruoxi Shi, Kaiming Kuang, Yinhao Zhu, Xuanlin Li, Shizhong Han, Hong Cai, Fatih Porikli, and Hao Su. OpenShape: Scaling up 3D shape representation towards open-world understanding. In Adv. Neural Inf. Process. Syst. (NeurIPS), volume 36, 2023
2023
-
[27]
Visual instruction tuning
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. Visual instruction tuning. InAdv. Neu- ral Inf. Process. Syst. (NeurIPS), volume 36, 2023. 24 Hand-Object Interaction in the Age of Large Foundation Models: Reconstruction, Generation, and Embodied Transfer
2023
-
[28]
Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou. Qwen-VL: A versatile vision-language model for understanding, localiza- tion, text reading, and beyond.arXiv preprint arXiv:2308.12966, 2023
Pith/arXiv arXiv 2023
-
[29]
DINOv2: Learning robust vi- sual features without supervision.arXiv preprint arXiv:2304.07193, 2023
Maxime Oquab, Timoth´ee Darcet, Th´eo Moutakanni, Huy V o, Marc Szafraniec, Vasil Khalidov, Pierre Fer- nandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, et al. DINOv2: Learning robust vi- sual features without supervision.arXiv preprint arXiv:2304.07193, 2023
Pith/arXiv arXiv 2023
-
[30]
High- resolution image synthesis with latent diffusion mod- els
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bj ¨orn Ommer. High- resolution image synthesis with latent diffusion mod- els. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2022
2022
-
[31]
Cogvideox: Text-to-video diffusion models with an expert transformer
Zhuoyi Yang, Jiayan Teng, Wendi Zheng, Ming Ding, Shiyu Huang, Jiazheng Xu, Yuanming Yang, Wenyi Hong, Xiaohan Zhang, Guanyu Feng, et al. Cogvideox: Text-to-video diffusion models with an expert transformer. InProc. Int. Conf. Learn. Repre- sent. (ICLR), 2025
2025
-
[32]
Objaverse: A universe of annotated 3D objects
Matt Deitke, Dustin Schwenk, Jordi Salvador, Luca Weihs, Oscar Michel, Eli VanderBilt, Ludwig Schmidt, Kiana Ehsani, Aniruddha Kembhavi, and Ali Farhadi. Objaverse: A universe of annotated 3D objects. InProc. IEEE/CVF Conf. Comput. Vis. Pat- tern Recognit. (CVPR), 2023
2023
-
[33]
Learning trans- ferable visual models from natural language supervi- sion
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sas- try, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning trans- ferable visual models from natural language supervi- sion. InProc. Int. Conf. Mach. Learn. (ICML), 2021
2021
-
[34]
Reconstructing hand-held objects in 3D from images and videos
Jane Wu, Georgios Pavlakos, Georgia Gkioxari, and Jitendra Malik. Reconstructing hand-held objects in 3D from images and videos. InProc. Int. Conf. 3D Vis. (3DV), 2026
2026
-
[35]
Ghost: Fast category- agnostic hand-object interaction reconstruction from RGB videos using gaussian splatting
Ahmed Tawfik Aboukhadra, Marcel Rogge, Nadia Robertini, Abdalla Arafa, Jameel Malik, Ahmed El- hayek, and Didier Stricker. Ghost: Fast category- agnostic hand-object interaction reconstruction from RGB videos using gaussian splatting. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2026
2026
-
[36]
Hand-held object reconstruction from RGB video with dynamic interaction
Shijian Jiang, Qi Ye, Rengan Xie, Yuchi Huo, and Jim- ing Chen. Hand-held object reconstruction from RGB video with dynamic interaction. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2025
2025
-
[37]
Reconstructing hands in 3D with transformers
Georgios Pavlakos, Dandan Shan, Ilija Radosavovic, Angjoo Kanazawa, David Fouhey, and Jitendra Malik. Reconstructing hands in 3D with transformers. In Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2024
2024
-
[38]
Wilor: End- to-end 3D hand localization and reconstruction in-the- wild
Rolandos Alexandros Potamias, Jinglei Zhang, Jiankang Deng, and Stefanos Zafeiriou. Wilor: End- to-end 3D hand localization and reconstruction in-the- wild. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2025
2025
-
[39]
ViTPose: Simple vision transformer baselines for human pose estimation
Yufei Xu, Jing Zhang, Qiming Zhang, and Dacheng Tao. ViTPose: Simple vision transformer baselines for human pose estimation. InAdv. Neural Inf. Process. Syst. (NeurIPS), 2022
2022
-
[40]
Latent action pretraining from videos
Seonghyeon Ye, Joel Jang, Byeongguk Jeon, Se June Joo, Jianwei Yang, Baolin Peng, Ajay Mandlekar, Reuben Tan, Yu-Wei Chao, Bill Yuchen Lin, et al. Latent action pretraining from videos. InProc. Int. Conf. Learn. Represent. (ICLR), 2025
2025
-
[41]
Ruihan Yang, Qinxi Yu, Yecheng Wu, Rui Yan, Borui Li, An-Chieh Cheng, Xueyan Zou, Yunhao Fang, Xuxin Cheng, Ri-Zhao Qiu, et al. Egovla: Learning vision-language-action models from egocentric hu- man videos.arXiv preprint arXiv:2507.12440, 2025
Pith/arXiv arXiv 2025
-
[42]
Dexmv: Imitation learning for dexterous manipula- tion from human videos
Yuzhe Qin, Yueh-Hua Wu, Shaowei Liu, Hanwen Jiang, Ruihan Yang, Yang Fu, and Xiaolong Wang. Dexmv: Imitation learning for dexterous manipula- tion from human videos. InProc. Eur. Conf. Comput. Vis. (ECCV), 2022
2022
-
[43]
Yangcen Liu, Shuo Cheng, Xinchen Yin, Woo Chul Shin, Alfred Cueva, Yiran Yang, Zhenyang Chen, Chuye Zhang, and Danfei Xu. Egoengine: From ego- centric human videos to high-fidelity dexterous robot demonstrations.arXiv preprint arXiv:2606.12604, 2026
Pith/arXiv arXiv 2026
-
[44]
Efficient annotation and learning for 3D hand pose estimation: A survey.Int
Takehiko Ohkawa, Ryosuke Furuta, and Yoichi Sato. Efficient annotation and learning for 3D hand pose estimation: A survey.Int. J. Comput. Vis., 131(12), 2023
2023
-
[45]
A survey of deep learning methods and datasets for hand pose estimation from hand-object interaction images.Comput
Taeyun Woo, Wonjung Park, Woohyun Jeong, and Jinah Park. A survey of deep learning methods and datasets for hand pose estimation from hand-object interaction images.Comput. Graph., 116, 2023
2023
-
[46]
Advances in vision-based deep learning methods for interacting hands reconstruction: A survey.Comput
Yu Miao and Yue Liu. Advances in vision-based deep learning methods for interacting hands reconstruction: A survey.Comput. Graph., 124, 2024. 25 Hand-Object Interaction in the Age of Large Foundation Models: Reconstruction, Generation, and Embodied Transfer
2024
-
[47]
An overview of learning-based dexterous grasping: recent advances and future directions.Artif
Xu Song, Yongyao Li, Yunfan Zhang, Yufei Liu, and Lei Jiang. An overview of learning-based dexterous grasping: recent advances and future directions.Artif. Intell. Rev., 58(10), 2025
2025
-
[48]
Arthoi: Taming foundation models for monocular 4D reconstruction of hand-articulated- object interactions
Zikai Wang, Zhilu Zhang, Yiqing Wang, Hui Li, and Wangmeng Zuo. Arthoi: Taming foundation models for monocular 4D reconstruction of hand-articulated- object interactions. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2026
2026
-
[49]
Javier Romero, Dimitrios Tzionas, and Michael J. Black. Embodied hands: modeling and capturing hands and bodies together.ACM Trans. Graph., 36 (6), 2017. doi: 10.1145/3130800.3130883
arXiv 2017
-
[50]
Nimble: a non-rigid hand model with bones and muscles.ACM Trans
Yuwei Li, Longwen Zhang, Zesong Qiu, Yingwenqi Jiang, Nianyi Li, Yuexin Ma, Yuyao Zhang, Lan Xu, and Jingyi Yu. Nimble: a non-rigid hand model with bones and muscles.ACM Trans. Graph., 41(4), 2022
2022
-
[51]
HOnnotate: A method for 3D an- notation of hand and object poses
Shreyas Hampali, Mahdi Rad, Markus Oberweger, and Vincent Lepetit. HOnnotate: A method for 3D an- notation of hand and object poses. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2020
2020
-
[52]
AlignSDF: Pose-aligned signed distance fields for hand-object reconstruction
Zerui Chen, Yana Hasson, Cordelia Schmid, and Ivan Laptev. AlignSDF: Pose-aligned signed distance fields for hand-object reconstruction. InProc. Eur. Conf. Comput. Vis. (ECCV), 2022
2022
-
[53]
HandNeRF: Learn- ing to reconstruct hand-object interaction scene from a single RGB image
Hongsuk Choi, Nikhil Chavan-Dafle, Jiacheng Yuan, V olkan Isler, and Hyunsoo Park. HandNeRF: Learn- ing to reconstruct hand-object interaction scene from a single RGB image. InProc. IEEE Int. Conf. Robot. Autom. (ICRA), 2024
2024
-
[54]
Simultaneous localization and mapping: part i.IEEE Robot
Hugh Durrant-Whyte and Tim Bailey. Simultaneous localization and mapping: part i.IEEE Robot. Autom. Mag., 13(2), 2006
2006
-
[55]
Seitz, and Richard Szeliski
Noah Snavely, Steven M. Seitz, and Richard Szeliski. Photo tourism: Exploring photo collections in 3D. In ACM SIGGRAPH 2006 Papers. ACM, 2006
2006
-
[56]
LatentHOI: On the generalizable hand object motion generation with latent hand diffusion
Muchen Li, Sammy Christen, Chengde Wan, Yu- jun Cai, Renjie Liao, Leonid Sigal, and Shugao Ma. LatentHOI: On the generalizable hand object motion generation with latent hand diffusion. In Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2025
2025
-
[57]
3D hand pose estimation in everyday egocentric images
Aditya Prakash, Ruisen Tu, Matthew Chang, and Saurabh Gupta. 3D hand pose estimation in everyday egocentric images. InProc. Eur. Conf. Comput. Vis. (ECCV), 2024
2024
-
[58]
DDF-HO: Hand-held object recon- struction via conditional directed distance field
Chenyangguang Zhang, Yan Di, Ruida Zhang, Guangyao Zhai, Fabian Manhardt, Federico Tombari, and Xiangyang Ji. DDF-HO: Hand-held object recon- struction via conditional directed distance field. In Adv. Neural Inf. Process. Syst. (NeurIPS), volume 36, 2023
2023
-
[59]
Ncrf: neural contact radiance fields for free-viewpoint rendering of hand-object interaction
Zhongqun Zhang, Jifei Song, Eduardo P´erez-Pellitero, Yiren Zhou, Hyung Jin Chang, and Ale ˇs Leonardis. Ncrf: neural contact radiance fields for free-viewpoint rendering of hand-object interaction. InProc. Int. Conf. 3D Vis. (3DV), 2024
2024
-
[60]
Gan- hand: Predicting human grasp affordances in multi- object scenes
Enric Corona, Albert Pumarola, Guillem Alenya, Francesc Moreno-Noguer, and Gr´egory Rogez. Gan- hand: Predicting human grasp affordances in multi- object scenes. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2020
2020
-
[61]
Artigrasp: Physically plausible synthesis of bi-manual dexterous grasping and articulation
Hui Zhang, Sammy Christen, Zicong Fan, Luocheng Zheng, Jemin Hwangbo, Jie Song, and Otmar Hilliges. Artigrasp: Physically plausible synthesis of bi-manual dexterous grasping and articulation. InProc. Int. Conf. 3D Vis. (3DV), 2024
2024
-
[62]
Deepsdf: Learning continuous signed distance functions for shape representation
Jeong Joon Park, Peter Florence, Julian Straub, Richard Newcombe, and Steven Lovegrove. Deepsdf: Learning continuous signed distance functions for shape representation. InProc. IEEE/CVF Conf. Com- put. Vis. Pattern Recognit. (CVPR), 2019
2019
-
[63]
A skeleton-driven neural occupancy representation for articulated hands
Korrawe Karunratanakul, Adrian Spurr, Zicong Fan, Otmar Hilliges, and Siyu Tang. A skeleton-driven neural occupancy representation for articulated hands. InProc. Int. Conf. 3D Vis. (3DV), 2021
2021
-
[64]
Srinivasan, Matthew Tancik, Jonathan T
Ben Mildenhall, Pratul P. Srinivasan, Matthew Tancik, Jonathan T. Barron, Ravi Ramamoorthi, and Ren Ng. NeRF: Representing scenes as neural radiance fields for view synthesis. InProc. Eur. Conf. Comput. Vis. (ECCV), 2020. doi: 10.1007/978-3-030-58452-8 \ 24
-
[65]
3D gaussian splat- ting for real-time radiance field rendering.ACM Trans
Bernhard Kerbl, Georgios Kopanas, Thomas Leimk¨uhler, and George Drettakis. 3D gaussian splat- ting for real-time radiance field rendering.ACM Trans. Graph., 42(4), 2023
2023
-
[66]
Affordpose: A large-scale dataset of hand- object interactions with affordance-driven hand pose
Juntao Jian, Xiuping Liu, Manyi Li, Ruizhen Hu, and Jian Liu. Affordpose: A large-scale dataset of hand- object interactions with affordance-driven hand pose. InProc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), 2023
2023
-
[67]
Th´eo Morales, Omid Taheri, and Gerard Lacey. A versatile and differentiable hand-object interaction 26 Hand-Object Interaction in the Age of Large Foundation Models: Reconstruction, Generation, and Embodied Transfer representation. InProc. IEEE/CVF Winter Conf. Appl. Comput. Vis. (WACV), 2025
2025
-
[68]
Black, and Dima Damen
Siddhant Bansal, Zhifan Zhu, Shashank Tripathi, Ji- ahe Zhao, Michael J. Black, and Dima Damen. To- wards in-the-wild egocentric 3D hand-object pose estimation. InProc. Eur. Conf. Comput. Vis. (ECCV), 2026
2026
-
[69]
HOI4D: A 4D egocentric dataset for category-level human-object interaction
Yunze Liu, Yun Liu, Che Jiang, Kangbo Lyu, Weikang Wan, Hao Shen, Boqiang Liang, Zhoujie Fu, He Wang, and Li Yi. HOI4D: A 4D egocentric dataset for category-level human-object interaction. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recog- nit. (CVPR), 2022
2022
-
[70]
Arctic: A dataset for dexterous bimanual hand-object manipulation
Zicong Fan, Omid Taheri, Dimitrios Tzionas, Muhammed Kocabas, Manuel Kaufmann, Michael J Black, and Otmar Hilliges. Arctic: A dataset for dexterous bimanual hand-object manipulation. In Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2023
2023
-
[71]
Handoccnet: Occlusion-robust 3D hand mesh estimation network
JoonKyu Park, Yeonguk Oh, Gyeongsik Moon, Hong- suk Choi, and Kyoung Mu Lee. Handoccnet: Occlusion-robust 3D hand mesh estimation network. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recog- nit. (CVPR), 2022
2022
-
[72]
Mobrecon: Mobile-friendly hand mesh reconstruction from monocular image
Xingyu Chen, Yufeng Liu, Yajiao Dong, Xiong Zhang, Chongyang Ma, Yanmin Xiong, Yuan Zhang, and Xiaoyan Guo. Mobrecon: Mobile-friendly hand mesh reconstruction from monocular image. In Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2022
2022
-
[73]
A simple baseline for ef- ficient hand mesh reconstruction
Zhishan Zhou, Shihao Zhou, Zhi Lv, Minqiang Zou, Yao Tang, and Jiajun Liang. A simple baseline for ef- ficient hand mesh reconstruction. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2024
2024
-
[74]
Model-based 3D hand reconstruction via self- supervised learning
Yujin Chen, Zhigang Tu, Di Kang, Linchao Bao, Ying Zhang, Xuefei Zhe, Ruizhi Chen, and Junsong Yuan. Model-based 3D hand reconstruction via self- supervised learning. InProc. IEEE/CVF Conf. Com- put. Vis. Pattern Recognit. (CVPR), 2021
2021
-
[75]
Keypoint fusion for RGB-D based 3D hand pose estimation
Xingyu Liu, Pengfei Ren, Yuanyuan Gao, Jingyu Wang, Haifeng Sun, Qi Qi, Zirui Zhuang, and Jianxin Liao. Keypoint fusion for RGB-D based 3D hand pose estimation. InProc. AAAI Conf. Artif. Intell., 2024
2024
-
[76]
Hope-net: A graph-based model for hand-object pose estimation
Bardia Doosti, Shujon Naha, Majid Mirbagheri, and David J Crandall. Hope-net: A graph-based model for hand-object pose estimation. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2020
2020
-
[77]
Keypoint transformer: Solv- ing joint identification in challenging hands and ob- ject interactions for accurate 3D pose estimation
Shreyas Hampali, Sayan Deb Sarkar, Mahdi Rad, and Vincent Lepetit. Keypoint transformer: Solv- ing joint identification in challenging hands and ob- ject interactions for accurate 3D pose estimation. In Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2022
2022
-
[78]
Thor-net: End-to-end graformer-based realistic two hands and object reconstruction with self-supervision
Ahmed Tawfik Aboukhadra, Jameel Malik, Ahmed El- hayek, Nadia Robertini, and Didier Stricker. Thor-net: End-to-end graformer-based realistic two hands and object reconstruction with self-supervision. InProc. IEEE/CVF Winter Conf. Appl. Comput. Vis. (WACV), 2023
2023
-
[79]
HOISDF: Constraining 3D hand- object pose estimation with global signed distance fields
Haozhe Qi, Chen Zhao, Mathieu Salzmann, and Alexander Mathis. HOISDF: Constraining 3D hand- object pose estimation with global signed distance fields. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2024
2024
-
[80]
Contactart: Learning 3D interaction priors for category-level ar- ticulated object and hand poses estimation
Zehao Zhu, Jiashun Wang, Yuzhe Qin, Deqing Sun, Varun Jampani, and Xiaolong Wang. Contactart: Learning 3D interaction priors for category-level ar- ticulated object and hand poses estimation. InProc. Int. Conf. 3D Vis. (3DV), 2024
2024
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.