Pith. sign in

REVIEW 3 major objections 3 minor 292 references

This survey argues that the scattered field of foundation-model-assisted hand-object interaction is unified by a taxonomy of eight types of prior knowledge in three families, and that every method can be understood by which prior it injects

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-04 01:18 UTC pith:FNALSTCS

load-bearing objection A useful organizing taxonomy for HOI+foundation models, but the uncertainty-mitigation claims are inferred rather than evidenced. the 3 major comments →

arxiv 2607.28394 v2 pith:FNALSTCS submitted 2026-07-30 cs.CV

Hand-Object Interaction in the Age of Large Foundation Models:Reconstruction, Generation, and Embodied Transfer

classification cs.CV
keywords hand-object interactionfoundation modelsHOI reconstructionHOI generationembodied transfertaxonomycomputer vision survey
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper is a systematic survey arguing that hand-object interaction (HOI) research in the foundation-model era is best understood not by architecture or dataset, but by what cross-domain knowledge a pretrained model contributes and where it enters the pipeline. It proposes a taxonomy of eight foundation-model priors in three families—geometric (shape retrieval, shape reconstruction, spatial reconstruction), semantic (semantic grounding, language reasoning), and visual (visual representation, image generation, video generation). It claims this taxonomy can organize methods across six HOI tasks in reconstruction and generation, and that tracing prior injection explains which of five recurring uncertainties each method reduces. It further claims that HOI-derived knowledge transfers to robot learning through five routes, and that current evaluation metrics measure geometry but not interaction correctness. A sympathetic reader would care because the survey gives the field a common vocabulary for comparing methods and designing integrated HOI systems.

Core claim

The paper's central claim is that the apparent fragmentation of HOI-plus-foundation-model work reflects a missing organizing dimension: what knowledge is introduced, where it enters, and which uncertainty it reduces. The authors define a deliberately narrow boundary—a method counts as foundation-model-prior only when an explicitly identified, large-scale general-purpose pretrained model contributes cross-domain knowledge through predictions, representations, transferred parameters, adaptation, or distillation—and on that basis catalog eight sub-priors in three families. They argue this taxonomy covers six HOI tasks (pose estimation, hand-held object reconstruction, dynamic reconstruction, gr

What carries the argument

The central organizing device is the eight-sub-prior taxonomy combined with a pipeline abstraction: each method is characterized by its prior source, the representation it injects, and the injection operator it uses. This vocabulary—initialization, regularization, conditioning, token fusion, score-guided regularization, retargeting—describes how foundation-model knowledge enters the HOI backbone and task head. The taxonomy is linked to a five-uncertainty model (shape, spatial, physical, semantic, dynamic), so each prior family is mapped to the specific ambiguities it mitigates. For embodied transfer, the machinery is a five-route diagram tracing HOI evidence through a transferred signal and

Load-bearing premise

The load-bearing premise is the paper's stipulative boundary for what counts as a foundation-model prior: a method is included only when an explicitly identified, large-scale general-purpose pretrained model contributes cross-domain knowledge through predictions, representations, transferred parameters, adaptation, or distillation, and task-specialized pretraining does not count.

What would settle it

A concrete test would be to re-run the survey's taxonomic tables under a broader inclusion rule that also counts task-specialized pretrained initializations as foundation priors; if the qualitative mapping between prior families and the five uncertainties still holds, the boundary is not load-bearing, but if the taxonomy's cells become crowded and the uncertainty mapping blurs, the paper's organizational claim is weakened. Alternatively, a single benchmark that jointly reports geometry, contact agreement, physical plausibility, and task success across the six tasks could settle whether current

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • HOI methods should be characterized and compared by the foundation-model priors they exploit, not only by architecture or training data.
  • Geometric, semantic, and visual priors are complementary, so multi-prior systems are the natural next step for reducing shape, spatial, physical, semantic, and dynamic uncertainty together.
  • Evaluation that reports only geometry or image fidelity is insufficient; contact agreement, physical plausibility, and functional task success must be reported alongside.
  • Embodied transfer is best viewed as a downstream consumer of HOI reconstruction and generation outputs, with concrete transfer routes from human evidence to robot policy.
  • The taxonomy provides a shared vocabulary that can make future HOI papers comparable and can guide the design of integrated, verifiable HOI systems.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The taxonomy's gatekeeping boundary is an author choice rather than an empirical result; re-running the survey's tables with a broader definition of foundation priors (for example, including task-specialized pretrained initializations) would shift coverage and primary/auxiliary assignments.
  • The eight sub-priors are not cleanly orthogonal in practice: the same vision-language model counts as a semantic grounding prior when used for localization and as a language reasoning prior when used for intent inference, suggesting the taxonomy is a lens for reading methods rather than a unique partition.
  • A concrete extension the paper leaves implicit is a joint benchmark that measures geometry, contact, physical plausibility, and task success together; such a benchmark would directly test the survey's claim that current metrics have blind spots.
  • The emerging line of action-conditioned video generation (world models) is identified as an inference pattern rather than an injection operator; editorially, this could mature into a fourth visual sub-prior or a new task category as interactive rollouts grow.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 3 minor

Summary. This survey organizes the hand-object interaction (HOI) literature around eight foundation-model sub-priors grouped into geometric, semantic, and visual families, and maps how these priors are represented, injected, and adapted across six HOI reconstruction/generation tasks and embodied transfer to robot learning. The paper claims to be the first systematic review of foundation-model priors for HOI, and supports the taxonomy with representative methods in Tables 3–5, a dataset/evaluation summary, and a live repository.

Significance. If the taxonomy and the uncertainty-mitigation mapping are accepted, the survey would provide a useful organizing principle for a rapidly growing but fragmented field. The manuscript is internally consistent, the gatekeeping definition of 'foundation-model prior' is explicit and applied transparently (e.g., excluding ViTPose initialization in HaMeR), and the paper repeatedly identifies evaluation blind spots. These are strengths. However, the central analytical output — the mapping from each prior to the HOI uncertainties it 'helps reduce' — is largely inferential and not backed by the surveyed papers' evidence, which weakens the strongest claim.

major comments (3)
  1. [Abstract, Table 3, Secs. 3.3/5.3/7.2.3] The abstract claims the survey reveals 'which HOI uncertainty it helps reduce,' and Table 3's Unc.↓ column operationalizes this. Yet the cited papers rarely ablate the foundation-model component against the listed uncertainty. The paper itself concedes: Sec. 3.3 states a reconstructed shape 'should not be interpreted as interaction evidence by itself'; Sec. 5.3 says 'image realism is not interaction correctness'; Sec. 7.2.3 notes physical metrics 'depend strongly on mesh quality, friction, contact modeling, and simulator settings.' The Unc.↓ assignments thus appear to be authors' inference, not literature evidence. Please either (a) reclassify the Unc.↓ column as a hypothesized mechanism with a clear caveat, or (b) for each Table 3 row, cite an ablation/experiment from the original paper that supports the assignment. This is load-bearing because the abstract frames the entire survey arou
  2. [Sec. 1] The manuscript claims 'the first systematic review' of foundation-model priors for HOI, but no search/selection methodology is provided: no databases, query terms, inclusion/exclusion criteria, screening process, or date cutoff are documented. Without this, the 'systematic' claim cannot be audited and the survey cannot be distinguished from an author-selected narrative review. Please add a methodology paragraph (or appendix) documenting the protocol, or soften the claim to 'first literature survey' / 'comprehensive review'.
  3. [Table 3 and Table 4] The representative method list and dataset tables rely on a large number of arXiv preprints and very recent 2026 venue entries (e.g., GeoHand arXiv:2605.17354, ScaleHP arXiv:2606.25619, several CVPR 2026 entries). Given the survey's cutoff is not stated, the 'first systematic' claim and the balanced coverage of the field are difficult to assess. State the literature cutoff date, and mark entries that are preprints or not yet peer-reviewed at that date. This is especially relevant because several Table 3 exemplars are from the authors' own group, and the selection criteria for 'representative' methods are not specified.
minor comments (3)
  1. [Fig. 7] The figure caption and labels refer to 'Sec. 4.2: Human-Data Pretraining', 'Sec. 4.3: Human-to-Robot Skill Transfer', and 'Sec. 4.4: HOI-to-Robot Data Engines', but the corresponding sections in the text are 6.2, 6.3, and 6.4. Update the figure numbering.
  2. [Sec. 6.2.1] The phrase 'frame-aligned action chunks' and '6DoF object trajectories' appear without prior definition in the HOI taxonomy; consider adding these to the interaction-representation list in Sec. 2.2.2 for terminological consistency.
  3. [Throughout] The text frequently uses 'systematically analyze' and 'systematic coverage' without a clear definition of systematicity. Align these terms with the (proposed) methodology, or use more neutral phrasing such as 'structured analysis.'

Circularity Check

0 steps flagged

No significant circularity: the taxonomy is stipulative and the uncertainty mapping is interpretive, but no result reduces to its inputs by construction.

full rationale

This is a survey, not a derivation. Its central deliverable is a taxonomy of foundation-model priors and a qualitative mapping from priors to five uncertainties. The boundary in Sec. 1 is explicitly stipulative: "we use a deliberately narrow boundary: a method is considered a foundation-model-prior method only when an explicitly identified, large-scale general-purpose pretrained model contributes cross-domain knowledge to the HOI pipeline." Classification decisions are therefore authorial choices, not circular deductions. The Unc.↓ column in Table 3 is interpretive attribution rather than a fitted/predicted quantity; the paper itself qualifies the evidence base in several places: Sec. 3.3 says "the reconstructed shape should not be interpreted as interaction evidence by itself," Sec. 5.3 says "image realism is not interaction correctness," Sec. 7.2.3 says penetration and force-closure measures "depend strongly on mesh quality, friction, contact modeling, and simulator settings," and Sec. 8.4 lists prior reliability as an open problem. These caveats weaken the empirical support for the mapping, but they do not make it circular: no quantity is fitted and then presented as a prediction, and no equation or definition forces the survey's conclusions. The self-citations (GeoHand, HandOS, and ScaleHP in Table 3, with author overlap; MoGe-2 also has overlapping authors) are visible, but these are externally falsifiable method papers used as exemplars, the taxonomy does not depend on them, and per rule 4 self-citation alone is not circularity. The central claim—that the fragmented HOI literature can be organized by eight foundation-model sub-priors in three families—is a transparent analytic frame with independent content.

Axiom & Free-Parameter Ledger

0 free parameters · 3 axioms · 0 invented entities

The survey rests on three declared or implicit premises: a stipulative boundary for what counts as a foundation-model prior; a five-way decomposition of HOI uncertainty; and an assumption that each method can be tagged with a single primary sub-prior. These are author choices rather than demonstrated empirical facts; they gate every subsequent table and figure. No free parameters or invented physical entities appear.

axioms (3)
  • ad hoc to paper A method counts as foundation-model-prior only if an explicitly identified, large-scale general-purpose pretrained model contributes cross-domain knowledge through predictions, representations, transferred parameters, adaptation, or distillation.
    Stipulative boundary defined in Sec. 1 ('To make this question precise...'); the entire selection and classification in Tables 3-5 and Figs. 4-6 depend on this inclusion rule, and it excludes e.g. HaMeR's ViTPose initialization.
  • domain assumption HOI failure modes decompose into five uncertainties: shape, spatial, physical, semantic, dynamic.
    Introduced in Sec. 1 and used by Fig. 3 and all prior-family sections to explain which priors mitigate which uncertainties; the mapping loses force if the decomposition is not exhaustive or orthogonal.
  • domain assumption Each method can be assigned a unique primary sub-prior and optional auxiliary sub-priors with a single injection operator.
    Table 3 tags each method with P/A and operator; the survey's analyses in Secs. 3-5 assume these assignments are unambiguously correct, but many methods combine multiple priors and the primary/auxiliary choice is judgment-based.

pith-pipeline@v1.3.0-alltime-deepseek · 50996 in / 13436 out tokens · 145726 ms · 2026-08-04T01:18:23.076967+00:00 · methodology

0 comments
read the original abstract

Hand-object interaction (HOI) modeling remains challenging because it requires joint reasoning about hand articulation, object geometry, contact, semantics, and dynamics under severe visual uncertainty. Foundation models introduce transferable prior knowledge learned from large-scale cross-domain data, offering new ways to address these challenges beyond task-specific data and models. However, the rapidly growing literature remains fragmented, and existing studies typically describe these methods simply as ``using large models'' without systematically characterizing what knowledge is introduced, where it enters the HOI pipeline, or which HOI uncertainty it helps reduce. This survey presents the first systematic review of foundation-model priors for HOI. We organize the literature into six HOI tasks spanning reconstruction and generation. More importantly, we establish a taxonomy of eight foundation-model sub-priors grouped into geometric, semantic, and visual families. Geometric priors encompass shape retrieval, shape reconstruction, and spatial reconstruction; semantic priors include semantic grounding and language reasoning; and visual priors cover visual representation, image generation, and video generation. Based on this taxonomy, we systematically analyze how different priors are represented, injected, and adapted across HOI pipelines and tasks. Beyond how foundation models empower HOI, we further examine how HOI-derived knowledge is used in robot learning, including human-data pretraining, human-to-robot skill transfer, and HOI-to-robot data generation. Finally, we summarize datasets and evaluation protocols, and discuss limitations and future directions toward more generalizable HOI systems. To support long-term progress, we curate a live repository that continuously aggregates emerging methods and benchmarks.

Figures

Figures reproduced from arXiv: 2607.28394 by Jiaolong Yang, Junzhi Yu, Lei Zhang, Luping Xiao, Shiyang Liu, Weiquan Lin, Xingyu Chen, Xu Tang, Yu Deng.

Figure 1
Figure 1. Figure 1: Overview of this survey. The center organizes six HOI tasks into reconstruction (pose estimation, object reconstruc [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Taxonomy roadmap of this survey. Three foundation-model prior families are decomposed into section-level [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Residual HOI uncertainties and corresponding [PITH_FULL_IMAGE:figures/full_fig_p008_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Injection mechanisms of geometric priors for HOI. [PITH_FULL_IMAGE:figures/full_fig_p010_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Injection mechanisms of semantic priors for HOI. [PITH_FULL_IMAGE:figures/full_fig_p013_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Injection mechanisms of visual priors for HOI. [PITH_FULL_IMAGE:figures/full_fig_p015_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Five routes by which HOI evidence becomes robot [PITH_FULL_IMAGE:figures/full_fig_p018_7.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

292 extracted references · 66 linked inside Pith

  1. [1]

    H+O: unified egocentric recognition of 3D hand- object poses and interactions

    Bugra Tekin, Federica Bogo, and Marc Pollefeys. H+O: unified egocentric recognition of 3D hand- object poses and interactions. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2019. doi: 10.1109/CVPR.2019.00464

  2. [2]

    Learning joint reconstruction of hands and manipulated objects

    Yana Hasson, Gul Varol, Dimitrios Tzionas, Igor Kalevatykh, Michael J Black, Ivan Laptev, and Cordelia Schmid. Learning joint reconstruction of hands and manipulated objects. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2019

  3. [3]

    CPF: Learning a contact potential field to model the hand-object interaction

    Lixin Yang, Xinyu Zhan, Kailin Li, Wenqiang Xu, Jiefeng Li, and Cewu Lu. CPF: Learning a contact potential field to model the hand-object interaction. InProc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), 2021

  4. [4]

    HOLD: Category-agnostic 3D re- construction of interacting hands and objects from video

    Zicong Fan, Maria Parelli, Maria Eleni Kadoglou, Xu Chen, Muhammed Kocabas, Michael J Black, and Otmar Hilliges. HOLD: Category-agnostic 3D re- construction of interacting hands and objects from video. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2024

  5. [5]

    What’s in your hands? 3D reconstruction of generic objects in hands

    Yufei Ye, Abhinav Gupta, and Shubham Tulsiani. What’s in your hands? 3D reconstruction of generic objects in hands. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2022

  6. [6]

    Grasping field: Learning implicit representations for human grasps

    Korrawe Karunratanakul, Jinlong Yang, Yan Zhang, Michael J Black, Krikamol Muandet, and Siyu Tang. Grasping field: Learning implicit representations for human grasps. InProc. Int. Conf. 3D Vis. (3DV), 2020

  7. [7]

    DUSt3R: Geomet- ric 3D vision made easy

    Shuzhe Wang, Vincent Leroy, Yohann Cabon, Boris Chidlovskii, and Jerome Revaud. DUSt3R: Geomet- ric 3D vision made easy. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2024

  8. [8]

    Samarth Brahmbhatt, Chengcheng Tang, Christo- pher D Twigg, Charles C Kemp, and James Hays. 23 Hand-Object Interaction in the Age of Large Foundation Models: Reconstruction, Generation, and Embodied Transfer Contactpose: A dataset of grasps with object con- tact and hand pose. InProc. Eur. Conf. Comput. Vis. (ECCV), 2020

  9. [9]

    S2contact: Graph-based network for 3D hand-object contact estimation with semi-supervised learning

    Tze Ho Elden Tse, Zhongqun Zhang, Kwang In Kim, Ales Leonardis, Feng Zheng, and Hyung Jin Chang. S2contact: Graph-based network for 3D hand-object contact estimation with semi-supervised learning. In Proc. Eur. Conf. Comput. Vis. (ECCV), 2022. doi: 10.1007/978-3-031-19769-7\ 33

  10. [10]

    Contactopt: Optimizing contact to improve grasps

    Patrick Grady, Chengcheng Tang, Christopher D Twigg, Minh V o, Samarth Brahmbhatt, and Charles C Kemp. Contactopt: Optimizing contact to improve grasps. InProc. IEEE/CVF Conf. Comput. Vis. Pat- tern Recognit. (CVPR), 2021

  11. [11]

    Deep- SimHO: Stable pose estimation for hand-object in- teraction via physics simulation

    Rong Wang, Wei Mao, and Hongdong Li. Deep- SimHO: Stable pose estimation for hand-object in- teraction via physics simulation. InAdv. Neural Inf. Process. Syst. (NeurIPS), volume 36, 2023

  12. [12]

    D- grasp: Physically plausible dynamic grasp synthesis for hand-object interactions

    Sammy Christen, Muhammed Kocabas, Emre Aksan, Jemin Hwangbo, Jie Song, and Otmar Hilliges. D- grasp: Physically plausible dynamic grasp synthesis for hand-object interactions. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2022

  13. [13]

    Grounding DINO: Marrying DINO with grounded pre-training for open-set object detection

    Shilong Liu, Zhaoyang Zeng, Tianhe Ren, Feng Li, Hao Zhang, Jie Yang, Qing Jiang, Chunyuan Li, Jian- wei Yang, Hang Su, et al. Grounding DINO: Marrying DINO with grounded pre-training for open-set object detection. InProc. Eur. Conf. Comput. Vis. (ECCV), 2024

  14. [14]

    Segment anything

    Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C Berg, Wan-Yen Lo, et al. Segment anything. InProc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), 2023

  15. [15]

    SemGrasp: Semantic grasp generation via language aligned discretization

    Kailin Li, Jingbo Wang, Lixin Yang, Cewu Lu, and Bo Dai. SemGrasp: Semantic grasp generation via language aligned discretization. InProc. Eur. Conf. Comput. Vis. (ECCV), 2024

  16. [16]

    Text2grasp: Synthesis of grasps by text prompts for object grasping parts

    Xiaoyun Chang and Yi Sun. Text2grasp: Synthesis of grasps by text prompts for object grasping parts. In Proc. Int. Symp. Neural Netw. (ISNN), 2025

  17. [17]

    Towards unconstrained joint hand-object re- construction from RGB videos

    Yana Hasson, G¨ul Varol, Cordelia Schmid, and Ivan Laptev. Towards unconstrained joint hand-object re- construction from RGB videos. InProc. Int. Conf. 3D Vis. (3DV), 2021

  18. [18]

    Grasp in gaussians: Fast monocular recon- struction of dynamic hand-object interactions.arXiv preprint arXiv:2604.12929, 2026

    Ayce Idil Aytekin, Xu Chen, Zhengyang Shen, Thabo Beeler, Helge Rhodin, Rishabh Dabral, and Christian Theobalt. Grasp in gaussians: Fast monocular recon- struction of dynamic hand-object interactions.arXiv preprint arXiv:2604.12929, 2026

  19. [19]

    EasyHOI: Unleashing the power of large models for reconstructing hand-object inter- actions in the wild

    Yumeng Liu, Xiaoxiao Long, Zemin Yang, Yuan Liu, Marc Habermann, Christian Theobalt, Yuexin Ma, and Wenping Wang. EasyHOI: Unleashing the power of large models for reconstructing hand-object inter- actions in the wild. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2025

  20. [20]

    MagicHOI: Leveraging 3D priors for accurate hand-object recon- struction from short monocular video clips

    Shibo Wang, Haonan He, Maria Parelli, Christoph Gebhardt, Zicong Fan, and Jie Song. MagicHOI: Leveraging 3D priors for accurate hand-object recon- struction from short monocular video clips. InProc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), 2025

  21. [21]

    Diffusion-guided reconstruction of everyday hand-object interaction clips

    Yufei Ye, Poorvi Hebbar, Abhinav Gupta, and Shub- ham Tulsiani. Diffusion-guided reconstruction of everyday hand-object interaction clips. InProc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), 2023

  22. [22]

    Hand-object interaction image gen- eration

    Hezhen Hu, Weilun Wang, Wengang Zhou, and Houqiang Li. Hand-object interaction image gen- eration. InAdv. Neural Inf. Process. Syst. (NeurIPS), volume 35, 2022

  23. [23]

    HOIDiffusion: Generating realistic 3D hand-object interaction data

    Mengqi Zhang, Yang Fu, Zheng Ding, Sifei Liu, Zhuowen Tu, and Xiaolong Wang. HOIDiffusion: Generating realistic 3D hand-object interaction data. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recog- nit. (CVPR), 2024

  24. [24]

    Svimo: Synchro- nized diffusion for video and motion generation in hand-object interaction scenarios

    Lingwei Dang, Ruizhi Shao, Hongwen Zhang, Wei Min, Yebin Liu, and Qingyao Wu. Svimo: Synchro- nized diffusion for video and motion generation in hand-object interaction scenarios. InAdv. Neural Inf. Process. Syst. (NeurIPS), volume 38, 2025

  25. [25]

    Instantmesh: Efficient 3D mesh generation from a single image with sparse- view large reconstruction models.arXiv preprint arXiv:2404.07191, 2024

    Jiale Xu, Weihao Cheng, Yiming Gao, Xintao Wang, Shenghua Gao, and Ying Shan. Instantmesh: Efficient 3D mesh generation from a single image with sparse- view large reconstruction models.arXiv preprint arXiv:2404.07191, 2024

  26. [26]

    OpenShape: Scaling up 3D shape representation towards open-world understanding

    Minghua Liu, Ruoxi Shi, Kaiming Kuang, Yinhao Zhu, Xuanlin Li, Shizhong Han, Hong Cai, Fatih Porikli, and Hao Su. OpenShape: Scaling up 3D shape representation towards open-world understanding. In Adv. Neural Inf. Process. Syst. (NeurIPS), volume 36, 2023

  27. [27]

    Visual instruction tuning

    Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. Visual instruction tuning. InAdv. Neu- ral Inf. Process. Syst. (NeurIPS), volume 36, 2023. 24 Hand-Object Interaction in the Age of Large Foundation Models: Reconstruction, Generation, and Embodied Transfer

  28. [28]

    Qwen-VL: A versatile vision-language model for understanding, localiza- tion, text reading, and beyond.arXiv preprint arXiv:2308.12966, 2023

    Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou. Qwen-VL: A versatile vision-language model for understanding, localiza- tion, text reading, and beyond.arXiv preprint arXiv:2308.12966, 2023

  29. [29]

    DINOv2: Learning robust vi- sual features without supervision.arXiv preprint arXiv:2304.07193, 2023

    Maxime Oquab, Timoth´ee Darcet, Th´eo Moutakanni, Huy V o, Marc Szafraniec, Vasil Khalidov, Pierre Fer- nandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, et al. DINOv2: Learning robust vi- sual features without supervision.arXiv preprint arXiv:2304.07193, 2023

  30. [30]

    High- resolution image synthesis with latent diffusion mod- els

    Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bj ¨orn Ommer. High- resolution image synthesis with latent diffusion mod- els. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2022

  31. [31]

    Cogvideox: Text-to-video diffusion models with an expert transformer

    Zhuoyi Yang, Jiayan Teng, Wendi Zheng, Ming Ding, Shiyu Huang, Jiazheng Xu, Yuanming Yang, Wenyi Hong, Xiaohan Zhang, Guanyu Feng, et al. Cogvideox: Text-to-video diffusion models with an expert transformer. InProc. Int. Conf. Learn. Repre- sent. (ICLR), 2025

  32. [32]

    Objaverse: A universe of annotated 3D objects

    Matt Deitke, Dustin Schwenk, Jordi Salvador, Luca Weihs, Oscar Michel, Eli VanderBilt, Ludwig Schmidt, Kiana Ehsani, Aniruddha Kembhavi, and Ali Farhadi. Objaverse: A universe of annotated 3D objects. InProc. IEEE/CVF Conf. Comput. Vis. Pat- tern Recognit. (CVPR), 2023

  33. [33]

    Learning trans- ferable visual models from natural language supervi- sion

    Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sas- try, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning trans- ferable visual models from natural language supervi- sion. InProc. Int. Conf. Mach. Learn. (ICML), 2021

  34. [34]

    Reconstructing hand-held objects in 3D from images and videos

    Jane Wu, Georgios Pavlakos, Georgia Gkioxari, and Jitendra Malik. Reconstructing hand-held objects in 3D from images and videos. InProc. Int. Conf. 3D Vis. (3DV), 2026

  35. [35]

    Ghost: Fast category- agnostic hand-object interaction reconstruction from RGB videos using gaussian splatting

    Ahmed Tawfik Aboukhadra, Marcel Rogge, Nadia Robertini, Abdalla Arafa, Jameel Malik, Ahmed El- hayek, and Didier Stricker. Ghost: Fast category- agnostic hand-object interaction reconstruction from RGB videos using gaussian splatting. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2026

  36. [36]

    Hand-held object reconstruction from RGB video with dynamic interaction

    Shijian Jiang, Qi Ye, Rengan Xie, Yuchi Huo, and Jim- ing Chen. Hand-held object reconstruction from RGB video with dynamic interaction. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2025

  37. [37]

    Reconstructing hands in 3D with transformers

    Georgios Pavlakos, Dandan Shan, Ilija Radosavovic, Angjoo Kanazawa, David Fouhey, and Jitendra Malik. Reconstructing hands in 3D with transformers. In Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2024

  38. [38]

    Wilor: End- to-end 3D hand localization and reconstruction in-the- wild

    Rolandos Alexandros Potamias, Jinglei Zhang, Jiankang Deng, and Stefanos Zafeiriou. Wilor: End- to-end 3D hand localization and reconstruction in-the- wild. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2025

  39. [39]

    ViTPose: Simple vision transformer baselines for human pose estimation

    Yufei Xu, Jing Zhang, Qiming Zhang, and Dacheng Tao. ViTPose: Simple vision transformer baselines for human pose estimation. InAdv. Neural Inf. Process. Syst. (NeurIPS), 2022

  40. [40]

    Latent action pretraining from videos

    Seonghyeon Ye, Joel Jang, Byeongguk Jeon, Se June Joo, Jianwei Yang, Baolin Peng, Ajay Mandlekar, Reuben Tan, Yu-Wei Chao, Bill Yuchen Lin, et al. Latent action pretraining from videos. InProc. Int. Conf. Learn. Represent. (ICLR), 2025

  41. [41]

    Egovla: Learning vision-language-action models from egocentric hu- man videos.arXiv preprint arXiv:2507.12440, 2025

    Ruihan Yang, Qinxi Yu, Yecheng Wu, Rui Yan, Borui Li, An-Chieh Cheng, Xueyan Zou, Yunhao Fang, Xuxin Cheng, Ri-Zhao Qiu, et al. Egovla: Learning vision-language-action models from egocentric hu- man videos.arXiv preprint arXiv:2507.12440, 2025

  42. [42]

    Dexmv: Imitation learning for dexterous manipula- tion from human videos

    Yuzhe Qin, Yueh-Hua Wu, Shaowei Liu, Hanwen Jiang, Ruihan Yang, Yang Fu, and Xiaolong Wang. Dexmv: Imitation learning for dexterous manipula- tion from human videos. InProc. Eur. Conf. Comput. Vis. (ECCV), 2022

  43. [43]

    Egoengine: From ego- centric human videos to high-fidelity dexterous robot demonstrations.arXiv preprint arXiv:2606.12604, 2026

    Yangcen Liu, Shuo Cheng, Xinchen Yin, Woo Chul Shin, Alfred Cueva, Yiran Yang, Zhenyang Chen, Chuye Zhang, and Danfei Xu. Egoengine: From ego- centric human videos to high-fidelity dexterous robot demonstrations.arXiv preprint arXiv:2606.12604, 2026

  44. [44]

    Efficient annotation and learning for 3D hand pose estimation: A survey.Int

    Takehiko Ohkawa, Ryosuke Furuta, and Yoichi Sato. Efficient annotation and learning for 3D hand pose estimation: A survey.Int. J. Comput. Vis., 131(12), 2023

  45. [45]

    A survey of deep learning methods and datasets for hand pose estimation from hand-object interaction images.Comput

    Taeyun Woo, Wonjung Park, Woohyun Jeong, and Jinah Park. A survey of deep learning methods and datasets for hand pose estimation from hand-object interaction images.Comput. Graph., 116, 2023

  46. [46]

    Advances in vision-based deep learning methods for interacting hands reconstruction: A survey.Comput

    Yu Miao and Yue Liu. Advances in vision-based deep learning methods for interacting hands reconstruction: A survey.Comput. Graph., 124, 2024. 25 Hand-Object Interaction in the Age of Large Foundation Models: Reconstruction, Generation, and Embodied Transfer

  47. [47]

    An overview of learning-based dexterous grasping: recent advances and future directions.Artif

    Xu Song, Yongyao Li, Yunfan Zhang, Yufei Liu, and Lei Jiang. An overview of learning-based dexterous grasping: recent advances and future directions.Artif. Intell. Rev., 58(10), 2025

  48. [48]

    Arthoi: Taming foundation models for monocular 4D reconstruction of hand-articulated- object interactions

    Zikai Wang, Zhilu Zhang, Yiqing Wang, Hui Li, and Wangmeng Zuo. Arthoi: Taming foundation models for monocular 4D reconstruction of hand-articulated- object interactions. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2026

  49. [49]

    Javier Romero, Dimitrios Tzionas, and Michael J. Black. Embodied hands: modeling and capturing hands and bodies together.ACM Trans. Graph., 36 (6), 2017. doi: 10.1145/3130800.3130883

  50. [50]

    Nimble: a non-rigid hand model with bones and muscles.ACM Trans

    Yuwei Li, Longwen Zhang, Zesong Qiu, Yingwenqi Jiang, Nianyi Li, Yuexin Ma, Yuyao Zhang, Lan Xu, and Jingyi Yu. Nimble: a non-rigid hand model with bones and muscles.ACM Trans. Graph., 41(4), 2022

  51. [51]

    HOnnotate: A method for 3D an- notation of hand and object poses

    Shreyas Hampali, Mahdi Rad, Markus Oberweger, and Vincent Lepetit. HOnnotate: A method for 3D an- notation of hand and object poses. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2020

  52. [52]

    AlignSDF: Pose-aligned signed distance fields for hand-object reconstruction

    Zerui Chen, Yana Hasson, Cordelia Schmid, and Ivan Laptev. AlignSDF: Pose-aligned signed distance fields for hand-object reconstruction. InProc. Eur. Conf. Comput. Vis. (ECCV), 2022

  53. [53]

    HandNeRF: Learn- ing to reconstruct hand-object interaction scene from a single RGB image

    Hongsuk Choi, Nikhil Chavan-Dafle, Jiacheng Yuan, V olkan Isler, and Hyunsoo Park. HandNeRF: Learn- ing to reconstruct hand-object interaction scene from a single RGB image. InProc. IEEE Int. Conf. Robot. Autom. (ICRA), 2024

  54. [54]

    Simultaneous localization and mapping: part i.IEEE Robot

    Hugh Durrant-Whyte and Tim Bailey. Simultaneous localization and mapping: part i.IEEE Robot. Autom. Mag., 13(2), 2006

  55. [55]

    Seitz, and Richard Szeliski

    Noah Snavely, Steven M. Seitz, and Richard Szeliski. Photo tourism: Exploring photo collections in 3D. In ACM SIGGRAPH 2006 Papers. ACM, 2006

  56. [56]

    LatentHOI: On the generalizable hand object motion generation with latent hand diffusion

    Muchen Li, Sammy Christen, Chengde Wan, Yu- jun Cai, Renjie Liao, Leonid Sigal, and Shugao Ma. LatentHOI: On the generalizable hand object motion generation with latent hand diffusion. In Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2025

  57. [57]

    3D hand pose estimation in everyday egocentric images

    Aditya Prakash, Ruisen Tu, Matthew Chang, and Saurabh Gupta. 3D hand pose estimation in everyday egocentric images. InProc. Eur. Conf. Comput. Vis. (ECCV), 2024

  58. [58]

    DDF-HO: Hand-held object recon- struction via conditional directed distance field

    Chenyangguang Zhang, Yan Di, Ruida Zhang, Guangyao Zhai, Fabian Manhardt, Federico Tombari, and Xiangyang Ji. DDF-HO: Hand-held object recon- struction via conditional directed distance field. In Adv. Neural Inf. Process. Syst. (NeurIPS), volume 36, 2023

  59. [59]

    Ncrf: neural contact radiance fields for free-viewpoint rendering of hand-object interaction

    Zhongqun Zhang, Jifei Song, Eduardo P´erez-Pellitero, Yiren Zhou, Hyung Jin Chang, and Ale ˇs Leonardis. Ncrf: neural contact radiance fields for free-viewpoint rendering of hand-object interaction. InProc. Int. Conf. 3D Vis. (3DV), 2024

  60. [60]

    Gan- hand: Predicting human grasp affordances in multi- object scenes

    Enric Corona, Albert Pumarola, Guillem Alenya, Francesc Moreno-Noguer, and Gr´egory Rogez. Gan- hand: Predicting human grasp affordances in multi- object scenes. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2020

  61. [61]

    Artigrasp: Physically plausible synthesis of bi-manual dexterous grasping and articulation

    Hui Zhang, Sammy Christen, Zicong Fan, Luocheng Zheng, Jemin Hwangbo, Jie Song, and Otmar Hilliges. Artigrasp: Physically plausible synthesis of bi-manual dexterous grasping and articulation. InProc. Int. Conf. 3D Vis. (3DV), 2024

  62. [62]

    Deepsdf: Learning continuous signed distance functions for shape representation

    Jeong Joon Park, Peter Florence, Julian Straub, Richard Newcombe, and Steven Lovegrove. Deepsdf: Learning continuous signed distance functions for shape representation. InProc. IEEE/CVF Conf. Com- put. Vis. Pattern Recognit. (CVPR), 2019

  63. [63]

    A skeleton-driven neural occupancy representation for articulated hands

    Korrawe Karunratanakul, Adrian Spurr, Zicong Fan, Otmar Hilliges, and Siyu Tang. A skeleton-driven neural occupancy representation for articulated hands. InProc. Int. Conf. 3D Vis. (3DV), 2021

  64. [64]

    Srinivasan, Matthew Tancik, Jonathan T

    Ben Mildenhall, Pratul P. Srinivasan, Matthew Tancik, Jonathan T. Barron, Ravi Ramamoorthi, and Ren Ng. NeRF: Representing scenes as neural radiance fields for view synthesis. InProc. Eur. Conf. Comput. Vis. (ECCV), 2020. doi: 10.1007/978-3-030-58452-8 \ 24

  65. [65]

    3D gaussian splat- ting for real-time radiance field rendering.ACM Trans

    Bernhard Kerbl, Georgios Kopanas, Thomas Leimk¨uhler, and George Drettakis. 3D gaussian splat- ting for real-time radiance field rendering.ACM Trans. Graph., 42(4), 2023

  66. [66]

    Affordpose: A large-scale dataset of hand- object interactions with affordance-driven hand pose

    Juntao Jian, Xiuping Liu, Manyi Li, Ruizhen Hu, and Jian Liu. Affordpose: A large-scale dataset of hand- object interactions with affordance-driven hand pose. InProc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), 2023

  67. [67]

    Th´eo Morales, Omid Taheri, and Gerard Lacey. A versatile and differentiable hand-object interaction 26 Hand-Object Interaction in the Age of Large Foundation Models: Reconstruction, Generation, and Embodied Transfer representation. InProc. IEEE/CVF Winter Conf. Appl. Comput. Vis. (WACV), 2025

  68. [68]

    Black, and Dima Damen

    Siddhant Bansal, Zhifan Zhu, Shashank Tripathi, Ji- ahe Zhao, Michael J. Black, and Dima Damen. To- wards in-the-wild egocentric 3D hand-object pose estimation. InProc. Eur. Conf. Comput. Vis. (ECCV), 2026

  69. [69]

    HOI4D: A 4D egocentric dataset for category-level human-object interaction

    Yunze Liu, Yun Liu, Che Jiang, Kangbo Lyu, Weikang Wan, Hao Shen, Boqiang Liang, Zhoujie Fu, He Wang, and Li Yi. HOI4D: A 4D egocentric dataset for category-level human-object interaction. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recog- nit. (CVPR), 2022

  70. [70]

    Arctic: A dataset for dexterous bimanual hand-object manipulation

    Zicong Fan, Omid Taheri, Dimitrios Tzionas, Muhammed Kocabas, Manuel Kaufmann, Michael J Black, and Otmar Hilliges. Arctic: A dataset for dexterous bimanual hand-object manipulation. In Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2023

  71. [71]

    Handoccnet: Occlusion-robust 3D hand mesh estimation network

    JoonKyu Park, Yeonguk Oh, Gyeongsik Moon, Hong- suk Choi, and Kyoung Mu Lee. Handoccnet: Occlusion-robust 3D hand mesh estimation network. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recog- nit. (CVPR), 2022

  72. [72]

    Mobrecon: Mobile-friendly hand mesh reconstruction from monocular image

    Xingyu Chen, Yufeng Liu, Yajiao Dong, Xiong Zhang, Chongyang Ma, Yanmin Xiong, Yuan Zhang, and Xiaoyan Guo. Mobrecon: Mobile-friendly hand mesh reconstruction from monocular image. In Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2022

  73. [73]

    A simple baseline for ef- ficient hand mesh reconstruction

    Zhishan Zhou, Shihao Zhou, Zhi Lv, Minqiang Zou, Yao Tang, and Jiajun Liang. A simple baseline for ef- ficient hand mesh reconstruction. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2024

  74. [74]

    Model-based 3D hand reconstruction via self- supervised learning

    Yujin Chen, Zhigang Tu, Di Kang, Linchao Bao, Ying Zhang, Xuefei Zhe, Ruizhi Chen, and Junsong Yuan. Model-based 3D hand reconstruction via self- supervised learning. InProc. IEEE/CVF Conf. Com- put. Vis. Pattern Recognit. (CVPR), 2021

  75. [75]

    Keypoint fusion for RGB-D based 3D hand pose estimation

    Xingyu Liu, Pengfei Ren, Yuanyuan Gao, Jingyu Wang, Haifeng Sun, Qi Qi, Zirui Zhuang, and Jianxin Liao. Keypoint fusion for RGB-D based 3D hand pose estimation. InProc. AAAI Conf. Artif. Intell., 2024

  76. [76]

    Hope-net: A graph-based model for hand-object pose estimation

    Bardia Doosti, Shujon Naha, Majid Mirbagheri, and David J Crandall. Hope-net: A graph-based model for hand-object pose estimation. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2020

  77. [77]

    Keypoint transformer: Solv- ing joint identification in challenging hands and ob- ject interactions for accurate 3D pose estimation

    Shreyas Hampali, Sayan Deb Sarkar, Mahdi Rad, and Vincent Lepetit. Keypoint transformer: Solv- ing joint identification in challenging hands and ob- ject interactions for accurate 3D pose estimation. In Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2022

  78. [78]

    Thor-net: End-to-end graformer-based realistic two hands and object reconstruction with self-supervision

    Ahmed Tawfik Aboukhadra, Jameel Malik, Ahmed El- hayek, Nadia Robertini, and Didier Stricker. Thor-net: End-to-end graformer-based realistic two hands and object reconstruction with self-supervision. InProc. IEEE/CVF Winter Conf. Appl. Comput. Vis. (WACV), 2023

  79. [79]

    HOISDF: Constraining 3D hand- object pose estimation with global signed distance fields

    Haozhe Qi, Chen Zhao, Mathieu Salzmann, and Alexander Mathis. HOISDF: Constraining 3D hand- object pose estimation with global signed distance fields. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2024

  80. [80]

    Contactart: Learning 3D interaction priors for category-level ar- ticulated object and hand poses estimation

    Zehao Zhu, Jiashun Wang, Yuzhe Qin, Deqing Sun, Varun Jampani, and Xiaolong Wang. Contactart: Learning 3D interaction priors for category-level ar- ticulated object and hand poses estimation. InProc. Int. Conf. 3D Vis. (3DV), 2024

Showing first 80 references.