REVIEW 5 major objections 5 minor 113 references
FoundHand: Large-Scale Domain-Specific Learning for Controllable Hand Image Generation
T0 review · 5 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read A diffusion model trained on 10 million hand images makes pose-controllable hand generation state of the art across six tasks.
desk verdict A useful hand dataset and multi-task generation model whose flagship gesture-transfer claim rests on a same-pose reconstruction metric that does not measure pose transfer. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is pair-based training with spatially aligned multi-modal conditioning. The model is a latent diffusion vision transformer (DiT) pretrained on ImageNet; the reference and target frames are encoded with a Stable Diffusion VAE, and the reference image latents, keypoint heatmaps, and hand masks are aligned and passed through a shared embedder, with 3D self-attention between the two frames. Target heatmaps are Gaussian heatmaps over 42 keypoints, 21 per hand, with zero channels for absent hands, so the model can handle single or dual hands and can be conditioned without masks at test time. A binary flag tells the model whether the frame pair comes from synchronized views (viewpoint change) or a temporal sequence (pose change), and stochastic conditioning on previously generated views is used at inference for novel view synthesis and video generation.
What would settle it
Take a single hand image with known multi-view ground truth, estimate 3D joints, project them into a target camera far outside the small cone used in the paper (for example, a 90-degree rotation), generate the view with FoundHand, and measure the error between the generated image's detected 2D keypoints and the ground-truth projected keypoints. If the error grows sharply with baseline or the generated hand becomes anatomically implausible, the sufficiency of 2D keypoints for viewpoint control is disproven.
Extended reading notes
Core claim
FoundHand's central discovery is that a diffusion model trained on a large, diverse collection of hand image pairs can treat 2D keypoint heatmaps as a universal representation for hand articulation and camera viewpoint. During training, the model sees pairs of frames, either consecutive frames from hand videos or synchronized views from multi-view captures, and learns to map a reference image plus target heatmaps to the target image. At inference, the same conditioning enables reposing, where reference appearance is preserved while the keypoints change; domain transfer, where the keypoints stay fixed but the reference drives style; and novel view synthesis, where 3D joints estimated from a single view are projected into new cameras and the model generates consistent views without explicit camera parameters. The paper reports state-of-the-art quantitative results on gesture transfer, domain-transfer aided 3D hand estimation, novel view synthesis on InterHand2.6M, and zero-shot video and hand-object interaction synthesis, outperforming task-specific baselines.
Load-bearing premise
The load-bearing premise is that 2D keypoints encode enough information about camera viewpoint and depth for the model to synthesize correct novel views, so a large-baseline viewpoint change or a pose with strong depth ambiguity could break the claim that no explicit camera parameters are needed.
Editorial extensions
If this is right
- Hand pose editing becomes a simple keypoint-painting interface: users specify target 2D keypoints and the model preserves appearance while articulating the hand.
- Domain transfer can be used as a data-augmentation tool: fine-tuning a 3D hand estimator on FoundHand-transferred images improves mesh recovery metrics on a target domain.
- Novel views can be generated from a single hand image without explicit 3D supervision, outperforming NeRF-based zero-shot view synthesis baselines on InterHand2.6M.
- Malformed hands produced by text-to-image models can be repaired zero-shot via Repaint-style inpainting conditioned only on keypoints.
- Hand and hand-object videos can be synthesized by conditioning each frame on the first frame plus recent generated frames, with no video-specific training.
Reading between the lines
- Because the viewpoint sampling is restricted to a small cone around the reference camera, the paper's novel-view evidence does not establish that 2D keypoints fully resolve large-baseline depth ambiguity; a testable extension is to quantify the view cone angle at which reprojection error degrades.
- The automatic MediaPipe and SAM annotations receive no reported quality-control filtering, so the dataset may contain noisy keypoint and mask labels; a useful extension would be to measure how generation quality changes when a subset is cleaned or human-verified.
- The same pair-based conditioning could be applied to other articulated structures, such as faces or bodies, where 2D landmarks plus a reference image might support reposing, style transfer, and view synthesis without 3D supervision.
- The stochastic conditioning used for video could be turned into a simple consistency metric: freeze the first frame and measure pixel-level drift across long generated sequences to test whether the model maintains long-term identity.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces FoundHand-10M, a dataset of over 10M hand images assembled from 12 existing video and multi-view datasets, with automatic MediaPipe 2D keypoints and SAM segmentation masks. It then presents FoundHand, a latent diffusion transformer that performs image-to-image translation conditioned on 2D keypoint heatmaps, an optional reference image, and a binary flag indicating whether the input pair is a temporal pose change or a synchronized viewpoint change. The claimed core capabilities are gesture transfer, domain transfer, and novel view synthesis, with zero-shot applications for fixing malformed hands, synthesizing motion-controlled video, and generating hand-object interaction video. The paper reports state-of-the-art quantitative results across six hand-related tasks and emphasizes that 2D keypoints encode both articulation and viewpoint, avoiding explicit camera parameters.
Significance. If the central claims are substantiated, the contribution is significant: a large unified hand dataset with a common annotation convention, a controllable two-frame diffusion formulation that avoids full video training, and evidence that 2D-keypoint conditioning can drive pose and appearance changes across diverse in-the-wild inputs. The strengths include the scale of the dataset (10M images from 12 sources), the tractable two-frame 3D self-attention design, and the breadth of qualitative demonstrations from photorealistic images to artistic styles and hand-object interactions. However, the quantitative support for the state-of-the-art claim is thin in several places, and the load-bearing evaluation of gesture transfer and the camera-parameter claim need revision before the results can be considered established.
major comments (5)
- [Section 5, Table 1] The quantitative protocol for gesture transfer sets the target pose equal to the reference pose and reports PSNR/SSIM/LPIPS/FID between the generated image and the reference. This measures pose-preserving reconstruction—essentially an identity or autoencoding test—rather than the ability to adopt a different target pose, and a model that copies the reference can score perfectly. Since gesture transfer is a core capability and one of the six tasks supporting the SOTA claim, the comparison needs a metric on genuinely different target poses, such as PCK/OKS of the generated keypoints against the requested target keypoints, or paired data with ground-truth target images.
- [Section 5, Table 1] The text reports FID "between reference and generated images." FID is defined for distributions, not image pairs; without a precise description of which image sets form the two distributions and how the same-pose reconstruction setup is used, the FID numbers cannot be interpreted or compared across methods. The authors should specify the reference and generated distributions, sample sizes, and whether the metric is computed across the full test set or on a per-image basis.
- [Section 4.3, Novel View Synthesis] The claim that 2D keypoints "eliminate the need for explicit camera parameters during training or inference" contradicts the described pipeline, which explicitly assumes a camera projection K to lift reference keypoints to 3D joints J via an off-the-shelf estimator and then projects J into target cameras. Thus camera parameters are assumed and used at inference time. The claim should be revised to what is actually true—for example, that camera parameters are not provided as network inputs during training—or the pipeline description must change. This matters because the NVS results are cited as evidence for the versatility of the 2D-keypoint representation.
- [Section 4.1, FoundHand-10M] The paper gives no annotation-quality control for the MediaPipe and SAM labels on the 10M images. Because the model conditions on these 2D keypoints and masks, annotation errors directly bound the achievable pose fidelity and the claimed "precise control." The authors should report a validation statistic on a subset (e.g., agreement with manual labels or a downstream pose-estimation metric) and describe any filtering steps used to remove low-quality annotations.
- [Section 5, Table 3(b)] The zero-shot video synthesis comparison is based on 12 in-the-wild videos with no standard errors, no significance tests, and no description of how the reported metrics are aggregated. Given the paper's claim of outperforming task-specific video baselines, this quantitative evidence is too thin; the authors should provide per-video results or a larger evaluation, and clarify whether the metric is computed on generated frames only.
minor comments (5)
- [Section 7] The dataset name is inconsistent between "FoundHand-10M" (Sections 1 and 4.1) and "FoundHand10M" (Section 7); please unify the spelling.
- [Figure 3] The caption reads "Ours w/ skeleton Ours," which appears to be a typo; the intended comparison is unclear.
- [Table 1] The caption describes the metrics as "identity consistency metrics," but FID and LPIPS are perceptual or distribution-level metrics; consider a more accurate label.
- [Table 3] Table 3 lists "Zero123 [44]" while reference [44] is titled "Zero-1-to-3"; please ensure the citation label matches the method being compared.
- [Section 3] The classifier-free guidance equation uses the convention ϵ̂θ = wϵθ(zτ;τ,c)+(1−w)ϵθ(zτ;τ,∅); the text should state the value of w used in the experiments and any dependence on the binary flag y.
Circularity Check
No circular derivation: FoundHand's claims are empirical and externally grounded; the same-pose gesture-transfer protocol is an evaluation weakness, not a circular step.
full rationale
FoundHand is an empirical systems paper. The dataset is assembled from twelve external datasets with MediaPipe and SAM annotations; the diffusion model is trained on an L2 denoising objective; and the headline results are comparisons against external baselines on external benchmarks. The statement that 2D keypoints encode articulation and viewpoint is a representation premise, not a derived result, and the NVS pipeline in Section 4.3 explicitly estimates 3D joints and projects them into target cameras, so the claim about eliminating camera parameters is an overstatement rather than a circular reduction. Self-citations (Genheld, DIVA-360, Manus, GeoDiffuser) appear only in related work and do not carry the central argument. The gesture-transfer quantitative protocol in Section 5, which sets the target pose equal to the reference pose and measures reconstruction quality, is self-referential as evidence for pose transfer, since a model that copies the reference would score well. However, this is an evaluation-design limitation, not circular reasoning: no fitted parameter is renamed as a prediction, and no equation or derivation reduces the claimed capability to its input by construction. The paper's central claims therefore do not exhibit circularity.
Assumptions & free parameters
free parameters (3)
- Classifier-free guidance scale w =
not reported
- Binary flag y for pose vs view transformation =
not reported
- Condition dropout probability =
not reported
assumptions (4)
- domain assumption 2D keypoints are projections of 3D hand joints, so they encode both articulation and camera viewpoint.
- domain assumption MediaPipe and SAM annotations on 10M images are accurate enough to serve as training signal.
- domain assumption Image pairs from video sequences and synchronized multiviews provide sufficient supervision to learn physically plausible pose and view transformations.
- domain assumption Horizontal flipping and hand swapping augmentation teach correct handedness transformation.
Cite this review
Pith. "Pith review of FoundHand: Large-Scale Domain-Specific Learning for Controllable Hand Image Generation." pith.science (2026). https://pith.science/paper/TWWF7BBQ
@misc{pith2026241202690,
author = {Pith},
title = {Pith review of: FoundHand: Large-Scale Domain-Specific Learning for Controllable Hand Image Generation},
year = {2026},
howpublished = {\url{https://pith.science/paper/TWWF7BBQ}},
note = {Machine review of arXiv:2412.02690}
}
read the original abstract
Despite remarkable progress in image generation models, generating realistic hands remains a persistent challenge due to their complex articulation, varying viewpoints, and frequent occlusions. We present FoundHand, a large-scale domain-specific diffusion model for synthesizing single and dual hand images. To train our model, we introduce FoundHand-10M, a large-scale hand dataset with 2D keypoints and segmentation mask annotations. Our insight is to use 2D hand keypoints as a universal representation that encodes both hand articulation and camera viewpoint. FoundHand learns from image pairs to capture physically plausible hand articulations, natively enables precise control through 2D keypoints, and supports appearance control. Our model exhibits core capabilities that include the ability to repose hands, transfer hand appearance, and even synthesize novel views. This leads to zero-shot capabilities for fixing malformed hands in previously generated images, or synthesizing hand video sequences. We present extensive experiments and evaluations that demonstrate state-of-the-art performance of our method.
Figures
Figures from the paper (5 more)
Reference graph
Works this paper leans on
-
[1]
https : / / www
Midjourney. https : / / www . midjourney . com / home. (Accessed on 11/12/2023). 2
2023
-
[2]
https : //www.youtube.com/watch?v=24yjRbBah3w ,
Why ai art struggles with hands - youtube. https : //www.youtube.com/watch?v=24yjRbBah3w , . (Accessed on 11/12/2023). 2
2023
-
[3]
https : //www.reddit.com/r/AskReddit/comments/ y8oe9l/why_ai_cant_draw_hands/ ,
Why ai can’t draw hands? : r/askreddit. https : //www.reddit.com/r/AskReddit/comments/ y8oe9l/why_ai_cant_draw_hands/ , . (Accessed on 11/12/2023). 2
2023
-
[4]
Pose with Style: Detail-Preserving Pose-Guided Image Synthesis with Conditional StyleGAN
Badour Albahar, Jingwan Lu, Jimei Yang, Zhixin Shu, Eli Shechtman, and Jia-Bin Huang. Pose with style: Detail- preserving pose-guided image synthesis with conditional stylegan. ArXiv, abs/2109.06166, 2021. 2
work page Pith review arXiv 2021
-
[5]
Renderdiffusion: Image diffusion for 3d reconstruction, in- painting and generation
Titas Anciukevi ˇcius, Zexiang Xu, Matthew Fisher, Paul Henderson, Hakan Bilen, Niloy J Mitra, and Paul Guerrero. Renderdiffusion: Image diffusion for 3d reconstruction, in- painting and generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 12608–12618, 2023. 5
2023
-
[6]
James Burgess, Kuan-Chieh Wang, and Serena Yeung. Viewpoint textual inversion: Unleashing novel view syn- thesis with pretrained 2d diffusion models. arXiv preprint arXiv:2309.07986, 2023. 5
work page Pith review arXiv 2023
-
[7]
Z. Cao, G. Hidalgo Martinez, T. Simon, S. Wei, and Y . A. Sheikh. Openpose: Realtime multi-person 2d pose estima- tion using part affinity fields. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2019. 3
2019
-
[8]
Efficient geometry-aware 3d generative adversarial networks
Eric R Chan, Connor Z Lin, Matthew A Chan, Koki Nagano, Boxiao Pan, Shalini De Mello, Orazio Gallo, Leonidas J Guibas, Jonathan Tremblay, Sameh Khamis, et al. Efficient geometry-aware 3d generative adversarial networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 16123– 16133, 2022. 2
2022
Show all 113 references
-
[9]
Chan, Koki Nagano, Matthew A
Eric R. Chan, Koki Nagano, Matthew A. Chan, Alexan- der W. Bergman, Jeong Joon Park, Axel Levy, Miika Ait- tala, Shalini De Mello, Tero Karras, and Gordon Wetzstein. GeNVS: Generative novel view synthesis with 3D-aware diffusion models. In arXiv, 2023. 5
2023
-
[10]
Narang, Karl Van Wyk, Umar Iqbal, Stan Birchfield, Jan Kautz, and Di- eter Fox
Yu-Wei Chao, Wei Yang, Yu Xiang, Pavlo Molchanov, Ankur Handa, Jonathan Tremblay, Yashraj S. Narang, Karl Van Wyk, Umar Iqbal, Stan Birchfield, Jan Kautz, and Di- eter Fox. DexYCB: A benchmark for capturing hand grasp- ing of objects. In IEEE/CVF Conference on Computer Vi- sio...
2021
-
[11]
Unpaired pose guided human image generation
Xu Chen, Jie Song, and Otmar Hilliges. Unpaired pose guided human image generation. In The IEEE Confer- ence on Computer Vision and Pattern Recognition (CVPR) Workshops, 2019. 2, 3
2019
-
[12]
4drecons: 4d neural implicit deformable objects reconstruction from a single rgb-d camera with geometrical and topological regu- larizations
Xiaoyan Cong, Haitao Yang, Liyan Chen, Kaifeng Zhang, Li Yi, Chandrajit Bajaj, and Qixing Huang. 4drecons: 4d neural implicit deformable objects reconstruction from a single rgb-d camera with geometrical and topological regu- larizations. arXiv preprint arXiv:2406.10167, 2024. 2
2024 arXiv
-
[13]
Rescaling egocentric vision: Collection, pipeline and chal- lenges for epic-kitchens-100
Dima Damen, Hazel Doughty, Giovanni Maria Farinella, Antonino Furnari, Evangelos Kazakos, Jian Ma, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, et al. Rescaling egocentric vision: Collection, pipeline and chal- lenges for epic-kitchens-100. International Journa...
2022
-
[14]
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In 2009 IEEE Conference on Computer Vision and Pattern Recognition, pages 248–255, 2009. 2, 3
2009
-
[15]
Disentangled and controllable face image genera- tion via 3d imitative-contrastive learning
Yu Deng, Jiaolong Yang, Dong Chen, Fang Wen, and Xin Tong. Disentangled and controllable face image genera- tion via 3d imitative-contrastive learning. 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 5153–5162, 2020. 2, 3
2020
-
[16]
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2018. 2
2018 arXiv
-
[17]
Diffusion models beat gans on image synthesis
Prafulla Dhariwal and Alexander Nichol. Diffusion models beat gans on image synthesis. Advances in neural informa- tion processing systems, 34:8780–8794, 2021. 2
2021
-
[18]
Scalable pre-training of large autoregressive image models, 2024
Alaaeldin El-Nouby, Michal Klein, Shuangfei Zhai, Miguel Angel Bautista, Alexander Toshev, Vaishaal Shankar, Joshua M Susskind, and Armand Joulin. Scalable pre-training of large autoregressive image models, 2024. 2
2024
-
[19]
Black, and Otmar Hilliges
Zicong Fan, Omid Taheri, Dimitrios Tzionas, Muhammed Kocabas, Manuel Kaufmann, Michael J. Black, and Otmar Hilliges. ARCTIC: A dataset for dexterous bimanual hand- object manipulation. In Proceedings IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2023. 2, 3
2023
-
[20]
Signllm: Sign languages production large language models
Sen Fang, Lei Wang, Ce Zheng, Yapeng Tian, and Chen Chen. Signllm: Sign languages production large language models. arXiv preprint arXiv:2405.10718, 2024. 3
2024 arXiv
-
[21]
Dreamoving: A human dance video genera- tion framework based on diffusion models
Mengyang Feng, Jinlin Liu, Kai Yu, Yuan Yao, Zheng Hui, Xiefan Guo, Xianhui Lin, Haolan Xue, Chen Shi, Xiaowen Li, et al. Dreamoving: A human dance video genera- tion framework based on diffusion models. arXiv preprint arXiv:2312.05107, 2023. 3
2023 arXiv
-
[22]
Dart: Articulated hand model with diverse accessories and rich textures
Daiheng Gao, Yuliang Xiu, Kailin Li, Lixin Yang, Feng Wang, Peng Zhang, Bang Zhang, Cewu Lu, and Ping Tan. Dart: Articulated hand model with diverse accessories and rich textures. Advances in Neural Information Processing Systems, 35:37055–37067, 2022. 2, 3
2022
-
[23]
Vivid-1-to-3: Novel view synthesis with video diffusion models, 2023
Jeong gi Kwak, Erqun Dong, Yuhe Jin, Hanseok Ko, Shweta Mahajan, and Kwang Moo Yi. Vivid-1-to-3: Novel view synthesis with video diffusion models, 2023. 5
2023
-
[24]
Generative adversarial nets
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. Generative adversarial nets. In Ad- vances in Neural Information Processing Systems, 2014. 2
2014
-
[25]
Ego4d: Around the world in 3,000 hours of egocentric video
Kristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis, Antonino Furnari, Rohit Girdhar, Jackson Hamburger, Hao Jiang, Miao Liu, Xingyu Liu, et al. Ego4d: Around the world in 3,000 hours of egocentric video. In Proceedings of the IEEE/CVF Conference on Computer Vi- sio...
2022
-
[26]
Effects of hand representations for typing in virtual reality
Jens Grubert, Lukas Witzani, Eyal Ofek, Michel Pahud, Matthias Kranz, and Per Ola Kristensson. Effects of hand representations for typing in virtual reality. In 2018 IEEE Conference on Virtual Reality and 3D User Interfaces (VR), pages 151–158. IEEE, 2018. 2
2018
-
[27]
Sparsectrl: Adding sparse controls to text-to-video diffusion models, 2023
Yuwei Guo, Ceyuan Yang, Anyi Rao, Maneesh Agrawala, Dahua Lin, and Bo Dai. Sparsectrl: Adding sparse controls to text-to-video diffusion models, 2023. 2
2023
-
[28]
Classifier-free diffusion guidance, 2022
Jonathan Ho and Tim Salimans. Classifier-free diffusion guidance, 2022. 3, 4
2022
-
[29]
Denoising dif- fusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising dif- fusion probabilistic models. Advances in neural informa- tion processing systems, 33:6840–6851, 2020. 2, 3
2020
-
[30]
Model-aware gesture-to-gesture transla- tion
Hezhen Hu, Weilun Wang, Wen gang Zhou, Weichao Zhao, and Houqiang Li. Model-aware gesture-to-gesture transla- tion. 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 16423–16432, 2021. 2
2021
-
[31]
Animate anyone: Consistent and controllable image-to-video synthesis for character animation
Li Hu, Xin Gao, Peng Zhang, Ke Sun, Bang Zhang, and Liefeng Bo. Animate anyone: Consistent and controllable image-to-video synthesis for character animation. arXiv preprint arXiv:2311.17117, 2023. 2, 3, 8, 9
2023 arXiv
-
[32]
Po-Hsiang Huang, Fu-En Yang, and Y . Wang. Learning identity-invariant motion representations for cross-id face reenactment. 2020 IEEE/CVF Conference on Computer Vi- sion and Pattern Recognition (CVPR) , pages 7082–7090,
2020
-
[33]
Hagrid – hand gesture recognition image dataset
Alexander Kapitanov, Karina Kvanchiani, Alexander Na- gaev, Roman Kraynov, and Andrei Makhliarchuk. Hagrid – hand gesture recognition image dataset. In Proceedings of the IEEE/CVF Winter Conference on Applications of Com- puter Vision (WACV), pages 4572–4581, 2024. 2, 3
2024
-
[34]
A style-based generator architecture for generative adversarial networks
Tero Karras, Samuli Laine, and Timo Aila. A style-based generator architecture for generative adversarial networks. 2019 IEEE/CVF Conference on Computer Vision and Pat- tern Recognition (CVPR), pages 4396–4405, 2018. 2
2019
-
[35]
3d gaussian splatting for real-time radiance field rendering
Bernhard Kerbl, Georgios Kopanas, Thomas Leimk ¨uhler, and George Drettakis. 3d gaussian splatting for real-time radiance field rendering. ACM Transactions on Graphics, 42(4), 2023. 5
2023
-
[36]
Sapiens: Foundation for human vi- sion models
Rawal Khirodkar, Timur Bagautdinov, Julieta Martinez, Su Zhaoen, Austin James, Peter Selednik, Stuart Anderson, and Shunsuke Saito. Sapiens: Foundation for human vi- sion models. In European Conference on Computer Vision, pages 206–228. Springer, 2025. 2
2025
-
[37]
Segment anything
Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C Berg, Wan-Yen Lo, et al. Segment anything. arXiv preprint arXiv:2304.02643, 2023. 2, 4
2023 arXiv
-
[38]
Graspdiffusion: Synthe- sizing realistic whole-body hand-object interaction
Patrick Kwon and Hanbyul Joo. Graspdiffusion: Synthe- sizing realistic whole-body hand-object interaction. arXiv preprint arXiv:2410.13911, 2024. 3
2024
-
[39]
Word-level deep sign language recognition from video: A new large-scale dataset and methods comparison
Dongxu Li, Cristian Rodriguez, Xin Yu, and Hongdong Li. Word-level deep sign language recognition from video: A new large-scale dataset and methods comparison. In Pro- ceedings of the IEEE/CVF winter conference on applica- tions of computer vision, pages 1459–1469, 2020. 2, 3
2020
-
[40]
Vihope: Visuotactile in-hand object 6d pose estimation with shape completion.IEEE Robotics and Automation Let- ters, 2023
Hongyu Li, Snehal Dikhale, Soshi Iba, and Nawid Jamali. Vihope: Visuotactile in-hand object 6d pose estimation with shape completion.IEEE Robotics and Automation Let- ters, 2023. 2
2023
-
[41]
Renderih: A large-scale synthetic dataset for 3d interacting hand pose estimation
Lijun Li, Linrui Tian, Xindi Zhang, Qi Wang, Bang Zhang, Liefeng Bo, Mengyuan Liu, and Chen Chen. Renderih: A large-scale synthetic dataset for 3d interacting hand pose estimation. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) , pages 20395–...
2023
-
[42]
Mesh graphormer
Kevin Lin, Lijuan Wang, and Zicheng Liu. Mesh graphormer. In Proceedings of the IEEE/CVF international conference on computer vision, pages 12939–12948, 2021. 3
2021
-
[43]
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll ´ar, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In Computer Vision–ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedin...
2014
-
[44]
Zero-1-to-3: Zero-shot one image to 3d object, 2023
Ruoshi Liu, Rundi Wu, Basile Van Hoorick, Pavel Tok- makov, Sergey Zakharov, and Carl V ondrick. Zero-1-to-3: Zero-shot one image to 3d object, 2023. 8
2023
-
[45]
Hoi4d: A 4d egocentric dataset for category-level human-object interaction
Yunze Liu, Yun Liu, Che Jiang, Kangbo Lyu, Weikang Wan, Hao Shen, Boqiang Liang, Zhoujie Fu, He Wang, and Li Yi. Hoi4d: A 4d egocentric dataset for category-level human-object interaction. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CV...
2022
-
[46]
Diva-360: The dynamic visual dataset for immersive neu- ral fields
Cheng-You Lu, Peisen Zhou, Angela Xing, Chandradeep Pokhariya, Arnab Dey, Ishaan Nikhil Shah, Rugved Ma- vidipalli, Dylan Hu, Andrew I Comport, Kefan Chen, et al. Diva-360: The dynamic visual dataset for immersive neu- ral fields. In Proceedings of the IEEE/CVF Conference on C...
2024
-
[47]
Handrefiner: Refining malformed hands in generated images by diffusion-based conditional inpaint- ing
Wenquan Lu, Yufei Xu, Jing Zhang, Chaoyue Wang, and Dacheng Tao. Handrefiner: Refining malformed hands in generated images by diffusion-based conditional inpaint- ing. arXiv preprint arXiv:2311.17957 , 2023. 2, 3, 7, 8, 9
2023 arXiv
-
[48]
Me- diapipe: A framework for building perception pipelines
Camillo Lugaresi, Jiuqiang Tang, Hadon Nash, Chris Mc- Clanahan, Esha Uboweja, Michael Hays, Fan Zhang, Chuo- Ling Chang, Ming Guang Yong, Juhyun Lee, et al. Me- diapipe: A framework for building perception pipelines. arXiv preprint arXiv:1906.08172, 2019. 2, 3, 4
1906 arXiv
-
[49]
Repaint: Inpainting using denoising diffusion probabilistic models,
Andreas Lugmayr, Martin Danelljan, Andres Romero, Fisher Yu, Radu Timofte, and Luc Van Gool. Repaint: Inpainting using denoising diffusion probabilistic models,
-
[50]
Pose guided person image gener- ation
Liqian Ma, Xu Jia, Qianru Sun, Bernt Schiele, Tinne Tuyte- laars, and Luc Van Gool. Pose guided person image gener- ation. In Advances in Neural Information Processing Sys- tems, pages 405–415, 2017. 2, 3
2017
-
[51]
Nerf: Representing scenes as neural radiance fields for view syn- thesis
Ben Mildenhall, Pratul P Srinivasan, Matthew Tancik, Jonathan T Barron, Ravi Ramamoorthi, and Ren Ng. Nerf: Representing scenes as neural radiance fields for view syn- thesis. Communications of the ACM, 65(1):99–106, 2021. 5, 7
2021
-
[52]
Genheld: Generating and editing handheld objects
Chaerin Min and Srinath Sridhar. Genheld: Generating and editing handheld objects. arXiv preprint arXiv:2406.05059,
-
[53]
Tsdf-sampling: Efficient sampling for neural surface field using truncated signed distance field
Chaerin Min, Sehyun Cha, Changhee Won, and Jongwoo Lim. Tsdf-sampling: Efficient sampling for neural surface field using truncated signed distance field. arXiv preprint arXiv:2311.17878, 2023. 5
2023 arXiv
-
[54]
Interhand2
Gyeongsik Moon, Shoou-I Yu, He Wen, Takaaki Shiratori, and Kyoung Mu Lee. Interhand2. 6m: A dataset and base- line for 3d interacting hand pose estimation from a single rgb image. In Computer Vision–ECCV 2020: 16th Euro- pean Conference, Glasgow, UK, August 23–28, 2020, Pro- c...
2020
-
[55]
A dataset of relighted 3D interacting hands
Gyeongsik Moon, Shunsuke Saito, Weipeng Xu, Rohan Joshi, Julia Buffalini, Harley Bellan, Nicholas Rosen, Jesse Richardson, Mize Mallorie, Philippe Bree, Tomas Simon, Bo Peng, Shubham Garg, Kevyn McPhail, and Takaaki Shi- ratori. A dataset of relighted 3D interacting hands. In ...
2023
-
[56]
Instant neural graphics primitives with a multiresolution hash encoding
Thomas M ¨uller, Alex Evans, Christoph Schied, and Alexander Keller. Instant neural graphics primitives with a multiresolution hash encoding. ACM Transactions on Graphics (ToG), 41(4):1–15, 2022. 5
2022
-
[57]
Han- diffuser: Text-to-image generation with realistic hand ap- pearances
Supreeth Narasimhaswamy, Uttaran Bhattacharya, Xiang Chen, Ishita Dasgupta, Saayan Mitra, and Minh Hoai. Han- diffuser: Text-to-image generation with realistic hand ap- pearances. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2468–...
2024
-
[58]
Han- diffuser: Text-to-image generation with realistic hand ap- pearances
Supreeth Narasimhaswamy, Uttaran Bhattacharya, Xiang Chen, Ishita Dasgupta, Saayan Mitra, and Minh Hoai. Han- diffuser: Text-to-image generation with realistic hand ap- pearances. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2468–...
2024
-
[59]
Im- proved denoising diffusion probabilistic models
Alexander Quinn Nichol and Prafulla Dhariwal. Im- proved denoising diffusion probabilistic models. In Inter- national Conference on Machine Learning , pages 8162–
-
[60]
Fsgan: Subject agnostic face swapping and reenactment
Yuval Nirkin, Yosi Keller, and Tal Hassner. Fsgan: Subject agnostic face swapping and reenactment. 2019 IEEE/CVF International Conference on Computer Vision (ICCV), pages 7183–7192, 2019. 2, 3
2019
-
[61]
Assemblyhands: Towards egocentric activity understanding via 3d hand pose estima- tion
Takehiko Ohkawa, Kun He, Fadime Sener, Tomas Hodan, Luan Tran, and Cem Keskin. Assemblyhands: Towards egocentric activity understanding via 3d hand pose estima- tion. In Proceedings of the IEEE/CVF conference on com- puter vision and pattern recognition , pages 12999–13008,
-
[62]
AssemblyHands: towards egocentric activity understanding via 3d hand pose esti- mation
Takehiko Ohkawa, Kun He, Fadime Sener, Tomas Hodan, Luan Tran, and Cem Keskin. AssemblyHands: towards egocentric activity understanding via 3d hand pose esti- mation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pages 12999–1300...
2023
-
[63]
Maxime Oquab, Timoth ´ee Darcet, Theo Moutakanni, Huy V . V o, Marc Szafraniec, Vasil Khalidov, Pierre Fernan- dez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, Russell Howes, Po-Yao Huang, Hu Xu, Vasu Sharma, Shang-Wen Li, Wojciech Galuba, Mike Rabbat, Mido Ass- ran, N...
2023
-
[64]
Atten- tionhand: Text-driven controllable hand image generation for 3d hand reconstruction in the wild
Junho Park, Kyeongbo Kong, and Suk-Ju Kang. Atten- tionhand: Text-driven controllable hand image generation for 3d hand reconstruction in the wild. arXiv preprint arXiv:2407.18034, 2024. 2, 3
2024 arXiv
-
[65]
Re- constructing hands in 3D with transformers
Georgios Pavlakos, Dandan Shan, Ilija Radosavovic, Angjoo Kanazawa, David Fouhey, and Jitendra Malik. Re- constructing hands in 3D with transformers. In CVPR,
-
[66]
Scalable diffusion mod- els with transformers
William Peebles and Saining Xie. Scalable diffusion mod- els with transformers. arXiv preprint arXiv:2212.09748 ,
-
[67]
Controlnext: Powerful and effi- cient control for image and video generation.arXiv preprint arXiv:2408.06070, 2024
Bohao Peng, Jian Wang, Yuechen Zhang, Wenbo Li, Ming- Chang Yang, and Jiaya Jia. Controlnext: Powerful and effi- cient control for image and video generation.arXiv preprint arXiv:2408.06070, 2024. 2, 3, 8, 9
2024 arXiv
-
[68]
Manus: Markerless grasp capture using articulated 3d gaussians
Chandradeep Pokhariya, Ishaan Nikhil Shah, Angela Xing, Zekun Li, Kefan Chen, Avinash Sharma, and Srinath Srid- har. Manus: Markerless grasp capture using articulated 3d gaussians. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2197...
2024
-
[69]
Movie gen: A cast of media foundation models
Adam Polyak, Amit Zohar, Andrew Brown, Andros Tjan- dra, Animesh Sinha, Ann Lee, Apoorv Vyas, Bowen Shi, Chih-Yao Ma, Ching-Yao Chuang, et al. Movie gen: A cast of media foundation models. arXiv preprint arXiv:2410.13720, 2024. 2
-
[70]
Dreamfusion: Text-to-3d using 2d diffusion
Ben Poole, Ajay Jain, Jonathan T Barron, and Ben Milden- hall. Dreamfusion: Text-to-3d using 2d diffusion. arXiv preprint arXiv:2209.14988, 2022. 7
2022 arXiv
-
[71]
Unicontrol: A unified diffu- sion model for controllable visual generation in the wild
Can Qin, Shu Zhang, Ning Yu, Yihao Feng, Xinyi Yang, Yingbo Zhou, Huan Wang, Juan Carlos Niebles, Caiming Xiong, Silvio Savarese, et al. Unicontrol: A unified diffu- sion model for controllable visual generation in the wild. arXiv preprint arXiv:2305.11147, 2023. 3, 5, 6, 7
2023 arXiv
-
[72]
Handcraft: Anatomically correct restoration of malformed hands in diffusion generated images
Zhen Qin, Yiqun Zhang, Yang Liu, and Dylan Campbell. Handcraft: Anatomically correct restoration of malformed hands in diffusion generated images. 2024. 2, 3, 7
2024
-
[73]
Learn- ing transferable visual models from natural language super- vision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learn- ing transferable visual models from natural language super- vision. In International conference on machine learning...
2021
-
[74]
Towards realistic generative 3d face models
Aashish Rai, Hiresh Gupta, Ayush Pandey, Francisco Vi- cente Carrasco, Shingo Jason Takagi, Amaury Aubel, Daeil Kim, Aayush Prakash, and Fernando De la Torre. Towards realistic generative 3d face models. In Proceedings of the IEEE/CVF Winter Conference on Applications of Compu...
2024
-
[76]
Zero-shot text-to-image generation
Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea V oss, Alec Radford, Mark Chen, and Ilya Sutskever. Zero-shot text-to-image generation. In Interna- tional Conference on Machine Learning, pages 8821–8831. PMLR, 2021. 2
2021
-
[77]
Hierarchical text-conditional image generation with clip latents
Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. Hierarchical text-conditional image generation with clip latents. arXiv preprint arXiv:2204.06125, 2022. 2, 3
2022 arXiv
-
[78]
Li, and Shan Liu
Yurui Ren, Gezhong Li, Yuanqi Chen, Thomas H. Li, and Shan Liu. Pirenderer: Controllable portrait image gener- ation via semantic neural rendering. 2021 IEEE/CVF In- ternational Conference on Computer Vision (ICCV), pages 13739–13748, 2021. 2, 3
2021
-
[80]
High-resolution image synthesis with latent diffusion models, 2021
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bj ¨orn Ommer. High-resolution image synthesis with latent diffusion models, 2021. 2, 3, 4
2021
-
[81]
Embodied hands: Modeling and capturing hands and bod- ies together
Javier Romero, Dimitrios Tzionas, and Michael J Black. Embodied hands: Modeling and capturing hands and bod- ies together. arXiv preprint arXiv:2201.02610, 2022. 2, 3, 7
2022 arXiv
-
[82]
Denton, Seyed Kamyar Seyed Ghasemipour, Burcu Karagol Ayan, Seyedeh Sara Mah- davi, Raphael Gontijo Lopes, Tim Salimans, Jonathan Ho, David J
Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily L. Denton, Seyed Kamyar Seyed Ghasemipour, Burcu Karagol Ayan, Seyedeh Sara Mah- davi, Raphael Gontijo Lopes, Tim Salimans, Jonathan Ho, David J. Fleet, and Mohammad Norouzi. Photorealistic text-to-image ...
2022 arXiv
-
[83]
Geodiffuser: Geometry-based im- age editing with diffusion models
Rahul Sajnani, Jeroen Vanbaar, Jie Min, Kapil Katyal, and Srinath Sridhar. Geodiffuser: Geometry-based im- age editing with diffusion models. arXiv preprint arXiv:2404.14403, 2024. 2
2024 arXiv
-
[84]
Zeronvs: Zero-shot 360-degree view synthesis from a single image
Kyle Sargent, Zizhang Li, Tanmay Shah, Charles Her- rmann, Hong-Xing Yu, Yunzhi Zhang, Eric Ryan Chan, Dmitry Lagun, Li Fei-Fei, Deqing Sun, et al. Zeronvs: Zero-shot 360-degree view synthesis from a single image. In Proceedings of the IEEE/CVF Conference on Computer Vision an...
2024
-
[85]
Laion-5b: An open large-scale dataset for train- ing next generation image-text models
Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Worts- man, et al. Laion-5b: An open large-scale dataset for train- ing next generation image-text models. Advances in Neural Inf...
2022
-
[86]
Zero123++: a single image to consis- tent multi-view diffusion base model
Ruoxi Shi, Hansheng Chen, Zhuoyang Zhang, Minghua Liu, Chao Xu, Xinyue Wei, Linghao Chen, Chong Zeng, and Hao Su. Zero123++: a single image to consis- tent multi-view diffusion base model. arXiv preprint arXiv:2310.15110, 2023. 7
-
[87]
Sangineto, St ´ephane Lathuili `ere, and N
Aliaksandr Siarohin, E. Sangineto, St ´ephane Lathuili `ere, and N. Sebe. Deformable gans for pose-based human im- age generation. 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3408–3416, 2017. 2, 3
2018
-
[88]
Constrained 6-dof grasp gener- ation on complex shapes for improved dual-arm manipula- tion, 2024
Gaurav Singh, Sanket Kalwar, Md Faizal Karim, Bi- pasha Sen, Nagamanikandan Govindan, Srinath Sridhar, and K Madhava Krishna. Constrained 6-dof grasp gener- ation on complex shapes for improved dual-arm manipula- tion, 2024. 2
2024
-
[89]
Constrained 6-dof grasp gener- ation on complex shapes for improved dual-arm manipula- tion
Gaurav Singh, Sanket Kalwar, Md Faizal Karim, Bi- pasha Sen, Nagamanikandan Govindan, Srinath Sridhar, and K Madhava Krishna. Constrained 6-dof grasp gener- ation on complex shapes for improved dual-arm manipula- tion. arXiv preprint arXiv:2404.04643, 2024. 2
2024 arXiv
-
[90]
The effectiveness of mae pre-pretraining for billion-scale pretraining
Mannat Singh, Quentin Duval, Kalyan Vasudev Alwala, Haoqi Fan, Vaibhav Aggarwal, Aaron Adcock, Armand Joulin, Piotr Doll ´ar, Christoph Feichtenhofer, Ross Gir- shick, Rohit Girdhar, and Ishan Misra. The effectiveness of mae pre-pretraining for billion-scale pretraining. In ICCV,
-
[91]
Weiss, Niru Mah- eswaranathan, and Surya Ganguli
Jascha Sohl-Dickstein, Eric A. Weiss, Niru Mah- eswaranathan, and Surya Ganguli. Deep unsupervised learning using nonequilibrium thermodynamics, 2015. 3
2015
-
[92]
Denoising diffusion implicit models
Jiaming Song, Chenlin Meng, and Stefano Ermon. Denoising diffusion implicit models. arXiv preprint arXiv:2010.02502, 2020. 2
2010 arXiv
-
[93]
Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole
Yang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole. Score- based generative modeling through stochastic differential equations, 2021. 3
2021
-
[94]
Controlling the world by sleight of hand, 2024
Sruthi Sudhakar, Ruoshi Liu, Basile Van Hoorick, Carl V ondrick, and Richard Zemel. Controlling the world by sleight of hand, 2024. 3, 5, 6, 7, 8, 9
2024
-
[95]
Dimensionx: Create any 3d and 4d scenes from a single image with con- trollable video diffusion, 2024
Wenqiang Sun, Shuo Chen, Fangfu Liu, Zilong Chen, Yueqi Duan, Jun Zhang, and Yikai Wang. Dimensionx: Create any 3d and 4d scenes from a single image with con- trollable video diffusion, 2024. 2
2024
-
[96]
Anycontrol: Create your artwork with versatile control on text-to-image generation, 2024
Yanan Sun, Yanchen Liu, Yinhao Tang, Wenjie Pei, and Kai Chen. Anycontrol: Create your artwork with versatile control on text-to-image generation, 2024. 3, 5, 6, 7
2024
-
[97]
Controllable 3d generative adversarial face model via disentangling shape and appearance
Fariborz Taherkhani, Aashish Rai, Quankai Gao, Shaunak Srivastava, Xuanbai Chen, Fernando De la Torre, Steven Song, Aayush Prakash, and Daeil Kim. Controllable 3d generative adversarial face model via disentangling shape and appearance. In Proceedings of the IEEE/CVF Win- ter ...
2023
-
[98]
Total generate: Cycle in cy- cle generative adversarial networks for generating human faces, hands, bodies, and natural scenes
Hao Tang and Nicu Sebe. Total generate: Cycle in cy- cle generative adversarial networks for generating human faces, hands, bodies, and natural scenes. Trans. Multi., 24: 2963–2974, 2022. 2, 3
2022
-
[99]
Hao Tang, Wei Wang, Dan Xu, Yan Yan, and N. Sebe. Ges- turegan for hand gesture-to-gesture translation in the wild. Proceedings of the 26th ACM international conference on Multimedia, 2018. 3, 5, 7
2018
-
[100]
Real-time continuous pose recovery of human hands using convolutional networks
Jonathan Tompson, Murphy Stein, Yann Lecun, and Ken Perlin. Real-time continuous pose recovery of human hands using convolutional networks. ACM Transactions on Graphics (ToG), 33(5):1–10, 2014. 4
2014
-
[101]
Realishuman: A two- stage approach for refining malformed human parts in gen- erated images, 2024
Benzhi Wang, Jingkai Zhou, Jingqi Bai, Yang Yang, Wei- hua Chen, Fan Wang, and Zhen Lei. Realishuman: A two- stage approach for refining malformed human parts in gen- erated images, 2024. 3, 7, 8, 9
2024
-
[102]
Hall, and Shimin Hu
Miao Wang, Guo-Ye Yang, Ruilong Li, Runze Liang, Song- Hai Zhang, Peter M. Hall, and Shimin Hu. Example-guided style-consistent image synthesis from semantic labeling. 2019 IEEE/CVF Conference on Computer Vision and Pat- tern Recognition (CVPR), pages 1495–1504, 2019. 2, 3
2019
-
[103]
Imagedream: Image-prompt multi-view diffusion for 3d generation
Peng Wang and Yichun Shi. Imagedream: Image-prompt multi-view diffusion for 3d generation. arXiv preprint arXiv:2312.02201, 2023. 7, 8
2023 arXiv
-
[104]
Dexgraspnet: A large-scale robotic dexterous grasp dataset for general ob- jects based on simulation
Ruicheng Wang, Jialiang Zhang, Jiayi Chen, Yinzhen Xu, Puhao Li, Tengyu Liu, and He Wang. Dexgraspnet: A large-scale robotic dexterous grasp dataset for general ob- jects based on simulation. In 2023 IEEE International Con- ference on Robotics and Automation (ICRA), pages 1135...
2023
-
[105]
Novel view synthesis with diffusion models,
Daniel Watson, William Chan, Ricardo Martin-Brualla, Jonathan Ho, Andrea Tagliasacchi, and Mohammad Norouzi. Novel view synthesis with diffusion models,
-
[106]
Srinivasan, Dor Verbin, Jonathan T
Rundi Wu, Ben Mildenhall, Philipp Henzler, Keunhong Park, Ruiqi Gao, Daniel Watson, Pratul P. Srinivasan, Dor Verbin, Jonathan T. Barron, Ben Poole, and Aleksander Holynski. Reconfusion: 3d reconstruction with diffusion priors. arXiv, 2023. 5
2023
-
[107]
Ef- fective whole-body pose estimation with two-stages distil- lation
Zhendong Yang, Ailing Zeng, Chun Yuan, and Yu Li. Ef- fective whole-body pose estimation with two-stages distil- lation. In Proceedings of the IEEE/CVF International Con- ference on Computer Vision, pages 4210–4220, 2023. 3
2023
-
[108]
Cogvideox: Text-to- video diffusion models with an expert transformer
Zhuoyi Yang, Jiayan Teng, Wendi Zheng, Ming Ding, Shiyu Huang, Jiazheng Xu, Yuanming Yang, Wenyi Hong, Xiaohan Zhang, Guanyu Feng, et al. Cogvideox: Text-to- video diffusion models with an expert transformer. arXiv preprint arXiv:2408.06072, 2024. 2
2024 arXiv
-
[109]
Affordance diffusion: Synthesizing hand-object inter- actions
Yufei Ye, Xueting Li, Abhinav Gupta, Shalini De Mello, Stan Birchfield, Jiaming Song, Shubham Tulsiani, and Sifei Liu. Affordance diffusion: Synthesizing hand-object inter- actions. In CVPR, 2023. 3
2023
-
[110]
Representation alignment for generation: Training diffu- sion transformers is easier than you think
Sihyun Yu, Sangkyung Kwak, Huiwon Jang, Jongheon Jeong, Jonathan Huang, Jinwoo Shin, and Saining Xie. Representation alignment for generation: Training diffu- sion transformers is easier than you think. arXiv preprint arXiv:2410.06940, 2024. 5
-
[111]
Hand1000: Generating realistic hands from text with only 1,000 images
Haozhuo Zhang, Bin Zhu, Yu Cao, and Yanbin Hao. Hand1000: Generating realistic hands from text with only 1,000 images. arXiv preprint arXiv:2408.15461, 2024. 3
2024 arXiv
-
[112]
Adding conditional control to text-to-image diffusion models
Lvmin Zhang, Anyi Rao, and Maneesh Agrawala. Adding conditional control to text-to-image diffusion models. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 3836–3847, 2023. 2, 3
2023
-
[113]
Learning anchor transformations for 3d garment animation
Fang Zhao, Zekun Li, Shaoli Huang, Junwu Weng, Tianfei Zhou, Guo-Sen Xie, Jue Wang, and Ying Shan. Learning anchor transformations for 3d garment animation. In Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 491–500, 2023. 2
2023
-
[114]
Sparsefusion: Dis- tilling view-conditioned diffusion for 3d reconstruction
Zhizhuo Zhou and Shubham Tulsiani. Sparsefusion: Dis- tilling view-conditioned diffusion for 3d reconstruction. In Proceedings of the IEEE/CVF Conference on Computer Vi- sion and Pattern Recognition , pages 12588–12597, 2023. 5
2023
-
[116]
Learn- ing to estimate 3d hand pose from single rgb im- ages
Christian Zimmermann and Thomas Brox. Learn- ing to estimate 3d hand pose from single rgb im- ages. Technical report, arXiv:1705.01389, 2017. https://arxiv.org/abs/1705.01389. 2
2017 arXiv
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.