REVIEW 4 major objections 6 minor 91 references
EgoAnimate: Generating Human Animations from Egocentric top-down Views
T0 review · 4 major / 6 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read EgoAnimate generates a fully animatable human avatar from a single egocentric top-down image by first synthesizing a frontal T-pose view with Stable Diffusion and ControlNet, then passing it to off-the-shelf avatar animation tools.
desk verdict New task formulation and a plausible proof-of-concept, but synthetic ground-truth targets keep the reported fidelity numbers from backing the realism claim. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the TopDown-to-Frontal (T2F) synthesis module: a Stable Diffusion latent diffusion model with a frozen VAE for encoding the top-down image, a ControlNet branch that injects an SMPL T-pose mask for structural pose control, and CLIP image embeddings projected and added through cross-attention to carry high-level semantics from the occluded input. The training objective combines the standard noise-prediction term with an LPIPS perceptual loss in image space, which the ablation studies show is needed for clothing-type preservation. This module is what bridges the domain gap; the avatar generation stage is deliberately assembled from frozen, publicly available tools so that the paper can isolate the effect of the frontal-view translation.
What would settle it
Collect paired top-down and true frontal footage for new subjects without any diffusion-model enhancement, run the trained T2F module, and compute the paper's own metrics against the true frontal frames; the central claim would be contradicted if lower-body clothing accuracy falls well below the reported 79% or if the PSNR and LPIPS gaps versus the enhanced targets are large and systematic.
Extended reading notes
Core claim
On its own terms, the paper claims that heavily occluded top-down egocentric views can be converted into frontal T-pose images that preserve the subject's clothing and body shape, and that those images are good enough for off-the-shelf avatar animation. The T2F module finetunes a latent diffusion model conditioned three ways: the VAE encoding of the top-down image, a ControlNet-injected SMPL pose mask, and CLIP embeddings of the input fused via cross-attention; it is trained with a diffusion noise-prediction loss plus an LPIPS perceptual loss. From the synthesized frontal image, the 3D branch uses MagicMan to generate 360-degree views and ExAvatar to build a gaussian-splatting avatar, while the 2D branch feeds the image directly to UniAnimate. Quantitative results, such as full-body PSNR 17.73, SSIM 0.8743, and LPIPS 0.0835, along with lower-body clothing accuracy of 79% and upper-body accuracy of 87% under the best configuration, are offered as evidence, together with out-of-distribution examples from Ego4D, GoPro footage, and UnrealEgo-RW.
Load-bearing premise
The argument assumes that frontal views enhanced by an off-the-shelf diffusion model are a valid proxy for what the person actually looks like from the front; if that enhancement changes clothing, body shape, or pose, the reported accuracy and avatar quality may reflect the enhancer's style rather than true reconstruction.
Editorial extensions
If this is right
- A user wearing only a head-mounted camera could get an animatable avatar without a separate front-facing camera, multi-view rig, or per-person training run.
- The same generated frontal image can drive different animation backends: UniAnimate for 2D animation or MagicMan plus ExAvatar for a 3D avatar, with the user study favoring UniAnimate while noting its longer runtime.
- The T2F module transfers to out-of-distribution egocentric sources it never saw in training, including Ego4D clips, internet GoPro footage, and UnrealEgo-RW frames.
- Because the face is synthesized rather than recovered, the pipeline suits body-and-clothing telepresence but not applications requiring identity-preserving facial appearance.
Reading between the lines
- If the diffusion-enhanced frontal targets systematically alter clothing or body shape, the reported PSNR, SSIM, LPIPS, and clothing-accuracy figures could mostly measure agreement with the enhancer's style; a real-frontal benchmark would separate those two readings.
- The conditioning recipe of VAE latent plus ControlNet structure plus CLIP semantics is not egocentric-specific, so the same module could be lifted to other sparse-view human synthesis tasks such as side-profile or selfie-to-frontal translation.
- Using short top-down video clips instead of a single frame could exploit temporal cues to resolve ambiguous clothing geometry like long coats or flowing garments, a testable extension the paper mentions but does not run.
- The user study's runtime-versus-quality inversion suggests that for real-time telepresence the practical choice may be a faster 2D animator, and that T2F should be optimized jointly with the downstream animator rather than judged by image metrics alone.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. EgoAnimate proposes a modular two-stage pipeline for creating an animatable full-body avatar from a single egocentric top-down image. The first stage is a TopDown-to-Frontal (T2F) module built on Stable Diffusion, ControlNet pose conditioning, and CLIP features; it is trained on a custom paired dataset of roughly 3,000 top-down frames and 300 manually curated frontal views enhanced with an off-the-shelf diffusion model. The second stage feeds the synthesized frontal image either into MagicMan+ExAvatar for a 3D Gaussian-splatting avatar or into image-to-video animators such as UniAnimate for 2D animation. The paper reports ablations over ControlNet, perceptual loss, VAE choice, and CLIP encoding, a clothing-type accuracy metric, a user study comparing four animation backbones, and qualitative results on Ego4D, UnrealEgo, and Internet footage. The main claimed contribution is that the T2F module produces realistic frontal views from top-down inputs, robust to pose variation and occlusion, and that the full pipeline generalizes to new identities without per-person training.
Significance. If the central claim holds, EgoAnimate addresses a genuinely under-explored problem: reconstructing animatable human avatars from a single egocentric HMD view without a multi-view capture studio. The design is pragmatic, reusing frozen pretrained components (Stable Diffusion, ControlNet, MagicMan, ExAvatar, UniAnimate) and requiring only a small finetuned T2F module, which is attractive for accessibility. The qualitative demonstrations on out-of-distribution inputs, including Ego4D and UnrealEgo, suggest the pipeline has real generalization potential. However, the quantitative evaluation is not yet convincing because all fidelity metrics are computed against frontal images that were themselves generated and enhanced by a diffusion model, not against real frontal captures. The paper also lacks any comparison with prior egocentric avatar methods. The claimed robustness and fidelity therefore remain plausible but not established. The modular formulation and the new dataset are useful contributions, but the evaluation needs substantial strengthening before the paper's main claims can be accepted.
major comments (4)
- [§3.3, §4.1, Tables 1 and 2] The quantitative evaluation is circular in an important sense. Section 3.3 states that frontal views were 'enhanced using an off-the-shelf image diffusion model [9]' and admits they are 'not exact ground-truth,' yet Table 1 (PSNR, SSIM, LPIPS) and Table 2 (clothing accuracy) compare all generated frontal images against exactly these DALL-E-enhanced targets. The model is trained to match this target distribution and then scored against the same distribution, so the reported numbers (e.g., PSNR 17.73, SSIM 0.8743, LPIPS 0.0835) do not measure fidelity to real frontal appearance. The paper should evaluate on a held-out set of real, unenhanced frontal captures, even if small, and should analyze whether the DALL-E enhancement systematically changes clothing, body shape, or pose.
- [§3.3 vs. §4.1] The dataset description is internally inconsistent. Section 3.3 says the dataset consists of approximately 3,000 top-down images and 300 frontal images, while Section 4.1 says training used 540 frontal images. Table 2 reports accuracy 'across 100 samples' while Section 4.1 says the test set contains 60 samples. The exact number of unique subjects, frontal images, paired top-down frames, and train/test splits must be clarified, otherwise the experiments are not reproducible.
- [§4.2] The avatar-animation evaluation relies entirely on a ranking-based user study with no ground-truth animations and no quantitative fidelity-to-reality metric. The paper itself states that there is an 'absence of ground-truth animations with matching clothing.' Consequently, the claim that the full pipeline 'generalizes to new identities' (Section 5) rests on qualitative inspection and subjective preference scores. The user study can justify the choice of UniAnimate among the tested animators, but it cannot validate the reconstruction accuracy of the T2F stage or of the final avatar against the real subject. A small quantitative evaluation on real captures, for example comparing body-shape or clothing-attribute consistency before and after animation, would materially strengthen this claim.
- [§2.3, §4] The paper positions itself against prior egocentric avatar works, EgoRenderer and EgoAvatar, but never compares with them empirically. Section 2.3 argues those methods require person-specific training, but no experiments demonstrate an advantage over them on any shared input. Since the paper's novelty claim is that it avoids multi-view person-specific training, a comparison under a common protocol, or at least a clear explanation of why such a comparison is infeasible, is needed. Without it, the superiority of the proposed approach over existing egocentric avatar methods is not evidenced.
minor comments (6)
- [§3.1] In the loss equation, the time step t is not defined in the text, and the notation \tilde{z}_t appears without explanation; please specify the noising process and the exact conditioning arguments of the noise predictor.
- [Figure 3 caption] The caption refers to '^' and an empty symbol for frozen and trainable components, but the symbols are missing in the text rendering; please ensure both markers are printed.
- [§1, §4.1] There are typos such as 'addes' (Section 1), 'furher' (Section 4.1), and 'hypothesis' used as a verb; a copyedit pass is needed.
- [§4.1] The subsection is titled 'Baselines,' but it describes component ablations and alternative encoders rather than comparisons with external baselines; consider renaming it to avoid confusion.
- [References] Some references appear duplicated (e.g., [2] and [4] are the same UnrealEgo paper; [48] and [49] appear to be the same Zero-1-to-3 paper in arXiv and CVPR forms). These should be consolidated.
- [§4.2] The user study description lacks details on the number of raters per video, inter-rater agreement, and the exact definition of Borda Score; adding these would improve reproducibility.
Circularity Check
The headline realism claim is validated by matching DALL-E-enhanced targets that also served as training labels, so the reported PSNR/SSIM/LPIPS numbers measure fit to a synthetic style rather than to true frontal appearance.
-
fitted input called prediction
[Section 3.3 'Ground Truth Enhancement'; Section 4.1 'Metrics' and Tables 1-2]
"Frontal views are sparse and visually enhanced using an off-the-shelf diffusion-based method [9] as a post-processing step. While not exact ground-truth, they provide sufficient visual supervision to guide plausible frontal reconstruction ... we quantify the correspondence between generated and ground-truth frontal images using PSNR (Peak Signal-to-Noise Ratio) for pixel-level accuracy, SSIM ... and LPIPS ..."
The 'ground-truth' frontal images used as supervision and as evaluation targets were produced by the same DALL-E enhancement process [9]. The training objective already includes LPIPS against these enhanced images (Section 3.1), and the quantitative claims in Tables 1 and 2 are computed against the same class of synthetic targets. PSNR/SSIM/LPIPS and clothing accuracy therefore measure how well the model imitates DALL-E's enhanced style, not how faithfully it reconstructs true frontal appearance. The central claim that the T2F module 'produces realistic frontal views' is operationalized as agreement with a generated target distribution, so the reported numbers are self-referential rather than evidence of real-world fidelity.
full rationale
No load-bearing self-citation or uniqueness-import argument appears; the architecture and downstream avatar pipeline (MagicMan/ExAvatar/UniAnimate) are external components used as drop-ins, and the out-of-distribution qualitative results give some independent support. The one substantive circularity is evaluative: the DALL-E-enhanced frontal images are both the training labels and the 'ground truth' for all quantitative metrics, so the realism claim is not independently tested against real frontal views. A secondary reproducibility inconsistency (Section 3.3 reports 300 frontal images; Section 4.1 says training used 540) further weakens the quantitative section but is not itself circularity.
Assumptions & free parameters
free parameters (4)
- lambda_perc =
0.2
- lambda_diff =
1.0
- frontal perturbation probability q =
not specified
- independent top-down rotation probability p =
not specified
assumptions (5)
- domain assumption A single top-down egocentric image contains enough information to infer the person's frontal appearance, including occluded regions such as pants and back.
- domain assumption DALL-E-enhanced frontal views are a valid ground truth for training and evaluation.
- domain assumption The Stable Diffusion latent space and ControlNet pose conditioning preserve clothing and body shape during frontal synthesis.
- domain assumption Off-the-shelf image-to-motion models can produce high-quality animation from a synthetic frontal T-pose image.
- domain assumption Top-down and frontal views can be matched by timestamp and body pose despite not having identical poses.
Cite this review
Pith. "Pith review of EgoAnimate: Generating Human Animations from Egocentric top-down Views." pith.science (2026). https://pith.science/paper/OVGRUULA
@misc{pith2026250709230,
author = {Pith},
title = {Pith review of: EgoAnimate: Generating Human Animations from Egocentric top-down Views},
year = {2026},
howpublished = {\url{https://pith.science/paper/OVGRUULA}},
note = {Machine review of arXiv:2507.09230}
}
read the original abstract
An ideal digital telepresence experience requires accurate replication of a person's body, clothing, and movements. To capture and transfer these movements into virtual reality, the egocentric (first-person) perspective can be adopted, which enables the use of a portable and cost-effective device without front-view cameras. However, this viewpoint introduces challenges such as occlusions and distorted body proportions. There are few works reconstructing human appearance from egocentric views, and none use a generative prior-based approach. Some methods create avatars from a single egocentric image during inference, but still rely on multi-view datasets during training. To our knowledge, this is the first study using a generative backbone to reconstruct animatable avatars from egocentric inputs. Based on Stable Diffusion, our method reduces training burden and improves generalizability. Inspired by methods such as SiTH and MagicMan, which perform 360-degree reconstruction from a frontal image, we introduce a pipeline that generates realistic frontal views from occluded top-down images using ControlNet and a Stable Diffusion backbone. Our goal is to convert a single top-down egocentric image into a realistic frontal representation and feed it into an image-to-motion model. This enables generation of avatar motions from minimal input, paving the way for more accessible and generalizable telepresence systems.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[9]
James Betker, Gabriel Goh, Li Jing, Tim Brooks, Jianfeng Wang, Linjie Li, Long Ouyang, Juntang Zhuang, Joyce Lee, Yufei Guo, et al . 2023. Improving im- age generation with better captions. Computer Science. https://cdn. openai. com/papers/dall-e-3. pdf (2023)
2023
-
[1]
Hiroyasu Akada, Jian Wang, Vladislav Golyanik, and Christian Theobalt. 2023. 3D Human Pose Perception from Egocentric Stereo Videos. 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2023), 767–776. https://api.semanticscholar.org/CorpusID:266725400
2023
-
[2]
Hiroyasu Akada, Jian Wang, Vladislav Golyanik, and Christian Theobalt. 2024. 3D Human Pose Perception from Egocentric Stereo Videos. In Computer Vision and Pattern Recognition (CVPR)
2024
-
[3]
Hiroyasu Akada, Jian Wang, Soshi Shimada, Masaki Takahashi, Christian Theobalt, and Vladislav Golyanik. 2022. UnrealEgo: A New Dataset for Ro- bust Egocentric 3D Human Motion Capture. InEuropean Conference on Computer Vision (ECCV)
2022
-
[4]
Hiroyasu Akada, Jian Wang, Soshi Shimada, Masaki Takahashi, Christian Theobalt, and Vladislav Golyanik. 2022. UnrealEgo: A New Dataset for Ro- bust Egocentric 3D Human Motion Capture. InEuropean Conference on Computer Vision. https://api.semanticscholar.org/CorpusID:251253160
2022
-
[5]
Magnor, Weipeng Xu, Christian Theobalt, and Gerard Pons-Moll
Thiemo Alldieck, Marcus A. Magnor, Weipeng Xu, Christian Theobalt, and Gerard Pons-Moll. 2018. Detailed Human Avatars from Monocular Video. 2018 Interna- tional Conference on 3D Vision (3DV) (2018), 98–109. https://api.semanticscholar. org/CorpusID:51929974
2018
-
[6]
Samaneh Azadi, Akbar Shah, Thomas Hayes, Devi Parikh, and Sonal Gupta
-
[7]
Omer Bar-Tal, Lior Yariv, Yaron Lipman, and Tali Dekel. 2023. MultiDiffusion: Fus- ing Diffusion Paths for Controlled Image Generation. In International Conference on Machine Learning. https://api.semanticscholar.org/CorpusID:256900756
2023
Show all 91 references
-
[8]
Miguel Barreda-Ángeles and Tilo Hartmann. 2022. Psychological benefits of using social virtual reality platforms during the covid-19 pandemic: The role of social and spatial presence. Computers in Human Behavior (2022)
2022
-
[10]
Bhattarai, Matthias Nießner, and Artem Sevastopolsky
Ananta R. Bhattarai, Matthias Nießner, and Artem Sevastopolsky. 2023. Tri- PlaneNet: An Encoder for EG3D Inversion. 2024 IEEE/CVF Winter Confer- ence on Applications of Computer Vision (W ACV) (2023), 3043–3053. https: //api.semanticscholar.org/CorpusID:257687746
2023
-
[11]
Tim Brooks, Aleksander Holynski, and Alexei A. Efros. 2023. InstructPix2Pix: Learning to Follow Image Editing Instructions. In CVPR
2023
-
[12]
Zhongang Cai, Daxuan Ren, Ailing Zeng, Zhengyu Lin, Tao Yu, Wenjia Wang, Xiangyu Fan, Yangmin Gao, Yifan Yu, Liang Pan, Fangzhou Hong, Mingyuan Zhang, Chen Change Loy, Lei Yang, and Ziwei Liu. 2022. HuMMan: Multi-Modal 4D Human Dataset for Versatile Sensing and Modeling. ArXiv...
2022 arXiv
-
[13]
Dan Casas and Marc Comino-Trinidad. 2023. SMPLitex: A Generative Model and Dataset for 3D Human Texture Estimation from Single Image. In British Machine Vision Conference (BMVC)
2023
-
[14]
Lin, Matthew Chan, Koki Nagano, Boxiao Pan, Shalini De Mello, Orazio Gallo, Leonidas J
Eric Chan, Connor Z. Lin, Matthew Chan, Koki Nagano, Boxiao Pan, Shalini De Mello, Orazio Gallo, Leonidas J. Guibas, Jonathan Tremblay, S. Khamis, Tero Kar- ras, and Gordon Wetzstein. 2021. Efficient Geometry-aware 3D Generative Adver- sarial Networks. 2022 IEEE/CVF Conference...
2021
-
[15]
Eric Chan, Marco Monteiro, Petr Kellnhofer, Jiajun Wu, and Gordon Wetzstein
-
[16]
Chan, Koki Nagano, Matthew A
Eric R. Chan, Koki Nagano, Matthew A. Chan, Alexander W. Bergman, Jeong Joon Park, Axel Levy, Miika Aittala, Shalini De Mello, Tero Karras, and Gordon Wet- zstein. 2023. Generative Novel View Synthesis with 3D-Aware Diffusion Models. In Proceedings of the IEEE/CVF Internationa...
2023
-
[17]
Jianchun Chen, Jian Wang, Yinda Zhang, Rohit Pandey, Thabo Beeler, Marc Habermann, and Christian Theobalt. 2024. EgoAvatar: Egocentric View-Driven and Photorealistic Full-body Avatars. In SIGGRAPH Asia
2024
-
[18]
Yuedong Chen, Haofei Xu, Chuanxia Zheng, Bohan Zhuang, Marc Pollefeys, Andreas Geiger, Tat-Jen Cham, and Jianfei Cai. 2024. MVSplat: Efficient 3D Gaussian Splatting from Sparse Multi-View Images. In European Conference on Computer Vision. https://api.semanticscholar.org/Corpus...
2024
-
[19]
Zheng Chen, Zhiqi Zhang, Junsong Yuan, Yi Xu, and Lantao Liu. 2024. Show Your Face: Restoring Complete Facial Images from Partial Observations for VR Meeting. 2024 IEEE/CVF Winter Conference on Applications of Computer Vision (W ACV)(2024), 8673–8682. https://api.semanticschol...
2024
-
[20]
Guillaume Couairon, Jakob Verbeek, Holger Schwenk, and Matthieu Cord. 2023. DiffEdit: Diffusion-based semantic image editing with mask guidance. In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023 . OpenReview.net. htt...
2023
-
[21]
Matt Deitke, Dustin Schwenk, Jordi Salvador, Luca Weihs, Oscar Michel, Eli VanderBilt, Ludwig Schmidt, Kiana Ehsani, Aniruddha Kembhavi, and Ali Farhadi
-
[22]
Elgharib, R
Mohamed A. Elgharib, R. MallikarjunB., Ayush Kumar Tewari, Hyeongwoo Kim, Wentao Liu, Hans-Peter Seidel, and Christian Theobalt. 2019. EgoFace: Egocentric Face Performance Capture and Videorealistic Reenactment.ArXiv abs/1905.10822 (2019). https://api.semanticscholar.org/Corpu...
2019 arXiv
-
[23]
Kristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis, Antonino Furnari, Rohit Girdhar, Jackson Hamburger, Hao Jiang, Miao Liu, Xingyu Liu, et al. 2022. Ego4d: Around the world in 3,000 hours of egocentric video. In Conference on computer vision and pattern recognition
2022
-
[24]
Jiatao Gu, Lingjie Liu, Peng Wang, and Christian Theobalt. 2021. StyleNeRF: A Style-based 3D-Aware Generator for High-resolution Image Synthesis. ArXiv abs/2110.08985 (2021). https://api.semanticscholar.org/CorpusID:239016913
2021 arXiv
-
[25]
Sang-Hun Han, Min-Gyu Park, Ju Hong Yoon, Ju-Mi Kang, Young-Jae Park, and Hae-Gon Jeon. 2023. High-fidelity 3D Human Digitization from Single 2K Reso- lution Images. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recog- nition (CVPR) (2023), 12869–12879. https://api.s...
2023
-
[26]
Xu He, Xiaoyu Li, Di Kang, Jiangnan Ye, Chaopeng Zhang, Liyang Chen, Xiangjun Gao, Han Zhang, Zhiyong Wu, and Haolin Zhuang. 2024. MagicMan: Genera- tive Novel View Synthesis of Humans with 3D-Aware Diffusion and Iterative Refinement. arXiv:2408.14211 [cs.CV]
2024 arXiv
-
[27]
Hsuan-I Ho, Jie Song, and Otmar Hilliges. 2024. SiTH: Single-view Textured Human Reconstruction with Image-Conditioned Diffusion. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)
2024
-
[28]
Jin Gyu Hong, Seung Young Noh, Hee-Kyung Lee, Won-Sik Cheong, and Ju Yong Chang. 2024. 3D Clothed Human Reconstruction from Sparse Multi-View Images. 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW) (2024), 677–687. https://api.semanticscho...
2024
-
[29]
Yicong Hong, Kai Zhang, Jiuxiang Gu, Sai Bi, Yang Zhou, Difan Liu, Feng Liu, Kalyan Sunkavalli, Trung Bui, and Hao Tan. 2023. LRM: Large Recon- struction Model for Single Image to 3D. ArXiv abs/2311.04400 (2023). https: //api.semanticscholar.org/CorpusID:265050698
2023 arXiv
-
[30]
Edward Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, and Weizhu Chen
J. Edward Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, and Weizhu Chen. 2021. LoRA: Low-Rank Adaptation of Large Language Models. ArXiv abs/2106.09685 (2021). https://api.semanticscholar.org/CorpusID: 235458009
2021 arXiv
-
[31]
Li Hu. 2024. Animate anyone: Consistent and controllable image-to-video syn- thesis for character animation. In Conference on Computer Vision and Pattern Recognition
2024
-
[32]
Hu, Kripasindhu Sarkar, Lingjie Liu, Matthias Zwicker, and Christian Theobalt
T. Hu, Kripasindhu Sarkar, Lingjie Liu, Matthias Zwicker, and Christian Theobalt
-
[33]
Mustafa Işık, Martin Rünz, Markos Georgopoulos, Taras Khakhulin, Jonathan Starck, Lourdes de Agapito, and Matthias Nießner. 2023. HumanRF: High-Fidelity Neural Radiance Fields for Humans in Motion. ACM Transactions on Graphics (TOG) 42 (2023), 1 – 12. https://api.semanticschol...
2023
-
[34]
Tianjian Jiang, Xu Chen, Jie Song, and Otmar Hilliges. 2022. InstantAvatar: Learning Avatars from Monocular Video in 60 Seconds.2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2022), 16922–16932. https: //api.semanticscholar.org/CorpusID:254877356
2022
-
[35]
Saragih, Shih-En Wei, Tenia Wang, Stephen Lombardi, Danielle Belko, Autumn Trimble, and Hernán Badino
Amin Jourabloo, Fernando De la Torre, Jason M. Saragih, Shih-En Wei, Tenia Wang, Stephen Lombardi, Danielle Belko, Autumn Trimble, and Hernán Badino. EgoAnimate: Generating Human Animations from Egocentric top-down Views arXiv Preprint, June 2025,
2025
-
[36]
Johanna Karras, Aleksander Holynski, Ting-Chun Wang, and Ira Kemelmacher- Shlizerman. 2023. Dreampose: Fashion image-to-video synthesis via stable diffu- sion. In International Conference on Computer Vision
2023
-
[37]
Bahjat Kawar, Shiran Zada, Oran Lang, Omer Tov, Huiwen Chang, Tali Dekel, Inbar Mosseri, and Michal Irani. 2023. Imagic: Text-Based Real Image Editing with Diffusion Models. In Conference on Computer Vision and Pattern Recognition 2023
2023
-
[38]
Bernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, and George Drettakis
-
[39]
Rawal Khirodkar, Timur Bagautdinov, Julieta Martinez, Su Zhaoen, Austin James, Peter Selednik, Stuart Anderson, and Shunsuke Saito. 2024. Sapiens: Foundation for Human Vision Models. arXiv preprint arXiv:2408.12569 (2024)
2024 arXiv
-
[40]
2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2021), 20291–20300
Robust Egocentric Photo-realistic Facial Expression Transfer for Virtual Re- ality. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2021), 20291–20300. https://api.semanticscholar.org/CorpusID:233209885
2021
-
[41]
Jeong-gi Kwak, Erqun Dong, Yuhe Jin, Hanseok Ko, Shweta Mahajan, and Kwang Moo Yi. 2024. Vivid-1-to-3: Novel view synthesis with video diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 6775–6785
2024
-
[42]
Youngjoong Kwon, Baole Fang, Yixing Lu, Haoye Dong, Cheng Zhang, Fran- cisco Vicente Carrasco, Albert Mosella-Montoro, Jianjin Xu, Shingo Takagi, Daeil Kim, Aayush Prakash, and Fernando De la Torre. 2024. Generalizable Human Gaussians for Sparse View Synthesis. ArXiv abs/2407....
2024 arXiv
-
[43]
Yushi Lan, Fangzhou Hong, Shuai Yang, Shangchen Zhou, Xuyi Meng, Bo Dai, Xingang Pan, and Chen Change Loy. 2024. LN3Diff: Scalable Latent Neural Fields Diffusion for Speedy 3D Generation. In European Conference on Computer Vision . https://api.semanticscholar.org/CorpusID:268531938
2024
-
[44]
ACM Transactions on Graphics 42, 4 (2023), 75:1–75:14
3D Gaussian Splatting for Real-Time Radiance Field Rendering. ACM Transactions on Graphics 42, 4 (2023), 75:1–75:14. doi:10.1145/3592990
2023 doi
-
[45]
Zhe Li, Zerong Zheng, Lizhen Wang, and Yebin Liu. 2024. Animatable Gaussians: Learning Pose-Dependent Gaussian Maps for High-Fidelity Human Avatar Model- ing. 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2024), 19711–19722. https://api.semanticsc...
2024
-
[46]
Tobias Kirschstein, Shenhan Qian, Simon Giebenhain, Tim Walter, and Matthias Nießner. 2023. NeRSemble: Multi-view Radiance Field Reconstruction of Human Heads. ACM Transactions on Graphics (TOG) 42 (2023), 1 – 14. https://api. semanticscholar.org/CorpusID:258480129
2023
-
[47]
MukundVarma, Zexiang Xu, and Hao Su
Minghua Liu, Chao Xu, Haian Jin, Ling Chen, T. MukundVarma, Zexiang Xu, and Hao Su. 2023. One-2-3-45: Any Single Image to 3D Mesh in 45 Seconds without Per-Shape Optimization. ArXiv abs/2306.16928 (2023). https://api. semanticscholar.org/CorpusID:259286991
2023 arXiv
-
[48]
Ruoshi Liu, Rundi Wu, Basile Van Hoorick, Pavel Tokmakov, Sergey Zakharov, and Carl Vondrick. 2023. Zero-1-to-3: Zero-shot One Image to 3D Object. 2023 IEEE/CVF International Conference on Computer Vision (ICCV) (2023), 9264–9275. https://api.semanticscholar.org/CorpusID:257631738
2023
-
[49]
Ruoshi Liu, Rundi Wu, Basile Van Hoorick, Pavel Tokmakov, Sergey Zakharov, and Carl Vondrick. 2023. Zero-1-to-3: Zero-shot One Image to 3D Object. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) . 9298–9309
2023
-
[50]
Yuheng Li, Haotian Liu, Qingyang Wu, Fangzhou Mu, Jianwei Yang, Jianfeng Gao, Chunyuan Li, and Yong Jae Lee. 2023. GLIGEN: Open-Set Grounded Text- to-Image Generation. 2023 IEEE/CVF Conference on Computer Vision and Pat- tern Recognition (CVPR) (2023), 22511–22521. https://api...
2023
-
[51]
Yuan Liu, Chu-Hsing Lin, Zijiao Zeng, Xiaoxiao Long, Lingjie Liu, Taku Ko- mura, and Wenping Wang. 2023. SyncDreamer: Generating Multiview-consistent Images from a Single-view Image. ArXiv abs/2309.03453 (2023). https://api. semanticscholar.org/CorpusID:261582503
2023 arXiv
-
[52]
Chen-Hsuan Lin, Jun Gao, Luming Tang, Towaki Takikawa, Xiaohui Zeng, Xun Huang, Karsten Kreis, Sanja Fidler, Ming-Yu Liu, and Tsung-Yi Lin. 2022. Magic3D: High-Resolution Text-to-3D Content Creation. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2...
2022
-
[53]
Yixing Lu, Junting Dong, Youngjoong Kwon, Qin Zhao, Bo Dai, and Fer- nando De la Torre. 2025. GAS: Generative Avatar Synthesis from a Single Image. arXiv:2502.06957 [cs.CV] https://arxiv.org/abs/2502.06957
2025 arXiv
-
[54]
Srinivasan, Matthew Tancik, Jonathan T
Ben Mildenhall, Pratul P. Srinivasan, Matthew Tancik, Jonathan T. Barron, Ravi Ramamoorthi, and Ren Ng. 2021. NeRF: representing scenes as neural radiance fields for view synthesis. Commun. ACM 65, 1 (Dec. 2021), 99–106. doi:10.1145/ 3503250
2021
-
[55]
Luvizon, Christian Theobalt, and Vladislav Golyanik
Christen Millerdurai, Hiroyasu Akada, Jian Wang, Diogo C. Luvizon, Christian Theobalt, and Vladislav Golyanik. 2024. EventEgo3D: 3D Human Motion Capture from Egocentric Event Streams. 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2024), 1186–1195....
2024
-
[56]
Xinqi Liu, Chenming Wu, Jialun Liu, Xing Liu, Jinbo Wu, Chen Zhao, Haocheng Feng, Errui Ding, and Jingdong Wang. 2024. GVA: Reconstructing Vivid 3D Gaussian Avatars from Monocular Videos. ArXiv abs/2402.16607 (2024). https: //api.semanticscholar.org/CorpusID:268032552
2024 arXiv
-
[57]
Chong Mou, Xintao Wang, Liangbin Xie, Jing Zhang, Zhongang Qi, Ying Shan, and Xiaohu Qie. 2023. T2I-Adapter: Learning Adapters to Dig out More Control- lable Ability for Text-to-Image Diffusion Models. InAAAI Conference on Artificial Intelligence. https://api.semanticscholar.o...
2023
-
[58]
Matthew Loper, Naureen Mahmood, Javier Romero, Gerard Pons-Moll, and Michael J. Black. 2015. SMPL: A Skinned Multi-Person Linear Model. ACM Trans. Graphics (Proc. SIGGRAPH Asia) 34, 6 (Oct. 2015), 248:1–248:16
2015
-
[59]
Barron, and Ben Mildenhall
Ben Poole, Ajay Jain, Jonathan T. Barron, and Ben Mildenhall. 2022. Dream- Fusion: Text-to-3D using 2D Diffusion. ArXiv abs/2209.14988 (2022). https: //api.semanticscholar.org/CorpusID:252596091
2022 arXiv
-
[60]
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. 2021. Learning Transferable Visual Models From Natural Language Supervision. arXiv preprint...
2021 arXiv
-
[61]
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. 2022. High-Resolution Image Synthesis With Latent Diffusion Models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 10684–10695
2022
-
[62]
Gyeongsik Moon, Takaaki Shiratori, and Shunsuke Saito. 2024. Expressive whole- body 3D gaussian avatar. In European Conference on Computer Vision
2024
-
[63]
Shunsuke Saito, Zeng Huang, Ryota Natsume, Shigeo Morishima, Angjoo Kanazawa, and Hao Li. 2019. PIFu: Pixel-Aligned Implicit Function for High- Resolution Clothed Human Digitization. 2019 IEEE/CVF International Conference on Computer Vision (ICCV) (2019), 2304–2314. https://ap...
2019
-
[64]
Alex Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam, Pamela Mishkin, Bob McGrew, Ilya Sutskever, and Mark Chen. 2021. GLIDE: Towards Photoreal- istic Image Generation and Editing with Text-Guided Diffusion Models. CoRR abs/2112.10741 (2021). arXiv:2112.10741 https://ar...
2021 arXiv
-
[65]
Yang Song, Liyue Shen, Lei Xing, and Stefano Ermon. 2021. Solving Inverse Problems in Medical Imaging with Score-Based Generative Models.arXiv preprint arXiv:2111.08005 (Nov 2021). arXiv:2111.08005 [eess.IV] https://arxiv.org/abs/ 2111.08005
2021 arXiv
-
[66]
Jiapeng Tang, Davide Davoli, Tobias Kirschstein, Liam Schoneveld, and Matthias Nießner. 2024. GAF: Gaussian Avatar Reconstruction from Monocular Videos via Multi-view Diffusion. ArXiv abs/2412.10209 (2024). https://api.semanticscholar. org/CorpusID:274762758
2024 arXiv
-
[67]
Jiaxiang Tang, Jiawei Ren, Hang Zhou, Ziwei Liu, and Gang Zeng. 2023. Dream- Gaussian: Generative Gaussian Splatting for Efficient 3D Content Creation.ArXiv abs/2309.16653 (2023). https://api.semanticscholar.org/CorpusID:263131552
2023 arXiv
-
[68]
Michał Rzeszewski and Leighton Evans. 2020. Virtual place during quarantine–a curious case of VRChat. Rozwój Regionalny i Polityka Regionalna (2020)
2020
-
[69]
Denis Tomè, Patrick Peluse, Lourdes de Agapito, and Hernán Badino. 2019. xR- EgoPose: Egocentric 3D Human Pose From an HMD Camera. 2019 IEEE/CVF International Conference on Computer Vision (ICCV) (2019), 7727–7737. https: //api.semanticscholar.org/CorpusID:198179395
2019
-
[70]
Yichun Shi, Peng Wang, Jianglong Ye, Mai Long, Kejie Li, and X. Yang. 2023. MVDream: Multi-view Diffusion for 3D Generation. ArXiv abs/2308.16512 (2023). https://api.semanticscholar.org/CorpusID:261395233
2023 arXiv
-
[71]
Xiang Wang, Shiwei Zhang, Changxin Gao, Jiayu Wang, Xiaoqiang Zhou, Yingya Zhang, Luxin Yan, and Nong Sang. 2025. UniAnimate: Taming Unified Video Dif- fusion Models for Consistent Human Image Animation.Science China Information Sciences (2025)
2025
-
[72]
Zhou Wang, Alan C Bovik, Hamid R Sheikh, and Eero P Simoncelli. 2004. Image quality assessment: from error visibility to structural similarity.IEEE Transactions on Image Processing 13, 4 (2004), 600–612
2004
-
[73]
Zhengyi Wang, Cheng Lu, Yikai Wang, Fan Bao, Chongxuan Li, Hang Su, and Jun Zhu. 2023. ProlificDreamer: High-Fidelity and Diverse Text-to-3D Generation with Variational Score Distillation. ArXiv abs/2305.16213 (2023). https://api. semanticscholar.org/CorpusID:258887357
2023 arXiv
-
[74]
Denis Tomè, Thiemo Alldieck, Patrick Peluse, Gerard Pons-Moll, Lourdes de Agapito, Hernán Badino, and Fernando de la Torre. 2020. SelfPose: 3D Ego- centric Pose Estimation From a Headset Mounted Camera. IEEE Transactions on Pattern Analysis and Machine Intelligence 45 (2020), ...
2020
-
[75]
Christopher Wewer, Kevin Raj, Eddy Ilg, Bernt Schiele, and Jan Eric Lenssen
-
[76]
Shuyuan Tu, Zhen Xing, Xintong Han, Zhi-Qi Cheng, Qi Dai, Chong Luo, and Zuxuan Wu. 2024. StableAnimator: High-Quality Identity-Preserving Human Image Animation. arXiv preprint arXiv:2411.17697 (2024)
2024 arXiv
-
[77]
Jiaxin Xie, Hao Ouyang, Jingtan Piao, Chenyang Lei, and Qifeng Chen. 2022. High- fidelity 3D GAN Inversion by Pseudo-multi-view Optimization. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2022), 321–331. https://api.semanticscholar.org/CorpusID:254044374
2022
-
[78]
Zhongcong Xu, Jianfeng Zhang, Jun Hao Liew, Hanshu Yan, Jia-Wei Liu, Chenxu Zhang, Jiashi Feng, and Mike Zheng Shou. 2024. MagicAnimate: Temporally Consistent Human Image Animation using Diffusion Model. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern ...
2024
-
[79]
Ziyang Yuan, Yiming Zhu, Yu Li, Hongyu Liu, and Chun Yuan. 2023. Make Encoder Great Again in 3D GAN Inversion through Geometry and Occlusion- Aware Encoding. 2023 IEEE/CVF International Conference on Computer Vision (ICCV) (2023), 2437–2447. https://api.semanticscholar.org/Cor...
2023
-
[80]
Daniel Watson, William Chan, Ricardo Martin-Brualla, Jonathan Ho, Andrea Tagliasacchi, and Mohammad Norouzi. 2022. Novel View Synthesis with Diffusion Models. arXiv preprint arXiv:2210.04628 (2022). Available at https://arxiv.org/ abs/2210.04628
2022 arXiv
-
[81]
Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shechtman, and Oliver Wang
-
[82]
Yuang Zhang, Jiaxi Gu, Li-Wen Wang, Han Wang, Junqi Cheng, Yuefeng Zhu, and Fangyuan Zou. 2024. MimicMotion: High-Quality Human Motion Video Gen- eration with Confidence-aware Pose Guidance. arXiv preprint arXiv:2406.19680 (2024)
2024 arXiv
-
[83]
Syed, Kawin Setsompop, and Akshay S
Tiange Xiang, Mahmut Yurt, Ali B. Syed, Kawin Setsompop, and Akshay S. Chaud- hari. 2023. DDM2: Self-Supervised Diffusion MRI Denoising with Generative Diffusion Models. ArXiv abs/2302.03018 (2023). https://api.semanticscholar.org/ CorpusID:256615188
2023 arXiv
-
[87]
Lvmin Zhang, Anyi Rao, and Maneesh Agrawala. 2023. Adding Conditional Con- trol to Text-to-Image Diffusion Models. In Proceedings of the IEEE/CVF Interna- tional Conference on Computer Vision (ICCV) . 3813–3824. doi:10.1109/ICCV51070. 2023.00346
2023
-
[91]
Shunyuan Zheng, Boyao Zhou, Ruizhi Shao, Boning Liu, Shengping Zhang, Liqiang Nie, and Yebin Liu. 2023. GPS-Gaussian: Generalizable Pixel-Wise 3D Gaussian Splatting for Real-Time Human Novel View Synthesis. 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (C...
2023
-
[2018]
The Unreasonable Effectiveness of Deep Features as a Perceptual Metric. In CVPR
-
[2020]
2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2020), 5795–5805
pi-GAN: Periodic Implicit Generative Adversarial Networks for 3D-Aware Image Synthesis. 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2020), 5795–5805. https://api.semanticscholar.org/CorpusID: 227247980
2020
-
[2021]
2021 IEEE/CVF International Conference on Computer Vision (ICCV) (2021), 14508– 14518
EgoRenderer: Rendering Human Avatars from Egocentric Camera Images. 2021 IEEE/CVF International Conference on Computer Vision (ICCV) (2021), 14508– 14518. https://api.semanticscholar.org/CorpusID:244037328
2021
-
[2022]
2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2022), 13142–13153
Objaverse: A Universe of Annotated 3D Objects. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2022), 13142–13153. https: //api.semanticscholar.org/CorpusID:254685588
2022
-
[2023]
In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV)
Make-An-Animation: Large-Scale Text-conditional 3D Human Motion Generation. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). 15039–15048
-
[2024]
In European Conference on Computer Vision
latentSplat: Autoencoding Variational Gaussians for Fast Generalizable 3D Reconstruction. In European Conference on Computer Vision . https://api. arXiv Preprint, June 2025, G. Kutay Türkoglu, Julian Tanke, Iheb Belgacem, and Lev Markhasin semanticscholar.org/CorpusID:268681424
2025
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.