REVIEW 4 major objections 5 minor 1 cited by
SimAvatar: Simulation-Ready Avatars with Layered Hair and Clothing
T0 review · 4 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read This paper claims that a text prompt can produce a fully simulation-ready 3D avatar with separate body, garment, and hair layers, each carrying 3D Gaussians for realistic appearance and each driven by a physics or neural simulator when…
desk verdict A solid layered avatar pipeline with a real new garment-diffusion component; the 'fully simulation-ready' claim outruns the evidence. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the layered geometry with per-layer Gaussian attachment. For meshes, each Gaussian stores its position, rotation, and scale in a face-local coordinate frame, so deforming the mesh automatically deforms the Gaussians. For hair, each line segment carries one long thin Gaussian whose parameters are computed directly from the segment endpoints. The garment mesh is produced by a VAE that encodes sampled surface points into a latent code and decodes an unsigned distance field, which is meshed to yield the clean open surface a cloth simulator needs. Appearance learning uses score distillation with separate implicit fields for body, garment, and hair, plus a regularization that forces opacity to decrease from hair root to tip.
What would settle it
Run the pipeline on prompts describing garment types absent from the training set (for example, hooded cloaks, saris, or asymmetric capes) and check whether the decoded meshes are free of holes and self-intersections and whether the garment simulator produces stable sequences without vertex divergence or interpenetration under novel poses.
Extended reading notes
Core claim
The central claim is that the reason earlier text-to-avatar systems produce avatars that look wrong in motion is representation choice: they entangle geometry in a single surface or in implicit fields that simulators cannot consume. SimAvatar's solution is to assign each body part the representation it needs: a parametric body mesh for skinning, a clean non-watertight garment mesh decoded from a learned latent diffusion model, and hair strands. Appearance is then layered on as 3D Gaussians, optimized separately for body, garment, and hair with text-prompt-specific score distillation and a hair opacity regularization that keeps strands connected. The paper reports that the resulting avatars are the first from a text prompt to be fully simulation-ready, animating with realistic cloth and hair dynamics rather than skinning artifacts.
Load-bearing premise
The text-conditioned garment diffusion model, trained on about 20,000 meshes, must generate clean, smooth, non-watertight garments that match arbitrary user prompts and remain valid inputs for the simulator; the paper validates this only qualitatively.
Editorial extensions
If this is right
- Text-generated avatars can be dropped into existing cloth and hair simulation pipelines without manual retopology or mesh cleanup.
- Novel pose sequences produce physically plausible dynamics, such as loose dresses following leg motion and hair flowing, instead of linear blend skinning artifacts.
- Because body, garment, and hair are separate layers, users can edit or recombine them independently.
- The text-conditioned garment diffusion model can be reused as a standalone component for language-driven garment design.
- With a neural simulator for garments, animation at inference time remains fast enough for interactive use.
Reading between the lines
- If the garment model generalizes beyond its roughly 20,000 training meshes, the same layered pipeline could extend to accessories and footwear, which the paper lists as currently entangled with body or garment layers.
- The hair opacity-gradient trick suggests a general recipe for attaching Gaussians to strand-like geometry; it could transfer to fur, grass, or bristle simulation, where broken transparent segments are also a problem.
- Because each layer is optimized against its own prompt, the body appearance field should be reusable across different garments and hairstyles given the same identity, enabling mix-and-match avatars without re-optimization.
- A direct quantitative test of simulation readiness would be measuring simulator stability (e.g., the fraction of prompts whose garment mesh simulates without self-intersection or divergence); the paper does not report such a metric.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. SimAvatar proposes a two-stage framework for text-driven generation of 3D human avatars with layered, simulation-ready geometry. The body is represented by SMPL, the garment by a mesh produced through a VAE plus latent diffusion model trained on ~20,000 GPG/CLOTH3D meshes, and the hair by strands from HAAR; 3D Gaussians are attached to all three layers and optimized with SDS and a hair-opacity regularizer. For animation, the body uses LBS while garments are simulated with HOOD and hair with a strand-based physics simulator. The paper claims this is 'the first to produce highly realistic, fully simulation-ready 3D avatars' and evaluates with qualitative comparisons, a user study, VQAScore, and CLIP-score.
Significance. If the central claim is supported, this is a meaningful advance: existing text-to-avatar methods either entangle all layers in a single geometry or use implicit representations that cannot be directly simulated, whereas SimAvatar's layered design with mesh- and strand-based geometry is a sensible route to combining diffusion-based generation with physics or neural simulation. The text-conditioned garment diffusion model is a useful contribution, and the qualitative results show realistic appearance with plausible wrinkles and hair motion. The paper also includes well-motivated ablations for the hair constraint, prompt engineering, and layer-wise training. However, the paper's headline claim of 'fully simulation-ready' is currently validated only qualitatively; the quantitative evidence is thin and partly circular. These issues are fixable with additional experiments and a more carefully scoped claim, so the work is promising but not yet fully supported.
major comments (4)
- [Sec. 3.2, Sec. 4.3, Appendix A] The central claim in the Abstract and Sec. 1 that avatars are 'fully simulation-ready' requires that every diffusion-generated garment mesh be a valid input to HOOD and remain stable under animation. Section 3.2 describes a VAE plus latent diffusion model trained on roughly 20,000 meshes, but Section 4.3 and Appendix A report only appearance preference and VQA/CLIP scores; there is no simulation success/failure rate across the 22 prompts, no mesh-quality metrics (such as non-manifold edges, self-intersection volume, Laplacian smoothness, or genus), and no failure analysis. HOOD is a learned GNN trained on existing cloth datasets, so out-of-distribution topologies or noisy UDF extractions could produce meshes that are visually plausible but numerically unstable in the simulator. I recommend adding a per-prompt simulation success/failure table and standard mesh-quality statistics, or explicitly scoping the claim to the garment types supported by the training data.
- [Sec. 5] The Conclusions state that simulating garments and hair sequentially 'can fail in certain cases, such as avatars wearing hoods.' This directly conflicts with the unqualified 'fully simulation-ready' claim in the Abstract and Introduction. The authors should either systematically characterize these failure cases and quantify their frequency, or revise the central claim to be conditional on supported garment types. As written, the admitted failure mode undercuts a load-bearing part of the contribution.
- [Sec. 4.3, Table 1] The user study is based on 18 users and 540 votes with no confidence intervals or significance testing, making the reported preference percentages hard to interpret. More importantly, the motion-preference comparison includes Fantasia3D, which the authors themselves state 'cannot be readily animated' in Sec. 4.2; including a non-animatable baseline in a motion-preference study inflates the reported preference. Please report confidence intervals or inter-rater agreement, and either remove Fantasia3D from the motion comparison or justify its inclusion as a deliberate reference point.
- [Sec. 4.1, Appendix A, Table 2] The VQAScore evaluation may be affected by circularity: VQAScore uses an image-to-text foundation model, and the CLOTH3D training data for the garment diffusion model were annotated with prompts generated by GPT-4V (Sec. 4.1). If the VQA evaluator shares the same model family, it may be biased toward the authors' training distribution. Please discuss this potential bias explicitly or use an independent evaluator or model family for the VQA score, since the quantitative claim of 'significantly higher' alignment depends on this metric.
minor comments (5)
- [Title and throughout] The running text contains 'A vatars' with an extra space (e.g., in the title and section headings); this should be corrected to 'Avatars'.
- [Sec. 4.2] The phrase 'start-of-the-art' should be 'state-of-the-art.'
- [Eq. (3)] The expression for the hair Gaussian rotation, 'ri = [1 + µi · di, u× di]', appears malformed: the first component '1 + µi · di' is dimensionally inconsistent, and the intended rotation representation (e.g., axis-angle or quaternion) is not clear. Please check the formula.
- [Appendix A, Table 2] In Table 2, the baseline HumanGaussians is abbreviated as 'HG' in the header but the abbreviation is not defined in the caption or surrounding text; please define it.
- [Figures 7 and 8] The column headers in Figures 7 and 8 appear concatenated (e.g., 'TADAFantasia3DTADAGAvatarHumanGaussians'); this is likely a rendering or formatting issue and should be fixed.
Circularity Check
No significant circularity: the system is assembled from independent learned generators and external simulators; no prediction reduces to its input by construction.
full rationale
SimAvatar is an empirical systems paper; its central claim, "fully simulation-ready avatars," is an engineering result supported by construction, qualitative comparisons, and user studies, not by a deduction that could collapse into its premises. I looked for places where an input is relabeled as a prediction. The garment mesh path (Sec. 3.2) trains a VAE on UDF fields and a latent diffusion model on the GPG/CLOTH3D latent codes; the generated mesh is decoded by MeshUDF and is never used as its own training signal. The appearance module (Sec. 3.3) uses SDS (Eq. 4) with a pretrained image diffusion model and defines Gaussian-to-mesh binding through Eq. 2 and Eq. 3; these are forward kinematic mappings, not fitted parameters that encode the evaluation targets. The simulation stage feeds generated meshes and strands to HOOD and a strand simulator and then transfers motion to the Gaussians; no simulation output is back-substituted as the definition of "simulation-ready." The same-group citations ([23] for the hair simulator and [90] for implicit-field Gaussian attributes) are published tools and empirical observations; the central claim does not reduce to them, and no uniqueness theorem is imported from the authors' prior work. The paper's own Sec. 5 limitation, that sequential hair/garment simulation "can fail in certain cases, such as avatars wearing hoods," is an acknowledged robustness gap, not a circular step. The VQAScore evaluation (Appendix A) uses a generic foundation VLM; the manuscript does not state that this VLM is the same GPT-4V instance used to annotate CLOTH3D (Sec. 4.1), so no evaluator-circularity can be established from the text. The lack of quantitative simulation-success or mesh-quality metrics is a support gap for the "simulation-ready" claim, but it is not circularity: the outputs are not equal to the inputs by construction.
Assumptions & free parameters
free parameters (4)
- lambda_grad =
0.0001
- lambda_KL =
0.1
- lambda_hair =
1.0
- gamma =
0.001
assumptions (5)
- domain assumption The Stable Diffusion-based text-to-image prior (Realistic Vision) provides sufficiently strong appearance supervision for body, hair, and garment layers via SDS.
- domain assumption HOOD neural cloth simulator can simulate the UDF-extracted garment meshes for arbitrary SMPL poses and body shapes.
- domain assumption HAAR and BodyShapeGPT, both external text-conditioned models, produce hair strands and SMPL body parameters consistent with the input prompt and with each other.
- domain assumption MeshUDF reconstruction from the predicted unsigned distance field yields a clean and simulation-ready mesh topology.
- standard math Score Distillation Sampling gradient (Eq. 4) is a valid approximation of the diffusion-model score.
Cite this review
Pith. "Pith review of SimAvatar: Simulation-Ready Avatars with Layered Hair and Clothing." pith.science (2026). https://pith.science/paper/NT37CRIM
@misc{pith2026241209545,
author = {Pith},
title = {Pith review of: SimAvatar: Simulation-Ready Avatars with Layered Hair and Clothing},
year = {2026},
howpublished = {\url{https://pith.science/paper/NT37CRIM}},
note = {Machine review of arXiv:2412.09545}
}
read the original abstract
We introduce SimAvatar, a framework designed to generate simulation-ready clothed 3D human avatars from a text prompt. Current text-driven human avatar generation methods either model hair, clothing, and the human body using a unified geometry or produce hair and garments that are not easily adaptable for simulation within existing simulation pipelines. The primary challenge lies in representing the hair and garment geometry in a way that allows leveraging established prior knowledge from foundational image diffusion models (e.g., Stable Diffusion) while being simulation-ready using either physics or neural simulators. To address this task, we propose a two-stage framework that combines the flexibility of 3D Gaussians with simulation-ready hair strands and garment meshes. Specifically, we first employ three text-conditioned 3D generative models to generate garment mesh, body shape and hair strands from the given text prompt. To leverage prior knowledge from foundational diffusion models, we attach 3D Gaussians to the body mesh, garment mesh, as well as hair strands and learn the avatar appearance through optimization. To drive the avatar given a pose sequence, we first apply physics simulators onto the garment meshes and hair strands. We then transfer the motion onto 3D Gaussians through carefully designed mechanisms for each body part. As a result, our synthesized avatars have vivid texture and realistic dynamic motion. To the best of our knowledge, our method is the first to produce highly realistic, fully simulation-ready 3D avatars, surpassing the capabilities of current approaches.
Figures
Figures from the paper (7 more)
Forward citations
Cited by 1 Pith paper
-
Detangled: A Framework for Creating, Editing, and Inferencing Feature Rich Hair Strands
A 5D texture parameterization plus centerline-based canonical space and supervised diffusion enables generation and texture transfer of feature-rich hair strands independent of style.
Reference graph
Works this paper leans on
-
[1]
https : / / huggingface
Realistic vision. https : / / huggingface . co / stablediffusionapi/realistic- vision- 51 ,
-
[2]
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. GPT-4 technical report. arXiv preprint arXiv:2303.08774, 2023. 7
arXiv 2023
-
[3]
Video based reconstruc- tion of 3d people models
Thiemo Alldieck, Marcus Magnor, Weipeng Xu, Christian Theobalt, and Gerard Pons-Moll. Video based reconstruc- tion of 3d people models. In CVPR, 2018. 3
2018
-
[4]
Learning to re- construct people in clothing from a single rgb camera
Thiemo Alldieck, Marcus Magnor, Bharat Lal Bhatnagar, Christian Theobalt, and Gerard Pons-Moll. Learning to re- construct people in clothing from a single rgb camera. In CVPR, 2019
2019
-
[5]
Tex2shape: Detailed full human body geometry from a single image
Thiemo Alldieck, Gerard Pons-Moll, Christian Theobalt, and Marcus Magnor. Tex2shape: Detailed full human body geometry from a single image. In ICCV, 2019. 3
2019
-
[6]
Bodyshapegpt: Smpl body shape manipulation with llms
Baldomero R ´Arbol and Dan Casas. Bodyshapegpt: Smpl body shape manipulation with llms. In ECCVW, 2024. 6
2024
-
[7]
Large steps in cloth sim- ulation
David Baraff and Andrew Witkin. Large steps in cloth sim- ulation. In Conference on Computer Graphics and Interac- tive Techniques, 1998. 3
1998
-
[8]
Discrete elastic rods
Mikl ´os Bergou, Max Wardetzky, Stephen Robinson, Basile Audoly, and Eitan Grinspun. Discrete elastic rods. In SIG- GRAPH, 2008. 4
2008
Show all 101 references
-
[9]
Discrete viscous threads
Mikl ´os Bergou, Basile Audoly, Etienne V ouga, Max Wardetzky, and Eitan Grinspun. Discrete viscous threads. ACM Transactions on graphics (TOG), 2010. 4
2010
-
[10]
Cloth3d: clothed 3d humans
Hugo Bertiche, Meysam Madadi, and Sergio Escalera. Cloth3d: clothed 3d humans. In ECCV, 2020. 3, 4, 7
2020
-
[11]
PBNS: Physically based neural simulator for unsupervised garment pose space deformation
Hugo Bertiche, Meysam Madadi, and Sergio Escalera. PBNS: Physically based neural simulator for unsupervised garment pose space deformation. ACM Transaction on Graphics (TOG), 2020
2020
-
[12]
Deepsd: Automatic deep skinning and pose space deformation for 3d garment animation
Hugo Bertiche, Meysam Madadi, Emilio Tylson, and Ser- gio Escalera. Deepsd: Automatic deep skinning and pose space deformation for 3d garment animation. In ICCV,
-
[13]
Es- timating cloth simulation parameters from video
Kiran S Bhat, Christopher D Twigg, Jessica K Hodgins, Pradeep Khosla, Zoran Popovic, and Steven M Seitz. Es- timating cloth simulation parameters from video. SIG- GRAPH, 2003. 3
2003
-
[14]
Multi-garment net: Learning to dress 3d people from images
Bharat Lal Bhatnagar, Garvita Tiwari, Christian Theobalt, and Gerard Pons-Moll. Multi-garment net: Learning to dress 3d people from images. In ICCV, 2019. 3
2019
-
[15]
Loopreg: Self-supervised learning of implicit surface correspondences, pose and shape for 3d human mesh registration
Bharat Lal Bhatnagar, Cristian Sminchisescu, Christian Theobalt, and Gerard Pons-Moll. Loopreg: Self-supervised learning of implicit surface correspondences, pose and shape for 3d human mesh registration. In NeurIPS, 2020. 3
2020
-
[16]
DreamAvatar: Text-and-Shape Guided 3D Human Avatar Generation via Diffusion Models
Yukang Cao, Yan-Pei Cao, Kai Han, Ying Shan, and Kwan- Yee K Wong. DreamAvatar: Text-and-Shape Guided 3D Human Avatar Generation via Diffusion Models. In CVPR,
-
[17]
Neural- ABC: Neural parametric models for articulated body with clothes
Honghu Chen, Yuxin Yao, and Juyong Zhang. Neural- ABC: Neural parametric models for articulated body with clothes. IEEE Transactions on Visualization and Computer Graphics, 2024. 3, 5
2024
-
[18]
Fanta- sia3D: Disentangling Geometry and Appearance for High- quality Text-to-3D Content Creation
Rui Chen, Yongwei Chen, Ningxin Jiao, and Kui Jia. Fanta- sia3D: Disentangling Geometry and Appearance for High- quality Text-to-3D Content Creation. In ICCV, 2023. 2, 4, 8, 13
2023
-
[19]
Stable but respon- sive cloth
Kwang-Jin Choi and Hyeong-Seok Ko. Stable but respon- sive cloth. In ACM SIGGRAPH Courses. 2005. 3
2005
-
[20]
Smplicit: Topology-aware generative model for clothed people
Enric Corona, Albert Pumarola, Guillem Alenya, Ger- ard Pons-Moll, and Francesc Moreno-Noguer. Smplicit: Topology-aware generative model for clothed people. In CVPR, 2021. 3
2021
-
[21]
https://openai.com/dall-e-2, 2022
Dalle2. https://openai.com/dall-e-2, 2022. 2
2022
-
[22]
Simple and scalable frictional contacts for thin nodal objects
Gilles Daviet. Simple and scalable frictional contacts for thin nodal objects. ACM Transactions on Graphics (TOG),
-
[23]
Interactive hair simulation on the gpu using admm
Gilles Daviet. Interactive hair simulation on the gpu using admm. In SIGGRAPH, 2023. 2, 3, 4
2023
-
[24]
Interactive hair simulation on the gpu using admm
Gilles Daviet. Interactive hair simulation on the gpu using admm. In SIGGRAPH, 2023. 4
2023
-
[25]
DrapeNet: Garment Generation and Self-Supervised Draping
Luca De Luigi, Ren Li, Benoit Guillard, Mathieu Salz- mann, and Pascal Fua. DrapeNet: Garment Generation and Self-Supervised Draping. In CVPR, 2023. 3, 5, 7
2023
-
[26]
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2018. 5
2018 arXiv
-
[27]
TELA: Text to layer-wise 3d clothed human generation
Junting Dong, Qi Fang, Zehuan Huang, Xudong Xu, Jingbo Wang, Sida Peng, and Bo Dai1. TELA: Text to layer-wise 3d clothed human generation. In ECCV, 2024. 2, 3
2024
-
[28]
A multi-scale model for simulating liquid-hair interactions
Yun (Raymond) Fei, Henrique Teles Maia, Christopher Batty, Changxi Zheng, and Eitan Grinspun. A multi-scale model for simulating liquid-hair interactions. ACM Trans- action on Graphics (TOG), 2017. 3
2017
-
[29]
Learning deformable tetrahedral meshes for 3d reconstruction
Jun Gao, Wenzheng Chen, Tommy Xiang, Alec Jacobson, Morgan McGuire, and Sanja Fidler. Learning deformable tetrahedral meshes for 3d reconstruction. InNeurIPS, 2020. 3
2020
-
[30]
HOOD: Hierarchical graphs for gener- alized modelling of clothing dynamics
Artur Grigorev, Bernhard Thomaszewski, Michael J Black, and Otmar Hilliges. HOOD: Hierarchical graphs for gener- alized modelling of clothing dynamics. In CVPR, 2023. 2, 3, 4
2023
-
[31]
Drape: Dressing any person
Peng Guan, Loretta Reiss, David A Hirshberg, Alexander Weiss, and Michael J Black. Drape: Dressing any person. SIGGRAPH, 2012. 3
2012
-
[32]
Meshudf: Fast and differentiable meshing of unsigned distance field networks
Benoit Guillard, Federico Stella, and Pascal Fua. Meshudf: Fast and differentiable meshing of unsigned distance field networks. In ECCV, 2022. 3, 5
2022
-
[33]
Dresscode: Autoregressively sewing and gen- erating garments from text guidance
Kai He, Kaixin Yao, Qixuan Zhang, Jingyi Yu, Lingjie Liu, and Lan Xu. Dresscode: Autoregressively sewing and gen- erating garments from text guidance. ACM Transactions on Graphics (TOG), 2024. 7
2024
-
[34]
AvatarCLIP: Zero-Shot Text-Driven Generation and Animation of 3D Avatars.SIG- GRAPH, 2022
Fangzhou Hong, Mingyuan Zhang, Liang Pan, Zhongang Cai, Lei Yang, and Ziwei Liu. AvatarCLIP: Zero-Shot Text-Driven Generation and Animation of 3D Avatars.SIG- GRAPH, 2022. 3
2022
-
[35]
Sag-free initialization for strand- based hybrid hair simulation
Jerry Hsu, Tongtong Wang, Zherong Pan, Xifeng Gao, Cem Yuksel, and Kui Wu. Sag-free initialization for strand- based hybrid hair simulation. ACM Transaction of Graph- ics, 2023. 3
2023
-
[36]
Human- liff: Layer-wise 3d human generation with diffusion model
Shoukang Hu, Fangzhou Hong, Tao Hu, Liang Pan, Haiyi Mei, Weiye Xiao, Lei Yang, and Ziwei Liu. Human- liff: Layer-wise 3d human generation with diffusion model. arXiv preprint, 2023. 2, 3
2023
-
[37]
Humannorm: Learning normal diffusion model for high-quality and realistic 3d human generation
Xin Huang, Ruizhi Shao, Qi Zhang, Hongwen Zhang, and Ying Feng. Humannorm: Learning normal diffusion model for high-quality and realistic 3d human generation. In CVPR, 2024. 2
2024
-
[38]
Dreamwaltz: Make a scene with complex 3d animatable avatars
Yukun Huang, Jianan Wang, Ailing Zeng, He Cao, Xi- anbiao Qi, Yukai Shi, Zheng-Jun Zha, and Lei Zhang. Dreamwaltz: Make a scene with complex 3d animatable avatars. In NeurIPS, 2023. 3
2023
-
[39]
TeCH: Text- guided Reconstruction of Lifelike Clothed Humans
Yangyi Huang, Hongwei Yi, Yuliang Xiu, Tingting Liao, Jiaxiang Tang, Deng Cai, and Justus Thies. TeCH: Text- guided Reconstruction of Lifelike Clothed Humans. In 3DV, 2024. 2
2024
-
[40]
ClipMatrix: Text-controlled creation of 3d textured meshes
Nikolay Jetchev. ClipMatrix: Text-controlled creation of 3d textured meshes. arXiv preprint arXiv:2307.05663, 2023. 2
2023 arXiv
-
[41]
Bcnet: Learning body and cloth shape from a single image
Boyi Jiang, Juyong Zhang, Yang Hong, Jinhao Luo, Ligang Liu, and Hujun Bao. Bcnet: Learning body and cloth shape from a single image. In ECCV, 2020. 3
2020
-
[42]
Avatar- Craft: Transforming text into neural human avatars with parameterized shape and pose control
Ruixiang Jiang, Can Wang, Jingbo Zhang, Menglei Chai, Mingming He, Dongdong Chen, and Jing Liao. Avatar- Craft: Transforming text into neural human avatars with parameterized shape and pose control. In ICCV, 2023. 2, 3
2023
-
[43]
Elucidating the design space of diffusion-based generative models
Tero Karras, Miika Aittala, Timo Aila, and Samuli Laine. Elucidating the design space of diffusion-based generative models. In NeurIPS, 2022. 5
2022
-
[44]
3d gaussian splatting for real-time radiance field rendering
Bernhard Kerbl, Georgios Kopanas, Thomas Leimk ¨uhler, and George Drettakis. 3d gaussian splatting for real-time radiance field rendering. SIGGRAPH, 2023. 6
2023
-
[45]
GALA: Generating animatable layered assets from a single scan
Taeksoo Kim, Byungjun Kim, Shunsuke Saito, and Han- byul Joo. GALA: Generating animatable layered assets from a single scan. In CVPR, 2024. 3
2024
-
[46]
Dreamhuman: Animatable 3d avatars from text
Nikos Kolotouros, Thiemo Alldieck, Andrei Zanfir, Ed- uard Gabriel Bazavan, Mihai Fieraru, and Cristian Smin- chisescu. Dreamhuman: Animatable 3d avatars from text. In NeurIPS, 2023. 2, 3, 8
2023
-
[47]
Generating datasets of 3d garments with sewing patterns
Maria Korosteleva and Sung-Hee Lee. Generating datasets of 3d garments with sewing patterns. In NeurIPS, 2021. 4
2021
-
[48]
Generating datasets of 3d garments with sewing patterns
Maria Korosteleva and Sung-Hee Lee. Generating datasets of 3d garments with sewing patterns. In NeurIPS, 2021. 7
2021
-
[49]
Diffcloth: Differentiable cloth simulation with dry fric- tional contact
Yifei Li, Tao Du, Kui Wu, Jie Xu, and Wojciech Matusik. Diffcloth: Differentiable cloth simulation with dry fric- tional contact. SIGGRAPH, 2022. 3
2022
-
[50]
Diffavatar: Simulation-ready garment optimization with differentiable simulation
Yifei Li, Hsiao-yu Chen, Egor Larionov, Nikolaos Sarafi- anos, Wojciech Matusik, and Tuur Stuyck. Diffavatar: Simulation-ready garment optimization with differentiable simulation. In CVPR, 2024. 3
2024
-
[51]
Lin, and Vladlen Koltun
Junbang Liang, Ming C. Lin, and Vladlen Koltun. Differ- entiable cloth simulation for inverse problems. In NeurIPS,
-
[52]
TADA! text to animatable digital avatars
Tingting Liao, Hongwei Yi, Yuliang Xiu, Jiaxiang Tang, Yangyi Huang, Justus Thies, and Michael J Black. TADA! text to animatable digital avatars. In 3DV, 2024. 2, 3, 4, 8, 13
2024
-
[53]
Magic3D: High- Resolution Text-to-3D Content Creation
Chen-Hsuan Lin, Jun Gao, Luming Tang, Towaki Takikawa, Xiaohui Zeng, Xun Huang, Karsten Kreis, Sanja Fidler, Ming-Yu Liu, and Tsung-Yi Lin. Magic3D: High- Resolution Text-to-3D Content Creation. In CVPR, 2023. 2
2023
-
[54]
Evaluating text-to-visual generation with image- to-text generation
Zhiqiu Lin, Deepak Pathak, Baiqi Li, Jiayao Li, Xide Xia, Graham Neubig, Pengchuan Zhang, and Deva Ra- manan. Evaluating text-to-visual generation with image- to-text generation. arXiv preprint arXiv:2404.01291, 2024. 13
2024 arXiv
-
[55]
Humangaus- sian: Text-driven 3d human generation with gaussian splat- ting
Xian Liu, Xiaohang Zhan, Jiaxiang Tang, Ying Shan, Gang Zeng, Dahua Lin, Xihui Liu, and Ziwei Liu. Humangaus- sian: Text-driven 3d human generation with gaussian splat- ting. In CVPR, 2024. 2, 6, 13
2024
-
[56]
Matthew Loper, Naureen Mahmood, Javier Romero, Ger- ard Pons-Moll, and Michael J. Black. SMPL: A skinned multi-person linear model. ToG, 34(6):248:1–248:16, 2015. 2, 3, 4, 6
2015
-
[57]
Gaussianhair: Hair modeling and rendering with light- aware gaussians
Haimin Luo, Min Ouyang, Zijun Zhao, Suyi Jiang, Long- wen Zhang, Qixuan Zhang, Wei Yang, Lan Xu, and Jingyi Yu. Gaussianhair: Hair modeling and rendering with light- aware gaussians. arXiv preprint arXiv:2402.10483, 2024. 4, 6
2024 arXiv
-
[58]
Learn- ing to dress 3d people in generative clothing
Qianli Ma, Jinlong Yang, Anurag Ranjan, Sergi Pujades, Gerard Pons-Moll, Siyu Tang, and Michael J Black. Learn- ing to dress 3d people in generative clothing. In CVPR,
-
[59]
Srinivasan, Matthew Tancik, Jonathan T
Ben Mildenhall, Pratul P. Srinivasan, Matthew Tancik, Jonathan T. Barron, Ravi Ramamoorthi, and Ren Ng. NeRF: Representing scenes as neural radiance fields for view synthesis. In ECCV, 2020. 2
2020
-
[60]
Predicting loose-fitting garment de- formations using bone-driven motion networks
Xiaoyu Pan, Jiaming Mai, Xinwei Jiang, Dongxue Tang, Jingxiang Li, Tianjia Shao, Kun Zhou, Xiaogang Jin, and Dinesh Manocha. Predicting loose-fitting garment de- formations using bone-driven motion networks. In SIG- GRAPH, 2022. 3
2022
-
[61]
Tailornet: Predicting clothing in 3d as a function of human pose, shape and garment style
Chaitanya Patel, Zhouyingcheng Liao, and Gerard Pons- Moll. Tailornet: Predicting clothing in 3d as a function of human pose, shape and garment style. In CVPR, 2020. 3
2020
-
[62]
Kry, and Kaleem Siddiqi
Emmanuel Piuze, Paul G. Kry, and Kaleem Siddiqi. Gen- eralized helicoids for modeling hair geometry. Computer Graphics Forum, 2011. 3
2011
-
[63]
ClothCap: Seamless 4d clothing capture and retar- geting
Gerard Pons-Moll, Sergi Pujades, Sonny Hu, and Michael J Black. ClothCap: Seamless 4d clothing capture and retar- geting. SIGGRAPH, 2017. 3
2017
-
[64]
DreamFusion: Text-to-3d using 2d diffusion
Ben Poole, Ajay Jain, Jonathan T Barron, and Ben Milden- hall. DreamFusion: Text-to-3d using 2d diffusion. InICLR,
-
[65]
Gaussianavatars: Photorealistic head avatars with rigged 3d gaussians
Shenhan Qian, Tobias Kirschstein, Liam Schoneveld, Da- vide Davoli, Simon Giebenhain, and Matthias Nießner. Gaussianavatars: Photorealistic head avatars with rigged 3d gaussians. In CVPR, 2024. 6
2024
-
[66]
DIG: Draping Implicit Garment over the Human Body
Li Ren, Benoit Guillard, Edoardo Remelli, and Pascal Fua. DIG: Draping Implicit Garment over the Human Body. In ACCV, 2022. 3
2022
-
[67]
Texture: Text-guided texturing of 3d shapes
Elad Richardson, Gal Metzer, Yuval Alaluf, Raja Giryes, and Daniel Cohen-Or. Texture: Text-guided texturing of 3d shapes. SIGGRAPH, 2023. 2
2023
-
[68]
High-resolution image synthesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bj ¨orn Ommer. High-resolution image synthesis with latent diffusion models. In CVPR, 2022. 2
2022
-
[69]
Photorealistic text-to-image diffusion models with deep language understanding
Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily L Denton, Kamyar Ghasemipour, Raphael Gontijo Lopes, Burcu Karagol Ayan, Tim Sali- mans, et al. Photorealistic text-to-image diffusion models with deep language understanding. In NeurIPS, 2022. 2
2022
-
[70]
Learning-based animation of clothing for virtual try-on
Igor Santesteban, Miguel A Otaduy, and Dan Casas. Learning-based animation of clothing for virtual try-on. Computer Graphics Forum, 38(2):355–366, 2019. 3
2019
-
[71]
Otaduy, and Dan Casas
Igor Santesteban, Nils Thuerey, Miguel A. Otaduy, and Dan Casas. Self-supervised collision handling via generative 3d garment models for virtual try-on. In CVPR, 2021. 3
2021
-
[72]
Self-Supervised Collision Handling via Generative 3D Garment Models for Virtual Try-On
Igor Santesteban, Nils Thuerey, Miguel A Otaduy, and Dan Casas. Self-Supervised Collision Handling via Generative 3D Garment Models for Virtual Try-On. In CVPR, 2021. 3
2021
-
[73]
SNUG: Self-Supervised Neural Dynamic Garments
Igor Santesteban, Miguel A Otaduy, and Dan Casas. SNUG: Self-Supervised Neural Dynamic Garments. In CVPR, 2022. 3
2022
-
[74]
Otaduy, Nils Thuerey, and Dan Casas
Igor Santesteban, Miguel A. Otaduy, Nils Thuerey, and Dan Casas. ULNeF: Untangled layered neural fields for mix- and-match virtual try-on. In NeurIPS, 2022. 3
2022
-
[75]
Garment3dgen: 3d garment stylization and texture generation
Nikolaos Sarafianos, Tuur Stuyck, Xiaoyu Xiang, Yilei Li, Jovan Popovic, and Rakesh Ranjan. Garment3dgen: 3d garment stylization and texture generation. arXiv preprint arXiv:2403.18816, 2024. 3
2024 arXiv
-
[76]
CT2Hair: High-fidelity 3d hair modeling using com- puted tomography
Yuefan Shen, Shunsuke Saito, Ziyan Wang, Olivier Maury, Chenglei Wu, Jessica Hodgins, Youyi Zheng, and Giljoo Nam. CT2Hair: High-fidelity 3d hair modeling using com- puted tomography. Transactions on Graphics, 2023. 3
2023
-
[77]
Neural Haircut: Prior-Guided Strand-Based Hair Reconstruction
Vanessa Sklyarova, Jenya Chelishev, Andreea Dogaru, Igor Medvedev, Victor Lempitsky, and Egor Zakharov. Neural Haircut: Prior-Guided Strand-Based Hair Reconstruction. In ICCV, 2023. 3
2023
-
[78]
Black, and Justus Thies
Vanessa Sklyarova, Egor Zakharov, Otmar Hilliges, Michael J. Black, and Justus Thies. Text-conditioned gener- ative model of 3d strand-based human hairstyles. In CVPR,
-
[79]
Deep- cloth: Neural garment representation for shape and style editing
Zhaoqi Su, Tao Yu, Yangang Wang, and Yebin Liu. Deep- cloth: Neural garment representation for shape and style editing. TPAMI, 2023. 3
2023
-
[80]
Dreamgaussian: Generative gaussian splat- ting for efficient 3d content creation
Jiaxiang Tang, Jiawei Ren, Hang Zhou, Ziwei Liu, and Gang Zeng. Dreamgaussian: Generative gaussian splat- ting for efficient 3d content creation. arXiv preprint arXiv:2309.16653, 2023. 2
2023 arXiv
-
[81]
Thomaszewski, Simon Pabst, and Wolfgang Straßer
B. Thomaszewski, Simon Pabst, and Wolfgang Straßer. Asynchronous cloth simulation. 2008. 3
2008
-
[82]
A simple approach to nonlinear tensile stiffness for accurate cloth simulation
Pascal V olino, Nadia Magnenat-Thalmann, and Francois Faure. A simple approach to nonlinear tensile stiffness for accurate cloth simulation. Transaction of Graphics (TOG),
-
[83]
O’Brien, and Ravi Ramamoorthi
Huamin Wang, James F. O’Brien, and Ravi Ramamoorthi. Data-driven elastic models for cloth: Modeling and mea- surement. In SIGGRAPH, 2011. 3
2011
-
[84]
Disentangled clothed avatar generation from text descriptions
Jionghao Wang, Yuan Liu, Zhiyang Dou, Zhengming Yu, Yongqing Liang, Xin Li, Wenping Wang, Rong Xie, and Li Song. Disentangled clothed avatar generation from text descriptions. arXiv preprint:2312.05295, 2023. 2, 3, 4, 6
2023 arXiv
-
[85]
Neus: Learning neural implicit surfaces by volume rendering for multi-view recon- struction
Peng Wang, Lingjie Liu, Yuan Liu, Christian Theobalt, Taku Komura, and Wenping Wang. Neus: Learning neural implicit surfaces by volume rendering for multi-view recon- struction. NeurIPS, 34:27171–27183, 2021. 2, 3
2021
-
[86]
Humancoser: Layered 3d hu- man generation via semantic-aware diffusion model
Yi Wang, Jian Ma, Ruizhi Shao, Qiao Feng, Yu-Kun Lai, Yebin Liu, and Kun Li. Humancoser: Layered 3d hu- man generation via semantic-aware diffusion model. arXiv preprint: 2312.05804, 2023. 3
2023 arXiv
-
[87]
Prolificdreamer: High- fidelity and diverse text-to-3d generation with variational score distillation
Zhengyi Wang, Cheng Lu, Yikai Wang, Fan Bao, Chongx- uan Li, Hang Su, and Jun Zhu. Prolificdreamer: High- fidelity and diverse text-to-3d generation with variational score distillation. In NeurIPS, 2023. 2
2023
-
[88]
Direct3d: Scalable image-to-3d generation via 3d latent diffusion transformer
Shuang Wu, Youtian Lin, Feihu Zhang, Yifei Zeng, Jingxi Xu, Philip Torr, Xun Cao, and Yao Yao. Direct3d: Scalable image-to-3d generation via 3d latent diffusion transformer. arXiv preprint arXiv:2405.14832, 2024. 4
2024 arXiv
-
[89]
Dream3d: Zero- shot text-to-3d synthesis using 3d shape prior and text-to- image diffusion models
Jiale Xu, Xintao Wang, Weihao Cheng, Yan-Pei Cao, Ying Shan, Xiaohu Qie, and Shenghua Gao. Dream3d: Zero- shot text-to-3d synthesis using 3d shape prior and text-to- image diffusion models. In Computer Vision and Pattern Recognition (CVPR), 2023. 2
2023
-
[90]
Gavatar: Ani- matable 3d gaussian avatars with implicit mesh learning
Ye Yuan, Xueting Li, Yangyi Huang, Shalini De Mello, Koki Nagano, Jan Kautz, and Umar Iqbal. Gavatar: Ani- matable 3d gaussian avatars with implicit mesh learning. In CVPR, 2024. 2, 3, 4, 6, 8, 13
2024
-
[91]
Hair meshes
Cem Yuksel, Scott Schaefer, and John Keyser. Hair meshes. ACM Trans. Graph., 2009. 3
2009
-
[92]
Point-based modeling of human clothing
Ilya Zakharkin, Kirill Mazur, Artur Grigorev, and Victor Lempitsky. Point-based modeling of human clothing. In CVPR, 2021. 3
2021
-
[93]
Human hair recon- struction with strand-aligned 3d gaussians
Egor Zakharov, Vanessa Sklyarova, Michael J Black, Giljoo Nam, Justus Thies, and Otmar Hilliges. Human hair recon- struction with strand-aligned 3d gaussians. In ECCV, 2024. 4, 6
2024
-
[94]
Avatarbooth: High-quality and customizable 3d human avatar generation
Yifei Zeng, Yuanxun Lu, Xinya Ji, Yao Yao, Hao Zhu, and Xun Cao. Avatarbooth: High-quality and customizable 3d human avatar generation. arXiv preprint: 2306.09864 ,
-
[95]
3dshape2vecset: A 3d shape representation for neural fields and generative diffusion models
Biao Zhang, Jiapeng Tang, Matthias Niessner, and Peter Wonka. 3dshape2vecset: A 3d shape representation for neural fields and generative diffusion models. ACM Trans- actions on Graphics (TOG), 2023. 4, 5
2023
-
[96]
Detailed, accurate, human shape estimation from clothed 3d scan sequences
Chao Zhang, Sergi Pujades, Michael J Black, and Gerard Pons-Moll. Detailed, accurate, human shape estimation from clothed 3d scan sequences. In CVPR, 2017. 3
2017
-
[97]
Hao Zhang, Yao Feng, Peter Kulits, Yandong Wen, Justus Thies, and Michael J. Black. Teca: Text-guided generation and editing of compositional 3d avatars. arXiv, 2023. 2, 3
2023
-
[98]
Avatarverse: High-quality & stable 3d avatar cre- ation from text and pose
Huichao Zhang, Bowen Chen, Hao Yang, Liao Qu, Xu Wang, Li Chen, Chao Long, Feida Zhu, Kang Du, and Min Zheng. Avatarverse: High-quality & stable 3d avatar cre- ation from text and pose. In AAAI, 2024. 2, 3
2024
-
[99]
Adding conditional control to text-to-image diffusion models
Lvmin Zhang, Anyi Rao, and Maneesh Agrawala. Adding conditional control to text-to-image diffusion models. In ICCV, 2023. 2, 3
2023
-
[100]
Clay: A controllable large-scale generative model for creating high-quality 3d assets
Longwen Zhang, Ziyu Wang, Qixuan Zhang, Qiwei Qiu, Anqi Pang, Haoran Jiang, Wei Yang, Lan Xu, and Jingyi Yu. Clay: A controllable large-scale generative model for creating high-quality 3d assets. ACM Transactions on Graphics (TOG), 2024. 4
2024
-
[101]
#$ (b) no !!
Xuanmeng Zhang, Jianfeng Zhang, Chacko Rohan, Hongyi Xu, Guoxian Song, Yi Yang, and Jiashi Feng. Getavatar: Generative textured meshes for animatable human avatars. In ICCV, 2023. 2 A. Quantitative Evaluations VQAScore We quantitatively compare our method against baseline meth...
2023
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.