REVIEW 3 major objections 7 minor 74 references
FreeCloth: Free-form Generation Enhances Challenging Clothed Human Modeling
T0 review · 3 major / 7 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read The paper claims that LBS-based warping fails for loose garments because canonicalizing points far from the body is ill-defined, and that a hybrid framework which generates those regions instead of warping them achieves state-of-the-art…
desk verdict Hybrid LBS + free-form generation is a genuinely new idea with strong benchmark results, but the fixed clothing-cut mask is an untested correctness condition that needs scrutiny. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the garment-specific clothing-cut map, a per-subject labeling computed by running the SAM segmentation model on rendered front and back normal maps of a single near-canonical frame and back-projecting the loose-clothing pixels onto the UV map; it decides which body points go to which branch. On the generated side, the workhorse is a modified SpareNet-style free-form generator whose pose encoder is PointNet++ applied separately to four leg parts (left and right upper and lower legs), pooled into a structure-aware pose code, and combined with a shared global garment code to synthesize posed-space points without any LBS transformation. The deformed side uses a PointNet++ pose encoder, barycentric-interpolated per-point pose and garment codes, and an LBS-based pose decoder predicting canonical-space displacement and normals. A Chamfer and normal reconstruction loss, displacement and code regularization, and a signed-distance collision loss tie the branches together.
What would settle it
Re-run FreeCloth on ReSynth test sequences at extreme poses, such as deep hip flexion, a high leg lift, or a strong twist, while computing a fresh clothing-cut map per frame with the same SAM procedure, and compare FID, point-cloud completeness, and seam artifacts against the fixed-map version; if fresh maps substantially eliminate tears and penetrations at out-of-distribution poses while the fixed map fails, the fixed-map assumption is the load-bearing weakness rather than the free-form generator.
Extended reading notes
Core claim
The central claim is that loose clothing is not a deformation problem but a generation problem. Existing LBS-based methods predict pose-dependent offsets in canonical space and then skin them to the body; when the garment is far from the body, canonicalization is ill-defined and the result splits into fragments or collapses into pant-like shapes. FreeCloth's discovery is that a fixed, garment-specific segmentation, computed once per subject with SAM on a near-canonical frame, can cleanly separate the body into points that should be replicated, deformed, or generated. The generated region is synthesized directly in posed space by a style-based point generator conditioned on part-wise pose features from the legs and a global garment code, bypassing LBS entirely. Merging the LBS-deformed near-body points with generated loose points, under a collision loss, yields continuous, wrinkle-rich skirts and dresses and eliminates the split-up, open-surface, and over-bent-wrinkle artifacts of prior point-based methods on the hardest cases.
Load-bearing premise
The load-bearing premise is that the once-computed per-subject clothing-cut map keeps correctly labeling where cloth is loose and where it is tight across all poses, so no re-segmentation is needed when the garment shifts.
Editorial extensions
If this is right
- On the five ReSynth loose-clothing subjects, FreeCloth records the best average FID of 37.75 and MSE of 2.61 among POP, SkiRT, FITE, and itself, and the perceptual study shows its largest advantage on the loosest skirts and dresses.
- Because free-form generation does not need a clothing template, a 2D positional map, or a continuous LBS field, the pipeline remains single-stage and end-to-end while still representing open surfaces such as the underside of a skirt.
- The generator's garment code supports multi-subject modeling and interpolation along skirt length and tightness, indicating that pose and garment style can be controlled as separate factors.
- Inference runs at 64.1 FPS with 10.83M parameters on an RTX 3090, so the fidelity gain over the strongest baseline is not bought with added compute.
- For garments whose loose parts are not lower-body, such as suit collars, the free-form generator still learns to represent those loose components, suggesting the hybrid idea generalizes beyond skirts.
Reading between the lines
- The fixed clothing-cut map is computed from one near-canonical frame, so the most direct stress test is pose generalization: if a pose shifts the hem or exposes the legs, the boundary between deformed and generated regions may tear, and the paper's own ablation shows that map is essential.
- The part-aware pose code uses only four leg parts, so the claimed expressiveness is explicitly tuned to lower-body garments; transferring the same free-form idea to sleeves, capes, or hoods would require choosing analogous semantic parts.
- The paper's argument that Chamfer distance mis-scores valid clothing states suggests distribution-based metrics like FID should become standard for this task, and one could test whether the method's advantage persists on real-world scans or video with a wider pose distribution.
- The authors' stated future combination with 3D Gaussian Splatting could turn the point-cloud garment into a textured avatar, but that integration would likely expose seam artifacts at the deformed and generated boundary that the current surfel-rendered evaluation may not fully reveal.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes FreeCloth, a hybrid framework for pose-dependent clothed human modeling that combines LBS-based deformation for near-body clothing regions with a novel free-form generator for loose regions such as skirts and dresses. A garment-specific clothing-cut map, computed once per subject via SAM segmentation on a single near-canonical frame, assigns each body surface point to unclothed, deformed, or generated categories. The method is trained end-to-end on the ReSynth dataset with Chamfer, normal, regularization, and collision losses, and evaluated on five subjects using FID, MSE, perceptual studies, and GPT-4o preference. The authors report state-of-the-art results on average, with notable gains on the loosest garments, and claim to be the first to use free-form generation for learning-based clothed human modeling.
Significance. If the results hold, the paper introduces a novel and plausible paradigm: replacing LBS for loose garments with free-form point generation, which can avoid the split and pant-like artifacts of LBS-only methods. The paper includes useful ablations (hybrid design, collision loss, part-aware encoding, clothing-cut map), a code release, an efficiency analysis showing real-time inference, and a perceptual study. The main strengths are the clear problem decomposition and the empirical evidence of visual improvements on challenging loose garments. However, the central assumption of a fixed clothing-cut map across poses is not stress-tested, and the quantitative claims rely on a small evaluation without statistical guarantees.
major comments (3)
- [Sec. 3.1, Supp. A.2] The clothing-cut map is computed from a single near-canonical frame and then fixed for all training and test poses. This is load-bearing because the map determines which body vertices are omitted and replaced by free-form points (Eq. 8), and the ablation in Fig. 7 shows the map is critical to the merger. The paper does not test whether the boundary between deformed and generated regions shifts with pose, for example when a skirt hem lifts or a slit opens. The failure analysis in Sec. B.8 reports only penetration and seam artifacts, not mask-shift failures. Please provide an experiment that measures the stability of the segmentation mask across the pose distribution (e.g., IoU of SAM segments per pose, or reconstruction error in regions near the garment boundary), and report whether the fixed map causes missing body parts or tears in extreme poses.
- [Table 2, Supp. A.1, Supp. B.6] The quantitative evaluation reports single-run FID/MSE values without error bars or significance tests, and it is limited to five subjects. Moreover, hyperparameters are selected using the evaluation subjects: Ng is set per garment type in Supp. A.1, and K=8 is chosen via ablation on a long dress in Supp. B.6, which is itself an evaluation subject. Please provide multiple seeds or confidence intervals for the reported metrics, and either perform hyperparameter selection on a held-out split or report the sensitivity of the results to K and Ng on the test set.
- [Table 2, abstract] The abstract claims 'state-of-the-art performance' and 'particularly in the most challenging cases,' but on the loosest subject (felice-004) FreeCloth has worse FID than FITE (42.41 vs 38.61), and on anna-001 it is also worse (39.63 vs 38.21). The average FID improvement over FITE is small (37.75 vs 39.02). Please clarify whether the SOTA claim refers to averages, and discuss these per-subject exceptions, or provide evidence that the differences are significant given the variance.
minor comments (7)
- [Introduction] The word 'skidding' in the first paragraph should be 'skinning.'
- [Sec. 3.4, Supp. A.3] Eq. (9) defines the loss weight as λ_cd for the Chamfer term, but Supp. A.3 lists it as λ_p; please unify the notation.
- [Sec. 3.4, Eq. (13)] The collision loss is named L_c in Eq. (13) but L_col in Eq. (9); use a single name throughout.
- [Sec. 3.2, Eq. (5)] The equality x_i^d = p_i + T_i · r_i^c = T_i · (p_i^c + r_i^c) is only valid if p_i = T_i · p_i^c; state this explicitly to avoid confusion.
- [Supp. B.2, Table B2] The subject ordering in Table B2 (felice-004, christine-027, janett-025) differs from Table 2 (felice-004, janett-025, christine-027); keep the ordering consistent to allow direct comparison.
- [Supp. A.5] The perceptual study asks participants to select a single best result; consider also reporting pairwise preferences or a statistical test to support the claim that 63.4% preference is significant.
- [Supp. A.1] The modified SpareNet omits the refiner and adversarial rendering; please clarify whether this modification affects the comparison to the original SpareNet and whether the generator could benefit from those components.
Circularity Check
No significant circularity: the paper's central claims are empirical and benchmarked against external baselines, with no derivation that assumes its own conclusion.
full rationale
FreeCloth is an empirical systems paper: its contribution is a hybrid architecture combining a garment-specific clothing-cut map, an LBS-based deformation branch, and a free-form point generator trained with reconstruction, normal, regularization, and collision losses against ground-truth scans from the public ReSynth dataset. The central SOTA claim is established by comparison with external baselines (POP, SkiRT, FITE) using FID, MSE, human preference, and GPT-4o preference scores; none of these evaluations reduce to a parameter fitted to the target metric. The clothing-cut map is computed once per subject by running SAM on a near-canonical frame and back-projecting the result to the UV map (Sec. 3.1 and Supp. A.2); this is a fixed preprocessing choice, not a prediction generated by the model, and the paper's own ablations (Fig. 7, B13) and failure analysis (Sec. B.8) explicitly expose its assumptions rather than hiding them. The paper contains no self-citation chain that carries the argument: the cited baselines and priors ([39], [40], [33], [70]) are works by other research groups, and no uniqueness theorem or author-imported constraint is invoked to force the design. The discussion rejecting Chamfer distance is an argument about metric validity supported by external evidence (DPF, FITE, and point-removal experiments), not a circular reduction of the method's output to its input. Every load-bearing step is either a supervised learning objective on external ground truth or a comparison against external methods, so no Eq.-to-Eq. or fit-as-prediction circularity can be exhibited. The fixed-map pose-dependence concern is a genuine correctness risk, but it is a limitation, not circularity.
Assumptions & free parameters
free parameters (5)
- Ng (number of generated points) =
32768 (long dresses), 16384 (skirts)
- K (number of surface patches in the generator) =
8
- Kb (number of body parts for pose encoding) =
4
- Collision loss threshold epsilon =
not specified
- Loss weights =
lambda_p=1e4, lambda_n=1.0, lambda_rd=2e3, lambda_rg=1, lambda_col=2e-2
assumptions (5)
- domain assumption SMPL-X LBS provides a valid deformation model for near-body clothing
- domain assumption The ReSynth dataset and its official split are representative for evaluating loose-clothing modeling
- domain assumption Rendering point clouds into multi-view normal maps and computing FID/MSE is a faithful proxy for 3D geometric quality
- ad hoc to paper SAM segmentation on a single canonical frame reliably identifies loose regions for all poses
- domain assumption The free-form generator can synthesize loose clothing geometry from part-based pose features and a garment code
Cite this review
Pith. "Pith review of FreeCloth: Free-form Generation Enhances Challenging Clothed Human Modeling." pith.science (2026). https://pith.science/paper/C4HAM6BL
@misc{pith2026241119942,
author = {Pith},
title = {Pith review of: FreeCloth: Free-form Generation Enhances Challenging Clothed Human Modeling},
year = {2026},
howpublished = {\url{https://pith.science/paper/C4HAM6BL}},
note = {Machine review of arXiv:2411.19942}
}
read the original abstract
Achieving realistic animated human avatars requires accurate modeling of pose-dependent clothing deformations. Existing learning-based methods heavily rely on the Linear Blend Skinning (LBS) of minimally-clothed human models like SMPL to model deformation. However, they struggle to handle loose clothing, such as long dresses, where the canonicalization process becomes ill-defined when the clothing is far from the body, leading to disjointed and fragmented results. To overcome this limitation, we propose FreeCloth, a novel hybrid framework to model challenging clothed humans. Our core idea is to use dedicated strategies to model different regions, depending on whether they are close to or distant from the body. Specifically, we segment the human body into three categories: unclothed, deformed, and generated. We simply replicate unclothed regions that require no deformation. For deformed regions close to the body, we leverage LBS to handle the deformation. As for the generated regions, which correspond to loose clothing areas, we introduce a novel free-form, part-aware generator to model them, as they are less affected by movements. This free-form generation paradigm brings enhanced flexibility and expressiveness to our hybrid framework, enabling it to capture the intricate geometric details of challenging loose clothing, such as skirts and dresses. Experimental results on the benchmark dataset featuring loose clothing demonstrate that FreeCloth achieves state-of-the-art performance with superior visual fidelity and realism, particularly in the most challenging cases.
Figures
Figures from the paper (4 more)
Reference graph
Works this paper leans on
-
[1]
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ah- mad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. Gpt-4 technical report. arXiv preprint arXiv:2303.08774 ,
-
[2]
Automatic rigging and anima- tion of 3d characters
Ilya Baran and Jovan Popovi´c. Automatic rigging and anima- tion of 3d characters. ACM Transactions on graphics (TOG), 26(3):72–es, 2007. 2, 3
2007
-
[3]
Shape reconstruction by learn- ing differentiable surface representations
Jan Bednarik, Shaifali Parashar, Erhan Gundogdu, Mathieu Salzmann, and Pascal Fua. Shape reconstruction by learn- ing differentiable surface representations. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4716–4725, 2020. 3
work page 2020
-
[4]
Multi-garment net: Learning to dress 3d people from images
Bharat Lal Bhatnagar, Garvita Tiwari, Christian Theobalt, and Gerard Pons-Moll. Multi-garment net: Learning to dress 3d people from images. In Proceedings of the IEEE/CVF international conference on computer vision , pages 5420– 5430, 2019. 3
work page 2019
-
[5]
Bharat Lal Bhatnagar, Cristian Sminchisescu, Christian Theobalt, and Gerard Pons-Moll. Loopreg: Self-supervised learning of implicit surface correspondences, pose and shape for 3d human mesh registration. Advances in Neural Infor- mation Processing Systems, 33:12909–12922, 2020. 3
work page 2020
-
[6]
Dynamic surface function networks for clothed human bodies
Andrei Burov, Matthias Nießner, and Justus Thies. Dynamic surface function networks for clothed human bodies. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 10754–10764, 2021. 3
work page 2021
-
[7]
Snarf: Differentiable forward skin- ning for animating non-rigid neural implicit shapes
Xu Chen, Yufeng Zheng, Michael J Black, Otmar Hilliges, and Andreas Geiger. Snarf: Differentiable forward skin- ning for animating non-rigid neural implicit shapes. In In- ternational Conference on Computer Vision (ICCV) , pages 11594–11604, 2021. 3, 4
work page 2021
-
[8]
Nasa neural articulated shape approxi- mation
Boyang Deng, John P Lewis, Timothy Jeruzalski, Gerard Pons-Moll, Geoffrey Hinton, Mohammad Norouzi, and An- drea Tagliasacchi. Nasa neural articulated shape approxi- mation. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceed- ings, Part VII 16, pages 612–628. Springer, 2020. 3
work page 2020
Show all 74 references
-
[9]
Better patch stitching for parametric surface reconstruc- tion
Zhantao Deng, Jan Bedna ˇrík, Mathieu Salzmann, and Pascal Fua. Better patch stitching for parametric surface reconstruc- tion. In 2020 International Conference on 3D Vision (3DV), pages 593–602. IEEE, 2020. 3
2020
-
[10]
Learning elemen- tary structures for 3d shape generation and matching
Theo Deprelle, Thibault Groueix, Matthew Fisher, Vladimir Kim, Bryan Russell, and Mathieu Aubry. Learning elemen- tary structures for 3d shape generation and matching. Ad- vances in Neural Information Processing Systems, 32, 2019. 3
2019
-
[11]
Avatar reshap- ing and automatic rigging using a deformable model
Andrew Feng, Dan Casas, and Ari Shapiro. Avatar reshap- ing and automatic rigging using a deformable model. InPro- ceedings of the 8th ACM SIGGRAPH Conference on Motion in Games, pages 57–64, 2015. 2
2015
-
[12]
Learning deformable tetrahedral meshes for 3d reconstruction
Jun Gao, Wenzheng Chen, Tommy Xiang, Alec Jacobson, Morgan McGuire, and Sanja Fidler. Learning deformable tetrahedral meshes for 3d reconstruction. Advances In Neu- ral Information Processing Systems, 33:9936–9947, 2020. 3
2020
-
[13]
A papier-mâché ap- proach to learning 3d surface generation
Thibault Groueix, Matthew Fisher, Vladimir G Kim, Bryan C Russell, and Mathieu Aubry. A papier-mâché ap- proach to learning 3d surface generation. In Proceedings of the IEEE conference on computer vision and pattern recog- nition, pages 216–224, 2018. 3
2018
-
[14]
Drape: Dressing any person
Peng Guan, Loretta Reiss, David A Hirshberg, Alexander Weiss, and Michael J Black. Drape: Dressing any person. ACM Transactions on Graphics (ToG), 31(4):1–10, 2012. 3, 5
2012
-
[15]
Garnet: A two-stream network for fast and accurate 3d cloth draping
Erhan Gundogdu, Victor Constantin, Amrollah Seifoddini, Minh Dang, Mathieu Salzmann, and Pascal Fua. Garnet: A two-stream network for fast and accurate 3d cloth draping. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 8739–8748, 2019. 3, 5
2019
-
[16]
Arch++: Animation-ready clothed human recon- struction revisited
Tong He, Yuanlu Xu, Shunsuke Saito, Stefano Soatto, and Tony Tung. Arch++: Animation-ready clothed human recon- struction revisited. In Proceedings of the IEEE/CVF interna- tional conference on computer vision , pages 11046–11056,
-
[17]
Gans trained by a two time-scale update rule converge to a local nash equilib- rium
Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. Gans trained by a two time-scale update rule converge to a local nash equilib- rium. Advances in neural information processing systems , 30, 2017. 6
2017
-
[18]
Learning to train a point cloud reconstruction network with- out matching
Tianxin Huang, Xuemeng Yang, Jiangning Zhang, Jinhao Cui, Hao Zou, Jun Chen, Xiangrui Zhao, and Yong Liu. Learning to train a point cloud reconstruction network with- out matching. In European Conference on Computer Vision, pages 179–194. Springer, 2022. 3
2022
-
[19]
Humannorm: Learning normal diffusion model for high-quality and realistic 3d hu- man generation
Xin Huang, Ruizhi Shao, Qi Zhang, Hongwen Zhang, Ying Feng, Yebin Liu, and Qing Wang. Humannorm: Learning normal diffusion model for high-quality and realistic 3d hu- man generation. In Proceedings of the IEEE/CVF Confer- ence on Computer Vision and Pattern Recognition , pages...
2024
-
[20]
Tech: Text-guided reconstruction of lifelike clothed humans
Yangyi Huang, Hongwei Yi, Yuliang Xiu, Tingting Liao, Ji- axiang Tang, Deng Cai, and Justus Thies. Tech: Text-guided reconstruction of lifelike clothed humans. arXiv preprint arXiv:2308.08545, 2023. 3
2023 arXiv
-
[21]
Deformable 3d gaussian splatting for animat- able human avatars
HyunJun Jung, Nikolas Brasch, Jifei Song, Eduardo Perez- Pellitero, Yiren Zhou, Zhihao Li, Nassir Navab, and Ben- jamin Busam. Deformable 3d gaussian splatting for animat- able human avatars. arXiv preprint arXiv:2312.15059, 2023. 3
2023 arXiv
-
[22]
In- vertible neural skinning
Yash Kant, Aliaksandr Siarohin, Riza Alp Guler, Menglei Chai, Jian Ren, Sergey Tulyakov, and Igor Gilitschenski. In- vertible neural skinning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 8715–8725, 2023. 3
2023
-
[23]
Physics-inspired upsampling for cloth sim- ulation in games
Ladislav Kavan, Dan Gerszewski, Adam W Bargteil, and Peter-Pike Sloan. Physics-inspired upsampling for cloth sim- ulation in games. In ACM SIGGRAPH 2011 papers , pages 1–10. 2011. 2 9
2011
-
[24]
Screened poisson sur- face reconstruction
Michael Kazhdan and Hugues Hoppe. Screened poisson sur- face reconstruction. ACM Transactions on Graphics (ToG), 32(3):1–13, 2013. 7
2013
-
[25]
3d gaussian splatting for real-time radiance field rendering
Bernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, and George Drettakis. 3d gaussian splatting for real-time radiance field rendering. ACM Transactions on Graphics, 42 (4):1–14, 2023. 3, 8
2023
-
[26]
Chupa: Carving 3d clothed humans from skinned shape priors using 2d diffusion probabilistic models
Byungjun Kim, Patrick Kwon, Kwangho Lee, Myunggi Lee, Sookwan Han, Daesik Kim, and Hanbyul Joo. Chupa: Carving 3d clothed humans from skinned shape priors using 2d diffusion probabilistic models. arXiv preprint arXiv:2305.11870, 2023. 6
2023 arXiv
-
[27]
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980 ,
-
[28]
Segment any- thing
Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer White- head, Alexander C Berg, Wan-Yen Lo, et al. Segment any- thing. In Proceedings of the IEEE/CVF International Con- ference on Computer Vision, pages 4015–4026, 202...
2023
-
[29]
Hugs: Human gaussian splats
Muhammed Kocabas, Jen-Hao Rick Chang, James Gabriel, Oncel Tuzel, and Anurag Ranjan. Hugs: Human gaussian splats. arXiv preprint arXiv:2311.17910, 2023. 3
2023 arXiv
-
[30]
Deepwrin- kles: Accurate and realistic clothing modeling
Zorah Lahner, Daniel Cremers, and Tony Tung. Deepwrin- kles: Accurate and realistic clothing modeling. In Proceed- ings of the European conference on computer vision (ECCV), pages 667–684, 2018. 3, 5
2018
-
[31]
Multi-layered unseen gar- ments draping network
Dohae Lee and In-Kwon Lee. Multi-layered unseen gar- ments draping network. arXiv preprint arXiv:2304.03492 ,
-
[32]
Ani- matable gaussians: Learning pose-dependent gaussian maps for high-fidelity human avatar modeling
Zhe Li, Zerong Zheng, Lizhen Wang, and Yebin Liu. Ani- matable gaussians: Learning pose-dependent gaussian maps for high-fidelity human avatar modeling. arXiv preprint arXiv:2311.16096, 2023. 3
2023 arXiv
-
[33]
Learning implicit templates for point-based clothed human modeling
Siyou Lin, Hongwen Zhang, Zerong Zheng, Ruizhi Shao, and Yebin Liu. Learning implicit templates for point-based clothed human modeling. In European Conference on Com- puter Vision, pages 210–228. Springer, 2022. 1, 2, 3, 5, 6, 7, 8, 4
2022
-
[34]
Neuroskinning: Automatic skin binding for production characters with deep graph networks
Lijuan Liu, Youyi Zheng, Di Tang, Yi Yuan, Changjie Fan, and Kun Zhou. Neuroskinning: Automatic skin binding for production characters with deep graph networks. ACM Transactions on Graphics (ToG), 38(4):1–12, 2019. 2
2019
-
[35]
Morphing and sampling network for dense point cloud completion
Minghua Liu, Lu Sheng, Sheng Yang, Jing Shao, and Shi- Min Hu. Morphing and sampling network for dense point cloud completion. In Proceedings of the AAAI conference on artificial intelligence, pages 11596–11603, 2020. 3
2020
-
[36]
Smpl: A skinned multi- person linear model
Matthew Loper, Naureen Mahmood, Javier Romero, Gerard Pons-Moll, and Michael J Black. Smpl: A skinned multi- person linear model. In Seminal Graphics Papers: Pushing the Boundaries, Volume 2, pages 851–866. 2023. 3
2023
-
[37]
Learn- ing to dress 3d people in generative clothing
Qianli Ma, Jinlong Yang, Anurag Ranjan, Sergi Pujades, Gerard Pons-Moll, Siyu Tang, and Michael J Black. Learn- ing to dress 3d people in generative clothing. InProceedings of the IEEE/CVF Conference on Computer Vision and Pat- tern Recognition, pages 6469–6478, 2020. 3
2020
-
[38]
Scale: Modeling clothed humans with a surface codec of articulated local elements
Qianli Ma, Shunsuke Saito, Jinlong Yang, Siyu Tang, and Michael J Black. Scale: Modeling clothed humans with a surface codec of articulated local elements. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pat- tern Recognition, pages 16082–16093, 2021. 2, 3
2021
-
[39]
The power of points for modeling humans in clothing
Qianli Ma, Jinlong Yang, Siyu Tang, and Michael J Black. The power of points for modeling humans in clothing. In International Conference on Computer Vision (ICCV), pages 10974–10984, 2021. 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12
2021
-
[40]
Neural point-based shape modeling of humans in challeng- ing clothing
Qianli Ma, Jinlong Yang, Michael J Black, and Siyu Tang. Neural point-based shape modeling of humans in challeng- ing clothing. In 2022 International Conference on 3D Vision (3DV), pages 679–689. IEEE, 2022. 2, 3, 4, 5, 6, 7, 8, 1
2022
-
[41]
Occupancy networks: Learning 3d reconstruction in function space
Lars Mescheder, Michael Oechsle, Michael Niemeyer, Se- bastian Nowozin, and Andreas Geiger. Occupancy networks: Learning 3d reconstruction in function space. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 4460–4470, 2019. 3
2019
-
[42]
Nerf: Representing scenes as neural radiance fields for view syn- thesis
Ben Mildenhall, Pratul P Srinivasan, Matthew Tancik, Jonathan T Barron, Ravi Ramamoorthi, and Ren Ng. Nerf: Representing scenes as neural radiance fields for view syn- thesis. Communications of the ACM , 65(1):99–106, 2021. 3
2021
-
[43]
Human gaussian splatting: Real-time rendering of animatable avatars
Arthur Moreau, Jifei Song, Helisa Dhamo, Richard Shaw, Yiren Zhou, and Eduardo Pérez-Pellitero. Human gaussian splatting: Real-time rendering of animatable avatars. arXiv preprint arXiv:2311.17113, 2023. 3
2023 arXiv
-
[44]
Dress-me-up: A dataset & method for self-supervised 3d garment retargeting
Shanthika Naik, Kunwar Singh, Astitva Srivastava, Dhawal Sirikonda, Amit Raj, Varun Jampani, and Avinash Sharma. Dress-me-up: A dataset & method for self-supervised 3d garment retargeting. arXiv preprint arXiv:2401.03108, 2024. 2
2024 arXiv
-
[45]
A layered model of human body and garment deformation
Alexandros Neophytou and Adrian Hilton. A layered model of human body and garment deformation. In 2014 2nd In- ternational Conference on 3D Vision, pages 171–178. IEEE,
2014
-
[46]
Ash: Animatable gaussian splats for efficient and photoreal human rendering
Haokai Pang, Heming Zhu, Adam Kortylewski, Christian Theobalt, and Marc Habermann. Ash: Animatable gaussian splats for efficient and photoreal human rendering. arXiv preprint arXiv:2312.05941, 2023. 3
2023 arXiv
-
[47]
Deepsdf: Learning con- tinuous signed distance functions for shape representation
Jeong Joon Park, Peter Florence, Julian Straub, Richard Newcombe, and Steven Lovegrove. Deepsdf: Learning con- tinuous signed distance functions for shape representation. In Proceedings of the IEEE/CVF conference on computer vi- sion and pattern recognition, pages 165–174, 2019. 3, 4
2019
-
[48]
Tailornet: Predicting clothing in 3d as a function of human pose, shape and garment style
Chaitanya Patel, Zhouyingcheng Liao, and Gerard Pons- Moll. Tailornet: Predicting clothing in 3d as a function of human pose, shape and garment style. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 7365–7375, 2020. 3, 5
2020
-
[49]
Expressive body capture: 3d hands, face, and body from a single image
Georgios Pavlakos, Vasileios Choutas, Nima Ghorbani, Timo Bolkart, Ahmed AA Osman, Dimitrios Tzionas, and Michael J Black. Expressive body capture: 3d hands, face, and body from a single image. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognitio...
2019
-
[50]
Neural body: 10 Implicit neural representations with structured latent codes for novel view synthesis of dynamic humans
Sida Peng, Yuanqing Zhang, Yinghao Xu, Qianqian Wang, Qing Shuai, Hujun Bao, and Xiaowei Zhou. Neural body: 10 Implicit neural representations with structured latent codes for novel view synthesis of dynamic humans. In Proceed- ings of the IEEE/CVF Conference on Computer Visio...
2021
-
[51]
Dynamic point fields
Sergey Prokudin, Qianli Ma, Maxime Raafat, Julien Valentin, and Siyu Tang. Dynamic point fields. arXiv preprint arXiv:2304.02626, 2023. 3, 6, 2
2023 arXiv
-
[52]
Pointnet: Deep learning on point sets for 3d classification and segmentation
Charles R Qi, Hao Su, Kaichun Mo, and Leonidas J Guibas. Pointnet: Deep learning on point sets for 3d classification and segmentation. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 652–660,
-
[53]
Pointnet++: Deep hierarchical feature learning on point sets in a metric space
Charles Ruizhongtai Qi, Li Yi, Hao Su, and Leonidas J Guibas. Pointnet++: Deep hierarchical feature learning on point sets in a metric space. Advances in neural information processing systems, 30, 2017. 4, 5, 1
2017
-
[54]
Unif: United neural implicit functions for clothed human reconstruction and animation
Shenhan Qian, Jiale Xu, Ziwei Liu, Liqian Ma, and Shenghua Gao. Unif: United neural implicit functions for clothed human reconstruction and animation. In European Conference on Computer Vision , pages 121–137. Springer,
-
[55]
Pifu: Pixel-aligned implicit function for high-resolution clothed human digitiza- tion
Shunsuke Saito, Zeng Huang, Ryota Natsume, Shigeo Mor- ishima, Angjoo Kanazawa, and Hao Li. Pifu: Pixel-aligned implicit function for high-resolution clothed human digitiza- tion. In Proceedings of the IEEE/CVF international confer- ence on computer vision, pages 2304–2314, 2019
2019
-
[56]
Scanimate: Weakly supervised learning of skinned clothed avatar networks
Shunsuke Saito, Jinlong Yang, Qianli Ma, and Michael J Black. Scanimate: Weakly supervised learning of skinned clothed avatar networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 2886–2897, 2021. 3, 4, 2
2021
-
[57]
Learning-based animation of clothing for virtual try-on
Igor Santesteban, Miguel A Otaduy, and Dan Casas. Learning-based animation of clothing for virtual try-on. In Computer Graphics Forum , pages 355–366. Wiley Online Library, 2019. 5
2019
-
[58]
Self-supervised collision handling via generative 3d garment models for virtual try-on
Igor Santesteban, Nils Thuerey, Miguel A Otaduy, and Dan Casas. Self-supervised collision handling via generative 3d garment models for virtual try-on. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 11763–11773, 2021. 6
2021
-
[59]
Towards multi-layered 3d garments animation
Yidi Shao, Chen Change Loy, and Bo Dai. Towards multi-layered 3d garments animation. arXiv preprint arXiv:2305.10418, 2023. 6
2023 arXiv
-
[60]
Deep marching tetrahedra: a hybrid repre- sentation for high-resolution 3d shape synthesis
Tianchang Shen, Jun Gao, Kangxue Yin, Ming-Yu Liu, and Sanja Fidler. Deep marching tetrahedra: a hybrid repre- sentation for high-resolution 3d shape synthesis. Advances in Neural Information Processing Systems , 34:6087–6101,
-
[61]
A-nerf: Articulated neural radiance fields for learn- ing human shape, appearance, and pose
Shih-Yang Su, Frank Yu, Michael Zollhöfer, and Helge Rhodin. A-nerf: Articulated neural radiance fields for learn- ing human shape, appearance, and pose. Advances in Neural Information Processing Systems, 34:12278–12291, 2021. 3
2021
-
[62]
Deep- cloth: Neural garment representation for shape and style editing
Zhaoqi Su, Tao Yu, Yangang Wang, and Yebin Liu. Deep- cloth: Neural garment representation for shape and style editing. IEEE Transactions on Pattern Analysis and Machine Intelligence, 45(2):1581–1593, 2022. 3
2022
-
[63]
Sizer: A dataset and model for parsing 3d clothing and learning size sensitive 3d clothing
Garvita Tiwari, Bharat Lal Bhatnagar, Tony Tung, and Ger- ard Pons-Moll. Sizer: A dataset and model for parsing 3d clothing and learning size sensitive 3d clothing. InComputer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part III 16...
2020
-
[64]
Hu- mannerf: Free-viewpoint rendering of moving people from monocular video
Chung-Yi Weng, Brian Curless, Pratul P Srinivasan, Jonathan T Barron, and Ira Kemelmacher-Shlizerman. Hu- mannerf: Free-viewpoint rendering of moving people from monocular video. In Proceedings of the IEEE/CVF con- ference on computer vision and pattern Recognition , pages 162...
2022
-
[65]
Density-aware chamfer distance as a com- prehensive metric for point cloud completion.arXiv preprint arXiv:2111.12702, 2021
Tong Wu, Liang Pan, Junzhe Zhang, Tai Wang, Ziwei Liu, and Dahua Lin. Density-aware chamfer distance as a com- prehensive metric for point cloud completion.arXiv preprint arXiv:2111.12702, 2021. 3
2021 arXiv
-
[66]
Style-based point generator with ad- versarial rendering for point cloud completion
Chulin Xie, Chuxin Wang, Bo Zhang, Hao Yang, Dong Chen, and Fang Wen. Style-based point generator with ad- versarial rendering for point cloud completion. In Proceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4619–4628, 2021. 5, 1
2021
-
[67]
Econ: Explicit clothed humans obtained from normals
Yuliang Xiu, Jinlong Yang, Xu Cao, Dimitrios Tzionas, and Michael J Black. Econ: Explicit clothed humans obtained from normals. arXiv preprint arXiv:2212.07422, 2022. 6
2022 arXiv
-
[68]
Icon: Implicit clothed humans obtained from nor- mals
Yuliang Xiu, Jinlong Yang, Dimitrios Tzionas, and Michael J Black. Icon: Implicit clothed humans obtained from nor- mals. In 2022 IEEE/CVF Conference on Computer Vi- sion and Pattern Recognition (CVPR), pages 13286–13296. IEEE, 2022. 3, 6
2022
-
[69]
Point-based modeling of human clothing
Ilya Zakharkin, Kirill Mazur, Artur Grigorev, and Victor Lempitsky. Point-based modeling of human clothing. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 14718–14727, 2021. 3
2021
-
[70]
Closet: Modeling clothed humans on continuous surface with explicit template decomposition
Hongwen Zhang, Siyou Lin, Ruizhi Shao, Yuxiang Zhang, Zerong Zheng, Han Huang, Yandong Guo, and Yebin Liu. Closet: Modeling clothed humans on continuous surface with explicit template decomposition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recog...
2023
-
[71]
Gps- gaussian: Generalizable pixel-wise 3d gaussian splatting for real-time human novel view synthesis
Shunyuan Zheng, Boyao Zhou, Ruizhi Shao, Boning Liu, Shengping Zhang, Liqiang Nie, and Yebin Liu. Gps- gaussian: Generalizable pixel-wise 3d gaussian splatting for real-time human novel view synthesis. arXiv preprint arXiv:2312.02155, 2023. 3
2023 arXiv
-
[72]
Open3d: A modern library for 3d data processing
Qian-Yi Zhou, Jaesik Park, and Vladlen Koltun. Open3d: A modern library for 3d data processing. arXiv preprint arXiv:1801.09847, 2018. 1 11 FreeCloth: Free-form Generation Enhances Challenging Clothed Human Modeling Supplementary Material In Sec. A, we elaborate on the impleme...
2018 arXiv
-
[73]
felice-004
To assess the geometric visual quality, we render the front and the back views at a high resolution of1024 × 1024. The deployed baseline models are discussed above in Sec. A.4. 1 beatrice-025 Front View Back View anna-001 janett-025 christine-027 felice-004 Figure A1. The segm...
-
[74]
seams". 7 SkiRT GTFITEOursPOP Figure B15. Additional Qualitative comparison between baselines and our model.The subject ID is “felice-004
Additionally, we report the FLOPs and parameter counts to quantify computational resource requirements. As shown in Tab. B4, our model has the smallest number of parame- ters and achieves real-time inference speed at64.1 FPS. No- tably, our model significantly outperforms SOTA...
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.