REVIEW 3 major objections 5 minor 82 references
Creating Your Editable 3D Photorealistic Avatar with Tetrahedron-constrained Gaussian Splatting
T0 review · 3 major / 5 minor · reviewed 2026-08-16 · deepseek-v4-flash
Pith's one-line read A hybrid tetrahedron-constrained Gaussian splatting representation lets users turn a monocular video into a photorealistic 3D avatar and edit it locally with text prompts or reference images.
desk verdict A genuinely new hybrid representation for localized 3D avatar editing with a sensible decoupled pipeline, but the headline empirical claim is undercut by a self-similarity FID protocol. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is TetGS, a hybrid representation in which each Gaussian is explicitly embedded in a father tetrahedron $t_k$ of a tetrahedral grid derived from a signed distance field. A signed distance field $\psi_g$ is evaluated at grid vertices, the Marching Tetrahedron algorithm extracts mesh vertices by linear interpolation of the SDF values on tetrahedral edges, and each Gaussian's position is set as $\mu = w_a v^M_{i1} + w_b v^M_{i2} + w_c v^M_{i3} + \tau n$, where the $v^M$ are the vertices of the mesh triangle it sits on, $w$ are fixed barycentric weights, and $\tau$ is a learnable displacement along the surface normal. The two identity-like maps $g_{v\to e}$ (mesh vertex to its tetrahedral edge) and $h_{f\to t}$ (mesh triangle to its father tetrahedron) carry the localization: they allow the method to freeze the SDF at vertices of preserved tetrahedra, update only editable vertices under a dual SDS loss, and re-initialize Gaussians in the edited tetrahedra by repeating the same extraction procedure. During texture generation, the editing Gaussians are temporarily degraded to 2D disks with fixed opacity and view-independent color, which stabilizes few-shot inpainting, and then their full covariance, opacity, and spherical-harmonic attributes are reactivated for the final refinement.
What would settle it
Record a subject in loose clothing, edit the garment to a tight one, and compare renderings of the preserved regions before and after editing. If the Marching Tetrahedron extraction produces inverted tetrahedra or reassigns preserved triangles to different father tetrahedra, the preserved regions will change or the edit will fail geometrically; the paper's own supplementary material identifies this loose-to-tight case as a limitation, so a concrete failure there would falsify the claim of reliable localized geometric adaptation.
Extended reading notes
Core claim
The central discovery is that binding Gaussian kernels to tetrahedra converts an otherwise discrete, unstable 3D Gaussian splatting optimization into a controllable deformation problem: updating the signed distance values of tetrahedral vertices moves both the extracted surface and the Gaussians embedded on it. The father-tetrahedron mapping $h_{f\to t}$ lets the method split the mesh into preserved and edited regions; frozen SDF values keep the preserved Gaussians unchanged, while editable vertices are optimized under a dual global-and-local SDS loss on normal renderings. After the geometry settles, the reallocated editing Gaussians are first restricted to 2D surfels and painted with few-shot normal-guided inpainted images, then their full 3D attributes are reactivated and refined with image-to-image augmented views. The conclusion states that this yields high-fidelity photorealistic 3D avatar editing with diverse identities and accessories, and the paper backs it with qualitative results and the reported metric improvements.
Load-bearing premise
The approach assumes that deforming the editable region never flips or reconnects the tetrahedral mesh so severely that the link between a surface triangle and its parent tetrahedron breaks; under large geometric changes, such as loose to tight clothing, that link can fail and the promised localized editing would break down.
Editorial extensions
If this is right
- A 40- to 50-second monocular video plus a text prompt or a reference garment photo is sufficient to produce an editable, photorealistic 360-degree avatar, removing the need for multi-camera capture or manual modeling.
- Because geometry and appearance are optimized separately, the method avoids the noise and blur that arise when Gaussian geometry and texture are trained together under generative guidance; the Fig. 7 and Tab. 2 ablations support this.
- Non-edited regions stay visually intact because their tetrahedron vertices are frozen and a surface-aware regularization keeps the edited surface from occluding them; the Fig. 8 ablations show that removing the partitioning or either SDS term breaks this behavior.
- Reference-image virtual try-on is achieved by substituting a try-on diffusion model for the normal inpainter and adding normal and mask supervision, so both garment style and its geometric design transfer to the avatar.
- Supplementary results show the same pipeline supports texture doodling, by painting on guidance images, and continuous editing, by applying edits sequentially.
Reading between the lines
- If the decoupling claim generalizes, the same tetrahedral scaffold could stabilize localized editing of any 3D Gaussian scene with a clean surface prior, not just human avatars; a direct test would be object-level editing with a mask-defined region on a generic reconstruction.
- The reported gains are computed on the paper's own dataset of ten videos against three baselines; a stronger external check would be a public multi-view human dataset with ground-truth geometry, since FID and CLIP measure image statistics and text alignment rather than whether the deformed geometry is correct.
- The loose-to-tight failure acknowledged in the supplementary suggests the real boundary is topological change in the tetrahedral grid; adding an estimated inner-body shape prior, as the authors suggest, would turn that failure mode into a testable fix.
- Because the texture stage relies on a normal-based inpainter, its view consistency is inherited from that model; a stress test would render the edited avatar from unevenly spaced elevations and inspect shoulders and back for flicker or duplicated detail, which the paper's evenly spaced 60-view evaluation may not fully expose.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a three-stage pipeline for creating editable 3D avatars from a monocular video, using a hybrid tetrahedron-constrained Gaussian Splatting (TetGS) representation. In the first stage, a neural SDF field is optimized on the video and converted into a tetrahedral grid with embedded Gaussians. The second stage performs localized spatial adaptation by explicitly partitioning tetrahedra into preserved and editable sets, updating the SDF of editable vertices under global and local SDS losses plus surface-aware regularizations. The third stage generates appearance by first learning a coarse texture under few-shot normal-based inpainting and then refining the appearance with an attribute-activation I2I refinement. The method supports both text-guided editing and reference-image-based virtual try-on. The authors claim that the resulting avatars are high-fidelity, photorealistic, and quantitatively superior to GaussianEditor, DGE, and TIP-Editor on their collected 10-subject dataset.
Significance. If the empirical claims are substantiated, the paper makes a useful contribution: it provides an accessible pipeline from a monocular video and a text/image prompt to an editable, locally controllable 3D avatar, addressing a real gap in 3DGS-based editing. The technical presentation is a strength: the TetGS representation is described with explicit equations (Eq. 1–7), the three-stage pipeline is clearly motivated, and the supplementary material provides architectural details and ablations (Tables 2–3, Figs. 7–9). The decoupling of geometry and appearance optimization is a plausible remedy for the instability noted in prior 3DGS editing. However, the central quantitative claim of superiority over baselines is not supported by the reported evaluation, as the FID protocol measures similarity to the method's own reconstructions rather than to real imagery. The paper would be significantly strengthened by a re-evaluation against genuine photorealistic references, statistical error bars, and a direct test of the tetrahedron-mapping robustness for large geometric edits.
major comments (3)
- [Sec. 4.2, Table 1] The FID metric is computed between edited renderings and the system's own reconstructed images, as stated in Sec. 4.2: "a lower FID indicates greater similarity to the reconstructed images." This is a self-similarity measure, not a photorealism or fidelity measure, and a degenerate no-op edit that leaves the reconstruction unchanged would achieve near-zero FID. The reported gap (115.95 vs. 194.54/201.82/258.27) therefore cannot be interpreted as evidence that the edited avatars are more photorealistic or higher-fidelity than the baselines. The quantitative comparison should be re-run against a fixed set of real photographs (e.g., held-out frames from the input video or real-captured images of people), or supplemented by a user study and by perceptual metrics that do not reference the method's own output.
- [Sec. 4.2, Table 1; Datasets paragraph] All quantitative metrics are computed on only 60 rendered images from a 10-subject dataset, with no error bars, no multiple seeds, and no statistical significance test. The claimed "significant improvement" in FID and the 26.28 vs. 22–23 CLIP gap may be within run-to-run variation; the paper does not provide the variance or confidence intervals needed to support the superiority claim. Additionally, DINO similarity is reported only for TIP-Editor and not for GaussianEditor or DGE, so the table does not provide a complete comparison across methods on the image-guided task. The authors should report mean and standard deviation over at least three independent runs per method and per metric, and should either complete the DINO column for all baselines or explain why it is inapplicable.
- [Sec. 3.1, Sec. 3.2.2, Supp. H] The robustness of the father-tetrahedron mapping hf→t (Sec. 3.1) under large geometry changes is load-bearing for the localized spatial adaptation module. When the SDF values at editable vertices are optimized as described in Sec. 3.2.2, the Marching Tetrahedron algorithm can produce inverted or topologically different tetrahedra; the paper only states that the reallocated Gaussians Gedit are initialized "by following the procedure described in Sec. 3.1" and does not analyze whether the mapping remains consistent or whether preserved tetrahedra Vkeep_T can be corrupted. The authors acknowledge in Supp. H that loose-to-tight garment edits are problematic; this suggests the mapping reliability is an edge case that deserves a specific diagnostic: e.g., report the fraction of tetrahedra that change sign, or show that the preserved-region geometry and appearance remain unchanged (measured by local reconstruction error) for a large deformation. Without such a test, the claim of "precise region localization, geometric adaptability" is not fully supported.
minor comments (5)
- [Sec. 4.2] The text says "For quantitative evaluation, we use Frechet Inception Distance (FID) to access the quality of edited images." The verb should be "assess," and the metric name should be spelled "Fr\'echet" with the accent.
- [Sec. 4.2, Table 1] The sentence "producing photorealistic results comparable to real-world individuals" is not backed by any metric that compares to real-world photographs; consider softening it or providing such evidence.
- [Sec. 4.2, Table 1 and Sec. 4.3, Table 2] The paper reports absolute FID/CLIP values without confidence intervals in both the main comparison and the ablation. Reporting error bars even for the ablation table would help assess stability, especially since the "w/o AA" variant is only 4.38 FID points above the full model.
- [Sec. 4.2, qualitative comparison] The qualitative comparisons in Fig. 6 are informative, but the text does not specify the training time or iteration budget given to each baseline. Since the paper emphasizes efficiency (e.g., 1.2 hours for spatial adaptation), a brief statement about GPU-hours per method would aid reproducibility and fairness.
- [Supp. A.3, Table 3] In the supplementary Table 3, the text "3DGS-30K Ours-7K 3DGS-7K . We report the averaged chamfer distance, PSNR,..." contains a garbled fragment that appears to be a formatting error; please fix it.
Circularity Check
Quantitative superiority claim rests partly on a self-referential FID benchmark; the core representation and editing pipeline are otherwise self-contained.
-
self definitional
[Section 4.2, Quantitative comparison and Table 1]
"For quantitative evaluation, we use Frechet Inception Distance (FID) [74] to access the quality of edited images, where a lower FID indicates greater similarity to the reconstructed images, reflecting higher rendering fidelity and realism. All metrics are computed on 60 rendered images captured evenly around the 3D avatar."
The FID reference set is the reconstruction produced by the paper's own instantiation stage, and the edited result is initialized from exactly that reconstruction. During localized spatial adaptation the method explicitly freezes Gkeep and inherits their attributes, so preserved regions remain identical to the reference by construction. Consequently a no-op edit that leaves the reconstruction unchanged would achieve a near-perfect FID under this definition, and the reported gap (115.95 vs. 194.54) measures self-consistency with the input reconstruction rather than photorealism or editing quality. The conclusion of 'photorealistic results comparable to real-world individuals' is therefore not supported by this metric, although the CLIP and DINO results provide independent, partial support.
full rationale
The core representational claim--TetGS embedding Gaussians inside DMTet tetrahedra with the father-tetrahedron mapping--is defined by explicit interpolation equations (Eq. 1 and the barycentric position formula) and is not derived from its own outputs. The spatial adaptation and texture generation stages are supervised by standard external diffusion objectives (SDS, normal inpainting, I2I refinement, IDM-VTON), and the self-citations ([1,2]) appear only as related work without carrying the argument. The only load-bearing circular element is the FID evaluation: quality is defined as similarity to the paper's own reconstruction while the editing pipeline is deliberately built to preserve that reconstruction in non-editing regions. This makes the headline quantitative superiority claim partially self-referential. The acknowledged limitation in Supp. H (loose-to-tight garments) is an edge-case robustness issue, not a circularity. Overall, the system is not circular in its derivation, but one of its three quantitative validations reduces partly by construction to its own input, so a moderate score is warranted.
Assumptions & free parameters
free parameters (5)
- Loss weights for Eq. (7) =
{lambda_G_SDS=0.5, lambda_L_SDS=0.5, lambda_sa=5000, lambda_nc=2000}
- SDS guidance scale and timestep range =
guidance scale 50; timestep sampled from U(0.02, 0.80) then annealed to U(0.02, 0.20)
- Tetrahedron grid resolution =
512^3 grids around a standard human body
- Gaussians per tetrahedron =
3 for faces larger than average, 1 otherwise
- Training iterations and learning rates =
SDF MLP Adam lr=1e-3; spatial adaptation AdamW lr=2e-5, 10000 iters; texture stages 20 min + 3 min
assumptions (5)
- domain assumption Marching Tetrahedron extraction and the father-tetrahedron mapping hf to t are reliable for converting SDF values into a stable mesh and Gaussian embedding.
- domain assumption A pre-trained normal-conditioned diffusion model provides valid gradient directions for geometry editing via SDS.
- domain assumption Monocular SDF reconstruction from NeuDA-style implicit fields with normal regularization yields an accurate enough surface for TetGS initialization.
- domain assumption The few-shot normal-based inpainter (ControlNetPlus) generates view-consistent, photorealistic guidance for appearance learning.
- ad hoc to paper Freezing the SDF values of vertices belonging to preserved tetrahedra is sufficient to keep the preserved Gaussians spatially fixed.
Cite this review
Pith. "Pith review of Creating Your Editable 3D Photorealistic Avatar with Tetrahedron-constrained Gaussian Splatting." pith.science (2026). https://pith.science/paper/NCZPBFSW
@misc{pith2026250420403,
author = {Pith},
title = {Pith review of: Creating Your Editable 3D Photorealistic Avatar with Tetrahedron-constrained Gaussian Splatting},
year = {2026},
howpublished = {\url{https://pith.science/paper/NCZPBFSW}},
note = {Machine review of arXiv:2504.20403}
}
read the original abstract
Personalized 3D avatar editing holds significant promise due to its user-friendliness and availability to applications such as AR/VR and virtual try-ons. Previous studies have explored the feasibility of 3D editing, but often struggle to generate visually pleasing results, possibly due to the unstable representation learning under mixed optimization of geometry and texture in complicated reconstructed scenarios. In this paper, we aim to provide an accessible solution for ordinary users to create their editable 3D avatars with precise region localization, geometric adaptability, and photorealistic renderings. To tackle this challenge, we introduce a meticulously designed framework that decouples the editing process into local spatial adaptation and realistic appearance learning, utilizing a hybrid Tetrahedron-constrained Gaussian Splatting (TetGS) as the underlying representation. TetGS combines the controllable explicit structure of tetrahedral grids with the high-precision rendering capabilities of 3D Gaussian Splatting and is optimized in a progressive manner comprising three stages: 3D avatar instantiation from real-world monocular videos to provide accurate priors for TetGS initialization; localized spatial adaptation with explicitly partitioned tetrahedrons to guide the redistribution of Gaussian kernels; and geometry-based appearance generation with a coarse-to-fine activation strategy. Both qualitative and quantitative experiments demonstrate the effectiveness and superiority of our approach in generating photorealistic 3D editable avatars.
Figures
Figures from the paper (10 more)
Reference graph
Works this paper leans on
-
[1]
3dtoonify: Creating your high-fidelity 3d stylized avatar easily from 2d portrait images,
Y . Men, H. Liu, Y . Yao, M. Cui, X. Xie, and Z. Lian, “3dtoonify: Creating your high-fidelity 3d stylized avatar easily from 2d portrait images,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 10 127–10 137
2024
-
[2]
En3d: An enhanced generative model for sculpting 3d humans from 2d synthetic data,
Y . Men, B. Lei, Y . Yao, M. Cui, Z. Lian, and X. Xie, “En3d: An enhanced generative model for sculpting 3d humans from 2d synthetic data,” in Proceedings of the IEEE/CVF Confer- ence on Computer Vision and Pattern Recognition, 2024, pp. 9981–9991. 1, 2
2024
-
[3]
Tip-editor: An accurate 3d editor following both text- prompts and image-prompts,
J. Zhuang, D. Kang, Y .-P. Cao, G. Li, L. Lin, and Y . Shan, “Tip-editor: An accurate 3d editor following both text- prompts and image-prompts,” ACM Transactions on Graph- ics (TOG), vol. 43, no. 4, pp. 1–12, 2024. 2, 3, 4, 6, 8, 14
work page 2024
-
[4]
Nerf: Representing scenes as neural radiance fields for view synthesis,
B. Mildenhall, P. P. Srinivasan, M. Tancik, J. T. Barron, R. Ramamoorthi, and R. Ng, “Nerf: Representing scenes as neural radiance fields for view synthesis,” Communications of the ACM, vol. 65, no. 1, pp. 99–106, 2021. 2
work page 2021
-
[5]
Customize your nerf: Adaptive source driven 3d scene editing via local-global iterative training,
R. He, S. Huang, X. Nie, T. Hui, L. Liu, J. Dai, J. Han, G. Li, and S. Liu, “Customize your nerf: Adaptive source driven 3d scene editing via local-global iterative training,” in Proceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 6966–6975. 3
work page 2024
-
[6]
Instant neural graphics primitives with a multiresolution hash encoding,
T. M ¨uller, A. Evans, C. Schied, and A. Keller, “Instant neural graphics primitives with a multiresolution hash encoding,” ACM transactions on graphics (TOG), vol. 41, no. 4, pp. 1– 15, 2022. 2
work page 2022
-
[7]
Mip-nerf: A multi- scale representation for anti-aliasing neural radiance fields,
J. T. Barron, B. Mildenhall, M. Tancik, P. Hedman, R. Martin-Brualla, and P. P. Srinivasan, “Mip-nerf: A multi- scale representation for anti-aliasing neural radiance fields,” in Proceedings of the IEEE/CVF international conference on computer vision, 2021, pp. 5855–5864
work page 2021
-
[8]
Mip-nerf 360: Unbounded anti-aliased neural radiance fields,
J. T. Barron, B. Mildenhall, D. Verbin, P. P. Srinivasan, and P. Hedman, “Mip-nerf 360: Unbounded anti-aliased neural radiance fields,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022, pp. 5470–
work page 2022
Show all 82 references
-
[9]
Zip-nerf: Anti-aliased grid-based neural radi- ance fields,
J. T. Barron, B. Mildenhall, D. Verbin, P. P. Srinivasan, and P. Hedman, “Zip-nerf: Anti-aliased grid-based neural radi- ance fields,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 19 697–19 705
2023
-
[10]
Ref-nerf: Structured view-dependent appearance for neural radiance fields,
D. Verbin, P. Hedman, B. Mildenhall, T. Zickler, J. T. Barron, and P. P. Srinivasan, “Ref-nerf: Structured view-dependent appearance for neural radiance fields,” in 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, 2022, pp. 5481–5490. 2, 4
2022
-
[11]
Geo-neus: Geometry- consistent neural implicit surfaces learning for multi-view reconstruction,
Q. Fu, Q. Xu, Y . S. Ong, and W. Tao, “Geo-neus: Geometry- consistent neural implicit surfaces learning for multi-view reconstruction,” Advances in Neural Information Processing Systems, vol. 35, pp. 3403–3416, 2022. 2
2022
-
[12]
Neus: Learning neural implicit surfaces by vol- ume rendering for multi-view reconstruction,
P. Wang, L. Liu, Y . Liu, C. Theobalt, T. Komura, and W. Wang, “Neus: Learning neural implicit surfaces by vol- ume rendering for multi-view reconstruction,”arXiv preprint arXiv:2106.10689, 2021. 2, 4
2021 arXiv
-
[13]
Neuda: Neural deformable anchor for high-fidelity implicit surface recon- struction,
B. Cai, J. Huang, R. Jia, C. Lv, and H. Fu, “Neuda: Neural deformable anchor for high-fidelity implicit surface recon- struction,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 8476–
2023
-
[14]
Neuralangelo: High-fidelity neural sur- face reconstruction,
Z. Li, T. M ¨uller, A. Evans, R. H. Taylor, M. Unberath, M.-Y . Liu, and C.-H. Lin, “Neuralangelo: High-fidelity neural sur- face reconstruction,” in Proceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition, 2023, pp. 8456–8465
2023
-
[15]
Neus2: Fast learning of neural implicit sur- faces for multi-view reconstruction,
Y . Wang, Q. Han, M. Habermann, K. Daniilidis, C. Theobalt, and L. Liu, “Neus2: Fast learning of neural implicit sur- faces for multi-view reconstruction,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 3295–3306
2023
-
[16]
Adaptive shells for efficient neural radiance field rendering,
Z. Wang, T. Shen, M. Nimier-David, N. Sharp, J. Gao, A. Keller, S. Fidler, T. M ¨uller, and Z. Gojcic, “Adaptive shells for efficient neural radiance field rendering,” arXiv preprint arXiv:2311.10091, 2023. 2
2023 arXiv
-
[17]
3d gaussian splatting for real-time radiance field rendering
B. Kerbl, G. Kopanas, T. Leimk ¨uhler, and G. Drettakis, “3d gaussian splatting for real-time radiance field rendering.” ACM Trans. Graph., vol. 42, no. 4, pp. 139–1, 2023. 1, 2, 6, 12
2023
-
[18]
3dgsr: Implicit surface reconstruction with 3d gaussian splatting,
X. Lyu, Y .-T. Sun, Y .-H. Huang, X. Wu, Z. Yang, Y . Chen, J. Pang, and X. Qi, “3dgsr: Implicit surface reconstruction with 3d gaussian splatting,” arXiv preprint arXiv:2404.00409, 2024
2024 arXiv
-
[19]
2d gaus- sian splatting for geometrically accurate radiance fields,
B. Huang, Z. Yu, A. Chen, A. Geiger, and S. Gao, “2d gaus- sian splatting for geometrically accurate radiance fields,” in ACM SIGGRAPH 2024 Conference Papers, 2024, pp. 1–11
2024
-
[20]
Sugar: Surface-aligned gaus- sian splatting for efficient 3d mesh reconstruction and high- quality mesh rendering,
A. Gu ´edon and V . Lepetit, “Sugar: Surface-aligned gaus- sian splatting for efficient 3d mesh reconstruction and high- quality mesh rendering,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 5354–5363. 2
2024
-
[21]
Mip- splatting: Alias-free 3d gaussian splatting,
Z. Yu, A. Chen, B. Huang, T. Sattler, and A. Geiger, “Mip- splatting: Alias-free 3d gaussian splatting,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pat- tern Recognition, 2024, pp. 19 447–19 456
2024
-
[22]
Gaussianpro: 3d gaussian splatting with progressive propagation,
K. Cheng, X. Long, K. Yang, Y . Yao, W. Yin, Y . Ma, W. Wang, and X. Chen, “Gaussianpro: 3d gaussian splatting with progressive propagation,” in Forty-first International Conference on Machine Learning, 2024. 2
2024
-
[23]
High-resolution image synthesis with latent diffusion models,
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Om- mer, “High-resolution image synthesis with latent diffusion models,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2022, pp. 10 684– 10 695. 1, 2
2022
-
[24]
Photorealistic text-to-image diffusion models with deep language understanding,
C. Saharia, W. Chan, S. Saxena, L. Li, J. Whang, E. L. Den- ton, K. Ghasemipour, R. Gontijo Lopes, B. Karagol Ayan, T. Salimans et al. , “Photorealistic text-to-image diffusion models with deep language understanding,” Advances in neural information processing systems, vol. 35...
2022
-
[25]
Dreambooth: Fine tuning text-to-image diffu- sion models for subject-driven generation,
N. Ruiz, Y . Li, V . Jampani, Y . Pritch, M. Rubinstein, and K. Aberman, “Dreambooth: Fine tuning text-to-image diffu- sion models for subject-driven generation,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2023, pp. 22 500–22 510. 3, 14
2023
-
[26]
Adding conditional control to text-to-image diffusion models,
L. Zhang, A. Rao, and M. Agrawala, “Adding conditional control to text-to-image diffusion models,” in Proceedings of the IEEE/CVF International Conference on Computer Vi- sion, 2023, pp. 3836–3847. 2
2023
-
[27]
Dream- fusion: Text-to-3d using 2d diffusion,
B. Poole, A. Jain, J. T. Barron, and B. Mildenhall, “Dream- fusion: Text-to-3d using 2d diffusion,” arXiv preprint arXiv:2209.14988, 2022. 1, 2, 5, 8, 14
2022 arXiv
-
[28]
Magic3d: High-resolution text-to-3d content creation,
C.-H. Lin, J. Gao, L. Tang, T. Takikawa, X. Zeng, X. Huang, K. Kreis, S. Fidler, M.-Y . Liu, and T.-Y . Lin, “Magic3d: High-resolution text-to-3d content creation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pat- tern Recognition, 2023, pp. 300–309
2023
-
[29]
Dreamhuman: Animatable 3d avatars from text,
N. Kolotouros, T. Alldieck, A. Zanfir, E. Bazavan, M. Fier- aru, and C. Sminchisescu, “Dreamhuman: Animatable 3d avatars from text,”Advances in Neural Information Process- ing Systems, vol. 36, 2024. 1, 2
2024
-
[30]
Dreamavatar: Text-and-shape guided 3d human avatar generation via diffusion models,
Y . Cao, Y .-P. Cao, K. Han, Y . Shan, and K.-Y . K. Wong, “Dreamavatar: Text-and-shape guided 3d human avatar generation via diffusion models,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 958–968
2024
-
[31]
Dreamwaltz: Make a scene with com- plex 3d animatable avatars,
Y . Huang, J. Wang, A. Zeng, H. Cao, X. Qi, Y . Shi, Z.-J. Zha, and L. Zhang, “Dreamwaltz: Make a scene with com- plex 3d animatable avatars,”Advances in Neural Information Processing Systems, vol. 36, 2024. 1
2024
-
[32]
Humannorm: Learning normal diffusion model for high-quality and realistic 3d human generation,
X. Huang, R. Shao, Q. Zhang, H. Zhang, Y . Feng, Y . Liu, and Q. Wang, “Humannorm: Learning normal diffusion model for high-quality and realistic 3d human generation,” in Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 4568–4577. ...
2024
-
[33]
Seeavatar: Photorealistic text- to-3d avatar generation with constrained geometry and ap- pearance,
Y . Xu, Z. Yang, and Y . Yang, “Seeavatar: Photorealistic text- to-3d avatar generation with constrained geometry and ap- pearance,” arXiv preprint arXiv:2312.08889, 2023
2023 arXiv
-
[34]
Barbie: Text to barbie-style 3d avatars,
X. Sun, Z. Zhang, Y . Tai, Q. Wang, H. Tang, Z. Yi, and J. Yang, “Barbie: Text to barbie-style 3d avatars,” arXiv preprint arXiv:2408.09126, 2024
2024
-
[35]
Tada! text to animatable digital avatars,
T. Liao, H. Yi, Y . Xiu, J. Tang, Y . Huang, J. Thies, and M. J. Black, “Tada! text to animatable digital avatars,” in 2024 International Conference on 3D Vision (3DV). IEEE, 2024, pp. 1508–1519. 1, 2
2024
-
[36]
Smpl: A skinned multi-person linear model,
M. Loper, N. Mahmood, J. Romero, G. Pons-Moll, and M. J. Black, “Smpl: A skinned multi-person linear model,” in Seminal Graphics Papers: Pushing the Boundaries, Volume 2, 2023, pp. 851–866. 2
2023
-
[37]
Deep marching tetrahedra: a hybrid representation for high- resolution 3d shape synthesis,
T. Shen, J. Gao, K. Yin, M.-Y . Liu, and S. Fidler, “Deep marching tetrahedra: a hybrid representation for high- resolution 3d shape synthesis,”Advances in Neural Informa- tion Processing Systems, vol. 34, pp. 6087–6101, 2021. 2, 3, 12
2021
-
[38]
Avatar- booth: High-quality and customizable 3d human avatar gen- eration,
Y . Zeng, Y . Lu, X. Ji, Y . Yao, H. Zhu, and X. Cao, “Avatar- booth: High-quality and customizable 3d human avatar gen- eration,” arXiv preprint arXiv:2306.09864, 2023. 1, 2
2023 arXiv
-
[39]
Puzzlea- vatar: Assembling 3d avatars from personal albums,
Y . Xiu, Y . Ye, Z. Liu, D. Tzionas, and M. J. Black, “Puzzlea- vatar: Assembling 3d avatars from personal albums,” arXiv preprint arXiv:2405.14869, 2024. 2
2024 arXiv
-
[40]
Dreamv- ton: Customizing 3d virtual try-on with personalized diffu- sion models,
Z. Xie, H. Dong, Y . Gao, Z. Ma, and X. Liang, “Dreamv- ton: Customizing 3d virtual try-on with personalized diffu- sion models,” in Proceedings of the 32nd ACM International Conference on Multimedia, 2024, pp. 10 784–10 793
2024
-
[41]
Tech: Text-guided reconstruction of lifelike clothed humans,
Y . Huang, H. Yi, Y . Xiu, T. Liao, J. Tang, D. Cai, and J. Thies, “Tech: Text-guided reconstruction of lifelike clothed humans,” in 2024 International Conference on 3D Vision (3DV). IEEE, 2024, pp. 1531–1542. 1, 2, 6, 13
2024
-
[42]
Lora: Low-rank adaptation of large language models,
E. J. Hu, Y . Shen, P. Wallis, Z. Allen-Zhu, Y . Li, S. Wang, L. Wang, and W. Chen, “Lora: Low-rank adaptation of large language models,” arXiv preprint arXiv:2106.09685 , 2021. 3, 14
2021 arXiv
-
[43]
Humangaussian: Text-driven 3d hu- man generation with gaussian splatting,
X. Liu, X. Zhan, J. Tang, Y . Shan, G. Zeng, D. Lin, X. Liu, and Z. Liu, “Humangaussian: Text-driven 3d hu- man generation with gaussian splatting,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 6646–6657. 3
2024
-
[44]
Gavatar: Animatable 3d gaussian avatars with implicit mesh learning,
Y . Yuan, X. Li, Y . Huang, S. De Mello, K. Nagano, J. Kautz, and U. Iqbal, “Gavatar: Animatable 3d gaussian avatars with implicit mesh learning,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 896–905
2024
-
[45]
Headstudio: Text to animatable head avatars with 3d gaussian splatting,
Z. Zhou, F. Ma, H. Fan, and Y . Yang, “Headstudio: Text to animatable head avatars with 3d gaussian splatting,” arXiv preprint arXiv:2402.06149, 2024. 3
2024 arXiv
-
[46]
Instruct-nerf2nerf: Editing 3d scenes with instructions,
A. Haque, M. Tancik, A. A. Efros, A. Holynski, and A. Kanazawa, “Instruct-nerf2nerf: Editing 3d scenes with instructions,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 19 740–19 750. 3, 8, 14
2023
-
[47]
Dreamedi- tor: Text-driven 3d scene editing with neural fields,
J. Zhuang, C. Wang, L. Lin, L. Liu, and G. Li, “Dreamedi- tor: Text-driven 3d scene editing with neural fields,” in SIG- GRAPH Asia 2023 Conference Papers, 2023, pp. 1–10. 3
2023
-
[48]
Consistdreamer: 3d-consistent 2d dif- fusion for high-fidelity scene editing,
J.-K. Chen, S. R. Bul `o, N. M¨uller, L. Porzi, P. Kontschieder, and Y .-X. Wang, “Consistdreamer: 3d-consistent 2d dif- fusion for high-fidelity scene editing,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 21 071–21 080
2024
-
[49]
Posterior distillation sam- pling,
J. Koo, C. Park, and M. Sung, “Posterior distillation sam- pling,” inProceedings of the IEEE/CVF Conference on Com- puter Vision and Pattern Recognition , 2024, pp. 13 352– 13 361
2024
-
[50]
Genn2n: Gener- ative nerf2nerf translation,
X. Liu, H. Xue, K. Luo, P. Tan, and L. Yi, “Genn2n: Gener- ative nerf2nerf translation,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 5105–5114. 3
2024
-
[51]
Instructpix2pix: Learning to follow image editing instructions,
T. Brooks, A. Holynski, and A. A. Efros, “Instructpix2pix: Learning to follow image editing instructions,” in Proceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 18 392–18 402. 3, 6, 14
2023
-
[52]
Gaus- sianeditor: Editing 3d gaussians delicately with text instruc- tions,
J. Wang, J. Fang, X. Zhang, L. Xie, and Q. Tian, “Gaus- sianeditor: Editing 3d gaussians delicately with text instruc- tions,” inProceedings of the IEEE/CVF Conference on Com- puter Vision and Pattern Recognition , 2024, pp. 20 902– 20 911. 2, 3, 6, 8, 14
2024
-
[53]
Dge: Direct gaussian 3d editing by consistent multi-view editing,
M. Chen, I. Laina, and A. Vedaldi, “Dge: Direct gaussian 3d editing by consistent multi-view editing,” arXiv preprint arXiv:2404.18929, 2024. 2, 3, 6, 8, 14
2024 arXiv
-
[54]
3dego: 3d editing on the go!
U. Khalid, H. Iqbal, A. Farooq, J. Hua, and C. Chen, “3dego: 3d editing on the go!” in European Conference on Computer Vision. Springer, 2025, pp. 73–89
2025
-
[55]
Gaus- sianvton: 3d human virtual try-on via multi-stage gaus- sian splatting editing with image prompting,
H. Chen, Y . Huang, H. Huang, X. Ge, and D. Shao, “Gaus- sianvton: 3d human virtual try-on via multi-stage gaus- sian splatting editing with image prompting,” arXiv preprint arXiv:2405.07472, 2024. 2
2024 arXiv
-
[56]
View-consistent 3d editing with gaussian splatting,
Y . Wang, X. Yi, Z. Wu, N. Zhao, L. Chen, and H. Zhang, “View-consistent 3d editing with gaussian splatting,” in Eu- ropean Conference on Computer Vision . Springer, 2025, pp. 404–420
2025
-
[57]
Gs-vton: Controllable 3d virtual try-on with gaussian splatting,
Y . Cao, M. Hadi, L. Pan, and Z. Liu, “Gs-vton: Controllable 3d virtual try-on with gaussian splatting,” arXiv preprint arXiv:2410.05259, 2024. 3
2024 arXiv
-
[58]
Improv- ing diffusion models for authentic virtual try-on in the wild,
Y . Choi, S. Kwak, K. Lee, H. Choi, and J. Shin, “Improv- ing diffusion models for authentic virtual try-on in the wild,” arXiv preprint arXiv:2403.05139, 2024. 6, 13
2024 arXiv
-
[59]
Texture: Text-guided texturing of 3d shapes,
E. Richardson, G. Metzer, Y . Alaluf, R. Giryes, and D. Cohen-Or, “Texture: Text-guided texturing of 3d shapes,” in ACM SIGGRAPH 2023 conference proceedings, 2023, pp. 1–11. 5, 8
2023
-
[60]
Paint3d: Paint anything 3d with lighting-less texture diffusion models,
X. Zeng, X. Chen, Z. Qi, W. Liu, Z. Zhao, Z. Wang, B. Fu, Y . Liu, and G. Yu, “Paint3d: Paint anything 3d with lighting-less texture diffusion models,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 4252–4262
2024
-
[61]
Hyperdreamer: Hyper-realistic 3d content gen- eration and editing from a single image,
T. Wu, Z. Li, S. Yang, P. Zhang, X. Pan, J. Wang, D. Lin, and Z. Liu, “Hyperdreamer: Hyper-realistic 3d content gen- eration and editing from a single image,” inSIGGRAPH Asia 2023 Conference Papers, 2023, pp. 1–10. 5, 8
2023
-
[62]
Structure-from-motion revisited,
J. L. Sch ¨onberger and J.-M. Frahm, “Structure-from-motion revisited,” in Conference on Computer Vision and Pattern Recognition (CVPR), 2016. 6
2016
-
[63]
Adam: A method for stochastic optimiza- tion,
D. P. Kingma, “Adam: A method for stochastic optimiza- tion,” arXiv preprint arXiv:1412.6980, 2014. 6
2014 arXiv
-
[64]
Decoupled weight decay regularization,
I. Loshchilov, “Decoupled weight decay regularization,” arXiv preprint arXiv:1711.05101, 2017. 6
2017 arXiv
-
[65]
Zero-shot text-to-image generation,
A. Ramesh, M. Pavlov, G. Goh, S. Gray, C. V oss, A. Rad- ford, M. Chen, and I. Sutskever, “Zero-shot text-to-image generation,” in International conference on machine learn- ing. Pmlr, 2021, pp. 8821–8831. 1
2021
-
[66]
Prolificdreamer: High-fidelity and diverse text-to-3d gener- ation with variational score distillation,
Z. Wang, C. Lu, Y . Wang, F. Bao, C. Li, H. Su, and J. Zhu, “Prolificdreamer: High-fidelity and diverse text-to-3d gener- ation with variational score distillation,” Advances in Neural Information Processing Systems, vol. 36, 2024. 1
2024
-
[67]
High-resolution image synthesis with latent diffu- sion models,
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Om- mer, “High-resolution image synthesis with latent diffu- sion models,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , June 2022, pp. 10 684–10 695
2022
-
[68]
Controlnetplus
xinsir6, “Controlnetplus.” 2024, https://github.com/xinsir6/ ControlNetPlus. 6, 13
2024
-
[69]
An efficient method of triangulat- ing equi-valued surfaces by using tetrahedral cells,
A. Doi and A. Koide, “An efficient method of triangulat- ing equi-valued surfaces by using tetrahedral cells,” IEICE TRANSACTIONS on Information and Systems, vol. 74, no. 1, pp. 214–224, 1991. 3
1991
-
[70]
Gpt-4 technical report,
J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat et al. , “Gpt-4 technical report,” arXiv preprint arXiv:2303.08774, 2023. 6, 13
2023 arXiv
-
[71]
Pifu: Pixel-aligned implicit function for high-resolution clothed human digitization,
S. Saito, Z. Huang, R. Natsume, S. Morishima, A. Kanazawa, and H. Li, “Pifu: Pixel-aligned implicit function for high-resolution clothed human digitization,” in Proceedings of the IEEE/CVF international conference on computer vision, 2019, pp. 2304–2314. 6
2019
-
[72]
Segment anything,
A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.-Y . Lo et al., “Segment anything,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 4015–4026. 4, 14
2023
-
[73]
Self-correction for hu- man parsing,
P. Li, Y . Xu, Y . Wei, and Y . Yang, “Self-correction for hu- man parsing,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 44, no. 6, pp. 3260–3271, 2020. 4
2020
-
[74]
Gans trained by a two time-scale update rule converge to a local nash equilibrium,
M. Heusel, H. Ramsauer, T. Unterthiner, B. Nessler, and S. Hochreiter, “Gans trained by a two time-scale update rule converge to a local nash equilibrium,” Advances in neural information processing systems, vol. 30, 2017. 7
2017
-
[75]
Learning transferable visual models from natural language supervision,
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark et al., “Learning transferable visual models from natural language supervision,” in International conference on machine learn- ing. PMLR, 2021, pp. 8748–8763. 7
2021
-
[76]
Dinov2: Learning robust visual features without su- pervision,
M. Oquab, T. Darcet, T. Moutakanni, H. V o, M. Szafraniec, V . Khalidov, P. Fernandez, D. Haziza, F. Massa, A. El-Nouby et al., “Dinov2: Learning robust visual features without su- pervision,” arXiv preprint arXiv:2304.07193, 2023. 7
2023 arXiv
-
[77]
Direct learning of mesh and appearance via 3d gaussian splatting,
A. Lin and J. Li, “Direct learning of mesh and appearance via 3d gaussian splatting,”arXiv preprint arXiv:2405.06945,
-
[78]
Games: Mesh-based adapting and modification of gaussian splatting,
J. Waczy ´nska, P. Borycki, S. Tadeja, J. Tabor, and P. Spurek, “Games: Mesh-based adapting and modification of gaussian splatting,” arXiv preprint arXiv:2402.01459, 2024. 5
2024 arXiv
-
[79]
Sdedit: Guided image synthesis and edit- ing with stochastic differential equations,
C. Meng, Y . He, Y . Song, J. Song, J. Wu, J.-Y . Zhu, and S. Ermon, “Sdedit: Guided image synthesis and edit- ing with stochastic differential equations,” arXiv preprint arXiv:2108.01073, 2021. 6
2021 arXiv
-
[80]
Implicit geometric regularization for learning shapes,
A. Gropp, L. Yariv, N. Haim, M. Atzmon, and Y . Lip- man, “Implicit geometric regularization for learning shapes,” arXiv preprint arXiv:2002.10099, 2020. 12
2002 arXiv
-
[81]
Delta denoising score,
A. Hertz, K. Aberman, and D. Cohen-Or, “Delta denoising score,” in Proceedings of the IEEE/CVF International Con- ference on Computer Vision, 2023, pp. 2328–2337. 14 Creating Your Editable 3D Photorealistic Avatar with Tetrahedron-constrained Gaussian Splatting Supplementary M...
2023
-
[512]
A w o m a n w e a r i n g a p i n k T - s h i r t w i t h r e d a c c e n t s
We set the global and local text promptsyG andyL as ”photo of a man/woman wearing a ... garment”and ”photo of a ... garment” , respectively. For calculating the geomet- ric guidance LG SDS andLL SDS , we use a publicly available normal-adapted Stable Diffusion V1.5 model [32] ...
Reviewed August 16, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.