REVIEW 5 major objections 7 minor 73 references
SAGA: Surface-Aligned Gaussian Avatar
T0 review · 5 major / 7 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read This paper claims that a two-stage adhere-then-detach mesh alignment turns 3D Gaussian Splatting into animatable human avatars that generalize to novel views and poses and support direct mesh extraction.
desk verdict SAGA is a credible two-stage Gaussian-mesh avatar with supportive ablations, but the 37%/13% claim doesn't survive contact with its own table, and the 'first direct mesh extraction' is really depth-fusion; worth reviewing, not ready as-is. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the bound-triangle relationship between a Gaussian and a mesh face, expressed through barycentric coordinates. A Gaussian's center is a weighted sum of the three vertices of its bound triangle, and its flat normal is aligned with the triangle normal; jointly optimizing barycentric coordinates and vertex positions lets Gaussians flow on the mesh instead of being pinned in place. In the Detached Stage, the same binding is retained as a soft regularizer: a position loss (minimum distance to the triangle) and a normal loss (absolute cosine distance) keep geometry on-mesh, while a retraction step in Stage 1 and a Walking-on-Mesh step in Stage 2 correct Gaussians that slide out of their triangle. This combination is what lets the mesh transfer its smooth, well-defined geometry and natural skinning weights to the Gaussians without sacrificing the expressivity of 3DGS.
What would settle it
Run SAGA on a monocular sequence of a person in loose clothing, then compare the rendered depth and the volumetric depth-fusion mesh against a synchronized multi-view scan; if the template mesh is far from the true surface, the alignment losses should pull the Gaussians to the wrong depth and produce visible artifacts in exactly the loose regions.
Extended reading notes
Core claim
On the paper's own terms, the central discovery is a two-stage surface-aligned Gaussian representation for monocularly reconstructed, animatable humans. In the Adhered Stage, each Gaussian center is a barycentric combination of the vertices of a bound triangle of the SMPL body mesh; the Gaussian is flattened along one axis and its normal is locked to the triangle normal, while both barycentric coordinates and mesh vertices are optimized so the Gaussian can flow over the surface. In the Detached Stage, the center is freed and reparameterized directly, but the Gaussian keeps its triangle binding and is pulled back toward that triangle by a position-alignment loss (minimum distance to the triangle, not just a plane) and an orientation-alignment loss (cosine distance between the Gaussian normal and the triangle normal). A Walking-on-Mesh routine replaces the bound triangle with the nearest adjacent one when a Gaussian drifts outside, keeping the regularization honest. The paper argues this adhere-then-detach schedule enforces well-defined geometry and frame-consistent deformation, and for the first time permits high-quality mesh extraction directly from deformable Gaussians learned from monocular video.
Load-bearing premise
Everything rests on the template body mesh staying close to the real clothed surface once the first stage is done, which is untested for loose clothing and unusual body shapes.
Editorial extensions
If this is right
- Animating an avatar into a new pose should stop producing the broken zippers, armpit fractures, and needle-like joint artifacts shown for unregularized Gaussians, because the mesh binding supplies smoother deformation and more natural skinning weights.
- A trained avatar should be usable as a mesh, not only as a renderer: fusing rendered depth maps should give a clean surface that can be retargeted, edited, or relit.
- The cost profile should stay practical for interactive use: about 12 minutes of training on a single GPU and real-time rendering at 60+ FPS at 512x512 resolution.
- Warm-starting with a brief strict-adhered stage should be enough to inherit the mesh's geometry discipline, avoiding the long training runs that fully rigid mesh-bound avatars require.
- Compared with a rigid fixed-on-mesh approach, the paper reports both higher rendering quality and a roughly 150x training speedup, showing that the two-stage release of constraints is what makes mesh alignment affordable.
Reading between the lines
- The same adhere-then-detach schedule should transfer to other articulated subjects, such as animals or deformable objects, whenever a coarse template mesh is available, because the mechanism is template regularization rather than human-specific skinning.
- A direct stress test the paper leaves open is loose or flowing clothing; if the template mesh deviates far from the true surface, the alignment losses would pull Gaussians to the wrong location, so an adaptive alignment weight based on estimated mesh-to-surface error would be a natural extension.
- The triangle-binding across frames could also serve as an explicit temporal smoothness prior: counting how often Walking-on-Mesh reassigns a Gaussian and penalizing frequent reassignment would regularize non-rigid motion beyond what the paper reports.
- Multi-view or depth supervision at test time could relax the dependence on the template mesh and let the same representation capture geometry on bodies that the parametric body prior fits poorly.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes SAGA, a two-stage surface-aligned Gaussian representation for animatable human avatars reconstructed from monocular video. In the first Adhered Stage, Gaussian centers are parameterized by barycentric coordinates on SMPL mesh triangles and jointly optimized with the mesh vertices, while the Gaussians are flattened and their normals fixed to the triangle normals, enforcing strict surface adherence while allowing the Gaussians to flow on the mesh. In the second Detached Stage, the Gaussians are freed but kept under a position-alignment loss (Eq. 12) and a normal-alignment loss (Eq. 13) relative to their bound triangles, and a Walking-on-Mesh strategy (Algorithm 1) re-binds Gaussians that drift outside their triangles. A non-rigid deformation module and a pose-conditioned colorization MLP handle motion and appearance change. Experiments on ZJU-MoCap, MonoCap, and PeopleSnapshot report state-of-the-art or competitive novel-view and novel-pose results with roughly 12-minute training and 60+ FPS rendering, and the paper claims the first direct high-quality mesh extraction from deformable Gaussians learned from monocular video.
Significance. If the results hold, the adhere-then-detach idea is a genuinely useful contribution to the monocular avatar literature: it offers a clean way to obtain the geometric regularization of mesh-bound Gaussians without the expressivity loss of rigid binding, and it is supported by a structured set of ablations. Concrete strengths include that the paper runs official code of all baselines under unified protocols, reports training and rendering efficiency concretely, provides per-component ablations (Tabs. 4-7), measures the overhead of the walking-on-mesh module (1.7 ms), and evaluates on three datasets. The evaluation is not circular: the mesh alignment is a regularizer rather than a fitted target, and the reported metrics are on held-out views and poses. However, the quantitative support is thinner than the text suggests in several places: the headline percentage improvements are not reproducible from the tables, the ablation of the central alignment loss is at noise level, and the robustness of the core assumption (that the SMPL mesh tracks the true clothed surface) is untested in the regime where it could fail. These issues are fixable within the manuscript's scope.
major comments (5)
- [6.3.1, Tab. 1] The sentence "our method can notably surpass them on challenging subject by 37% and 13% respectively" is not reproducible from Tab. 1. For every subject and metric in that table, the relative improvement of SAGA over InstantNVR is at most approximately 29% (LPIPS on subject 394: (40.00-28.48)/40.00) and over GauHuman at most approximately 12% (LPIPS on subject 387: (38.72-34.20)/38.72), and on PSNR GauHuman actually beats SAGA on three of the six subjects (377, 387, 394). The authors should either identify the exact computation behind the 37% and 13% figures or correct the claim, since as written the headline margin is not supported by the reported numbers.
- [6.5.2, Tab. 5] The central quantitative evidence for the Gaussian-Mesh Alignment Regularization, one of the paper's main contributions, is a mean PSNR gain of 0.03 dB (31.08 vs 31.11), an SSIM gain of 0.0002, and an LPIPS gain of 0.0000 over six subjects. These margins are at or below typical run-to-run variation, and the loss weights entering L_geo (lambda_pa and lambda_na in Sec. 4.2.2) are not reported anywhere in the manuscript. Please provide per-subject results, the exact weights, and ideally multiple seeds or a statistical test; without these, the ablation cannot support the statement in Sec. 6.5.2 that the alignment loss "effectively regularizes the detached Gaussians" on the strength of the quantitative evidence alone.
- [4.2.2, Eqs. (12)-(13)] A load-bearing assumption of the two-stage scheme is that the optimized SMPL mesh remains a faithful proxy for the clothed surface during the Detached Stage, and this assumption is untested in the regime where it can fail. All ZJU-MoCap and PeopleSnapshot subjects wear tight-fitting clothing, yet Eqs. (12)-(13) pull every detached Gaussian toward its bound triangle; if a garment deviates strongly from body shape (skirts, baggy trousers, dresses), these losses would bias Gaussians toward the template rather than the true surface, and Walking-on-Mesh only re-binds within the same template's adjacent triangles. The paper should quantify the mesh-surface discrepancy (for example, silhouette error or distance of the learned canonical mesh to the reconstructed depth) and/or add a subject with loose clothing. I do not see a circularity problem here, since the mesh acts as a regularizer and the metrics are on held-out views and poses; the issue is empirical coverage of the proxy assumption, not the logic of the method. The small effect of the alignment loss in Tab. 5 in the tested regime makes the risk of a negative effect in the untested regime plausible rather than merely hypothetical.
- [6.5.1, Tab. 4] The stage-ablation protocols are underspecified, which limits what Tab. 4 can establish. For the "w/o adhered stage" row it is not stated how and where the Gaussians are initialized, nor when the Detached Stage begins (iteration 0 or iteration 3k). For the "w/o detached stage" row it is not stated whether training continues to 15k iterations under the adhered parameterization or stops at 3k, and since the non-rigid deformation module is active only in the Detached Stage, the ablation varies stage duration jointly with module activation. Please specify the exact training schedule and initialization used for each row.
- [6.6] The claim of "direct high-quality mesh extraction" and the "first successful attempt" (abstract, Sec. 1, Sec. 7) is supported only by qualitative depth visualizations and TSDF-fusion examples (Figs. 16-17), with no quantitative geometry metric. Moreover, because the representation is built on a deformed SMPL template, the extracted mesh inherits the SMPL prior by construction, so the improvement over 3DGS-Avatar in Fig. 17 may largely reflect that prior rather than the proposed alignment mechanism itself. Please add a quantitative geometry comparison (for example, multi-view depth consistency, silhouette IoU, or distance to a reference obtained from the 22 held-out ZJU-MoCap views) or temper the claim accordingly.
minor comments (7)
- [Throughout] Typographical and naming issues: "neglectable" (Sec. 6.5.3) should be "negligible"; "OutOfTrianlge" in Algorithm 1; "representaion" in the Tab. 4 caption; "SpattingAvatar" vs "SplattingAvatar" should be unified (Sec. 6.2 and Fig. 8 caption); and "Peoplesnapshot" vs "PeopleSnapshot" capitalization is inconsistent.
- [Eqs. (12), (17)] Eq. (12) indexes the sum from i=0 to N while Gaussians are indexed i=1..N elsewhere. In Eq. (17), after clamping the barycentric coordinates to [0,1], the coordinates are not renormalized to sum to 1; the normalization step should be stated explicitly.
- [Sec. 4.4, Eq. (24)] The per-frame latent vector psi in Eq. (24) is described as a "per-frame latent vector" but it is not explained how it is obtained (a learnable embedding per training frame?) or how it behaves for novel poses; please clarify.
- [Sec. 6.3.2, Tab. 3] The 42% LPIPS gain over SplattingAvatar on female-3-casual is arithmetically consistent with Tab. 3, but the explanation that PSNR/SSIM "favor the smoothed blurry results" is asserted without supporting evidence; either provide a quantitative demonstration (for example, a patch-level analysis) or soften the claim.
- [Sec. 6.4] Novel pose synthesis is one of the two headline tasks but is evaluated only qualitatively (Figs. 9, 10, 12); adding a quantitative metric on the AIST++/AMASS animations would make the generalization claim testable.
- [Sec. 5.2] Key hyperparameters (epsilon in Eq. (7), the loss weights lambda_mask, lambda_LPIPS, lambda_lap, lambda_normal, and the weights of the alignment losses, as well as the Gaussian count N) are deferred to the supplementary, which is not included in this submission; please include them in the main text or make the supplement available.
- [Tab. 2] In Tab. 2, 3DGSAvatar achieves better LPIPS than SAGA on Olek (0.0115 vs 0.0116) and Vlad (0.0158 vs 0.0165), so the caption claim "outperforms the comparison methods on most subjects" should acknowledge this exception or be rephrased per metric.
Circularity Check
No significant circularity: the two-stage mesh alignment is a regularizer, novel view and pose results are evaluated on held-out data, and the geometry extraction claim explicitly relies on the stated SMPL prior rather than on a fitted prediction.
full rationale
The paper's load-bearing claim is that aligning Gaussians with an SMPL mesh in an adhere-then-detach scheme improves novel view and pose synthesis. This is an empirical method proposal, not a derivation that reduces to its inputs. The alignment losses (Eqs. 12 and 13) are regularizers added to the appearance loss; Gaussian centers and rotations remain free variables optimized against image reconstruction on training frames, while the quantitative comparisons (Tables 1-3) and novel-pose evaluations (Figs. 9-10) use held-out views and out-of-distribution poses from AIST++ and AMASS. No reported metric is a fitted parameter renamed as a prediction: the ablation in Table 5 shows a small but real effect of the alignment loss on held-out metrics, so the improvement is not forced by construction. The extracted-mesh claim in Sec. 6.6 does inherit smoothness and template shape from the SMPL mesh, but the paper is explicit about this design ('leveraging the mesh as a geometry regularizer'), making it a stated prior rather than a hidden circular step. There are no load-bearing self-citations: the method builds on external works such as SMPL, 3DGS, and prior Gaussian-avatar methods, none of which are by the present authors. No uniqueness theorem is imported from the authors' own prior work, and no equation sets the predicted quantity equal to an input by definition. The paper is self-contained with respect to circularity; the unsupported 'first successful attempt' novelty claim, if challenged, would be a correctness or literature issue, not a circularity issue.
Assumptions & free parameters
free parameters (4)
- Gaussian flatness scale epsilon (s0 = epsilon) =
not specified numerically, 'epsilon << s1,s2'
- Stage schedule: 3k adhered / 12k detached iterations =
3000 / 12000
- Loss weights: lambda_lap, lambda_normal, lambda_mask, lambda_LPIPS =
not reported in main text; delegated to supplementary
- Initial learning rate for non-rigid and colorization MLPs =
1e-3 with 0.1 exponential decay
assumptions (5)
- domain assumption SMPL is a sufficiently accurate coarse proxy for the real clothed human surface in canonical space.
- domain assumption Mesh alignment regularizes monocular dynamic reconstruction and improves novel view/pose generalization.
- domain assumption Gaussian flattening with the smallest-scale axis as the surface normal adequately represents local surface orientation.
- standard math Linear blend skinning with SMPL pose parameters describes articulated human deformation.
- domain assumption Projection of a Gaussian center onto its bound triangle (Eq. 11) is differentiable and valid for detached Gaussians.
Cite this review
Pith. "Pith review of SAGA: Surface-Aligned Gaussian Avatar." pith.science (2026). https://pith.science/paper/ITKPO4WK
@misc{pith2026241200845,
author = {Pith},
title = {Pith review of: SAGA: Surface-Aligned Gaussian Avatar},
year = {2026},
howpublished = {\url{https://pith.science/paper/ITKPO4WK}},
note = {Machine review of arXiv:2412.00845}
}
read the original abstract
This paper presents a Surface-Aligned Gaussian representation for creating animatable human avatars from monocular videos,aiming at improving the novel view and pose synthesis performance while ensuring fast training and real-time rendering. Recently,3DGS has emerged as a more efficient and expressive alternative to NeRF, and has been used for creating dynamic human avatars. However,when applied to the severely ill-posed task of monocular dynamic reconstruction, the Gaussians tend to overfit the constantly changing regions such as clothes wrinkles or shadows since these regions cannot provide consistent supervision, resulting in noisy geometry and abrupt deformation that typically fail to generalize under novel views and poses.To address these limitations, we present SAGA,i.e.,Surface-Aligned Gaussian Avatar,which aligns the Gaussians with a mesh to enforce well-defined geometry and consistent deformation, thereby improving generalization under novel views and poses. Unlike existing strict alignment methods that suffer from limited expressive power and low realism,SAGA employs a two-stage alignment strategy where the Gaussians are first adhered on while then detached from the mesh, thus facilitating both good geometry and high expressivity. In the Adhered Stage, we improve the flexibility of Adhered-on-Mesh Gaussians by allowing them to flow on the mesh, in contrast to existing methods that rigidly bind Gaussians to fixed location. In the second Detached Stage, we introduce a Gaussian-Mesh Alignment regularization, which allows us to unleash the expressivity by detaching the Gaussians but maintain the geometric alignment by minimizing their location and orientation offsets from the bound triangles. Finally, since the Gaussians may drift outside the bound triangles during optimization, an efficient Walking-on-Mesh strategy is proposed to dynamically update the bound triangles.
Figures
Figures from the paper (12 more)
Reference graph
Works this paper leans on
-
[1]
Fusion4d: Real- time performance capture of challenging scenes,
M. Dou, S. Khamis, Y . Degtyarev, P. Davidson, S. R. Fanello, A. Kowdle, S. O. Escolano, C. Rhemann, D. Kim, J. Taylor et al., “Fusion4d: Real- time performance capture of challenging scenes,” ACM Transactions on Graphics (ToG), vol. 35, no. 4, pp. 1–13, 2016
work page 2016
-
[2]
Holoportation: Virtual 3d teleportation in real-time,
S. Orts-Escolano, C. Rhemann, S. Fanello, W. Chang, A. Kowdle, Y . Degtyarev, D. Kim, P. L. Davidson, S. Khamis, M. Dou et al. , “Holoportation: Virtual 3d teleportation in real-time,” in Proceedings of the 29th annual symposium on user interface software and technology , 2016, pp. 741–754
work page 2016
-
[3]
The relightables: V olumetric performance capture of humans with realistic relighting,
K. Guo, P. Lincoln, P. Davidson, J. Busch, X. Yu, M. Whalen, G. Harvey, S. Orts-Escolano, R. Pandey, J. Dourgarian et al. , “The relightables: V olumetric performance capture of humans with realistic relighting,” ACM Transactions on Graphics (ToG), vol. 38, no. 6, pp. 1–19, 2019
work page 2019
-
[4]
High-quality streamable free- viewpoint video,
A. Collet, M. Chuang, P. Sweeney, D. Gillett, D. Evseev, D. Calabrese, H. Hoppe, A. Kirk, and S. Sullivan, “High-quality streamable free- viewpoint video,” ACM Transactions on Graphics (ToG), vol. 34, no. 4, pp. 1–13, 2015
work page 2015
-
[5]
State of the art on neural rendering,
A. Tewari, O. Fried, J. Thies, V . Sitzmann, S. Lombardi, K. Sunkavalli, R. Martin-Brualla, T. Simon, J. Saragih, M. Nießner et al., “State of the art on neural rendering,” in Computer Graphics Forum , vol. 39, no. 2. Wiley Online Library, 2020, pp. 701–727
work page 2020
-
[6]
NeRF: Representing scenes as neural radiance fields for view synthesis,
B. Mildenhall, P. P. Srinivasan, M. Tancik, J. T. Barron, R. Ramamoorthi, and R. Ng, “NeRF: Representing scenes as neural radiance fields for view synthesis,” in European Conference on Computer Vision, 2020, pp. 405– 421
work page 2020
-
[7]
S. Peng, Y . Zhang, Y . Xu, Q. Wang, Q. Shuai, H. Bao, and X. Zhou, “Neural body: Implicit neural representations with structured latent codes for novel view synthesis of dynamic humans,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 9054–9063
work page 2021
-
[8]
HumanNeRF: Free-viewpoint rendering of moving people from monocular video,
C.-Y . Weng, B. Curless, P. P. Srinivasan, J. T. Barron, and I. Kemelmacher-Shlizerman, “HumanNeRF: Free-viewpoint rendering of moving people from monocular video,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , June 2022, pp. 16 210–16 220
work page 2022
Show all 73 references
-
[9]
Animatable implicit neural representations for creating realistic avatars from videos,
S. Peng, Z. Xu, J. Dong, Q. Wang, S. Zhang, Q. Shuai, H. Bao, and X. Zhou, “Animatable implicit neural representations for creating realistic avatars from videos,” IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024
2024
-
[10]
Neural actor: Neural free-view synthesis of human actors with pose control,
L. Liu, M. Habermann, V . Rudnev, K. Sarkar, J. Gu, and C. Theobalt, “Neural actor: Neural free-view synthesis of human actors with pose control,” ACM Transactions on Graphics (TOG) , vol. 40, no. 6, pp. 1– 16, 2021
2021
-
[11]
Neuman: Neural human radiance field from a single video,
W. Jiang, K. M. Yi, G. Samei, O. Tuzel, and A. Ranjan, “Neuman: Neural human radiance field from a single video,” in Computer Vision–ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23–27, 2022, Proceedings, Part XXXII. Springer, 2022, pp. 402–418
2022
-
[12]
Instant neural graphics primitives with a multiresolution hash encoding,
T. M ¨uller, A. Evans, C. Schied, and A. Keller, “Instant neural graphics primitives with a multiresolution hash encoding,” ACM Trans. Graph. , vol. 41, no. 4, pp. 102:1–102:15, Jul. 2022. [Online]. Available: https://doi.org/10.1145/3528223.3530127
2022
-
[13]
Tensorf: Tensorial radiance fields,
A. Chen, Z. Xu, A. Geiger, J. Yu, and H. Su, “Tensorf: Tensorial radiance fields,” in European Conference on Computer Vision (ECCV), 2022
2022
-
[14]
Direct voxel grid optimization: Super- fast convergence for radiance fields reconstruction,
C. Sun, M. Sun, and H.-T. Chen, “Direct voxel grid optimization: Super- fast convergence for radiance fields reconstruction,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 5459–5469
2022
-
[15]
Plenoxels: Radiance fields without neural networks,
S. Fridovich-Keil, A. Yu, M. Tancik, Q. Chen, B. Recht, and A. Kanazawa, “Plenoxels: Radiance fields without neural networks,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 5501–5510
2022
-
[16]
Baking neural radiance fields for real-time view synthesis,
P. Hedman, P. P. Srinivasan, B. Mildenhall, J. T. Barron, and P. Debevec, “Baking neural radiance fields for real-time view synthesis,”ICCV, 2021
2021
-
[17]
PlenOctrees for real-time rendering of neural radiance fields,
A. Yu, R. Li, M. Tancik, H. Li, R. Ng, and A. Kanazawa, “PlenOctrees for real-time rendering of neural radiance fields,” in ICCV, 2021
2021
-
[18]
Learning neural volu- metric representations of dynamic humans in minutes,
C. Geng, S. Peng, Z. Xu, H. Bao, and X. Zhou, “Learning neural volu- metric representations of dynamic humans in minutes,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 8759–8770
2023
-
[19]
Instantavatar: Learning avatars from monocular video in 60 seconds,
T. Jiang, X. Chen, J. Song, and O. Hilliges, “Instantavatar: Learning avatars from monocular video in 60 seconds,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 16 922–16 932
2023
-
[20]
3d gaussian splatting for real-time radiance field rendering,
B. Kerbl, G. Kopanas, T. Leimk ¨uhler, and G. Drettakis, “3d gaussian splatting for real-time radiance field rendering,” ACM Transactions on Graphics (ToG), vol. 42, no. 4, pp. 1–14, 2023
2023
-
[21]
3dgs-avatar: Animatable avatars via deformable 3d gaussian splatting,
Z. Qian, S. Wang, M. Mihajlovic, A. Geiger, and S. Tang, “3dgs-avatar: Animatable avatars via deformable 3d gaussian splatting,” arXiv preprint arXiv:2312.09228, 2023
2023 arXiv
-
[22]
Gauhuman: Articulated gaussian splatting from monocular human videos,
S. Hu, T. Hu, and Z. Liu, “Gauhuman: Articulated gaussian splatting from monocular human videos,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 20 418–20 431
2024
-
[23]
Drivable 3d gaussian avatars,
W. Zielonka, T. Bagautdinov, S. Saito, M. Zollh ¨ofer, J. Thies, and J. Romero, “Drivable 3d gaussian avatars,” 2023
2023
-
[24]
Hugs: Human gaussian splats,
M. Kocabas, R. Chang, J. Gabriel, O. Tuzel, and A. Ranjan, “Hugs: Human gaussian splats,” 2023. [Online]. Available: https: //arxiv.org/abs/2311.17910
2023 arXiv
-
[25]
Gart: Gaussian articulated template models,
J. Lei, Y . Wang, G. Pavlakos, L. Liu, and K. Daniilidis, “Gart: Gaussian articulated template models,” in Proceedings of the IEEE/CVF Confer- ence on Computer Vision and Pattern Recognition , 2024, pp. 19 876– 19 887
2024
-
[26]
Gps-gaussian: Generalizable pixel-wise 3d gaussian splatting for real- time human novel view synthesis,
S. Zheng, B. Zhou, R. Shao, B. Liu, S. Zhang, L. Nie, and Y . Liu, “Gps-gaussian: Generalizable pixel-wise 3d gaussian splatting for real- time human novel view synthesis,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 19 680–19 690
2024
-
[27]
Animatable gaussians: Learning pose-dependent gaussian maps for high-fidelity human avatar modeling,
Z. Li, Z. Zheng, L. Wang, and Y . Liu, “Animatable gaussians: Learning pose-dependent gaussian maps for high-fidelity human avatar modeling,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 19 711–19 722
2024
-
[28]
Sugar: Surface-aligned gaussian splatting for efficient 3d mesh reconstruction and high-quality mesh rendering,
A. Gu ´edon and V . Lepetit, “Sugar: Surface-aligned gaussian splatting for efficient 3d mesh reconstruction and high-quality mesh rendering,” arXiv preprint arXiv:2311.12775, 2023
2023 arXiv
-
[29]
Scaffold- gs: Structured 3d gaussians for view-adaptive rendering,
T. Lu, M. Yu, L. Xu, Y . Xiangli, L. Wang, D. Lin, and B. Dai, “Scaffold- gs: Structured 3d gaussians for view-adaptive rendering,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogni- tion (CVPR), June 2024, pp. 20 654–20 664
2024
-
[30]
Gaussianpro: 3d gaussian splatting with progressive propa- gation,
K. Cheng, X. Long, K. Yang, Y . Yao, W. Yin, Y . Ma, W. Wang, and X. Chen, “Gaussianpro: 3d gaussian splatting with progressive propa- gation,” in Forty-first International Conference on Machine Learning , 2024
2024
-
[31]
Deepcap: Monocular human performance capture using weak supervi- sion,
M. Habermann, W. Xu, M. Zollhofer, G. Pons-Moll, and C. Theobalt, “Deepcap: Monocular human performance capture using weak supervi- sion,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020, pp. 5052–5063
2020
-
[32]
Scanimate: Weakly supervised learning of skinned clothed avatar networks,
S. Saito, J. Yang, Q. Ma, and M. J. Black, “Scanimate: Weakly supervised learning of skinned clothed avatar networks,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2021, pp. 2886–2897
2021
-
[33]
Real-time deep dynamic characters,
M. Habermann, L. Liu, W. Xu, M. Zollhoefer, G. Pons-Moll, and C. Theobalt, “Real-time deep dynamic characters,” ACM Transactions on Graphics (ToG), vol. 40, no. 4, pp. 1–16, 2021
2021
-
[34]
Arch: Animatable reconstruction of clothed humans,
Z. Huang, Y . Xu, C. Lassner, H. Li, and T. Tung, “Arch: Animatable reconstruction of clothed humans,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , June 2020
2020
-
[35]
Gomavatar: Efficient animatable human modeling from monocular video using gaussians-on-mesh,
J. Wen, X. Zhao, Z. Ren, A. G. Schwing, and S. Wang, “Gomavatar: Efficient animatable human modeling from monocular video using gaussians-on-mesh,” arXiv preprint arXiv:2404.07991, 2024
2024 arXiv
-
[36]
Smpl: A skinned multi-person linear model,
M. Loper, N. Mahmood, J. Romero, G. Pons-Moll, and M. J. Black, “Smpl: A skinned multi-person linear model,” ACM transactions on graphics (TOG), vol. 34, no. 6, pp. 1–16, 2015
2015
-
[37]
Animatable neural radiance fields for modeling dynamic human bodies,
S. Peng, J. Dong, Q. Wang, S. Zhang, Q. Shuai, X. Zhou, and H. Bao, “Animatable neural radiance fields for modeling dynamic human bodies,” in IEEE/CVF International Conference on Computer Vision , 2021, pp. 14 314–14 323
2021
-
[38]
Neural articulated radiance field,
A. Noguchi, X. Sun, S. Lin, and T. Harada, “Neural articulated radiance field,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021, pp. 5762–5772
2021
-
[39]
A-nerf: Articulated neural radiance fields for learning human shape, appearance, and pose,
S.-Y . Su, F. Yu, M. Zollh ¨ofer, and H. Rhodin, “A-nerf: Articulated neural radiance fields for learning human shape, appearance, and pose,” Advances in Neural Information Processing Systems, vol. 34, pp. 12 278– 12 291, 2021
2021
-
[40]
Surface-aligned neural radiance fields for controllable 3d human synthesis,
T. Xu, Y . Fujita, and E. Matsumoto, “Surface-aligned neural radiance fields for controllable 3d human synthesis,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 15 883–15 892
2022
-
[41]
Neural human performer: Learning generalizable radiance fields for human performance render- ing,
Y . Kwon, D. Kim, D. Ceylan, and H. Fuchs, “Neural human performer: Learning generalizable radiance fields for human performance render- ing,” Advances in Neural Information Processing Systems, vol. 34, 2021
2021
-
[42]
Mps-nerf: Gener- alizable 3d human rendering from multiview images,
X. Gao, J. Yang, J. Kim, S. Peng, Z. Liu, and X. Tong, “Mps-nerf: Gener- alizable 3d human rendering from multiview images,”IEEE Transactions on Pattern Analysis and Machine Intelligence, pp. 1–12, 2022
2022
-
[43]
Doublefield: Bridging the neural surface and radiance fields for high- fidelity human reconstruction and rendering,
R. Shao, H. Zhang, H. Zhang, M. Chen, Y .-P. Cao, T. Yu, and Y . Liu, “Doublefield: Bridging the neural surface and radiance fields for high- fidelity human reconstruction and rendering,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 20...
2022
-
[44]
Deliffas: Deformable light fields for fast avatar synthesis,
Y . Kwon, L. Liu, H. Fuchs, M. Habermann, and C. Theobalt, “Deliffas: Deformable light fields for fast avatar synthesis,” Advances in Neural Information Processing Systems, vol. 36, 2024
2024
-
[45]
Drivable volumetric avatars using texel-aligned features,
E. Remelli, T. Bagautdinov, S. Saito, C. Wu, T. Simon, S.-E. Wei, K. Guo, Z. Cao, F. Prada, J. Saragih et al., “Drivable volumetric avatars using texel-aligned features,” in ACM SIGGRAPH 2022 Conference Proceedings, 2022, pp. 1–9
2022
-
[46]
Avatarrex: Real-time expressive full-body avatars,
Z. Zheng, X. Zhao, H. Zhang, B. Liu, and Y . Liu, “Avatarrex: Real-time expressive full-body avatars,” ACM Transactions on Graphics (TOG) , vol. 42, no. 4, pp. 1–19, 2023
2023
-
[47]
Dynamic 3d gaus- sians: Tracking by persistent dynamic view synthesis,
J. Luiten, G. Kopanas, B. Leibe, and D. Ramanan, “Dynamic 3d gaus- sians: Tracking by persistent dynamic view synthesis,” arXiv preprint arXiv:2308.09713, 2023
2023 arXiv
-
[48]
4d gaussian splatting: Towards efficient novel view synthesis for dynamic scenes,
Y . Duan, F. Wei, Q. Dai, Y . He, W. Chen, and B. Chen, “4d gaussian splatting: Towards efficient novel view synthesis for dynamic scenes,” arXiv preprint arXiv:2402.03307, 2024
2024 arXiv
-
[49]
Neural parametric gaussians for monocular non-rigid object reconstruction,
D. Das, C. Wewer, R. Yunus, E. Ilg, and J. E. Lenssen, “Neural parametric gaussians for monocular non-rigid object reconstruction,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 10 715–10 725
2024
-
[50]
Deformable 3d gaussians for high-fidelity monocular dynamic scene reconstruction,
Z. Yang, X. Gao, W. Zhou, S. Jiao, Y . Zhang, and X. Jin, “Deformable 3d gaussians for high-fidelity monocular dynamic scene reconstruction,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 20 331–20 341
2024
-
[51]
Hifi4g: High-fidelity human performance rendering via compact gaussian splatting,
Y . Jiang, Z. Shen, P. Wang, Z. Su, Y . Hong, Y . Zhang, J. Yu, and L. Xu, “Hifi4g: High-fidelity human performance rendering via compact gaussian splatting,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 19 734–19 745
2024
-
[52]
Gaussianavatar: Towards realistic human avatar modeling from a single video via animatable 3d gaussians,
L. Hu, H. Zhang, Y . Zhang, B. Zhou, B. Liu, S. Zhang, and L. Nie, “Gaussianavatar: Towards realistic human avatar modeling from a single video via animatable 3d gaussians,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 634– 644
2024
-
[53]
Splatarmor: Articulated gaussian splatting for animatable humans from monocular rgb videos,
R. Jena, G. S. Iyer, S. Choudhary, B. Smith, P. Chaudhari, and J. Gee, “Splatarmor: Articulated gaussian splatting for animatable humans from monocular rgb videos,” arXiv preprint arXiv:2311.10812, 2023
2023 arXiv
-
[54]
Uv gaussians: Joint learning of mesh deformation and gaussian textures for human avatar modeling,
Y . Jiang, Q. Liao, X. Li, L. Ma, Q. Zhang, C. Zhang, Z. Lu, and Y . Shan, “Uv gaussians: Joint learning of mesh deformation and gaussian textures for human avatar modeling,” arXiv preprint arXiv:2403.11589, 2024
2024 arXiv
-
[55]
Ash: Animatable gaussian splats for efficient and photoreal human rendering,
H. Pang, H. Zhu, A. Kortylewski, C. Theobalt, and M. Habermann, “Ash: Animatable gaussian splats for efficient and photoreal human rendering,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 1165–1175
2024
-
[56]
Image-to-image translation with conditional adversarial networks,
P. Isola, J.-Y . Zhu, T. Zhou, and A. A. Efros, “Image-to-image translation with conditional adversarial networks,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2017, pp. 1125– 1134
2017
-
[57]
Dynamic point fields,
S. Prokudin, Q. Ma, M. Raafat, J. Valentin, and S. Tang, “Dynamic point fields,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Oct. 2023, pp. 7964–7976
2023
-
[58]
Splattingavatar: Realistic real-time human avatars with mesh- embedded gaussian splatting,
Z. Shao, Z. Wang, Z. Li, D. Wang, X. Lin, Y . Zhang, M. Fan, and Z. Wang, “Splattingavatar: Realistic real-time human avatars with mesh- embedded gaussian splatting,” in Proceedings of the IEEE/CVF Confer- ence on Computer Vision and Pattern Recognition, 2024, pp. 1606–1616
2024
-
[59]
Haha: Highly articulated gaussian human avatars with textured mesh prior,
D. Svitov, P. Morerio, L. Agapito, and A. Del Bue, “Haha: Highly articulated gaussian human avatars with textured mesh prior,” arXiv preprint arXiv:2404.01053, 2024
2024 arXiv
-
[60]
The phong surface: Efficient 3d model fitting using lifted optimization,
J. Shen, T. J. Cashman, Q. Ye, T. Hutton, T. Sharp, F. Bogo, A. Fitzgib- bon, and J. Shotton, “The phong surface: Efficient 3d model fitting using lifted optimization,” in Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part I...
2020
-
[61]
User-specific hand modeling from monocular depth sequences,
J. Taylor, R. Stebbing, V . Ramakrishna, C. Keskin, J. Shotton, S. Izadi, A. Hertzmann, and A. Fitzgibbon, “User-specific hand modeling from monocular depth sequences,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2014, pp. 644–651
2014
-
[62]
Efficient and precise interactive hand tracking through joint, continuous optimization of pose and correspondences,
J. Taylor, L. Bordeaux, T. Cashman, B. Corish, C. Keskin, T. Sharp, E. Soto, D. Sweeney, J. Valentin, B. Luff et al. , “Efficient and precise interactive hand tracking through joint, continuous optimization of pose and correspondences,” ACM Transactions on Graphics (ToG) , vol...
2016
-
[63]
Ewa splatting,
M. Zwicker, H. Pfister, J. Van Baar, and M. Gross, “Ewa splatting,” IEEE Transactions on Visualization and Computer Graphics, vol. 8, no. 3, pp. 223–238, 2002
2002
-
[64]
Mesh-based gaussian splatting for real-time large-scale deformation,
L. Gao, J. Yang, B.-T. Zhang, J.-M. Sun, Y .-J. Yuan, H. Fu, and Y .-K. Lai, “Mesh-based gaussian splatting for real-time large-scale deformation,” arXiv preprint arXiv:2402.04796, 2024
2024 arXiv
-
[65]
Deformable 3d gaussians for high-fidelity monocular dynamic scene reconstruction,
Z. Yang, X. Gao, W. Zhou, S. Jiao, Y . Zhang, and X. Jin, “Deformable 3d gaussians for high-fidelity monocular dynamic scene reconstruction,” arXiv preprint arXiv:2309.13101, 2023
2023 arXiv
-
[66]
Leap: Learning articulated occupancy of people,
M. Mihajlovic, Y . Zhang, M. J. Black, and S. Tang, “Leap: Learning articulated occupancy of people,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2021, pp. 10 461–10 471
2021
-
[67]
The unreasonable effectiveness of deep features as a perceptual metric,
R. Zhang, P. Isola, A. A. Efros, E. Shechtman, and O. Wang, “The unreasonable effectiveness of deep features as a perceptual metric,” in CVPR, 2018
2018
-
[68]
Adam: A method for stochastic optimization,
D. P. Kingma, “Adam: A method for stochastic optimization,” arXiv preprint arXiv:1412.6980, 2014
2014 arXiv
-
[69]
Video based reconstruction of 3d people models,
T. Alldieck, M. Magnor, W. Xu, C. Theobalt, and G. Pons-Moll, “Video based reconstruction of 3d people models,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Jun 2018, pp. 8387– 8397, CVPR Spotlight Paper
2018
-
[70]
Animatable neural radiance fields from monocular rgb videos,
J. Chen, Y . Zhang, D. Kang, X. Zhe, L. Bao, X. Jia, and H. Lu, “Animatable neural radiance fields from monocular rgb videos,” arXiv preprint arXiv:2106.13629, 2021
2021 arXiv
-
[71]
Ai choreographer: Music conditioned 3d dance generation with aist++,
R. Li, S. Yang, D. A. Ross, and A. Kanazawa, “Ai choreographer: Music conditioned 3d dance generation with aist++,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2021, pp. 13 401–13 412
2021
-
[72]
Amass: Archive of motion capture as surface shapes,
N. Mahmood, N. Ghorbani, N. F. Troje, G. Pons-Moll, and M. J. Black, “Amass: Archive of motion capture as surface shapes,” in Proceedings of the IEEE/CVF international conference on computer vision, 2019, pp. 5442–5451
2019
-
[73]
A volumetric method for building complex models from range images,
B. Curless and M. Levoy, “A volumetric method for building complex models from range images,” in Proceedings of the 23rd annual confer- ence on Computer graphics and interactive techniques , 1996, pp. 303– 312. 8 B IOGRAPHY SECTION Ronghan Chen Ronghan Chen is currently a Ph.D...
1996
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.