Pith. sign in

REVIEW 4 major objections 6 minor 1 cited by

Sequential Gaussian Avatars with Hierarchical Motion Context

T0 review · 4 major / 6 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read Conditioning Gaussian avatar deformation on both skeleton residuals and per-vertex velocities captures loose-clothing motion at real-time speed.

desk verdict Solid incremental 3DGS avatar work; the vertex-velocity condition is plausible, but the paper never shows it beats a same-capacity pose-only control. read the letter →

arxiv 2411.16768 v2 pith:5H5BWLQZ submitted 2024-11-25 cs.CV

classification cs.CV
keywords 3DGaussianSplattinganimatablehumanavatarnon-rigiddeformationhierarchicalmotioncontextSMPLtemplatevelocitiesmulti-scaletemporalsamplingnovelviewsynthesisposesequenceconditioning
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that pose-only conditioning is the bottleneck for animatable Gaussian avatars: a single frame's pose cannot represent the many appearances the same pose can take during complex motion, especially in loose clothing. SeqAvatar adds a hierarchical motion context with two levels: coarse skeleton motion, obtained from differences of SMPL poses across frames, and fine-grained point-wise motion, obtained from finite-difference velocities of SMPL template vertices warped by linear blend skinning. Both conditions are sampled at several temporal scales and over neighbouring template vertices, then fed to the MLP that predicts each Gaussian's non-rigid deformation. The claim is that this restores real-time rendering while matching or surpassing slower NeRF-based temporal models and beating pose-only 3DGS baselines on DNA-Rendering, I3D-Human, and ZJU-MoCap.

What carries the argument

The load-bearing object is the vertex motion template field $F_{\mathcal{V}}$: a per-SMPL-vertex velocity array computed by warping the template to observation space with linear blend skinning and taking finite differences $V_t = (T^o_t - T^o_{t-s})/s$. Each canonical Gaussian looks up the $ au$ nearest template vertices' velocities, encodes them with an MLP, and concatenates the result with the coarse skeleton motion embedding obtained from pose residuals $\Delta P_t = \delta(P_t, P_{t-s})$. Spatio-temporal multi-scale sampling varies the interval $s$ across a set $S = \{s_0, s_0+\Delta s, \dots\}$, giving the deformation MLP both global body movement and local region motion over several temporal windows. This is what lets one network resolve the same pose at different moments into different non-rigid deformations.

What would settle it

Train the same pipeline twice on a video of a garment whose motion is dominated by inertia or external forces, such as a flaring skirt during a spin or a coat blowing in wind: once with SMPL-derived vertex velocities and once with vertex velocities taken from dense surface tracking of the real garment. If the second does not outperform the first, or if the first shows no gain over a pose-only baseline, the claim that fine vertex motion is the load-bearing improvement fails.

Watch

Extended reading notes

Core claim

The central claim is that an explicit Gaussian representation can carry motion conditions at two granularities, and that the fine granularity is what recovers appearance details far from the skeleton. The coarse condition is a sequence of per-joint rotation differences between adjacent frames, encoded into a 32-dimensional embedding. The fine condition is a motion template field that stores, for every SMPL vertex, its finite-difference velocity under linear blend skinning; each Gaussian samples its $ au$ nearest template vertices' velocities and encodes them into a 96-dimensional embedding. A spatio-temporal multi-scale sampling strategy builds these embeddings from several time intervals so that long-term motion trends and inter-frame details enter together. The non-rigid MLP then predicts position, scale, and rotation offsets for each Gaussian from its position, pose, and the two motion embeddings, and the deformed Gaussians are warped by LBS and splatted at real-time rates.

Load-bearing premise

The method's fine-grained motion signal comes from the coarse SMPL body model, not from tracking the actual clothing, so everything rests on the assumption that SMPL vertex motion faithfully represents how loose garments and local surfaces really move.

Editorial extensions

If this is right

  • Adding the coarse skeleton-motion condition alone improves over pose-only non-rigid deformation, and adding the fine vertex-motion condition improves further, which supports the hierarchy as the cause of the gains.
  • Multi-scale sampling over several time intervals gives better robustness than any single interval, so temporal context should be gathered at multiple scales rather than at one fixed step.
  • The method keeps real-time rendering while matching or beating the slower NeRF-style temporal baseline, so temporal conditioning does not have to sacrifice interactivity.
  • Novel-pose and out-of-distribution animations trained on one sequence render plausible results on poses from unseen sequences, indicating that the motion conditions generalize rather than merely memorize training poses.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Replacing the SMPL-derived velocity proxy with dense surface tracking or physics-based simulation is the clear next experiment; if it further sharpens garments, the coarse template is the bottleneck rather than the conditioning principle.
  • The same velocity-conditioned deformation could apply to any explicit representation with a template and skinning weights, including animals or clothed characters whose pose-to-appearance map is similarly ambiguous.
  • One could drive avatars from video by feeding observed vertex velocities directly into the deformation MLP, allowing re-enactment and animation to share the same machinery without first regressing pose.
  • Because the vertex motion template is a cheap pre-computation, the conditioning could be extended to higher-resolution templates or denser velocity fields without affecting the real-time inference cost.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. This paper proposes SeqAvatar, a 3D Gaussian Splatting (3DGS) framework for animatable human avatars from multi-view video. The method augments the standard pose-conditioned non-rigid deformation network with two additional motion conditions: a coarse skeleton-motion embedding ΔP computed as differences of SMPL joint rotations over a temporal window (Eqs. 7-8), and a fine vertex-motion embedding f_V obtained by finite differences of LBS-warped SMPL template vertex positions, aggregated to each Gaussian via KNN (Eqs. 10-13). A spatio-temporal multi-scale sampling strategy (STMS, Eq. 14) varies the time interval to combine long-term trends with inter-frame details. Experiments on DNA-Rendering, I3D-Human, and ZJU-MoCap report improved PSNR/SSIM/LPIPS over 3DGS-Avatar, GART, GauHuman, and the NeRF-based Dyco, at roughly 62 FPS versus Dyco's 0.7 FPS.

Significance. The work has practical value: it is a simple and modular motion-conditioning scheme that demonstrably improves 3DGS avatar rendering quality on loose-clothing, complex-motion benchmarks (Tables 1-2) while preserving real-time rendering (Table 8). Credit is due for the multi-dataset evaluation, the component-wise ablations (Tables 4, 6, 7), the efficiency analysis, and the candid limitation statement in Section 6. The conceptual novelty is modest: the fine motion condition is a spatially distributed re-encoding of pose history rather than an observation of true surface motion, and the incremental gains over pose-only conditioning are small (on the order of 0.1-0.3 dB). The significance of the central claim therefore rests on whether the reported gains are attributable to the motion semantics of the new conditions rather than to added model capacity; the control experiments requested below will determine this.

major comments (4)
  1. [§4.1, Eqs. (10)-(13); Table 4] The stress-test concern about the fine vertex-motion condition lands. V is a deterministic function of the SMPL pose history, the fixed template T, and the fixed skinning weights W: V_t is computed by LBS-warping the template (Eqs. 10-11), and the KNN sampling in Eq. (12) uses canonical Gaussian positions only, so V contains no measured garment or surface deformation beyond what the pose history implies. The observed gain from adding V (Table 4, (c)->(d): PSNR 31.89->32.01, LPIPS 32.17->31.23) is small and could be produced by the added parameters of E_knn and E_V rather than by the semantic content of the velocities, a possibility the paper itself leaves open in Section 6 ('our local velocity cues are derived from the coarse SMPL model rather than dense surface tracking'). Because the abstract credits 'fine-grained vertex motions' as the basis for superiority over pose-only 3DGS avatars, the authors should add a control that keeps capacity but destroys the spatial semantics of V, e.g., randomly permuting vertex velocities across the template before KNN sampling, or replacing f_V with a per-Gaussian learned latent of the same dimension (R^96). If the improvement in row (d) persists under either control, the proposed fine-motion mechanism is not supported.
  2. [Tables 1-5] All quantitative results are single-run and no variance estimates are reported. This matters because several differences that carry the paper's claims are small: adding V in Table 4 changes PSNR by +0.12 dB, and on ZJU-MoCap (Table 5) SeqAvatar trails GauHuman on PSNR (31.02 vs 31.04) and SSIM (0.9619 vs 0.9620). Without standard deviations over multiple seeds, these margins cannot be distinguished from run-to-run noise, and the ablative support for the central contribution is correspondingly weak. Please report mean and standard deviation over at least three runs for the headline comparisons and for the ablation rows in Table 4.
  3. [§5.2] The protocol for extending the monocular baselines 3DGS-Avatar, GART, and GauHuman to multi-view input is not described; the text only states that they are extended 'under the same settings.' The fairness of the comparison depends on details such as training views, iterations, adaptive densification settings, loss weights, and whether pose refinement is enabled. Without this information, the consistent gains in Tables 1-3 cannot be fully attributed to the proposed method rather than to differences in baseline tuning. Please document the adaptation protocol in the supplement.
  4. [Abstract; Table 5] The abstract's unconditional statement that SeqAvatar 'significantly outperforms 3DGS-based approaches' is stronger than the ZJU-MoCap evidence supports: in Table 5 the method does not exceed GauHuman on PSNR (31.02 vs 31.04) or SSIM (0.9619 vs 0.9620), and is better only on LPIPS (28.89 vs 31.81). The advantage appears to be dataset-dependent, being pronounced on the loose-clothing, complex-motion DNA-Rendering and I3D-Human sets but at parity on the controlled ZJU-MoCap set. Please calibrate the claim to this scope, e.g., by stating that the method significantly outperforms 3DGS baselines on complex-motion datasets and is comparable on controlled settings.
minor comments (6)
  1. [Eq. (25); §11] Loss weights are denoted λ1, λ2, λ3 in Eq. (25) but reported as λ0, λ1, λ2 in Section 11; please unify the notation and state which weight corresponds to which loss term for each dataset.
  2. [Figs. 1-2] The figures use a calligraphic symbol for vertex velocity while the text uses V and F_V, and the delta notation in the caption of Fig. 1 is not defined; please harmonize the notation between figures and equations.
  3. [§5.2, Tables 1-3] The baseline set is inconsistent across tables: GART appears only in Table 1, and Dyco appears only on I3D-Human and ZJU-MoCap; please state whether the same baselines were evaluated on all datasets and give the reason if not, e.g., availability of pretrained models or data splits.
  4. [§4.2, Eq. (12)] The KNN neighborhoods are computed from canonical Gaussian positions x_i, which are optimized and may drift from the SMPL template vertices; please state whether the neighborhoods are recomputed during optimization or fixed once at initialization, since this affects the meaning of the 'local region' being sampled.
  5. [§12.1 (Supp.)] There is a typo in the supplementary material: '3DGS-Avaar' should be '3DGS-Avatar'.
  6. [Table 8] The caption reads 'Mehods' and the table does not state the dataset, resolution, and hardware details beyond 'a single 4090 GPU'; please specify these details so the 62 FPS figure is reproducible.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the motion-condition inputs are frozen precomputed template kinematics, not fitted outputs, and the reported gains are measured against external baselines.

full rationale

SeqAvatar is an empirical fitting pipeline, not a derivation, so most circularity patterns do not apply. The coarse skeleton motion fΔP (Eqs. 7-8) and the fine vertex motion fV (Eqs. 10-13) are both input conditions computed from the frozen SMPL template and pose history: Vt=(LBS(T,Bt,W)-LBS(T,Bt-s,W))/s. This makes V a deterministic spatial re-encoding of the same pose sequence, and one could question whether it carries information beyond ΔP; the paper itself concedes in Sec. 6 that "our local velocity cues are derived from the coarse SMPL model rather than dense surface tracking." However, that is a feature-information concern, not circularity: V is not fitted to the target images, is not renamed as a prediction, and the non-rigid MLP output (δx, δs, δr) is not equivalent to V by construction. The paper also explicitly avoids a circular velocity design by rejecting Gaussian-position velocities in favor of the frozen SMPL template field. The reported improvements in Tables 1-4 are empirical comparisons against external baselines (Dyco, 3DGS-Avatar, GART, GauHuman) under shared data splits. The only author-overlap issue is that Dyco is a same-group baseline, but it is used as a comparison method rather than as a load-bearing justification, so it does not raise the circularity score. No circular step can be exhibited from the paper's equations or citations.

Assumptions & free parameters 4 free parameters · 4 assumptions · 0 invented entities

The method's central contribution is a conditioning scheme built on SMPL-derived motion signals. The key assumptions are that SMPL provides accurate geometry and that vertex velocities are a valid proxy for true surface motion; both are standard in the field but neither is independently verified in this paper. The main free parameters are sampling schedules and KNN counts chosen via ablations on the same benchmarks used for the final results.

free parameters (4)
  • Sampling scale set S = {24,33,42} (I3D-Human); {1,3,5} (DNA-Rendering)
    Chosen by hand; ablation in Table 6 shows S={24,33,42} gives the best average PSNR on I3D-Human, so the final configuration is tuned on the same benchmark.
  • Sequence length L per scale = 8
    Number of sampled frames per scale; fixed without an ablation in the main text.
  • KNN vertex count tau = 8
    Ablation in Table 7 shows tau=8 is best on I3D-Human; chosen after inspection of the same benchmark.
  • Loss weights lambda0, lambda1, lambda2 = 1.0, 0.01/0.1, 0.01/0.1 depending on dataset
    Standard weights set differently per dataset; not justified by independent validation.
assumptions (4)
  • domain assumption SMPL(-X) template and skinning weights accurately represent the human body and its deformation.
    All Gaussian positions and vertex velocities are initialized from SMPL and warped with LBS; if SMPL registration is wrong, both coarse and fine motion conditions are wrong. Invoked in Section 4.1 and Eq. 10.
  • standard math LBS is an adequate rigid-motion model for the body, with non-rigid residuals learned by an MLP.
    Section 3 Eq. 1 and Section 4.3 Eq. 21; standard in the field but an assumption about body deformation.
  • domain assumption The finite-difference velocity of SMPL template vertices (Eq. 11) is a suitable proxy for true surface motion of garments and local regions.
    This is the load-bearing premise for the fine vertex motion condition; the paper itself flags it as a limitation in Section 6.
  • domain assumption Per-frame SMPL pose estimates are accurate enough, after refinement, to derive meaningful motion context.
    Pose refinement MLP is introduced in Section 4.3, but residual pose errors propagate to both Delta P and V.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Sequential Gaussian Avatars with Hierarchical Motion Context." pith.science (2026). https://pith.science/paper/5H5BWLQZ

@misc{pith2026241116768,
  author       = {Pith},
  title        = {Pith review of: Sequential Gaussian Avatars with Hierarchical Motion Context},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/5H5BWLQZ}},
  note         = {Machine review of arXiv:2411.16768}
}
read the original abstract

The emergence of neural rendering has significantly advanced the rendering quality of 3D human avatars, with the recently popular 3DGS technique enabling real-time performance. However, SMPL-driven 3DGS human avatars still struggle to capture fine appearance details due to the complex mapping from pose to appearance during fitting. In this paper, we propose SeqAvatar, which excavates the explicit 3DGS representation to better model human avatars based on a hierarchical motion context. Specifically, we utilize a coarse-to-fine motion conditions that incorporate both the overall human skeleton and fine-grained vertex motions for non-rigid deformation. To enhance the robustness of the proposed motion conditions, we adopt a spatio-temporal multi-scale sampling strategy to hierarchically integrate more motion clues to model human avatars. Extensive experiments demonstrate that our method significantly outperforms 3DGS-based approaches and renders human avatars orders of magnitude faster than the latest NeRF-based models that incorporate temporal context, all while delivering performance that is at least comparable or even superior. Project page: https://zezeaaa.github.io/projects/SeqAvatar/

Figures

Figures reproduced from arXiv: 2411.16768 by the authors.

Figure 1
Figure 1. Illustration of the Hierarchical Motion Context Design. We model the motion-dependent appearance variations with both coarse skeleton condition ∆P and fine-grained point-wise velocity condition V. ∆P describes the overall human skeleton motion, which is derived from the difference between the human poses at adjacent frames. V is the point-wise velocity that indicates finer-grained motion in local regions. To achieve… view at source ↗
Figure 2
Figure 2. Overview of the proposed method. We first initialize canonical Gaussian positions x with SMPL template vertexes. For each Gaussian, we derive both coarse skeleton motion condition f∆P and fine-grained vertex motion condition fV that sampled from the vertex motion template FV (points in different colors represent different motions). Based on such hierarchical motion information, we utilize an MLP Enon−rigid to better… view at source ↗
Figure 3
Figure 3. Novel View Qualitative Results on DNA-Rendering. We zoom into the local region and compute the error maps compared with ground truth images. The results show that our method achieves competitive results on both the overall and local region qualities. method outperforms previous state-of-the-art human mod￾eling methods across all metrics. To demonstrate the visual improvement, we compare the quality of rendering imag… view at source ↗
Figures from the paper (6 more)
Figure 4
Figure 4. Figure 4: Novel Pose Qualitative Results on I3D-Human. We compare the proposed method with previous SOTA approaches. The rendered images and error maps demonstrate our robustness on novel pose rendering. mations, allowing our method to capture detailed appear￾ance variations cau…
Figure 5
Figure 5. Figure 5: Qualitative Ablation Results. We compare the visual influence of different components on ID2 1 scene of I3D-Human dataset. Ours Dyco Target Pose ID3_1 ID1_2 ID2_1 ID1_1 Ours GauHuman Target Pose 2_0019_10 2_0051_09 I3D-Human DNA-Rendering 2_0813_05 1_0206_04 [PITH_FUL…
Figure 6
Figure 6. Figure 6: Out-of-distribution Pose Animation. drawn from an unseen sequence exhibiting different human motions. The results suggest that our method is capable of generalizing to these significant pose changes, even though such poses were not observed during training [PITH_FULL_…
Figure 7
Figure 7. Figure 7: Novel View Qualitative Results on DNA-Rendering. We provide both the complete image and localized error map comparisons against other methods for a more comprehensive visualization. 11. Details of Optimization In our experiments, the number |S| (in Eq. (14)) of se￾quen…
Figure 8
Figure 8. Figure 8: Novel View Qualitative Results on I3D-Human [PITH_FULL_IMAGE:figures/full_fig_p015_8.png]
Figure 9
Figure 9. Figure 9: Qualitative Results on ZJU-MoCap. 12.2. Visual Comparisons On DNA-Rendering [PITH_FULL_IMAGE:figures/full_fig_p015_9.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Proxy-GS: Unified Occlusion Priors for Training and Inference in Structured 3D Gaussian Splatting

    cs.CV 2025-09 conditional novelty 5.0 of 10

    A proxy mesh rendered through hardware rasterization provides a cheap occlusion depth prior that culls hidden anchors at inference and guides densification at training, giving Octree-GS-like MLP splatting a 3 to 4x sp...

Reference graph

Works this paper leans on

90 extracted references · 49 canonical work pages · cited by 1 Pith paper

  1. [1]

    Mip-NeRF: A Multiscale Representation for Anti-aliasing Neural Radiance Fields

    Jonathan T Barron, Ben Mildenhall, Matthew Tancik, Peter Hedman, Ricardo Martin-Brualla, and Pratul P Srinivasan. Mip-NeRF: A Multiscale Representation for Anti-aliasing Neural Radiance Fields. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 5855– 5864, 2021. 2

  2. [2]

    Mip-nerf 360: Unbounded anti-aliased neural radiance fields

    Jonathan T Barron, Ben Mildenhall, Dor Verbin, Pratul P Srinivasan, and Peter Hedman. Mip-nerf 360: Unbounded anti-aliased neural radiance fields. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5470–5479, 2022. 2

  3. [3]

    Hexplane: A Fast Representa- tion for Dynamic Scenes

    Ang Cao and Justin Johnson. Hexplane: A Fast Representa- tion for Dynamic Scenes. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 130–141, 2023. 2

  4. [4]

    Tensorf: Tensorial radiance fields

    Anpei Chen, Zexiang Xu, Andreas Geiger, Jingyi Yu, and Hao Su. Tensorf: Tensorial radiance fields. In European Conference on Computer Vision , pages 333–350. Springer,

  5. [5]

    Animatable neural radiance fields from monocular rgb videos

    Jianchuan Chen, Ying Zhang, Di Kang, Xuefei Zhe, Lin- chao Bao, Xu Jia, and Huchuan Lu. Animatable neural radiance fields from monocular rgb videos. arXiv preprint arXiv:2106.13629, 2021. 2

  6. [6]

    Snarf: Differentiable forward skinning for animating non-rigid neural implicit shapes

    Xu Chen, Yufeng Zheng, Michael J Black, Otmar Hilliges, and Andreas Geiger. Snarf: Differentiable forward skinning for animating non-rigid neural implicit shapes. In Proceed- ings of the IEEE/CVF International Conference on Com- puter Vision, pages 11594–11604, 2021

  7. [7]

    Fast-snarf: A fast deformer for articulated neural fields

    Xu Chen, Tianjian Jiang, Jie Song, Max Rietmann, Andreas Geiger, Michael J Black, and Otmar Hilliges. Fast-snarf: A fast deformer for articulated neural fields. IEEE Transac- tions on Pattern Analysis and Machine Intelligence, 45(10): 11796–11809, 2023

  8. [8]

    Within the Dynamic Context: Inertia-aware 3D Human Modeling with Pose Sequence

    Yutong Chen, Yifan Zhan, Zhihang Zhong, Wei Wang, Xiao Sun, Yu Qiao, and Yinqiang Zheng. Within the dynamic con- text: Inertia-aware 3d human modeling with pose sequence. arXiv preprint arXiv:2403.19160, 2024. 2, 3, 5, 7, 8, 1

Show all 90 references
  1. [9]

    Dna-rendering: A diverse neural actor repository for high-fidelity human-centric rendering

    Wei Cheng, Ruixiang Chen, Siming Fan, Wanqi Yin, Keyu Chen, Zhongang Cai, Jingbo Wang, Yang Gao, Zhengming Yu, Zhengyu Lin, et al. Dna-rendering: A diverse neural actor repository for high-fidelity human-centric rendering. In Proceedings of the IEEE/CVF International Conferenc...

  2. [10]

    Depth-supervised nerf: Fewer views and faster train- ing for free

    Kangle Deng, Andrew Liu, Jun-Yan Zhu, and Deva Ra- manan. Depth-supervised nerf: Fewer views and faster train- ing for free. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12882– 12891, 2022. 2

  3. [11]

    Fast and robust multi-person 3d pose estimation and tracking from multiple views

    Junting Dong, Qi Fang, Wen Jiang, Yurou Yang, Hujun Bao, and Xiaowei Zhou. Fast and robust multi-person 3d pose estimation and tracking from multiple views. In T-PAMI,

  4. [12]

    Plenoxels: Radiance fields without neural networks

    Sara Fridovich-Keil, Alex Yu, Matthew Tancik, Qinhong Chen, Benjamin Recht, and Angjoo Kanazawa. Plenoxels: Radiance fields without neural networks. In CVPR, pages 5501–5510, 2022. 2

  5. [13]

    K-planes: Explicit Radiance Fields in Space, Time, and Appearance

    Sara Fridovich-Keil, Giacomo Meanti, Frederik Rahbæk Warburg, Benjamin Recht, and Angjoo Kanazawa. K-planes: Explicit Radiance Fields in Space, Time, and Appearance. In Proceedings of the IEEE/CVF Conference on Computer Vi- sion and Pattern Recognition, pages 12479–12488, 2023. 2

  6. [14]

    Dynamic neural radiance fields for monocular 4d facial avatar reconstruction

    Guy Gafni, Justus Thies, Michael Zollhofer, and Matthias Nießner. Dynamic neural radiance fields for monocular 4d facial avatar reconstruction. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 8649–8658, 2021. 2

  7. [15]

    V4D: V oxel for 4D Novel View Synthesis

    Wanshui Gan, Hongbin Xu, Yi Huang, Shifeng Chen, and Naoto Yokoya. V4D: V oxel for 4D Novel View Synthesis. IEEE Transactions on Visualization and Computer Graphics,

  8. [16]

    Dynamic View Synthesis from Dynamic Monocular Video

    Chen Gao, Ayush Saraf, Johannes Kopf, and Jia-Bin Huang. Dynamic View Synthesis from Dynamic Monocular Video. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 5712–5721, 2021. 2

  9. [17]

    Neural novel actor: Learning a generalized animatable neural representation for human actors

    Qingzhe Gao, Yiming Wang, Libin Liu, Lingjie Liu, Chris- tian Theobalt, and Baoquan Chen. Neural novel actor: Learning a generalized animatable neural representation for human actors. IEEE Transactions on Visualization and Com- puter Graphics, 2023. 2

  10. [18]

    Fastnerf: High-fidelity neural rendering at 200fps

    Stephan J Garbin, Marek Kowalski, Matthew Johnson, Jamie Shotton, and Julien Valentin. Fastnerf: High-fidelity neural rendering at 200fps. In ICCV, pages 14346–14355, 2021. 2

  11. [19]

    Learning neural volumetric representations of dy- namic humans in minutes

    Chen Geng, Sida Peng, Zhen Xu, Hujun Bao, and Xiaowei Zhou. Learning neural volumetric representations of dy- namic humans in minutes. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 8759–8770, 2023. 2

  12. [20]

    Humans in 4d: Re- constructing and tracking humans with transformers

    Shubham Goel, Georgios Pavlakos, Jathushan Rajasegaran, Angjoo Kanazawa, and Jitendra Malik. Humans in 4d: Re- constructing and tracking humans with transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 14783–14794, 2023. 2

  13. [21]

    Gaussianavatar: Towards realistic human avatar modeling from a single video via animatable 3d gaussians

    Liangxiao Hu, Hongwen Zhang, Yuxiang Zhang, Boyao Zhou, Boning Liu, Shengping Zhang, and Liqiang Nie. Gaussianavatar: Towards realistic human avatar modeling from a single video via animatable 3d gaussians. In Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pat...

  14. [22]

    Gauhuman: Articu- lated gaussian splatting from monocular human videos

    Shoukang Hu, Tao Hu, and Ziwei Liu. Gauhuman: Articu- lated gaussian splatting from monocular human videos. In Proceedings of the IEEE/CVF Conference on Computer Vi- sion and Pattern Recognition, pages 20418–20431, 2024. 1, 2, 3, 5, 6, 7, 8

  15. [23]

    Surmo: Surface- based 4d motion modeling for dynamic human rendering

    Tao Hu, Fangzhou Hong, and Ziwei Liu. Surmo: Surface- based 4d motion modeling for dynamic human rendering. In Proceedings of the IEEE/CVF Conference on Computer Vi- sion and Pattern Recognition, pages 6550–6560, 2024. 2 9

  16. [24]

    D-tensorf: Tenso- rial Radiance Fields for Dynamic Scenes

    Hankyu Jang and Daeyoung Kim. D-tensorf: Tenso- rial Radiance Fields for Dynamic Scenes. arXiv preprint arXiv:2212.02375, 2022. 2

  17. [25]

    Splatarmor: Articulated gaussian splatting for animat- able humans from monocular rgb videos

    Rohit Jena, Ganesh Subramanian Iyer, Siddharth Choud- hary, Brandon Smith, Pratik Chaudhari, and James Gee. Splatarmor: Articulated gaussian splatting for animat- able humans from monocular rgb videos. arXiv preprint arXiv:2311.10812, 2023. 2

  18. [26]

    Deformable 3d gaussian splatting for animat- able human avatars

    HyunJun Jung, Nikolas Brasch, Jifei Song, Eduardo Perez- Pellitero, Yiren Zhou, Zhihao Li, Nassir Navab, and Ben- jamin Busam. Deformable 3d gaussian splatting for animat- able human avatars. arXiv preprint arXiv:2312.15059, 2023. 2

  19. [27]

    3d gaussian splatting for real-time radiance field rendering

    Bernhard Kerbl, Georgios Kopanas, Thomas Leimk ¨uhler, and George Drettakis. 3d gaussian splatting for real-time radiance field rendering. ACM Transactions on Graphics, 42 (4), 2023. 1, 2, 3, 5

  20. [28]

    Geo- metric modeling in shape space

    Martin Kilian, Niloy J Mitra, and Helmut Pottmann. Geo- metric modeling in shape space. In ACM SIGGRAPH 2007 papers, pages 64–es. 2007. 1

  21. [29]

    Hugs: Human gaussian splats

    Muhammed Kocabas, Jen-Hao Rick Chang, James Gabriel, Oncel Tuzel, and Anurag Ranjan. Hugs: Human gaussian splats. In Proceedings of the IEEE/CVF conference on com- puter vision and pattern recognition , pages 505–515, 2024. 2

  22. [30]

    Neural human performer: Learning generalizable ra- diance fields for human performance rendering

    Youngjoong Kwon, Dahun Kim, Duygu Ceylan, and Henry Fuchs. Neural human performer: Learning generalizable ra- diance fields for human performance rendering. Advances in Neural Information Processing Systems, 34:24741–24752,

  23. [31]

    Gart: Gaussian articulated template mod- els

    Jiahui Lei, Yufu Wang, Georgios Pavlakos, Lingjie Liu, and Kostas Daniilidis. Gart: Gaussian articulated template mod- els. In Proceedings of the IEEE/CVF Conference on Com- puter Vision and Pattern Recognition , pages 19876–19887,

  24. [32]

    Hu- man101: Training 100+ fps human gaussians in 100s from 1 view

    Mingwei Li, Jiachen Tao, Zongxin Yang, and Yi Yang. Hu- man101: Training 100+ fps human gaussians in 100s from 1 view. arXiv preprint arXiv:2312.15258, 2023

  25. [33]

    Gaussianbody: Clothed human re- construction via 3d gaussian splatting

    Mengtian Li, Shengxiang Yao, Zhifeng Xie, Keyu Chen, and Yu-Gang Jiang. Gaussianbody: Clothed human re- construction via 3d gaussian splatting. arXiv preprint arXiv:2401.09720, 2024. 2

  26. [34]

    Neural 3D Video Synthesis from Multi-view Video

    Tianye Li, Mira Slavcheva, Michael Zollhoefer, Simon Green, Christoph Lassner, Changil Kim, Tanner Schmidt, Steven Lovegrove, Michael Goesele, Richard Newcombe, et al. Neural 3D Video Synthesis from Multi-view Video. In Proceedings of the IEEE/CVF Conference on Computer Vision...

  27. [35]

    Neural Scene Flow Fields for Space-time View Synthesis of Dynamic Scenes

    Zhengqi Li, Simon Niklaus, Noah Snavely, and Oliver Wang. Neural Scene Flow Fields for Space-time View Synthesis of Dynamic Scenes. In Proceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition, pages 6498–6508, 2021. 2

  28. [36]

    Ani- matable gaussians: Learning pose-dependent gaussian maps for high-fidelity human avatar modeling

    Zhe Li, Zerong Zheng, Lizhen Wang, and Yebin Liu. Ani- matable gaussians: Learning pose-dependent gaussian maps for high-fidelity human avatar modeling. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 19711–19722, 2024. 2

  29. [37]

    Barf: Bundle-adjusting neural radiance fields

    Chen-Hsuan Lin, Wei-Chiu Ma, Antonio Torralba, and Si- mon Lucey. Barf: Bundle-adjusting neural radiance fields. In ICCV, pages 5741–5751, 2021. 2

  30. [38]

    High-fidelity and real-time novel view synthesis for dynamic scenes

    Haotong Lin, Sida Peng, Zhen Xu, Tao Xie, Xingyi He, Hu- jun Bao, and Xiaowei Zhou. High-fidelity and real-time novel view synthesis for dynamic scenes. In SIGGRAPH Asia Conference Proceedings, 2023. 2

  31. [39]

    Neural sparse voxel fields

    Lingjie Liu, Jiatao Gu, Kyaw Zaw Lin, Tat-Seng Chua, and Christian Theobalt. Neural sparse voxel fields. pages 15651– 15663, 2020. 2

  32. [40]

    Gea: Reconstructing expressive 3d gaussian avatar from monocular video

    Xinqi Liu, Chenming Wu, Xing Liu, Jialun Liu, Jinbo Wu, Chen Zhao, Haocheng Feng, Errui Ding, and Jingdong Wang. Gea: Reconstructing expressive 3d gaussian avatar from monocular video. arXiv preprint arXiv:2402.16607 ,

  33. [41]

    Animatable 3d gaussian: Fast and high- quality reconstruction of multiple human avatars

    Yang Liu, Xiang Huang, Minghan Qin, Qinwei Lin, and Haoqian Wang. Animatable 3d gaussian: Fast and high- quality reconstruction of multiple human avatars. arXiv preprint arXiv:2311.16482, 2023. 2

  34. [42]

    Robust Dynamic Radi- ance Fields

    Yu-Lun Liu, Chen Gao, Andreas Meuleman, Hung-Yu Tseng, Ayush Saraf, Changil Kim, Yung-Yu Chuang, Jo- hannes Kopf, and Jia-Bin Huang. Robust Dynamic Radi- ance Fields. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 13–23,

  35. [43]

    Smpl: A skinned multi- person linear model

    Matthew Loper, Naureen Mahmood, Javier Romero, Gerard Pons-Moll, and Michael J Black. Smpl: A skinned multi- person linear model. ACM Transactions on Graphics, 34(6),

  36. [44]

    Nerf: Representing scenes as neural radiance fields for view syn- thesis

    Ben Mildenhall, Pratul P Srinivasan, Matthew Tancik, Jonathan T Barron, Ravi Ramamoorthi, and Ren Ng. Nerf: Representing scenes as neural radiance fields for view syn- thesis. In European Conference on Computer Vision, pages 405–421, 2020. 2

  37. [45]

    Human gaussian splatting: Real-time rendering of animatable avatars

    Arthur Moreau, Jifei Song, Helisa Dhamo, Richard Shaw, Yiren Zhou, and Eduardo P ´erez-Pellitero. Human gaussian splatting: Real-time rendering of animatable avatars. In Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 788–798, 2024. 2

  38. [46]

    Instant neural graphics primitives with a mul- tiresolution hash encoding

    Thomas M ¨uller, Alex Evans, Christoph Schied, and Alexan- der Keller. Instant neural graphics primitives with a mul- tiresolution hash encoding. ACM Transactions on Graphics (ToG), 41(4):1–15, 2022. 2

  39. [47]

    Bun- dle adjusted gaussian avatars deblurring

    Muyao Niu, Yifan Zhan, Qingtian Zhu, Zhuoxiao Li, Wei Wang, Zhihang Zhong, Xiao Sun, and Yinqiang Zheng. Bun- dle adjusted gaussian avatars deblurring. arXiv preprint arXiv:2411.16758, 2024. 2

  40. [48]

    Anicrafter: Customizing realistic human-centric animation via avatar- background conditioning in video diffusion models

    Muyao Niu, Mingdeng Cao, Yifan Zhan, Qingtian Zhu, Mingze Ma, Jiancheng Zhao, Yanhong Zeng, Zhihang Zhong, Xiao Sun, and Yinqiang Zheng. Anicrafter: Customizing realistic human-centric animation via avatar- background conditioning in video diffusion models. arXiv preprint arXi...

  41. [49]

    Nerfies: Deformable Neural Radiance Fields

    Keunhong Park, Utkarsh Sinha, Jonathan T Barron, Sofien Bouaziz, Dan B Goldman, Steven M Seitz, and Ricardo Martin-Brualla. Nerfies: Deformable Neural Radiance Fields. In Proceedings of the IEEE/CVF International Con- ference on Computer Vision, pages 5865–5874, 2021. 2 10

  42. [50]

    Temporal Interpola- tion Is All You Need for Dynamic Neural Radiance Fields

    Sungheon Park, Minjung Son, Seokhwan Jang, Young Chun Ahn, Ji-Yeon Kim, and Nahyup Kang. Temporal Interpola- tion Is All You Need for Dynamic Neural Radiance Fields. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4212–4221, 2023. 2

  43. [51]

    Georgios Pavlakos, Vasileios Choutas, Nima Ghorbani, Timo Bolkart, Ahmed A. A. Osman, Dimitrios Tzionas, and Michael J. Black. Expressive body capture: 3D hands, face, and body from a single image. In Proceedings IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), pa...

  44. [52]

    Structure consistent gaussian splatting with matching prior for few-shot novel view synthesis

    Rui Peng, Wangze Xu, Luyang Tang, Jianbo Jiao, Ronggang Wang, et al. Structure consistent gaussian splatting with matching prior for few-shot novel view synthesis. Advances in Neural Information Processing Systems, 37:97328–97352,

  45. [53]

    Neural body: Implicit neural representations with structured latent codes for novel view synthesis of dynamic humans

    Sida Peng, Yuanqing Zhang, Yinghao Xu, Qianqian Wang, Qing Shuai, Hujun Bao, and Xiaowei Zhou. Neural body: Implicit neural representations with structured latent codes for novel view synthesis of dynamic humans. In CVPR,

  46. [54]

    Dynamic point fields

    Sergey Prokudin, Qianli Ma, Maxime Raafat, Julien Valentin, and Siyu Tang. Dynamic point fields. In Proceed- ings of the IEEE/CVF International Conference on Com- puter Vision, pages 7964–7976, 2023. 1

  47. [55]

    D-neRF: Neural Radiance Fields for Dynamic Scenes

    Albert Pumarola, Enric Corona, Gerard Pons-Moll, and Francesc Moreno-Noguer. D-neRF: Neural Radiance Fields for Dynamic Scenes. In Proceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition, pages 10318–10327, 2021. 2

  48. [56]

    3dgs-avatar: Animatable avatars via deformable 3d gaussian splatting

    Zhiyin Qian, Shaofei Wang, Marko Mihajlovic, Andreas Geiger, and Siyu Tang. 3dgs-avatar: Animatable avatars via deformable 3d gaussian splatting. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5020–5030, 2024. 1, 2, 3, 5, 6, 7, 8

  49. [57]

    Tensor4d: Efficient Neural 4D Decomposition for High-fidelity Dynamic Reconstruc- tion and Rendering

    Ruizhi Shao, Zerong Zheng, Hanzhang Tu, Boning Liu, Hongwen Zhang, and Yebin Liu. Tensor4d: Efficient Neural 4D Decomposition for High-fidelity Dynamic Reconstruc- tion and Rendering. In Proceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition, pages...

  50. [58]

    Disentangled generation and aggregation for robust radiance fields

    Shihe Shen, Huachen Gao, Wangze Xu, Rui Peng, Luyang Tang, Kaiqiang Xiong, Jianbo Jiao, and Ronggang Wang. Disentangled generation and aggregation for robust radiance fields. In European Conference on Computer Vision, pages 218–236. Springer, 2024. 2

  51. [59]

    Novel view synthesis of human interactions from sparse multi-view videos

    Qing Shuai, Chen Geng, Qi Fang, Sida Peng, Wenhao Shen, Xiaowei Zhou, and Hujun Bao. Novel view synthesis of human interactions from sparse multi-view videos. In SIG- GRAPH Conference Proceedings, 2022. 2

  52. [60]

    Direct voxel grid optimization: Super-fast convergence for radiance fields reconstruction

    Cheng Sun, Min Sun, and Hwann-Tzong Chen. Direct voxel grid optimization: Super-fast convergence for radiance fields reconstruction. In CVPR, pages 5459–5469, 2022. 2

  53. [61]

    Sparf: Neural radiance fields from sparse and noisy poses

    Prune Truong, Marie-Julie Rakotosaona, Fabian Manhardt, and Federico Tombari. Sparf: Neural radiance fields from sparse and noisy poses. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 4190–4200, 2023. 2

  54. [62]

    Stableanimator: High- quality identity-preserving human image animation

    Shuyuan Tu, Zhen Xing, Xintong Han, Zhi-Qi Cheng, Qi Dai, Chong Luo, and Zuxuan Wu. Stableanimator: High- quality identity-preserving human image animation. In Pro- ceedings of the Computer Vision and Pattern Recognition Conference, pages 21096–21106, 2025. 2

  55. [63]

    Sparsenerf: Distilling depth ranking for few-shot novel view synthesis

    Guangcong Wang, Zhaoxi Chen, Chen Change Loy, and Zi- wei Liu. Sparsenerf: Distilling depth ranking for few-shot novel view synthesis. InProceedings of the IEEE/CVF inter- national conference on computer vision , pages 9065–9076,

  56. [64]

    Arah: Animatable volume rendering of articulated hu- man sdfs

    Shaofei Wang, Katja Schwarz, Andreas Geiger, and Siyu Tang. Arah: Animatable volume rendering of articulated hu- man sdfs. In European conference on computer vision, pages 1–19. Springer, 2022. 2

  57. [65]

    Unianimate: Taming unified video diffusion mod- els for consistent human image animation

    Xiang Wang, Shiwei Zhang, Changxin Gao, Jiayu Wang, Xiaoqiang Zhou, Yingya Zhang, Luxin Yan, and Nong Sang. Unianimate: Taming unified video diffusion mod- els for consistent human image animation. arXiv preprint arXiv:2406.01188, 2024. 2

  58. [66]

    Image Quality Assessment: from Error Visibility to Structural Similarity

    Zhou Wang, Alan C Bovik, Hamid R Sheikh, and Eero P Si- moncelli. Image Quality Assessment: from Error Visibility to Structural Similarity. IEEE transactions on image pro- cessing, 13(4):600–612, 2004. 3, 5

  59. [67]

    Nerf–: Neural radiance fields without known camera parameters

    Zirui Wang, Shangzhe Wu, Weidi Xie, Min Chen, and Victor Adrian Prisacariu. Nerf–: Neural radiance fields without known camera parameters. arXiv preprint arXiv:2102.07064, 2021. 2

  60. [68]

    Hu- mannerf: Free-viewpoint rendering of moving people from monocular video

    Chung-Yi Weng, Brian Curless, Pratul P Srinivasan, Jonathan T Barron, and Ira Kemelmacher-Shlizerman. Hu- mannerf: Free-viewpoint rendering of moving people from monocular video. In Proceedings of the IEEE/CVF con- ference on computer vision and pattern Recognition , pages 162...

  61. [69]

    Swift4d: Adaptive divide-and-conquer gaussian splatting for compact and efficient reconstruction of dynamic scene.arXiv preprint arXiv:2503.12307, 2025

    Jiahao Wu, Rui Peng, Zhiyan Wang, Lu Xiao, Luyang Tang, Jinbo Yan, Kaiqiang Xiong, and Ronggang Wang. Swift4d: Adaptive divide-and-conquer gaussian splatting for compact and efficient reconstruction of dynamic scene.arXiv preprint arXiv:2503.12307, 2025. 2

  62. [70]

    Space-time Neural Irradiance Fields for Free- viewpoint Video

    Wenqi Xian, Jia-Bin Huang, Johannes Kopf, and Changil Kim. Space-time Neural Irradiance Fields for Free- viewpoint Video. In Proceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition, pages 9421–9431, 2021. 2

  63. [71]

    Mvpgs: Excavating multi-view priors for gaussian splatting from sparse input views

    Wangze Xu, Huachen Gao, Shihe Shen, Rui Peng, Jianbo Jiao, and Ronggang Wang. Mvpgs: Excavating multi-view priors for gaussian splatting from sparse input views. In European Conference on Computer Vision, pages 203–220. Springer, 2024. 2

  64. [72]

    4k4d: Real-time 4d view synthesis at 4k resolution

    Zhen Xu, Sida Peng, Haotong Lin, Guangzhao He, Jiaming Sun, Yujun Shen, Hujun Bao, and Xiaowei Zhou. 4k4d: Real-time 4d view synthesis at 4k resolution. InCVPR, 2024. 2

  65. [73]

    Magicanimate: Temporally consistent human im- age animation using diffusion model

    Zhongcong Xu, Jianfeng Zhang, Jun Hao Liew, Hanshu Yan, Jia-Wei Liu, Chenxu Zhang, Jiashi Feng, and Mike Zheng Shou. Magicanimate: Temporally consistent human im- age animation using diffusion model. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Reco...

  66. [74]

    Instant gaussian stream: Fast and generalizable streaming of dy- namic scene reconstruction via gaussian splatting

    Jinbo Yan, Rui Peng, Zhiyan Wang, Luyang Tang, Jiayu Yang, Jie Liang, Jiahao Wu, and Ronggang Wang. Instant gaussian stream: Fast and generalizable streaming of dy- namic scene reconstruction via gaussian splatting. In Pro- ceedings of the Computer Vision and Pattern Recogniti...

  67. [75]

    Tomie: Towards modular growth in en- hanced smpl skeleton for 3d human with animatable gar- ments

    Yifan Zhan, Qingtian Zhu, Muyao Niu, Mingze Ma, Jiancheng Zhao, Zhihang Zhong, Xiao Sun, Yu Qiao, and Yinqiang Zheng. Tomie: Towards modular growth in en- hanced smpl skeleton for 3d human with animatable gar- ments. arXiv preprint arXiv:2410.08082, 2024. 2

  68. [76]

    R3-avatar: Record and retrieve temporal code- book for reconstructing photorealistic human avatars

    Yifan Zhan, Wangze Xu, Qingtian Zhu, Muyao Niu, Mingze Ma, Yifei Liu, Zhihang Zhong, Xiao Sun, and Yinqiang Zheng. R3-avatar: Record and retrieve temporal code- book for reconstructing photorealistic human avatars. arXiv preprint arXiv:2503.12751, 2025. 2

  69. [77]

    The Unreasonable Effectiveness of Deep Features as a Perceptual Metric

    Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shecht- man, and Oliver Wang. The Unreasonable Effectiveness of Deep Features as a Perceptual Metric. In Proceedings of the IEEE conference on computer vision and pattern recogni- tion, pages 586–595, 2018. 5

  70. [78]

    Gps- gaussian: Generalizable pixel-wise 3d gaussian splatting for real-time human novel view synthesis

    Shunyuan Zheng, Boyao Zhou, Ruizhi Shao, Boning Liu, Shengping Zhang, Liqiang Nie, and Yebin Liu. Gps- gaussian: Generalizable pixel-wise 3d gaussian splatting for real-time human novel view synthesis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Re...

  71. [79]

    Physa- vatar: Learning the physics of dressed 3d avatars from visual observations

    Yang Zheng, Qingqing Zhao, Guandao Yang, Wang Yi- fan, Donglai Xiang, Florian Dubost, Dmitry Lagun, Thabo Beeler, Federico Tombari, Leonidas Guibas, et al. Physa- vatar: Learning the physics of dressed 3d avatars from visual observations. arXiv preprint arXiv:2404.04421, 2024. 2, 8

  72. [80]

    Fsgs: Real-time few-shot view synthesis using gaussian splatting

    Zehao Zhu, Zhiwen Fan, Yifan Jiang, and Zhangyang Wang. Fsgs: Real-time few-shot view synthesis using gaussian splatting. In European conference on computer vision, pages 145–163. Springer, 2024. 2

  73. [81]

    Driv- able 3d gaussian avatars

    Wojciech Zielonka, Timur Bagautdinov, Shunsuke Saito, Michael Zollh ¨ofer, Justus Thies, and Javier Romero. Driv- able 3d gaussian avatars. arXiv preprint arXiv:2311.08581,

  74. [83]

    We follow Dyco [8] to derive each joint’s rotation variation ∆θ between adja- cent frames as its skeleton motion representation

    Details of Skeleton Motion ∆P The pose P is represented as the rotation relative to the par- ent node of all joints: P = {θ1, θ2, ..., θK}, (26) where θj ∈ R3 describes the j-th joint’s relative rotation with its parent in axis-angle form [43]. We follow Dyco [8] to derive eac...

  75. [84]

    We use images at a resolution of 512 ×512 for experiments and follow Dyco[8] for both training and testing splits

    Datasets I3D-Human Dataset [8] . We use images at a resolution of 512 ×512 for experiments and follow Dyco[8] for both training and testing splits. DNA-Rendering Dataset [9]. We use images with a reso- lution of 512 × 612 for training and 750 × 1024 for testing in our experime...

  76. [85]

    (17) describe the skele- ton and local region motions, respectively

    Motion Embeddings for Non-Rigid MLP The embeddings f∆P and fV in Eq. (17) describe the skele- ton and local region motions, respectively. For each frame, all Gaussian primitives {Gi} share the same skeleton motion embedding f∆P ∈ R32 and we con- cat it with Gi’s point-wise mot...

  77. [86]

    Details of Loss Functions Lcolor is the L1 loss between the rendered image I and the ground truth Igt: Lcolor = |I − Igt|. (30) We use Lssim to constrain the structure similarity between the rendered image and the ground truth, which is given by Lssim = 1 − SSIM(I, Igt), (31) ...

  78. [87]

    (14)) of se- quence sampled in different scales is set to 3, and each sequence length L (in Eq

    Details of Optimization In our experiments, the number |S| (in Eq. (14)) of se- quence sampled in different scales is set to 3, and each sequence length L (in Eq. (6)) is set to 8. We use the τ = 8 nearest SMPL template vertexes’ velocities to get each Gaussian primitive’s loc...

  79. [88]

    Experiments on ZJU-MoCap ZJU-MoCap [53] is a common benchmark dataset in hu- man avatar, which is mainly collected under controlled speeds and tight-fitting garments

    Additional Results 12.1. Experiments on ZJU-MoCap ZJU-MoCap [53] is a common benchmark dataset in hu- man avatar, which is mainly collected under controlled speeds and tight-fitting garments. We use 6 sequences (377, 386, 387, 392, 393, 394) for experiments following the datas...

  80. [89]

    6 shows the quantitative met- rics with different sequential sampling steps

    Additional Ablations Different Sample Steps Tab. 6 shows the quantitative met- rics with different sequential sampling steps. The results demonstrate that combining all coarse-to-fine temporal mo- tions as the condition leads to more robust non-rigid defor- mation and better p...

  81. [90]

    Computational Efficiency

    Computational Cost Analysis Table 8. Computational Efficiency. Mehods Train Time FPS Train Mem. Dyco 6 h 0.7 20 GB 3DGS-Avatar 25 m 25 6 GB Ours (3k iter) 5 m 62 8 GB We provide computational cost analysis for experiments on ZJU-Mocap (conducted on a single 4090 GPU). Tab. 8 s...

  82. [2023]

    2 12 Sequential Gaussian Avatars with Hierarchical Motion Context Supplementary Material

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.