Pith. sign in

REVIEW 3 major objections 5 minor 90 references

FATE: Full-head Gaussian Avatar with Textural Editing from Monocular Video

T0 review · 3 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read A single monocular video is enough to rebuild a complete 360-degree, animatable, texture-editable head avatar.

desk verdict Solid frontal reconstruction with credible ablations, but the 360-degree full-head claim rests on unquantified SphereHead/PTI pseudo-GT and needs a much harder look. read the letter →

arxiv 2411.15604 v2 pith:SOK7OXGA submitted 2024-11-23 cs.CV

classification cs.CV
keywords monocularheadavatar3DGaussiansplattingUV-spaceparameterizationsampling-baseddensificationneuralbakingtextureeditingfull-headcompletiongenerativepriorinversion
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper claims that a single, casually recorded monocular video is enough to reconstruct a complete $360^\circ$-renderable, animatable 3D head avatar — a first, to the authors' knowledge. It attacks the two standard bottlenecks in one pipeline: sampling-based densification replaces the gradient-threshold rules of vanilla 3DGS to spread Gaussians more efficiently, and a neural-baking stage converts the discrete splats into continuous UV attribute maps so the avatar can be edited as easily as a mesh texture. To finish the unseen back and sides of the head, a universal completion framework pulls priors from a pretrained full-head generative model and cross-trains hallucinated views with the real video frames. If the claims hold, a single phone video could feed games, film, and telepresence with an editable avatar, and the reported numbers place FATE at state of the art while using far fewer Gaussians than its main rivals.

What carries the argument

Three mechanisms carry the argument. (1) Sampling-based densification: instead of a threshold $\tau_{pos}$, the gradient magnitude $|\partial L / \partial \mu|$ of each Gaussian is used as an importance metric for multinomial sampling over the UV-bound face $f_i$; a new splat is born at uniformly drawn barycentric coordinates and inherits the sampled splat's attributes, so Gaussian count stays controlled while the distribution adapts. (2) Neural baking: a BakeNet (a U-Net) acts as a low-pass pre-filter $H$ between a multi-channel Gaussian noise map $F$ and a bilinear kernel $B$, yielding a continuous attribute function $f(p) = (F * H * B)(p)$ in UV space; sampled outputs replace the point-wise attributes in a second training stage, making the texture maps smooth and directly editable. (3) Full-head completion: the avatar is rendered on a horizontal camera orbit at neutral pose, valid views are kept by facial-landmark confidence, GFPGAN aligns image quality, multi-view Pivotal Tuning Inversion fits the subject into SphereHead, and the synthesized rear and side images are inverse-transformed, matted, and cross-trained with real frames as pseudo-ground-truth. The FLAME-based UV binding of Gaussians to mesh faces is the shared substrate that makes densification and baking meaningful.

What would settle it

Take a talking-head video with virtually no head turns (e.g., a fixed frontal interview clip), run FATE, and render the completed avatar at $\pm 90^\circ$ yaw; the paper's own failure analysis (Fig. 14) predicts identity changes or visible seams at those angles. A quantitative version compares side/rear renders against a multi-view or light-stage ground truth across subjects with varied hair and head shapes, measuring identity similarity and geometric error to map exactly when the $360^\circ$ claim breaks.

Watch

Extended reading notes

Core claim

FATE's central discovery is that the two long-standing obstacles of monocular head reconstruction — incomplete geometry and an inefficient, discrete Gaussian representation — can be solved in one coherent system. The paper argues that density should be controlled by sampling: treating the per-Gaussian gradient magnitude as an importance weight for multinomial sampling over binding triangles yields a better positional distribution than threshold-based cloning and splitting, producing the reported quality with roughly half the Gaussian count of the nearest competitor. It then argues that a discrete Gaussian avatar can be made editable by neural baking, where a U-Net pre-filter plus bilinear kernel maps the splat attributes onto continuous UV attribute maps, so color, opacity, scale, rotation, and offset become paint-able textures. Finally, to obtain the $360^\circ$ capability, the completion framework renders the neutral avatar on an orbit, restores image quality with GFPGAN, inverts the subject into SphereHead via multi-view Pivotal Tuning Inversion, and cross-trains the synthesized side and rear views with the real data. The authors report overall PSNR 28.37 and LPIPS 0.0586 — state of the art in their comparison — and claim to be the first to deliver an animatable, $360^\circ$ full-head monocular reconstruction.

Load-bearing premise

The side-and-back views are only as trustworthy as the generative prior: the completion framework assumes that SphereHead, after GFPGAN quality alignment and multi-view PTI inversion, produces credible pseudo-ground-truth for the unseen rear head — something the paper itself reports failing for long-haired subjects and for videos with almost no side views, where junction artifacts or identity changes appear.

Editorial extensions

If this is right

  • A single monocular portrait video is sufficient to obtain an animatable, $360^\circ$-renderable full-head avatar, removing the multi-view capture rig for consumer applications.
  • Sampling-based densification yields state-of-the-art or comparable rendering with far fewer Gaussians than prior UV-based methods (about 49k vs 72k for GaussianAvatars on INSTA), which shortens training and raises frame rates.
  • After neural baking, appearance edits — stickers, style transfers, recoloring — become direct texture-map operations instead of costly per-avatar optimization with diffusion models.
  • The completion framework is method-agnostic: the authors show it can be grafted onto other monocular reconstruction baselines to give them plausible side and rear views as well.
  • Neural baking involves a tunable trade-off between rendering fidelity and texture quality; the paper documents regularization options (e.g., rotation regularization, baking appearance only) to move along that trade-off.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • In our reading, the $360^\circ$ claim is bounded by the generative prior's coverage: the completion stage may quietly replace identity rather than reconstruct it whenever the subject's hair or head shape sits far from SphereHead's training distribution, so FATE's true frontier is where its completion framework stops being trustworthy.
  • The sampling-based densification mechanism is not head-specific; any surface-parameterized Gaussian representation that suffers from gradient-threshold densification explosions could adopt the same importance-sampling rule.
  • Baking into continuous UV maps hints at a converged asset pipeline: once attributes live on texture maps, the avatar could in principle be exported like a conventional textured mesh and edited in standard software — a step the paper does not take.
  • A natural testable extension is to measure completion quality as a function of the input video's angular coverage; we would expect a sharp drop in identity preservation below some minimum head-turn angle, consistent with the paper's reported failures.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper introduces FATE, a monocular 3D head avatar reconstruction method that combines sampling-based densification for Gaussian splats, a neural baking stage that converts discrete Gaussian attributes into continuous UV-space maps for editing, and a completion framework that uses SphereHead with GFPGAN-based quality alignment and PTI inversion to synthesize pseudo-ground-truth for non-frontal views. The authors report state-of-the-art frontal-view rendering metrics across four datasets, fewer Gaussians than prior methods, and qualitative results for editing and 360-degree completion.

Significance. If the frontal reconstruction and editing results are taken at face value, FATE contributes a practical UV-embedded Gaussian representation with efficient densification and a direct texture-editing workflow, backed by per-subject quantitative comparisons and ablations. The proposed completion framework is also potentially useful as a plug-in for other monocular head avatar methods. However, the headline claim of being the first 360-degree full-head monocular reconstruction is not quantitatively supported: the completion step relies on pseudo-ground-truth from a generative model, and the paper itself documents failure cases. The significance of the full-head claim therefore remains uncertain pending additional evaluation.

major comments (3)
  1. [Sec. 3.4, Figs. 8/17, Sec. 10, Fig. 14] The 360-degree full-head claim is not supported by quantitative evaluation. All reported metrics in Tables 1, 8, and 9 are computed on frontal test frames near the training camera distribution; the completed side and rear views are only shown qualitatively (Figs. 8, 17, 18). Since the completion framework is the sole source of supervision for non-frontal views and is based on SphereHead/PTI pseudo-ground-truth, the paper should provide a quantitative evaluation on held-out side/rear views (e.g., from subjects with multi-view ground truth or reserved side-view frames), an identity-similarity check, and a user study. This is particularly important because Sec. 10 and Fig. 14 document junction artifacts and identity changes for long-haired subjects and for videos with almost no side views.
  2. [Sec. 3.4, Sec. 10] The statement in Sec. 3.4 that 'some identity changes caused by GFPGAN in the frontal view are acceptable' conflicts with the goal of faithful full-head reconstruction. GFPGAN is applied to the real frontal frames that drive the PTI inversion, and the inverted generator is then used to create pseudo-data for unseen views. If the frontal view is altered, the identity of the completed avatar may drift; Fig. 14(b) shows an identity change in the side view. The paper should quantify identity preservation (e.g., face recognition cosine similarity between the real subject and the completed avatar) and report how often completion fails, rather than treating these as isolated examples.
  3. [Sec. 9, Table 4] The neural baking trade-off is reported on a single subject and the conclusions are not consistent: 'Bake App. Only' improves LPIPS but worsens PSNR and produces noisy textures, while the attribute regularization results show little sensitivity to λV. Because the paper's title and contributions emphasize textural editing, the editing capability needs more than qualitative figures; a quantitative comparison of editing quality (e.g., consistency of edited regions across views, speed, or a user study) is needed to substantiate that baking enables editing 'with the same ease and efficacy as mesh textures.'
minor comments (5)
  1. [Sec. 3.4] The completion framework is described as having 'three steps' but lists coordinate alignment, image quality alignment, and inversion/finetuning; clarify the naming. Also, the exact loss weights and scheduling for cross-training between real and pseudo-images are not given in the main text or supplementary.
  2. [Sec. 7.4] The choice to use only the latter half of the 30 orbit images as pseudo-data is an ad-hoc decision with no ablation; please justify or study this choice.
  3. [Table 4] The trade-off conclusions are based on a single subject ('bala case'); report results on more subjects to support the general claim.
  4. [Abstract/Introduction] The claim of being 'the first animatable and 360° full-head monocular reconstruction method' should explicitly differentiate from single-image 360-degree generative models such as PanoHead and SphereHead, which also produce full heads.
  5. [Throughout] There are typos and minor wording issues, including 'unvoidable' in Sec. 1, 'We also study to improve' in Sec. 3.3, and inconsistent capitalization of 'Full-head' versus 'full-head'.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: FATE's frontal reconstruction is benchmarked externally, and the full-head completion is an explicit SphereHead prior rather than a disguised self-derivation.

full rationale

The core monocular reconstruction (Sec. 3.1 and 3.2) is trained against real monocular frames and compared with external baselines (GaussianAvatars, FlashAvatar, MonoGaussianAvatar, SplattingAvatar) on standard metrics in Tables 1, 8, and 9. The ablations in Table 3 isolate the contributions of sampling-based densification, learnable blendshapes, and baking strategy, so these components are not fitted to the quantities they are used to predict. Neural baking (Sec. 3.3) reparameterizes the already-learned Gaussian avatar into continuous UV maps and is evaluated by rendering quality on the same monocular view distribution; it is a compression/editing tool, not a derivation that reduces to its own output. The full-head completion framework (Sec. 3.4) transparently uses SphereHead, an external pretrained full-head generative model, to synthesize pseudo-ground-truth for side and rear views; this is an explicit prior injection rather than a hidden equivalence, and the paper honestly documents its failure modes in Sec. 10, including junction artifacts and identity changes. The self-citations in Sec. 3.3 (refs [75,76]) support a design analogy for why convolutional inductive bias might yield continuous attribute maps, but they are accompanied by external citations [37,39] and are not load-bearing for the central reconstruction claim. Because every load-bearing comparison is against external benchmarks or real monocular supervision, and the completion step is an acknowledged external prior rather than a quantity derived from itself, no circular step can be exhibited from the paper's equations or descriptions.

Assumptions & free parameters 7 free parameters · 5 assumptions · 0 invented entities

The central claims rest on a set of hand-tuned hyperparameters (loss weights, densification interval, opacity thresholds) and on domain assumptions about the representational sufficiency of FLAME, zero-order SH lighting, and the transferability of SphereHead priors. No new physical entities are introduced.

free parameters (7)
  • Loss weights lambda1, lambda2, lambda3, lambda4 = 0.1, 100, 100, 0.1
    Hand-tuned in Eq. 15 to balance image, perceptual, mesh, and scale losses.
  • Densification interval = 1k Gaussians every 3k iterations
    Chosen empirically to control Gaussian growth during training.
  • Pruning threshold = opacity 5e-3
    Opacity threshold for pruning unsuitable Gaussians, stated in Sec. 7.2.
  • Opacity reset period = every 6k iterations
    Resets opacity to prevent over-transparent Gaussians.
  • r (aspect ratio limit in Lscale) = not stated
    Hyperparameter in Eq. 13 limiting the max/min scale ratio; no value is given in the main text or appendix.
  • Camera radius adjustment = 2.7 to 3.2
    Adjusted in completion to keep the portrait within the viewing frustum (Sec. 7.4).
  • PTI optimization iterations = 200 (latent) + 200 (fine-tune)
    Inversion steps in the completion framework (Sec. 7.4).
assumptions (5)
  • domain assumption The FLAME parametric head model provides sufficient template geometry for the frontal and rear head.
    The entire reconstruction relies on FLAME plus learnable blendshapes; if the template is inaccurate for the rear head, the Gaussian positions are biased and the completion inherits the error.
  • domain assumption Zero-degree spherical harmonics suffice for consistent lighting.
    Assumed in the implementation details (Sec. 4.1); rendering quality degrades under non-uniform lighting, as acknowledged in the limitations.
  • domain assumption The pretrained SphereHead model can generate plausible side and rear head appearance for the specific subject after PTI inversion.
    Core premise of the completion framework; the paper documents failure cases (Sec. 10, Fig. 14) where this assumption breaks.
  • domain assumption GFPGAN restoration improves domain alignment without corrupting identity.
    Used in image quality alignment; the paper admits identity changes are acceptable in the frontal view, which implies a tacit acceptance of this risk.
  • ad hoc to paper Using only the latter half of the 30 orbit images as pseudo-data is sufficient to avoid alignment artifacts.
    A design choice in Sec. 7.4 made to reduce artifacts from PTI alignment, without a systematic study of how many images are needed.

how reviews work

0 comments
Cite this review

Pith. "Pith review of FATE: Full-head Gaussian Avatar with Textural Editing from Monocular Video." pith.science (2026). https://pith.science/paper/SOK7OXGA

@misc{pith2026241115604,
  author       = {Pith},
  title        = {Pith review of: FATE: Full-head Gaussian Avatar with Textural Editing from Monocular Video},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/SOK7OXGA}},
  note         = {Machine review of arXiv:2411.15604}
}
abstract

Reconstructing high-fidelity, animatable 3D head avatars from effortlessly captured monocular videos is a pivotal yet formidable challenge. Although significant progress has been made in rendering performance and manipulation capabilities, notable challenges remain, including incomplete reconstruction and inefficient Gaussian representation. To address these challenges, we introduce FATE, a novel method for reconstructing an editable full-head avatar from a single monocular video. FATE integrates a sampling-based densification strategy to ensure optimal positional distribution of points, improving rendering efficiency. A neural baking technique is introduced to convert discrete Gaussian representations into continuous attribute maps, facilitating intuitive appearance editing. Furthermore, we propose a universal completion framework to recover non-frontal appearance, culminating in a 360$^\circ$-renderable 3D head avatar. FATE outperforms previous approaches in both qualitative and quantitative evaluations, achieving state-of-the-art performance. To the best of our knowledge, FATE is the first animatable and 360$^\circ$ full-head monocular reconstruction method for a 3D head avatar.

Figures

Figures reproduced from arXiv: 2411.15604 by the authors.

Figure 1
Figure 1. From a monocular portrait video input, we propose FATE to reconstruct an animatable 3D head avatar, which enables Gaussian [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Pipeline. In Stage I, we perform sampling-based densification in Sec. 3.2 in the UV space and train a Gaussian head avatar using the preprocessed monocular video dataset. The obtained head avatar can optionally use full-head completion in Sec 3.4 to recover non-frontal regions. In Stage II, given the learned head avatar, we construct a continuous function f(p) in the UV space using U-Net H and bilinear kernel B, bak… view at source ↗
Figure 3
Figure 3. 3DGS in Monocular Video. (a) In monocular recon￾struction, since the sides of the head avatar are rarely supervised, Gaussians tend to grow towards the direction of the rendering cam￾era. (b) This potentially results in position gradient visualizations during training, showing that most of the facial region displays distributions exceeding the threshold τpos. monocular videos exacerbates this issue. As shown in Fig.… view at source ↗
Figures from the paper (16 more)
Figure 5
Figure 5. Figure 5: Baked Results Visualization. We visualize the color texture map produced by neural baking on different subjects. convolutional operations incorporated into the generative model. On further analysis, we argue that the inductive bi￾ases of the CNN contribute to local smo…
Figure 6
Figure 6. Figure 6: Completion Framework. A universal framework is proposed to complete the side and rear appearance under monocular settings. onal projection, which differs from monocular video-based reconstruction. Therefore, establishing model space trans￾formation and enhancing the qu…
Figure 7
Figure 7. Figure 7: Monocular Reconstruction Results. Our method is more effective at capturing fine structure and high-frequency details (e.g. loose strands of hair, lip creases, and stubble in the facial area.). More reconstructed subjects are shown in supplementary materials [PITH_FUL…
Figure 8
Figure 8. Figure 8: Full-head Completion Results. The first row shows the side and back views rendered in our method without completion, and the second row shows the result after completion [PITH_FULL_IMAGE:figures/full_fig_p007_8.png]
Figure 9
Figure 9. Figure 9: Texture Editing Results. We show the effects of simply and effectively editing the baked texture map [PITH_FULL_IMAGE:figures/full_fig_p007_9.png]
Figure 10
Figure 10. Figure 10: BakeNet Architecture. We adopt a U-Net architecture as the backbone of BakeNet, leveraging its ability to construct rep￾resentations across various frequency bands from noise. 7.3. Neural Baking We use a simple U-Net [47] as shown in [PITH_FULL_IMAGE:figures/full_fig…
Figure 11
Figure 11. Figure 11: Incomplete Inversion Issues. In typical inversion op￾timization, the neck and top of the head of the portrait often fall outside the frame, as shown in (a). We obtained the result shown in (b) by adjusting the camera-to-object distance. (see [PITH_FULL_IMAGE:figures/…
Figure 12
Figure 12. Figure 12: Neural Baking Trade-off. We visualize the color texture maps produced by neural baking under different settings and the results after editing with a checking sticker. Bake Appearance Only We only use neural baking to ob￾tain texture maps for color and opacity, while t…
Figure 13
Figure 13. Figure 13: Neural Baking Failure. For long hair subjects, as in (a), direct neural baking will damage the fine geometry of the Gaussians composing the hair as in (b) [PITH_FULL_IMAGE:figures/full_fig_p016_13.png]
Figure 14
Figure 14. Figure 14: Full-head Completion Failure. Since the PTI results still differ from the real avatar, artifacts appear at the junction, as shown in the red box in (a). And for avatars with almost no side view in the training data, as shown in (b), it is difficult to estimate the exa…
Figure 15
Figure 15. Figure 15: Robustness to Imperfect Poses We add noise to camera translation to simulate less well-processed datasets. Note that 1 mm in the figure approximately corresponds to 1 cm in the real world [PITH_FULL_IMAGE:figures/full_fig_p017_15.png]
Figure 16
Figure 16. Figure 16: More Reconstructed Results. Our method excels at capturing fine structures and preserving high-frequency details (e.g., eyebrows, hair strands, eyeglass frames, and pupil colors.). 7 [PITH_FULL_IMAGE:figures/full_fig_p018_16.png]
Figure 17
Figure 17. Figure 17: More Full-head Completion Results. Odd rows display the results under novel views without applying the Full-head comple￾tion framework, while even rows show the results after completion. Our completion framework significantly enhances rendering quality under large vie…
Figure 18
Figure 18. Figure 18: Universal Completion Results. Odd rows display the results under novel views without applying the Full-head completion framework, while even rows show the results after completion. Our completion framework applies to various monocular reconstruction methods. 9 [PITH_…
Figure 19
Figure 19. Figure 19: Cross-reenactment Results. We use the expression and pose sequences from the driving source to animate different subjects, enabling the transfer of dynamic facial expressions and poses across various avatars [PITH_FULL_IMAGE:figures/full_fig_p021_19.png]
Figure 20
Figure 20. Figure 20: Editing Results. In (a), we show several results of directly editing the texture map by adding stickers, such as anime portraits, rainbows, kisses, mustaches, and logos. In (b), we present the results of applying style transfer to the texture map. 10 [PITH_FULL_IMAGE…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

90 extracted references · 69 canonical work pages

  1. [1]

    Gaussian shell maps for efficient 3d human generation

    Rameen Abdal, Wang Yifan, Zifan Shi, Yinghao Xu, Ryan Po, Zhengfei Kuang, Qifeng Chen, Dit-Yan Yeung, and Gor- don Wetzstein. Gaussian shell maps for efficient 3d human generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 9441–9451, 2024. 2

  2. [2]

    Ogras, and Linjie Luo

    Sizhe An, Hongyi Xu, Yichun Shi, Guoxian Song, Umit Y . Ogras, and Linjie Luo. Panohead: Geometry-aware 3d full- head synthesis in 360deg. In CVPR, pages 20950–20959,

  3. [3]

    Geneavatar: Generic expression-aware vol- umetric head avatar editing from a single image

    Chong Bao, Yinda Zhang, Yuan Li, Xiyu Zhang, Bang- bang Yang, Hujun Bao, Marc Pollefeys, Guofeng Zhang, and Zhaopeng Cui. Geneavatar: Generic expression-aware vol- umetric head avatar editing from a single image. In CVPR,

  4. [4]

    A morphable model for the synthesis of 3d faces

    V olker Blanz and Thomas Vetter. A morphable model for the synthesis of 3d faces. In Proceedings of the 26th Annual Conference on Computer Graphics and Interactive Tech- niques, page 187–194, USA, 1999. ACM Press/Addison- Wesley Publishing Co. 2

  5. [5]

    Tim Brooks, Aleksander Holynski, and Alexei A. Efros. In- structpix2pix: Learning to follow image editing instructions. In CVPR, 2023. 2

  6. [6]

    Buehler, Abhimitra Meka, Gengyan Li, Thabo Beeler, and Otmar Hilliges

    Marcel C. Buehler, Abhimitra Meka, Gengyan Li, Thabo Beeler, and Otmar Hilliges. Varitex: Variational neural face textures. In CVPR, 2021. 3

  7. [7]

    Preface: A data-driven volumetric prior for few-shot ultra high-resolution face synthesis

    Marcel C B ¨uhler, Kripasindhu Sarkar, Tanmay Shah, Gengyan Li, Daoye Wang, Leonhard Helminger, Ser- gio Orts-Escolano, Dmitry Lagun, Otmar Hilliges, Thabo Beeler, et al. Preface: A data-driven volumetric prior for few-shot ultra high-resolution face synthesis. InICCV, pages 3402–3413, 2023. 3

  8. [8]

    Facewarehouse: A 3d facial expression database for visual computing

    Chen Cao, Yanlin Weng, Shun Zhou, Yiying Tong, and Kun Zhou. Facewarehouse: A 3d facial expression database for visual computing. IEEE TVCG, 20(3):413–425, 2014. 2

Show all 90 references
  1. [9]

    Authentic volumetric avatars from a phone scan

    Chen Cao, Tomas Simon, Jin Kyu Kim, Gabe Schwartz, Michael Zollhoefer, Shun-Suke Saito, Stephen Lombardi, Shih-En Wei, Danielle Belko, Shoou-I Yu, Yaser Sheikh, and Jason Saragih. Authentic volumetric avatars from a phone scan. ACM TOG, 41(4), 2022. 3

  2. [10]

    pi-gan: Periodic implicit generative ad- versarial networks for 3d-aware image synthesis

    Eric Chan, Marco Monteiro, Petr Kellnhofer, Jiajun Wu, and Gordon Wetzstein. pi-gan: Periodic implicit generative ad- versarial networks for 3d-aware image synthesis. In CVPR,

  3. [11]

    Chan, Connor Z

    Eric R. Chan, Connor Z. Lin, Matthew A. Chan, Koki Nagano, Boxiao Pan, Shalini De Mello, Orazio Gallo, Leonidas Guibas, Jonathan Tremblay, Sameh Khamis, Tero Karras, and Gordon Wetzstein. Efficient geometry-aware 3D generative adversarial networks. In CVPR, 2022. 3

  4. [12]

    Tensorf: Tensorial radiance fields

    Anpei Chen, Zexiang Xu, Andreas Geiger, Jingyi Yu, and Hao Su. Tensorf: Tensorial radiance fields. In ECCV, 2022. 2

  5. [13]

    Monogaus- sianavatar: Monocular gaussian point-based head avatar

    Yufan Chen, Lizhen Wang, Qijing Li, Hongjiang Xiao, Shengping Zhang, Hongxun Yao, and Yebin Liu. Monogaus- sianavatar: Monocular gaussian point-based head avatar. In ACM SIGGRAPH 2024 Conference Papers, 2024. 2, 3, 7, 8

  6. [14]

    Black, and Timo Bolkart

    Radek Danecek, Michael J. Black, and Timo Bolkart. EMOCA: Emotion driven monocular face capture and ani- mation. In CVPR, pages 20311–20322, 2022. 1, 2

  7. [15]

    The light stages and their applications to pho- toreal digital actors

    Paul Debevec. The light stages and their applications to pho- toreal digital actors. ACM TOG, 2(4):1–6, 2012. 1

  8. [16]

    Accurate 3d face reconstruction with weakly-supervised learning: From single image to image set

    Yu Deng, Jiaolong Yang, Sicheng Xu, Dong Chen, Yunde Jia, and Xin Tong. Accurate 3d face reconstruction with weakly-supervised learning: From single image to image set. In CVPRW, 2019. 2

  9. [17]

    Bakedavatar: Baking neural fields for real- time head avatar synthesis

    Hao-Bin Duan, Miao Wang, Jin-Chuan Shi, Xu-Chuan Chen, and Yan-Pei Cao. Bakedavatar: Baking neural fields for real- time head avatar synthesis. ACM TOG, 42(6), 2023. 2

  10. [18]

    Black, and Timo Bolkart

    Yao Feng, Haiwen Feng, Michael J. Black, and Timo Bolkart. Learning an animatable detailed 3D face model from in-the-wild images. ACM TOG, 40(8), 2021. 1, 2

  11. [19]

    Dynamic neural radiance fields for monocular 4d facial avatar reconstruction

    Guy Gafni, Justus Thies, Michael Zollh ¨ofer, and Matthias Nießner. Dynamic neural radiance fields for monocular 4d facial avatar reconstruction. In CVPR, pages 8649–8658,

  12. [20]

    Dynamic neural radiance fields for monocular 4d facial avatar reconstruction

    Guy Gafni, Justus Thies, Michael Zollh ¨ofer, and Matthias Nießner. Dynamic neural radiance fields for monocular 4d facial avatar reconstruction. In CVPR, pages 8645–8654,

  13. [21]

    Reconstructing personalized se- mantic facial nerf models from monocular video.ACM TOG, 41(6), 2022

    Xuan Gao, Chenglai Zhong, Jun Xiang, Yang Hong, Yudong Guo, and Juyong Zhang. Reconstructing personalized se- mantic facial nerf models from monocular video.ACM TOG, 41(6), 2022. 2

  14. [22]

    Mani-gs: Gaussian splatting manipulation with triangular mesh

    Xiangjun Gao, Xiaoyu Li, Yiyu Zhuang, Qi Zhang, Wenbo Hu, Chaopeng Zhang, Yao Yao, Ying Shan, and Long Quan. Mani-gs: Gaussian splatting manipulation with triangular mesh. arXiv preprint arXiv:2405.17811, 2024. 2

  15. [23]

    Portrait video editing em- powered by multimodal generative priors

    Xuan Gao, Haiyao Xiao, Chenglai Zhong, Shimin Hu, Yudong Guo, and Juyong Zhang. Portrait video editing em- powered by multimodal generative priors. In ACM SIG- GRAPH Asia, 2024. 2

  16. [24]

    Neu- ral head avatars from monocular rgb videos

    Philip-William Grassal, Malte Prinzler, Titus Leistner, Carsten Rother, Matthias Nießner, and Justus Thies. Neu- ral head avatars from monocular rgb videos. InCVPR, pages 18653–18664, 2022. 2

  17. [25]

    Towards fast, accurate and stable 3d dense face alignment

    Jianzhu Guo, Xiangyu Zhu, Yang Yang, Fan Yang, Zhen Lei, and Stan Z Li. Towards fast, accurate and stable 3d dense face alignment. In ECCV, 2020. 6

  18. [26]

    The re- lightables: V olumetric performance capture of humans with realistic relighting

    Kaiwen Guo, Peter Lincoln, Philip Davidson, Jay Busch, Xueming Yu, Matt Whalen, Geoff Harvey, Sergio Orts- Escolano, Rohit Pandey, Jason Dourgarian, et al. The re- lightables: V olumetric performance capture of humans with realistic relighting. ACM TOG, 38(6):1–19, 2019. 1

  19. [27]

    Emotalk3d: High-fidelity free-view synthesis of emotional 3d talking head

    Qianyun He, Xinya Ji, Yicheng Gong, Yuanxun Lu, Zhengyu Diao, Linjia Huang, Yao Yao, Siyu Zhu, Zhan Ma, Songchen Xu, Xiaofei Wu, Zixiao Zhang, Xun Cao, and Hao Zhu. Emotalk3d: High-fidelity free-view synthesis of emotional 3d talking head. In ECCV, 2024. 2, 8, 5

  20. [28]

    Headnerf: A real-time nerf-based parametric head model

    Yang Hong, Bo Peng, Haiyao Xiao, Ligang Liu, and Juyong Zhang. Headnerf: A real-time nerf-based parametric head model. In CVPR, 2022. 3

  21. [29]

    Perceptual losses for real-time style transfer and super-resolution

    Justin Johnson, Alexandre Alahi, and Li Fei-Fei. Perceptual losses for real-time style transfer and super-resolution. In ECCV, 2016. 3 9

  22. [30]

    Perceptual losses for real-time style transfer and super-resolution

    Justin Johnson, Alexandre Alahi, and Li Fei-Fei. Perceptual losses for real-time style transfer and super-resolution. In ECCV, pages 694–711, 2016. 1

  23. [31]

    A style-based generator architecture for generative adversarial networks

    Tero Karras, Samuli Laine, and Timo Aila. A style-based generator architecture for generative adversarial networks. In CVPR, 2019. 3

  24. [32]

    Zhanghan Ke, Jiayu Sun, Kaican Li, Qiong Yan, and Ryn- son W.H. Lau. Modnet: Real-time trimap-free portrait mat- ting via objective decomposition. In AAAI, 2022. 6

  25. [33]

    3d gaussian splatting for real-time radiance field rendering

    Bernhard Kerbl, Georgios Kopanas, Thomas Leimk ¨uhler, and George Drettakis. 3d gaussian splatting for real-time radiance field rendering. ACM TOG, 42(4), 2023. 1, 4

  26. [34]

    Davis E. King. Dlib - a toolkit for machine learning and computer vision, 2009. 6, 2

  27. [35]

    Kingma and Jimmy Ba

    Diederik P. Kingma and Jimmy Ba. Adam: A method for stochastic optimization, 2017. 1

  28. [36]

    Nersemble: Multi-view radi- ance field reconstruction of human heads

    Tobias Kirschstein, Shenhan Qian, Simon Giebenhain, Tim Walter, and Matthias Nießner. Nersemble: Multi-view radi- ance field reconstruction of human heads. ACM TOG, 42(4),

  29. [37]

    Gghead: Fast and generalizable 3d gaussian heads.ACM SIGGRAPH Asia,

    Tobias Kirschstein, Simon Giebenhain, Jiapeng Tang, Markos Georgopoulos, and Matthias Nießner. Gghead: Fast and generalizable 3d gaussian heads.ACM SIGGRAPH Asia,

  30. [38]

    Spherehead: Stable 3d full-head synthesis with spherical tri-plane representation,

    Heyuan Li, Ce Chen, Tianhao Shi, Yuda Qiu, Sizhe An, Guanying Chen, and Xiaoguang Han. Spherehead: Stable 3d full-head synthesis with spherical tri-plane representation,

  31. [39]

    Uravatar: Universal relightable gaussian codec avatars

    Junxuan Li, Chen Cao, Gabriel Schwartz, Rawal Khirodkar, Christian Richardt, Tomas Simon, Yaser Sheikh, and Shun- suke Saito. Uravatar: Universal relightable gaussian codec avatars. In ACM SIGGRAPH Asia, 2024. 5

  32. [40]

    Tianye Li, Timo Bolkart, Michael. J. Black, Hao Li, and Javier Romero. Learning a model of facial shape and ex- pression from 4D scans. ACM TOG, 36(6):194:1–194:17,

  33. [41]

    3d gaussian blendshapes for head avatar animation

    Shengjie Ma, Yanlin Weng, Tianjia Shao, and Kun Zhou. 3d gaussian blendshapes for head avatar animation. In ACM SIGGRAPH 2024 Conference Papers, 2024. 3

  34. [42]

    Instant neural graphics primitives with a mul- tiresolution hash encoding

    Thomas M ¨uller, Alex Evans, Christoph Schied, and Alexan- der Keller. Instant neural graphics primitives with a mul- tiresolution hash encoding. ACM TOG, 41(4):102:1–102:15,

  35. [43]

    Pytorch: An imperative style, high-performance deep learning library

    Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Rai- son, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, L...

  36. [44]

    Georgios Pavlakos, Vasileios Choutas, Nima Ghorbani, Timo Bolkart, Ahmed A. A. Osman, Dimitrios Tzionas, and Michael J. Black. Expressive body capture: 3D hands, face, and body from a single image. In CVPR, pages 10975– 10985, 2019. 8

  37. [45]

    Gaus- sianavatars: Photorealistic head avatars with rigged 3d gaus- sians

    Shenhan Qian, Tobias Kirschstein, Liam Schoneveld, Davide Davoli, Simon Giebenhain, and Matthias Nießner. Gaus- sianavatars: Photorealistic head avatars with rigged 3d gaus- sians. CVPR, 2024. 2, 3, 7, 8

  38. [46]

    Pivotal tuning for latent-based editing of real im- ages

    Daniel Roich, Ron Mokady, Amit H Bermano, and Daniel Cohen-Or. Pivotal tuning for latent-based editing of real im- ages. ACM TOG, 2021. 6, 1, 2

  39. [47]

    U- net: Convolutional networks for biomedical image segmen- tation

    Olaf Ronneberger, Philipp Fischer, and Thomas Brox. U- net: Convolutional networks for biomedical image segmen- tation. In Medical Image Computing and Computer-Assisted Intervention – MICCAI 2015 , pages 234–241, Cham, 2015. Springer International Publishing. 5, 2

  40. [48]

    Relightable gaussian codec avatars

    Shunsuke Saito, Gabriel Schwartz, Tomas Simon, Junxuan Li, and Giljoo Nam. Relightable gaussian codec avatars. In CVPR, 2024. 2, 3

  41. [49]

    Graf: Generative radiance fields for 3d-aware image synthesis

    Katja Schwarz, Yiyi Liao, Michael Niemeyer, and Andreas Geiger. Graf: Generative radiance fields for 3d-aware image synthesis. In NeurIPS, 2020. 3

  42. [50]

    SplattingAvatar: Realistic Real-Time Human Avatars with Mesh-Embedded Gaussian Splatting

    Zhijing Shao, Zhaolong Wang, Zhuang Li, Duotun Wang, Xiangru Lin, Yu Zhang, Mingming Fan, and Zeyu Wang. SplattingAvatar: Realistic Real-Time Human Avatars with Mesh-Embedded Gaussian Splatting. In CVPR, 2024. 2, 3, 7, 8

  43. [51]

    Texttoon: Real-time text toonify head avatar from single video

    Luchuan Song, Lele Chen, Celong Liu, Pinxin Liu, and Chenliang Xu. Texttoon: Real-time text toonify head avatar from single video. In ACM SIGGRAPH Asia, 2024. 2

  44. [52]

    Tri 2-plane: V olumetric avatar reconstruction with feature pyramid

    Luchuan Song, Pinxin Liu, Lele Chen, Guojun Yin, and Chenliang Xu. Tri 2-plane: V olumetric avatar reconstruction with feature pyramid. ECCV, 2024. 2

  45. [53]

    Next3d: Genera- tive neural texture rasterization for 3d-aware head avatars

    Jingxiang Sun, Xuan Wang, Lizhen Wang, Xiaoyu Li, Yong Zhang, Hongwen Zhang, and Yebin Liu. Next3d: Genera- tive neural texture rasterization for 3d-aware head avatars. In CVPR, 2023. 3

  46. [54]

    Morf: Morphable radiance fields for multiview neural head modeling

    Daoye Wang, Prashanth Chandran, Gaspard Zoss, Derek Bradley, and Paulo Gotardo. Morf: Morphable radiance fields for multiview neural head modeling. In ACM SIG- GRAPH 2022 Conference Proceedings , New York, NY , USA, 2022. Association for Computing Machinery. 3

  47. [55]

    Uar-nvc: A unified autoregressive framework for memory-efficient neural video compression, 2025

    Jia Wang, Xinfeng Zhang, Gai Zhang, Jun Zhu, Lv Tang, and Li Zhang. Uar-nvc: A unified autoregressive framework for memory-efficient neural video compression, 2025. 2

  48. [56]

    To- wards real-world blind face restoration with generative facial prior

    Xintao Wang, Yu Li, Honglun Zhang, and Ying Shan. To- wards real-world blind face restoration with generative facial prior. In CVPR, 2021. 6

  49. [57]

    High-fidelity 3d face gener- ation from natural language descriptions

    Menghua Wu, Hao Zhu, Linjia Huang, Yiyu Zhuang, Yuanxun Lu, and Xun Cao. High-fidelity 3d face gener- ation from natural language descriptions. In Proceedings of IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023. 3

  50. [58]

    Gaussian head & shoulders: High fidelity neural upper body avatars with anchor gaussian guided texture warping, 2024

    Tianhao Wu, Jing Yang, Zhilin Guo, Jingyi Wan, Fangcheng Zhong, and Cengiz Oztireli. Gaussian head & shoulders: High fidelity neural upper body avatars with anchor gaussian guided texture warping, 2024. 4

  51. [59]

    Flashavatar: High-fidelity head avatar with efficient gaussian embedding

    Jun Xiang, Xuan Gao, Yudong Guo, and Juyong Zhang. Flashavatar: High-fidelity head avatar with efficient gaussian embedding. In CVPR, 2024. 2, 3, 7, 8, 1 10

  52. [60]

    Physgaussian: Physics- integrated 3d gaussians for generative dynamics

    Tianyi Xie, Zeshun Zong, Yuxing Qiu, Xuan Li, Yutao Feng, Yin Yang, and Chenfanfu Jiang. Physgaussian: Physics- integrated 3d gaussians for generative dynamics. arXiv preprint arXiv:2311.12198, 2023. 6

  53. [61]

    Avatarmav: Fast 3d head avatar reconstruction using motion-aware neural voxels

    Yuelang Xu, Lizhen Wang, Xiaochen Zhao, Hongwen Zhang, and Yebin Liu. Avatarmav: Fast 3d head avatar reconstruction using motion-aware neural voxels. In ACM SIGGRAPH 2023 Conference Proceedings, 2023. 2

  54. [62]

    Gaussian head avatar: Ultra high-fidelity head avatar via dynamic gaussians

    Yuelang Xu, Benwang Chen, Zhe Li, Hongwen Zhang, Lizhen Wang, Zerong Zheng, and Yebin Liu. Gaussian head avatar: Ultra high-fidelity head avatar via dynamic gaussians. In CVPR, 2024. 2, 3

  55. [63]

    Dialoguenerf: towards realistic avatar face-to- face conversation video generation.Visual Intelligence, 2(1): 24, 2024

    Yichao Yan, Zanwei Zhou, Zi Wang, Jingnan Gao, and Xi- aokang Yang. Dialoguenerf: towards realistic avatar face-to- face conversation video generation.Visual Intelligence, 2(1): 24, 2024. 3

  56. [64]

    Towards practical capture of high-fidelity relightable avatars

    Haotian Yang, Mingwu Zheng, Wanquan Feng, Haibin Huang, Yu-Kun Lai, Pengfei Wan, Zhongyuan Wang, and Chongyang Ma. Towards practical capture of high-fidelity relightable avatars. In ACM SIGGRAPH Asia, pages 1–11,

  57. [65]

    Bisenet: Bilateral segmentation network for real-time semantic segmentation

    Changqian Yu, Jingbo Wang, Chao Peng, Changxin Gao, Gang Yu, and Nong Sang. Bisenet: Bilateral segmentation network for real-time semantic segmentation. InECCV, page 334–349. Springer-Verlag, 2018. 1

  58. [66]

    Efros, Eli Shecht- man, and Oliver Wang

    Richard Zhang, Phillip Isola, Alexei A. Efros, Eli Shecht- man, and Oliver Wang. The unreasonable effectiveness of deep features as a perceptual metric. In CVPR, 2018. 8

  59. [67]

    Realcompo: Balancing realism and com- positionality improves text-to-image diffusion models.arXiv preprint arXiv:2402.12908, 2024

    Xinchen Zhang, Ling Yang, Yaqi Cai, Zhaochen Yu, Kaini Wang, Jiake Xie, Ye Tian, Minkai Xu, Yong Tang, Yujiu Yang, and Bin Cui. Realcompo: Balancing realism and com- positionality improves text-to-image diffusion models.arXiv preprint arXiv:2402.12908, 2024. 2

  60. [68]

    Iter- comp: Iterative composition-aware feedback learning from model gallery for text-to-image generation

    Xinchen Zhang, Ling Yang, Guohao Li, Yaqi Cai, Jiake Xie, Yong Tang, Yujiu Yang, Mengdi Wang, and Bin Cui. Iter- comp: Iterative composition-aware feedback learning from model gallery for text-to-image generation. arXiv preprint arXiv:2410.07171, 2024. 2

  61. [69]

    Psavatar: A point-based shape model for real- time head avatar animation with 3d gaussian splatting, 2024

    Zhongyuan Zhao, Zhenyu Bao, Qing Li, Guoping Qiu, and Kanglin Liu. Psavatar: A point-based shape model for real- time head avatar animation with 3d gaussian splatting, 2024. 2, 3, 4

  62. [70]

    B¨uhler, Xu Chen, Michael J

    Yufeng Zheng, Victoria Fern ´andez Abrevaya, Marcel C. B¨uhler, Xu Chen, Michael J. Black, and Otmar Hilliges. I M Avatar: Implicit morphable head avatars from videos. In CVPR, 2022. 2, 8, 1

  63. [71]

    Black, and Otmar Hilliges

    Yufeng Zheng, Wang Yifan, Gordon Wetzstein, Michael J. Black, and Otmar Hilliges. Pointavatar: Deformable point- based head avatars from videos. In CVPR, 2023. 8

  64. [72]

    Facescape: 3d facial dataset and bench- mark for single-view 3d face reconstruction

    Hao Zhu, Haotian Yang, Longwei Guo, Yidi Zhang, Yanru Wang, Mingkai Huang, Menghua Wu, Qiu Shen, Ruigang Yang, and Xun Cao. Facescape: 3d facial dataset and bench- mark for single-view 3d face reconstruction. IEEE TPAMI,

  65. [73]

    Mofanerf: Morphable facial neural radiance field

    Yiyu Zhuang, Hao Zhu, Xusen Sun, and Xun Cao. Mofanerf: Morphable facial neural radiance field. In ECCV, 2022. 3

  66. [74]

    Neai: A pre-convoluted representation for plug-and-play neural ambient illumina- tion

    Yiyu Zhuang, Qi Zhang, Xuan Wang, Hao Zhu, Ying Feng, Xiaoyu Li, Ying Shan, and Xun Cao. Neai: A pre-convoluted representation for plug-and-play neural ambient illumina- tion. arXiv preprint arXiv:2304.08757, 2023. 2

  67. [75]

    To- wards native generative model for 3d head avatar, 2024

    Yiyu Zhuang, Yuxiao He, Jiawei Zhang, Yanwen Wang, Ji- ahe Zhu, Yao Yao, Siyu Zhu, Xun Cao, and Hao Zhu. To- wards native generative model for 3d head avatar, 2024. 3, 5

  68. [76]

    Idol: Instant photorealistic 3d human creation from a single image

    Yiyu Zhuang, Jiaxi Lv, Hao Wen, Qing Shuai, Ailing Zeng, Hao Zhu, Shifeng Chen, Yujiu Yang, Xun Cao, and Wei Liu. Idol: Instant photorealistic 3d human creation from a single image. arXiv preprint arXiv:2412.14963, 2024. 5

  69. [77]

    Instant volumetric head avatars

    Wojciech Zielonka, Timo Bolkart, and Justus Thies. Instant volumetric head avatars. CVPR, pages 4574–4584, 2022. 2, 8

  70. [78]

    Towards metrical reconstruction of human faces

    Wojciech Zielonka, Timo Bolkart, and Justus Thies. Towards metrical reconstruction of human faces. In ECCV, 2022. 1, 2, 8

  71. [79]

    Ewa splatting

    Matthias Zwicker, Hanspeter Pfister, Jeroen van Baar, and Markus Gross. Ewa splatting. IEEE TVCG, 8(3):223–238,

  72. [81]

    To ensure that Σ is positive semi-definite, the covari- ance matrix is further decomposed into a rotation matrix R and a scaling matrix S: Σ = RSST RT

    Preliminary 3D Gaussian Splatting 3D Gaussian Splatting [33] is a point-based volume rendering method that models each primitive as a Gaussian kernel, formalized as follows: G (x) = e− 1 2 (x−µ)T Σ−1(x−µ), (16) where µ is Gaussian position and Σ is 3D covariance ma- trix. To e...

  73. [82]

    Datasets We used a total of 20 monocular portrait videos for our ex- periments

    Implementation Details 7.1. Datasets We used a total of 20 monocular portrait videos for our ex- periments. For 10 datasets with DECA-based preprocess- ing, we optimize the DECA-predicted FLAME coefficients during training and testing in line with IMAvatar [70]. For the test-t...

  74. [84]

    Monocular Results We provide the quantitative results for each subject in Tab

    Additional Results 8.1. Monocular Results We provide the quantitative results for each subject in Tab. 8 and Tab. 9, and more qualitative results are presented in Fig. 16. Our method demonstrates superior performance across multiple datasets. Other methods, such as FlashAvatar...

  75. [85]

    8, 9, neural baking causes certain metric degradation compared to the avatars optimized in a point-wise manner

    Neural Baking Trade-off As reported in the main content and Tab. 8, 9, neural baking causes certain metric degradation compared to the avatars optimized in a point-wise manner. We found that this is because convolutional neural networks (CNN) struggle to fit the complex distri...

  76. [86]

    As mentioned in Sec

    Failure Case and Limitation Our neural baking and full-head completion still have lim- itations. As mentioned in Sec. 9, since CNN is tricky to construct Gaussian geometry, neural baking may fail for in- tricate geometry. For instance, in the case of the woman 4 Figure 13. Neu...

  77. [87]

    We further evaluate the differences be- tween our method and GaussianAvatars when the camera translation is imperfect

    Noisy Pose Simulation To train head avatars from monocular videos, we require frame-by-frame RGB images along with the corresponding tracked coefficients. We further evaluate the differences be- tween our method and GaussianAvatars when the camera translation is imperfect. We ...

  78. [88]

    6, we supplement the training time and rendering FPS under identical hardware conditions

    Computational Efficiency In Tab. 6, we supplement the training time and rendering FPS under identical hardware conditions. Our method out- performs other UV space-based methods (FA, SA) regard- ing shorter training time and higher FPS. Compared to GA, our method achieves compa...

  79. [89]

    7, we introduce more ablation settings on two representative datasets and further report the number of Gaussian in different settings

    More Ablations As shown in Tab. 7, we introduce more ablation settings on two representative datasets and further report the number of Gaussian in different settings. We additionally conduct ex- periments as w/o densify∗, where ∗ indicates that Gaussians are removed based on o...

  80. [90]

    Data from consenting subjects will be made publicly available

    Ethics We used four subjects from EmoTalk3D [27], with all par- ticipants signing the consent for using their videos in this re- search and publication. Data from consenting subjects will be made publicly available. Our method generates realistic and animatable head avatars, e...

  81. [512]

    The decoder then reduces the number of chan- nels back to 64, and the final convolutional layer adjusts the output channels to 11

    The first convolutional layer increases the number of channels to 64, and the encoder of the U-Net processes the channels up to 1024, doubling the number of channels at each layer. The decoder then reduces the number of chan- nels back to 64, and the final convolutional layer ...

  82. [2002]

    1 11 FATE: Full-head Gaussian Avatar with Textural Editing from Monocular Video Supplementary Material This supplementary material provides additional imple- mentation details and experimental results. In Sec. 6, we introduce the preliminaries related to 3DGS and PTI. Sec. 7 d...

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.