REVIEW 3 cited by
HeadStudio: Text to Animatable Head Avatars with 3D Gaussian Splatting
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
abstract
Creating digital avatars from textual prompts has long been a desirable yet challenging task. Despite the promising results achieved with 2D diffusion priors, current methods struggle to create high-quality and consistent animated avatars efficiently. Previous animatable head models like FLAME have difficulty in accurately representing detailed texture and geometry. Additionally, high-quality 3D static representations face challenges in semantically driving with dynamic priors. In this paper, we introduce \textbf{HeadStudio}, a novel framework that utilizes 3D Gaussian splatting to generate realistic and animatable avatars from text prompts. Firstly, we associate 3D Gaussians with animatable head prior model, facilitating semantic animation on high-quality 3D representations. To ensure consistent animation, we further enhance the optimization from initialization, distillation, and regularization to jointly learn the shape, texture, and animation. Extensive experiments demonstrate the efficacy of HeadStudio in generating animatable avatars from textual prompts, exhibiting appealing appearances. The avatars are capable of rendering high-quality real-time ($\geq 40$ fps) novel views at a resolution of 1024. Moreover, These avatars can be smoothly driven by real-world speech and video. We hope that HeadStudio can enhance digital avatar creation and gain popularity in the community. Code is at: https://github.com/ZhenglinZhou/HeadStudio.
Forward citations
Cited by 3 Pith papers
-
GeoAvatar: Adaptive Geometrical Gaussian Splatting for 3D Head Avatar
GeoAvatar improves 3D head avatar quality by adaptively regulating Gaussian offsets per facial region, adding a detailed mouth structure with part-wise deformation, and releasing a new expressive monocular dataset, Dy...
-
Category-Aware 3D Object Composition with Disentangled Texture and Shape Multi-view Diffusion
C33D blends a 3D model with an object category by generating a fused front view, then using texture and shape multi-view diffusion plus adaptive inversion to reconstruct a novel, consistent 3D model.
-
PlantDreamer: Achieving Realistic 3D Plant Models with Diffusion-Guided Gaussian Splatting
A diffusion-guided Gaussian splatting pipeline generates realistic 3D plants from L-system meshes or point clouds and beats GaussianDreamer on masked PSNR for bean, kale and mint.
Discussion (0). Continue with ORCID to comment.