REVIEW 12 cited by
Drivable 3D Gaussian Avatars
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We present Drivable 3D Gaussian Avatars (D3GA), a multi-layered 3D controllable model for human bodies that utilizes 3D Gaussian primitives embedded into tetrahedral cages. The advantage of using cages compared to commonly employed linear blend skinning (LBS) is that primitives like 3D Gaussians are naturally re-oriented and their kernels are stretched via the deformation gradients of the encapsulating tetrahedron. Additional offsets are modeled for the tetrahedron vertices, effectively decoupling the low-dimensional driving poses from the extensive set of primitives to be rendered. This separation is achieved through the localized influence of each tetrahedron on 3D Gaussians, resulting in improved optimization. Using the cage-based deformation model, we introduce a compositional pipeline that decomposes an avatar into layers, such as garments, hands, or faces, improving the modeling of phenomena like garment sliding. These parts can be conditioned on different driving signals, such as keypoints for facial expressions or joint-angle vectors for garments and the body. Our experiments on two multi-view datasets with varied body shapes, clothes, and motions show higher-quality results. They surpass PSNR and SSIM metrics of other SOTA methods using the same data while offering greater flexibility and compactness.
Forward citations
Cited by 12 Pith papers
-
Instant Expressive Gaussian Head Avatars at Over 100 FPS
A single-photo avatar encoder with per-Gaussian feature-space deformation animates faces at 107 FPS with expression quality competitive with diffusion models.
-
Pippo: High-Resolution Multi-View Humans from a Single Image
A single-image multi-view diffusion transformer generates 1K-resolution turnaround views of humans, with attention biasing for many views and a new reprojection-error metric.
-
Deblur-Avatar: Animatable Avatars from Motion-Blurred Monocular Videos
Deblur-Avatar reconstructs sharp, animatable human avatars from motion-blurred monocular video by optimizing SMPL start and end poses and averaging rendered virtual frames inside 3D Gaussian Splatting.
-
SqueezeMe: Mobile-Ready Distillation of Gaussian Full-Body Avatars
Pose-corrective networks for Gaussian avatars are distilled into shared linear layers, giving three full-body avatars real-time animation and rendering on a Quest 3 headset.
-
HDGS: Textured 2D Gaussian Splatting for Enhanced Scene Rendering
HDGS improves 2D Gaussian splatting with per-surfel texture maps, per-ray sorting, Fisher pruning, and five-ray frustum sampling for sharper detail and reduced aliasing.
-
One Shot, One Talk: Whole-body Talking Avatar from a Single Image
A single photo becomes an animatable whole-body talking avatar by training a coupled 3DGS-mesh model on diffusion-generated pseudo-videos with perceptual supervision.
-
DevilSight: Augmenting Monocular Human Avatar Reconstruction through a Virtual Perspective
A monocular human avatar reconstruction method generates pseudo back-view videos with a fine-tuned diffusion model and uses them as extra training data for a 3D Gaussian avatar.
-
Wavelet-GS: 3D Gaussian Splatting with Wavelet Decomposition
Wavelet-GS splits a 3D point cloud into low- and high-frequency wavelet parts, trains each with its own strategy, plus a relight module, reporting gains over prior 3DGS variants on four datasets.
-
Sequential Gaussian Avatars with Hierarchical Motion Context
A 3D Gaussian avatar model that conditions non-rigid deformation on hierarchical skeleton and vertex motion reaches state-of-the-art rendering quality on three human-capture datasets.
-
DyGASR: Dynamic Generalized Exponential Splatting with Surface Alignment for Accelerated 3D Mesh Reconstruction
DyGASR reconstructs 3D meshes faster and with less memory by replacing Gaussians with generalized exponential splats, adding SuGaR-style surface alignment, and training at progressively higher resolutions.
-
SAT: Supervisor Regularization and Animation Augmentation for Two-process Monocular Texture 3D Human Reconstruction
A two-stage Gaussian-splatting framework with supervisor feature regularization and online animation augmentation improves monocular textured 3D human reconstruction on CustomHuman and THuman3.0.
-
Exploring Dynamic Novel View Synthesis Technologies for Cinematography
A review of dynamic novel view synthesis for cinematography, accompanied by a self-made montage using Nerfacto, 4D-GS, and SC-GS.
Discussion (0). Continue with ORCID to comment.