REVIEW 8 cited by
Real-time Neural Radiance Talking Portrait Synthesis via Audio-spatial Decomposition
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
While dynamic Neural Radiance Fields (NeRF) have shown success in high-fidelity 3D modeling of talking portraits, the slow training and inference speed severely obstruct their potential usage. In this paper, we propose an efficient NeRF-based framework that enables real-time synthesizing of talking portraits and faster convergence by leveraging the recent success of grid-based NeRF. Our key insight is to decompose the inherently high-dimensional talking portrait representation into three low-dimensional feature grids. Specifically, a Decomposed Audio-spatial Encoding Module models the dynamic head with a 3D spatial grid and a 2D audio grid. The torso is handled with another 2D grid in a lightweight Pseudo-3D Deformable Module. Both modules focus on efficiency under the premise of good rendering quality. Extensive experiments demonstrate that our method can generate realistic and audio-lips synchronized talking portrait videos, while also being highly efficient compared to previous methods.
Forward citations
Cited by 8 Pith papers
-
ViDS: Video Diffusion Shader using 3D Face Tracking
Fine-tuning a video diffusion model on dense 3DMM normal maps from Pixel3DMM lets a single portrait photo be animated with a driving video's expressions and pose, surpassing landmark- and latent-based portrait animati...
-
HM-Talker: Hybrid Motion Modeling for High-Fidelity Talking Head Synthesis
HM-Talker merges explicit facial-structure cues with implicit audio-driven motion features via stochastic feature pairing to improve talking-head realism and lip-sync.
-
Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation
MF-Talk, a mask-free and identity-reference-free three-stage pipeline, improves visual quality and identity preservation in talking-face generation while remaining competitive on lip-sync.
-
D^3-Talker: Dual-Branch Decoupled Deformation Fields for Few-Shot 3D Talking Head Synthesis
D3-Talker splits face deformation into a general, identity-agnostic branch driven by a facial motion prior and a personalized branch driven by audio, improving few-shot talking head lip sync.
-
M2DAO-Talker: Harmonizing Multi-granular Motion Decoupling and Alternating Optimization for Talking-head Generation
M2DAO-Talker decouples rigid, facial, and oral motion in 3D Gaussian Splatting and uses alternating optimization to reach state-of-the-art scores on small talking-head benchmarks.
-
Few-Shot Identity Adaptation for 3D Talking Heads via Global Gaussian Field
A shared global Gaussian field plus identity embeddings lets a 3D talking head model adapt to new speakers with a few seconds of footage while improving quality over prior per-identity models.
-
GGTalker: Talking Head Systhesis with Generalizable Gaussian Priors and Identity-Specific Adaptation
GGTalker combines large-scale audio-to-expression and expression-to-texture priors with rapid per-identity fine-tuning to create high-quality 3D talking heads from a short video.
-
SyncTalk++: High-Fidelity and Efficient Synchronized Talking Heads Synthesis Using Gaussian Splatting
SyncTalk++ synthesizes speech-driven talking-head videos via 3D Gaussian Splatting and reports state-of-the-art synchronization and quality at up to 101 FPS.
Discussion (0). Continue with ORCID to comment.