Pith. sign in

REVIEW 8 cited by

Real-time Neural Radiance Talking Portrait Synthesis via Audio-spatial Decomposition

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2211.12368 v1 pith:7353FRLU submitted 2022-11-22 cs.CV

classification cs.CV
keywords talkinggridportraitaudio-spatialdynamicefficientmodulenerf
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

While dynamic Neural Radiance Fields (NeRF) have shown success in high-fidelity 3D modeling of talking portraits, the slow training and inference speed severely obstruct their potential usage. In this paper, we propose an efficient NeRF-based framework that enables real-time synthesizing of talking portraits and faster convergence by leveraging the recent success of grid-based NeRF. Our key insight is to decompose the inherently high-dimensional talking portrait representation into three low-dimensional feature grids. Specifically, a Decomposed Audio-spatial Encoding Module models the dynamic head with a 3D spatial grid and a 2D audio grid. The torso is handled with another 2D grid in a lightweight Pseudo-3D Deformable Module. Both modules focus on efficiency under the premise of good rendering quality. Extensive experiments demonstrate that our method can generate realistic and audio-lips synchronized talking portrait videos, while also being highly efficient compared to previous methods.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ViDS: Video Diffusion Shader using 3D Face Tracking

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Fine-tuning a video diffusion model on dense 3DMM normal maps from Pixel3DMM lets a single portrait photo be animated with a driving video's expressions and pose, surpassing landmark- and latent-based portrait animati...

  2. HM-Talker: Hybrid Motion Modeling for High-Fidelity Talking Head Synthesis

    cs.CV 2025-08 conditional novelty 6.0 of 10

    HM-Talker merges explicit facial-structure cues with implicit audio-driven motion features via stochastic feature pairing to improve talking-head realism and lip-sync.

  3. Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation

    cs.CV 2025-07 conditional novelty 6.0 of 10

    MF-Talk, a mask-free and identity-reference-free three-stage pipeline, improves visual quality and identity preservation in talking-face generation while remaining competitive on lip-sync.

  4. D^3-Talker: Dual-Branch Decoupled Deformation Fields for Few-Shot 3D Talking Head Synthesis

    cs.CV 2025-08 conditional novelty 5.0 of 10

    D3-Talker splits face deformation into a general, identity-agnostic branch driven by a facial motion prior and a personalized branch driven by audio, improving few-shot talking head lip sync.

  5. M2DAO-Talker: Harmonizing Multi-granular Motion Decoupling and Alternating Optimization for Talking-head Generation

    cs.CV 2025-07 conditional novelty 5.0 of 10

    M2DAO-Talker decouples rigid, facial, and oral motion in 3D Gaussian Splatting and uses alternating optimization to reach state-of-the-art scores on small talking-head benchmarks.

  6. Few-Shot Identity Adaptation for 3D Talking Heads via Global Gaussian Field

    cs.CV 2025-06 conditional novelty 5.0 of 10

    A shared global Gaussian field plus identity embeddings lets a 3D talking head model adapt to new speakers with a few seconds of footage while improving quality over prior per-identity models.

  7. GGTalker: Talking Head Systhesis with Generalizable Gaussian Priors and Identity-Specific Adaptation

    cs.CV 2025-06 conditional novelty 5.0 of 10

    GGTalker combines large-scale audio-to-expression and expression-to-texture priors with rapid per-identity fine-tuning to create high-quality 3D talking heads from a short video.

  8. SyncTalk++: High-Fidelity and Efficient Synchronized Talking Heads Synthesis Using Gaussian Splatting

    cs.CV 2025-06 conditional novelty 5.0 of 10

    SyncTalk++ synthesizes speech-driven talking-head videos via 3D Gaussian Splatting and reports state-of-the-art synchronization and quality at up to 101 FPS.

Pith tools