Pith. sign in

REVIEW 2 cited by

GSTalker: Real-time Audio-Driven Talking Face Generation via Deformable Gaussian Splatting

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.19040 v1 pith:WCEXBCHJ submitted 2024-04-29 cs.CV

classification cs.CV
keywords gaussiandeformationgstalkertrainingaudio-drivenfieldreal-timerendering
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

We present GStalker, a 3D audio-driven talking face generation model with Gaussian Splatting for both fast training (40 minutes) and real-time rendering (125 FPS) with a 3$\sim$5 minute video for training material, in comparison with previous 2D and 3D NeRF-based modeling frameworks which require hours of training and seconds of rendering per frame. Specifically, GSTalker learns an audio-driven Gaussian deformation field to translate and transform 3D Gaussians to synchronize with audio information, in which multi-resolution hashing grid-based tri-plane and temporal smooth module are incorporated to learn accurate deformation for fine-grained facial details. In addition, a pose-conditioned deformation field is designed to model the stabilized torso. To enable efficient optimization of the condition Gaussian deformation field, we initialize 3D Gaussians by learning a coarse static Gaussian representation. Extensive experiments in person-specific videos with audio tracks validate that GSTalker can generate high-fidelity and audio-lips synchronized results with fast training and real-time rendering speed.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. PD-GS: Phoneme-Driven 3DGS for Audio-Driven Talking Heads

    cs.AI 2026-08 conditional novelty 6.0 of 10

    PD-GS fuses ASR-aligned phoneme tokens with HuBERT audio features via a learned gate in a 3DGS talker, improving lip landmark distance (LMD 2.66 on HDTF) over audio-only baselines.

  2. SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers

    cs.CV 2025-06 conditional novelty 5.0 of 10

    An audio-conditioned video diffusion transformer that animates portraits from image, video, text, and audio inputs with a sliding-window fusion for long videos.

Pith tools