Pith. sign in

REVIEW 1 cited by

Talking-head Generation with Rhythmic Head Motion

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2007.08547 v1 pith:GOPKVDLN submitted 2020-07-16 cs.CV cs.GR

classification cs.CVcs.GR
keywords headmotionvideoachievesembeddinggeneratemodulemovements
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

When people deliver a speech, they naturally move heads, and this rhythmic head motion conveys prosodic information. However, generating a lip-synced video while moving head naturally is challenging. While remarkably successful, existing works either generate still talkingface videos or rely on landmark/video frames as sparse/dense mapping guidance to generate head movements, which leads to unrealistic or uncontrollable video synthesis. To overcome the limitations, we propose a 3D-aware generative network along with a hybrid embedding module and a non-linear composition module. Through modeling the head motion and facial expressions1 explicitly, manipulating 3D animation carefully, and embedding reference images dynamically, our approach achieves controllable, photo-realistic, and temporally coherent talking-head videos with natural head movements. Thoughtful experiments on several standard benchmarks demonstrate that our method achieves significantly better results than the state-of-the-art methods in both quantitative and qualitative comparisons. The code is available on https://github.com/ lelechen63/Talking-head-Generation-with-Rhythmic-Head-Motion.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. GaussianSpeech: Audio-Driven Gaussian Avatars

    cs.CV 2024-11 conditional novelty 6.0 of 10

    A transformer-based sequence model drives a lightweight 3D Gaussian avatar from audio, producing synchronized, photorealistic talking-head animations with a new 16-camera dataset.

Pith tools