REVIEW 5 cited by
Audio2Head: Audio-driven One-shot Talking-head Generation with Natural Head Motion
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We propose an audio-driven talking-head method to generate photo-realistic talking-head videos from a single reference image. In this work, we tackle two key challenges: (i) producing natural head motions that match speech prosody, and (ii) maintaining the appearance of a speaker in a large head motion while stabilizing the non-face regions. We first design a head pose predictor by modeling rigid 6D head movements with a motion-aware recurrent neural network (RNN). In this way, the predicted head poses act as the low-frequency holistic movements of a talking head, thus allowing our latter network to focus on detailed facial movement generation. To depict the entire image motions arising from audio, we exploit a keypoint based dense motion field representation. Then, we develop a motion field generator to produce the dense motion fields from input audio, head poses, and a reference image. As this keypoint based representation models the motions of facial regions, head, and backgrounds integrally, our method can better constrain the spatial and temporal consistency of the generated videos. Finally, an image generation network is employed to render photo-realistic talking-head videos from the estimated keypoint based motion fields and the input reference image. Extensive experiments demonstrate that our method produces videos with plausible head motions, synchronized facial expressions, and stable backgrounds and outperforms the state-of-the-art.
Forward citations
Cited by 5 Pith papers
-
Proxy Avatar Meets Low-Rank Caching: Real-Time One-Shot Emotion-Controllable Portrait Animation
A cascade of a single-identity Gaussian emotion proxy, a one-shot diffusion retargeting model, and low-rank appearance caching enables real-time one-shot portrait animation with emotion control.
-
Who is a Better Talker: Subjective and Objective Quality Assessment for AI-Generated Talking Heads
A new large dataset and the FSCD model improve automated quality scoring of AI-generated talking-head videos, beating 15 baselines in correlation with human ratings.
-
Long-Term TalkingFace Generation via Motion-Prior Conditional Diffusion Model
A motion-prior diffusion model with archived-frame memory improves identity, lip-sync, and head-motion consistency in long talking-face videos.
-
SyncTalk++: High-Fidelity and Efficient Synchronized Talking Heads Synthesis Using Gaussian Splatting
SyncTalk++ synthesizes speech-driven talking-head videos via 3D Gaussian Splatting and reports state-of-the-art synchronization and quality at up to 101 FPS.
-
NTIRE 2025 XGC Quality Assessment Challenge: Methods and Results
All 19 valid entries in the NTIRE 2025 XGC quality assessment challenge outperformed their track baselines at predicting human quality scores for user-generated video, AI-generated video, and talking heads.
Discussion (0). Continue with ORCID to comment.