REVIEW 13 cited by
VASA-1: Lifelike Audio-Driven Talking Faces Generated in Real Time
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We introduce VASA, a framework for generating lifelike talking faces with appealing visual affective skills (VAS) given a single static image and a speech audio clip. Our premiere model, VASA-1, is capable of not only generating lip movements that are exquisitely synchronized with the audio, but also producing a large spectrum of facial nuances and natural head motions that contribute to the perception of authenticity and liveliness. The core innovations include a holistic facial dynamics and head movement generation model that works in a face latent space, and the development of such an expressive and disentangled face latent space using videos. Through extensive experiments including evaluation on a set of new metrics, we show that our method significantly outperforms previous methods along various dimensions comprehensively. Our method not only delivers high video quality with realistic facial and head dynamics but also supports the online generation of 512x512 videos at up to 40 FPS with negligible starting latency. It paves the way for real-time engagements with lifelike avatars that emulate human conversational behaviors.
Forward citations
Cited by 13 Pith papers
-
ViDS: Video Diffusion Shader using 3D Face Tracking
Fine-tuning a video diffusion model on dense 3DMM normal maps from Pixel3DMM lets a single portrait photo be animated with a driving video's expressions and pose, surpassing landmark- and latent-based portrait animati...
-
Conversational Human Audio-visual Talking Dialogue Generation
CHAT generates mutually responsive dyadic audio-visual dialogue clips from a single text prompt and yields a 50k synthetic pre-training set that improves facial reaction models on REACT 2024.
-
Identity Deepfake Threats to Biometric Authentication Systems: Public and Expert Perspectives
A UK-based mixed-methods study finds a wide gap between public trust in biometrics and expert concern about deepfake spoofing, and proposes a tri-layer mitigation framework.
-
Low-Rank Head Avatar Personalization with Registers
A Register Module, a learnable 3D feature space rigged to a 3DMM mesh, improves LoRA-based personalization of head avatars by teaching the model to focus on identity-specific DINOv2 features during adaptation.
-
Silence is Golden: Leveraging Adversarial Examples to Nullify Audio Control in LDM-based Talking-Head Generation
Silencer adds a nearly invisible disturbance to portraits that makes LDM-based talking-head models keep the mouth silent, and it survives several image-purification countermeasures.
-
Exploring Timeline Control for Facial Motion Generation
A diffusion model generates natural facial motions from user-specified multi-track timelines, using TICC-based frame-level action interval annotation for training and evaluation.
-
Pippo: High-Resolution Multi-View Humans from a Single Image
A single-image multi-view diffusion transformer generates 1K-resolution turnaround views of humans, with attention biasing for many views and a new reprojection-error metric.
-
EDTalk++: Full Disentanglement for Controllable Talking Head Synthesis
EDTalk++ disentangles talking-head video into four orthogonal motion banks (mouth, pose, eyes, expression) and drives them from either video or audio inputs.
-
MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation
MoDiT, a diffusion transformer conditioned on 3DMM coefficients and Wav2Lip references, produces talking-head videos with improved same-identity lip sync and more natural blinks in its reported benchmarks.
-
FixTalk: Taming Identity Leakage for High-Quality Talking Head Generation in Extreme Cases
FixTalk adds two modules to a real-time GAN talking-head model, decoupling identity from motion to stop identity leakage while using a memory to recover details and reduce artifacts.
-
Multi-View Face and Gesture Animation with Dynamic Gaussians
Combining separate face and hand models with a parametric body and Gaussian splatting enables multi-view-consistent upper-body avatars that can be re-animated with new expressions and gestures.
-
Wan-Streamer v0.2: Higher Resolution, Same Latency
Wan-Streamer v0.2 upgrades native-streaming audio-visual interaction to 640×368 at 25 FPS with unchanged ~200 ms model-side latency via a single-GPU thinker and multi-GPU Ulysses-style performer.
-
Human Motion Video Generation: A Survey
A comprehensive survey with a five-phase pipeline model for human motion video generation, covering over 200 papers and adding a new benchmark comparison of nine pose-guided methods.
Discussion (0). Continue with ORCID to comment.