REVIEW 19 cited by
LatentSync: Taming Audio-Conditioned Latent Diffusion Models for Lip Sync with SyncNet Supervision
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
End-to-end audio-conditioned latent diffusion models (LDMs) have been widely adopted for audio-driven portrait animation, demonstrating their effectiveness in generating lifelike and high-resolution talking videos. However, direct application of audio-conditioned LDMs to lip-synchronization (lip-sync) tasks results in suboptimal lip-sync accuracy. Through an in-depth analysis, we identified the underlying cause as the "shortcut learning problem", wherein the model predominantly learns visual-visual shortcuts while neglecting the critical audio-visual correlations. To address this issue, we explored different approaches for integrating SyncNet supervision into audio-conditioned LDMs to explicitly enforce the learning of audio-visual correlations. Since the performance of SyncNet directly influences the lip-sync accuracy of the supervised model, the training of a well-converged SyncNet becomes crucial. We conducted the first comprehensive empirical studies to identify key factors affecting SyncNet convergence. Based on our analysis, we introduce StableSyncNet, with an architecture designed for stable convergence. Our StableSyncNet achieved a significant improvement in accuracy, increasing from 91% to 94% on the HDTF test set. Additionally, we introduce a novel Temporal Representation Alignment (TREPA) mechanism to enhance temporal consistency in the generated videos. Experimental results show that our method surpasses state-of-the-art lip-sync approaches across various evaluation metrics on the HDTF and VoxCeleb2 datasets.
Forward citations
Cited by 19 Pith papers
-
TryOnCrafter: Unleashing Camera Trajectories for Realistic Video Virtual Try-on via a Renderable 4D Try-on Proxy
TryOnCrafter is the first DiT-based framework for camera-controllable video virtual try-on via a renderable 4D try-on proxy distilled from 2D priors into 3DGS avatar animated with SMPL-X.
-
Tell me Habibi, is it Real or Fake?
ArEnAV, the first large-scale Arabic-English code-switched audio-visual deepfake dataset, makes current state-of-the-art detectors fail much more than on monolingual data.
-
MultiRef-Compass: Towards Comprehensive Evaluation of Multi-Reference-to-Audio-Video Generation
MultiRef-Compass is a 350-sample benchmark and 14-metric protocol for multi-reference-to-audio-video generation; current models still fail at reference binding and audio-visual consistency.
-
Conversational Human Audio-visual Talking Dialogue Generation
CHAT generates mutually responsive dyadic audio-visual dialogue clips from a single text prompt and yields a 50k synthetic pre-training set that improves facial reaction models on REACT 2024.
-
SyncBreaker:Stage-Aware Multimodal Adversarial Attacks on Audio-Driven Talking Head Generation
SyncBreaker jointly attacks image and audio streams with Multi-Interval Sampling and Cross-Attention Fooling to degrade speech-driven talking head generation more than single-modality baselines.
-
HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning
HuMo uses a two-stage training scheme and a face-focus trick to generate human videos that follow text, keep a reference person's identity, and sync speech to audio, beating several single-task systems on benchmarks.
-
DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation
A single-DiT portrait animation framework with style and emotion branches plus parallel audio-style cross-attention claims faster, controllable talking-head generation without a Reference Net.
-
JOLT3D: Joint Learning of Talking Heads and 3DMM Parameters with Application to Lip-Sync
JOLT3D jointly trains a 3DMM reconstruction network with a talking head generator, then uses FACS mouth blendshapes from a diffusion model to lip-sync videos while preserving the original chin contour.
-
SpeakerVid-5M: A Large-Scale High-Quality Dataset for Audio-Visual Dyadic Interactive Human Generation
SpeakerVid-5M provides 5.2 million audio-visual human clips (8,743 hours) with rich annotations and a dyadic interaction benchmark for training interactive virtual humans.
-
Video Editing for Audio-Visual Dubbing
EdiDub treats visual dubbing as content-aware video editing: a diffusion model is conditioned on the original video and inverted latents preserve context while re-synchronizing lips to new speech.
-
Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation
MultiTalk is the first framework to generate multi-person conversational videos from multi-stream audio, using Label Rotary Position Embedding to bind each voice to the correct person.
-
HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters
HunyuanVideo-Avatar is an audio-driven video generator that enables emotion-controllable and multi-character animation by injecting character images, routing audio via face masks, and transferring emotion from referen...
-
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey
The paper offers the first focused review of MLLM-based video translation organized by a three-role taxonomy of Semantic Reasoner, Expressive Performer, and Visual Synthesizer, plus open challenges.
-
Few-Shot Identity Adaptation for 3D Talking Heads via Global Gaussian Field
A shared global Gaussian field plus identity embeddings lets a 3D talking head model adapt to new speakers with a few seconds of footage while improving quality over prior per-identity models.
-
SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers
An audio-conditioned video diffusion transformer that animates portraits from image, video, text, and audio inputs with a sliding-window fusion for long videos.
-
FaceEditTalker: Controllable Talking Head Generation with Facial Attribute Editing
A single framework can edit predefined facial attributes in audio-synchronized talking head videos while preserving identity and lip-sync quality.
-
Preview WB-DH: Towards Whole Body Digital Human Bench for the Generation of Whole-body Talking Avatar Videos
The paper previews a claimed 2M-clip multimodal benchmark for whole-body talking avatar video generation, with standard metrics and an initial evaluation of eight open-source models.
-
KLASSify to Verify: Audio-Visual Deepfake Detection Using SSL-based Audio and Handcrafted Visual Features
A challenge entry combining Wav2Vec-AASIST audio scores with lightweight handcrafted-feature video scores via calibration and maxout reports 92.78% AUC on AV-Deepfake1M++ testA.
-
JWB-DH-V1: Benchmark for Joint Whole-Body Talking Avatar and Speech Generation Version 1
A paper announcing a large-scale whole-body talking avatar benchmark and evaluation protocol, but with insufficient details to verify the dataset or the joint audio-video evaluation.
Discussion (0). Continue with ORCID to comment.