REVIEW 32 cited by
Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The field of portrait image animation, driven by speech audio input, has experienced significant advancements in the generation of realistic and dynamic portraits. This research delves into the complexities of synchronizing facial movements and creating visually appealing, temporally consistent animations within the framework of diffusion-based methodologies. Moving away from traditional paradigms that rely on parametric models for intermediate facial representations, our innovative approach embraces the end-to-end diffusion paradigm and introduces a hierarchical audio-driven visual synthesis module to enhance the precision of alignment between audio inputs and visual outputs, encompassing lip, expression, and pose motion. Our proposed network architecture seamlessly integrates diffusion-based generative models, a UNet-based denoiser, temporal alignment techniques, and a reference network. The proposed hierarchical audio-driven visual synthesis offers adaptive control over expression and pose diversity, enabling more effective personalization tailored to different identities. Through a comprehensive evaluation that incorporates both qualitative and quantitative analyses, our approach demonstrates obvious enhancements in image and video quality, lip synchronization precision, and motion diversity. Further visualization and access to the source code can be found at: https://fudan-generative-vision.github.io/hallo.
Forward citations
Cited by 32 Pith papers
-
EmpaAva: An Open-source Agentic 3D-Avatar Empathetic Live Chatbot
EmpaAva is an open-source, LLM-orchestrated 3D avatar chatbot that perceives user affect from speech and video, plans empathetic replies, and delivers them with synchronized emotional speech and facial motion.
-
Geometry-guided Emotion Modulation for Controllable and Photorealistic Emotional Talking Face Generation
A diffusion-based talking-face generator uses 3D blendshape coefficients to continuously control the emotion intensity of generated facial expressions.
-
ViDS: Video Diffusion Shader using 3D Face Tracking
Fine-tuning a video diffusion model on dense 3DMM normal maps from Pixel3DMM lets a single portrait photo be animated with a driving video's expressions and pose, surpassing landmark- and latent-based portrait animati...
-
SyncBreaker:Stage-Aware Multimodal Adversarial Attacks on Audio-Driven Talking Head Generation
SyncBreaker jointly attacks image and audio streams with Multi-Interval Sampling and Cross-Attention Fooling to degrade speech-driven talking head generation more than single-modality baselines.
-
STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits
STARCaster is a 2D spatio-temporal video diffusion model that unifies identity-conditioned audio-driven portrait animation and novel-view synthesis without explicit 3D reconstruction.
-
Text2Lip: Progressive Lip-Synced Talking Face Generation from Text via Viseme-Guided Rendering
A text-only talking face generator maps text to visemes, hallucinates pseudo-audio, and renders lip-synced video; the paper's headline metrics are partially contradicted by its own tables.
-
X-NeMo: Expressive Neural Motion Reenactment via Disentangled Latent Attention
X-NeMo trains a 1D identity-agnostic motion descriptor end-to-end with a diffusion model, enabling zero-shot portrait animation with improved identity and expression fidelity.
-
DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation
A single-DiT portrait animation framework with style and emotion branches plus parallel audio-style cross-attention claims faster, controllable talking-head generation without a Reference Net.
-
JOLT3D: Joint Learning of Talking Heads and 3DMM Parameters with Application to Lip-Sync
JOLT3D jointly trains a 3DMM reconstruction network with a talking head generator, then uses FACS mouth blendshapes from a diffusion model to lip-sync videos while preserving the original chin contour.
-
MagicAnime: A Hierarchically Annotated, Multimodal and Multitasking Dataset with Benchmarks for Cartoon Animation Generation
MagicAnime is a 400k-clip multimodal cartoon dataset with hierarchical annotations and benchmarks for image-to-video, pose-driven, face reenactment, and audio-driven animation generation.
-
Think-Before-Draw: Decomposing Emotion Semantics & Fine-Grained Controllable Expressive Talking Head Generation
A two-stage text-guidance framework, Think-Before-Draw, uses chain-of-thought prompting to convert emotion labels into facial muscle descriptions and then progressively guides a diffusion model from coarse emotion to ...
-
ARIG: Autoregressive Interactive Head Generation for Real-time Conversations
ARIG introduces a real-time, frame-wise autoregressive head generation framework with diffusion-based continuous motion prediction, improving interactive realism over clip-wise methods.
-
Silence is Golden: Leveraging Adversarial Examples to Nullify Audio Control in LDM-based Talking-Head Generation
Silencer adds a nearly invisible disturbance to portraits that makes LDM-based talking-head models keep the mouth silent, and it survives several image-purification countermeasures.
-
Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation
MultiTalk is the first framework to generate multi-person conversational videos from multi-stream audio, using Label Rotary Position Embedding to bind each voice to the correct person.
-
Exploring Timeline Control for Facial Motion Generation
A diffusion model generates natural facial motions from user-specified multi-track timelines, using TICC-based frame-level action interval annotation for training and evaluation.
-
HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters
HunyuanVideo-Avatar is an audio-driven video generator that enables emotion-controllable and multi-character animation by injecting character images, routing audio via face masks, and transferring emotion from referen...
-
Long-Term TalkingFace Generation via Motion-Prior Conditional Diffusion Model
A motion-prior diffusion model with archived-frame memory improves identity, lip-sync, and head-motion consistency in long talking-face videos.
-
TaoMate: Anchor-Guided Memory Bridging Evolving and Reference States for Real-Time Audio-Video Digital Human Generation
Anchor-guided fixed-capacity multimodal memory plus stage-parallel few-step denoising stabilizes long-form causal audio-video digital humans at real-time speed.
-
FA-LAM: Focus-Aware Large Avatar Model for One-Shot 4D Animatable Gaussian Head
A one-shot model for animatable 3D/4D Gaussian head reconstruction that adds attention regularization, decoupled reconstruction-animation training, and autoregressive visibility-gated fusion, reporting consistent metr...
-
InfinityHuman: Towards Long-Term Audio-Driven Human
A coarse-to-fine audio-driven animation framework that uses pose-guided refinement and hand-specific reward learning to generate long, identity-stable talking videos.
-
InfiniteTalk: Audio-driven Video Generation for Sparse-Frame Video Dubbing
Sparse-frame dubbing with adjacent-chunk keyframe sampling lets a streaming audio-video model produce full-body motion synchronized to new audio while preserving identity and camera motion.
-
EDTalk++: Full Disentanglement for Controllable Talking Head Synthesis
EDTalk++ disentangles talking-head video into four orthogonal motion banks (mouth, pose, eyes, expression) and drives them from either video or audio inputs.
-
MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation
MoDiT, a diffusion transformer conditioned on 3DMM coefficients and Wav2Lip references, produces talking-head videos with improved same-identity lip sync and more natural blinks in its reported benchmarks.
-
MoDA: Multi-modal Diffusion Architecture for Talking Head Generation
MoDA uses flow matching in a compact face-motion space with a progressively fused multi-modal transformer to generate expressive, lip-synced talking-head videos from a single image and audio.
-
FixTalk: Taming Identity Leakage for High-Quality Talking Head Generation in Extreme Cases
FixTalk adds two modules to a real-time GAN talking-head model, decoupling identity from motion to stop identity leakage while using a memory to recover details and reduce artifacts.
-
MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation
MirrorMe adapts the LTX video diffusion transformer to generate real-time, high-fidelity audio-driven halfbody animations with identity preservation and hand pose control.
-
GGTalker: Talking Head Systhesis with Generalizable Gaussian Priors and Identity-Specific Adaptation
GGTalker combines large-scale audio-to-expression and expression-to-texture priors with rapid per-identity fine-tuning to create high-quality 3D talking heads from a short video.
-
SyncTalk++: High-Fidelity and Efficient Synchronized Talking Heads Synthesis Using Gaussian Splatting
SyncTalk++ synthesizes speech-driven talking-head videos via 3D Gaussian Splatting and reports state-of-the-art synchronization and quality at up to 101 FPS.
-
Speaking images. A novel framework for the automated self-description of artworks
A four-stage open-source AI pipeline turns a digitized artwork into a short video where a depicted person animates and narrates the scene.
-
FaceEditTalker: Controllable Talking Head Generation with Facial Attribute Editing
A single framework can edit predefined facial attributes in audio-synchronized talking head videos while preserving identity and lip-sync quality.
-
Human Motion Video Generation: A Survey
A comprehensive survey with a five-phase pipeline model for human motion video generation, covering over 200 papers and adding a new benchmark comparison of nine pose-guided methods.
-
LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models
Using consistency distillation, INT8 quantization, and pipeline parallelism, the LLIA system generates portrait video from audio at 78 FPS, with 140 ms initial latency.
Discussion (0). Sign in to comment.