REVIEW 31 cited by
AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
In this study, we propose AniPortrait, a novel framework for generating high-quality animation driven by audio and a reference portrait image. Our methodology is divided into two stages. Initially, we extract 3D intermediate representations from audio and project them into a sequence of 2D facial landmarks. Subsequently, we employ a robust diffusion model, coupled with a motion module, to convert the landmark sequence into photorealistic and temporally consistent portrait animation. Experimental results demonstrate the superiority of AniPortrait in terms of facial naturalness, pose diversity, and visual quality, thereby offering an enhanced perceptual experience. Moreover, our methodology exhibits considerable potential in terms of flexibility and controllability, which can be effectively applied in areas such as facial motion editing or face reenactment. We release code and model weights at https://github.com/scutzzj/AniPortrait
Forward citations
Cited by 31 Pith papers
-
Instant Expressive Gaussian Head Avatars at Over 100 FPS
A single-photo avatar encoder with per-Gaussian feature-space deformation animates faces at 107 FPS with expression quality competitive with diffusion models.
-
Geometry-guided Emotion Modulation for Controllable and Photorealistic Emotional Talking Face Generation
A diffusion-based talking-face generator uses 3D blendshape coefficients to continuously control the emotion intensity of generated facial expressions.
-
LeapTalk: Breaking the Latency-Quality Trade-off in Talking Head Generation
LeapTalk distills a multi-step diffusion teacher into a one-step Brownian-bridge student and reports stable streaming talking-head generation at up to 200 FPS.
-
ViDS: Video Diffusion Shader using 3D Face Tracking
Fine-tuning a video diffusion model on dense 3DMM normal maps from Pixel3DMM lets a single portrait photo be animated with a driving video's expressions and pose, surpassing landmark- and latent-based portrait animati...
-
SyncBreaker:Stage-Aware Multimodal Adversarial Attacks on Audio-Driven Talking Head Generation
SyncBreaker jointly attacks image and audio streams with Multi-Interval Sampling and Cross-Attention Fooling to degrade speech-driven talking head generation more than single-modality baselines.
-
AHOY! Animatable Humans under Occlusion from YouTube Videos with Gaussian Splatting and Video Diffusion Priors
Identity-finetuned video diffusion plus RF-Inversion can supply multi-view body supervision that lets 3D Gaussian avatars be completed and animated from heavily occluded monocular video.
-
STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits
STARCaster is a 2D spatio-temporal video diffusion model that unifies identity-conditioned audio-driven portrait animation and novel-view synthesis without explicit 3D reconstruction.
-
FantasyTalking2: Timestep-Layer Adaptive Preference Optimization for Audio-Driven Portrait Animation
A three-part system, Talking-Critic, Talking-NSQ, and TLPO, aligns diffusion portrait animation models to human preferences and improves lip-sync, motion naturalness, and visual quality.
-
LaVieID: Local Autoregressive Diffusion Transformers for Identity-Preserving Video Creation
LaVieID improves identity-preserving text-to-video by routing local facial parts into early DiT blocks and autoregressively refining denoised video tokens in temporal chunks.
-
PoseGuard: Pose-Guided Generation with Safety Guardrails
PoseGuard degrades output quality of pose-guided video generators for unsafe poses while preserving fidelity for benign poses, using LoRA-based safety alignment.
-
Text2Lip: Progressive Lip-Synced Talking Face Generation from Text via Viseme-Guided Rendering
A text-only talking face generator maps text to visemes, hallucinates pseudo-audio, and renders lip-synced video; the paper's headline metrics are partially contradicted by its own tables.
-
X-NeMo: Expressive Neural Motion Reenactment via Disentangled Latent Attention
X-NeMo trains a 1D identity-agnostic motion descriptor end-to-end with a diffusion model, enabling zero-shot portrait animation with improved identity and expression fidelity.
-
DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation
A single-DiT portrait animation framework with style and emotion branches plus parallel audio-style cross-attention claims faster, controllable talking-head generation without a Reference Net.
-
MagicAnime: A Hierarchically Annotated, Multimodal and Multitasking Dataset with Benchmarks for Cartoon Animation Generation
MagicAnime is a 400k-clip multimodal cartoon dataset with hierarchical annotations and benchmarks for image-to-video, pose-driven, face reenactment, and audio-driven animation generation.
-
HairShifter: Consistent and High-Fidelity Video Hair Transfer via Anchor-Guided Animation
HairShifter transfers a reference hairstyle onto a person throughout a video by animating a high-quality anchor frame and using a gated decoder that preserves non-hair regions.
-
Democratizing High-Fidelity Co-Speech Gesture Video Generation
Co-speech gesture video is generated by first predicting 2D skeleton motion from audio via feature-concatenated diffusion, then rendering with an off-the-shelf video model, supported by a new 405-hour public dataset.
-
Audio-Sync Video Generation with Multi-Stream Temporal Control
MTV splits audio into speech, effects, and music to separately drive lip sync, event timing, and visual mood in video generation, trained on a new 392K-clip dataset.
-
Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation
MultiTalk is the first framework to generate multi-person conversational videos from multi-stream audio, using Label Rotary Position Embedding to bind each voice to the correct person.
-
Exploring Timeline Control for Facial Motion Generation
A diffusion model generates natural facial motions from user-specified multi-track timelines, using TICC-based frame-level action interval annotation for training and evaluation.
-
Long-Term TalkingFace Generation via Motion-Prior Conditional Diffusion Model
A motion-prior diffusion model with archived-frame memory improves identity, lip-sync, and head-motion consistency in long talking-face videos.
-
InfinityHuman: Towards Long-Term Audio-Driven Human
A coarse-to-fine audio-driven animation framework that uses pose-guided refinement and hand-specific reward learning to generate long, identity-stable talking videos.
-
InfiniteTalk: Audio-driven Video Generation for Sparse-Frame Video Dubbing
Sparse-frame dubbing with adjacent-chunk keyframe sampling lets a streaming audio-video model produce full-body motion synchronized to new audio while preserving identity and camera motion.
-
EDTalk++: Full Disentanglement for Controllable Talking Head Synthesis
EDTalk++ disentangles talking-head video into four orthogonal motion banks (mouth, pose, eyes, expression) and drives them from either video or audio inputs.
-
MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation
MoDiT, a diffusion transformer conditioned on 3DMM coefficients and Wav2Lip references, produces talking-head videos with improved same-identity lip sync and more natural blinks in its reported benchmarks.
-
FixTalk: Taming Identity Leakage for High-Quality Talking Head Generation in Extreme Cases
FixTalk adds two modules to a real-time GAN talking-head model, decoupling identity from motion to stop identity leakage while using a memory to recover details and reduce artifacts.
-
SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers
An audio-conditioned video diffusion transformer that animates portraits from image, video, text, and audio inputs with a sliding-window fusion for long videos.
-
Speaking images. A novel framework for the automated self-description of artworks
A four-stage open-source AI pipeline turns a digitized artwork into a short video where a depicted person animates and narrates the scene.
-
FaceEditTalker: Controllable Talking Head Generation with Facial Attribute Editing
A single framework can edit predefined facial attributes in audio-synchronized talking head videos while preserving identity and lip-sync quality.
-
Human Motion Video Generation: A Survey
A comprehensive survey with a five-phase pipeline model for human motion video generation, covering over 200 papers and adding a new benchmark comparison of nine pose-guided methods.
-
LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models
Using consistency distillation, INT8 quantization, and pipeline parallelism, the LLIA system generates portrait video from audio at 78 FPS, with 140 ms initial latency.
- Controllable and Expressive One-Shot Video Head Swapping
Discussion (0). Sign in to comment.