Pith. sign in

REVIEW 41 cited by

LivePortrait: Efficient Portrait Animation with Stitching and Retargeting Control

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.03168 v2 pith:U7UULCU6 submitted 2024-07-03 cs.CV

classification cs.CV
keywords animationcontrollabilityframeworkgenerationliveportraitportraitbettercomputational
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Portrait Animation aims to synthesize a lifelike video from a single source image, using it as an appearance reference, with motion (i.e., facial expressions and head pose) derived from a driving video, audio, text, or generation. Instead of following mainstream diffusion-based methods, we explore and extend the potential of the implicit-keypoint-based framework, which effectively balances computational efficiency and controllability. Building upon this, we develop a video-driven portrait animation framework named LivePortrait with a focus on better generalization, controllability, and efficiency for practical usage. To enhance the generation quality and generalization ability, we scale up the training data to about 69 million high-quality frames, adopt a mixed image-video training strategy, upgrade the network architecture, and design better motion transformation and optimization objectives. Additionally, we discover that compact implicit keypoints can effectively represent a kind of blendshapes and meticulously propose a stitching and two retargeting modules, which utilize a small MLP with negligible computational overhead, to enhance the controllability. Experimental results demonstrate the efficacy of our framework even compared to diffusion-based methods. The generation speed remarkably reaches 12.8ms on an RTX 4090 GPU with PyTorch. The inference code and models are available at https://github.com/KwaiVGI/LivePortrait

Discussion (0). Sign in to comment.

Forward citations

Cited by 41 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. InteracVid: Building a Real Interactive Audio-Visual Response Dataset from Live-Chat Videos

    cs.CV 2026-08 conditional novelty 7.0 of 10

    InteracVid delivers 454K livestream-derived context-query-response triplets, pairing real or LLM-reconstructed chat triggers with real audio-video reactions, and shows fine-tuning gains on genuine queries.

  2. Instant Expressive Gaussian Head Avatars at Over 100 FPS

    cs.CV 2025-12 conditional novelty 7.0 of 10

    A single-photo avatar encoder with per-Gaussian feature-space deformation animates faces at 107 FPS with expression quality competitive with diffusion models.

  3. FantasyPortrait: Enhancing Multi-Character Portrait Animation with Expression-Augmented Diffusion Transformers

    cs.CV 2025-07 conditional novelty 7.0 of 10

    A DiT-based portrait animation model transfers implicit facial expressions to one or more characters using a masked cross-attention mechanism, supported by a new multi-face dataset and benchmark.

  4. Proxy Avatar Meets Low-Rank Caching: Real-Time One-Shot Emotion-Controllable Portrait Animation

    cs.CV 2026-08 conditional novelty 6.0 of 10

    A cascade of a single-identity Gaussian emotion proxy, a one-shot diffusion retargeting model, and low-rank appearance caching enables real-time one-shot portrait animation with emotion control.

  5. TongueReenact: Geometry-Anchored Tongue Synthesis for Face Reenactment

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A two-stage pipeline (Gaussian reenactment plus geometry-anchored masked diffusion) transfers cross-identity tongue dynamics, roughly doubling tongue-specific metrics over prior reenactment baselines.

  6. LeapTalk: Breaking the Latency-Quality Trade-off in Talking Head Generation

    cs.CV 2026-07 conditional novelty 6.0 of 10

    LeapTalk distills a multi-step diffusion teacher into a one-step Brownian-bridge student and reports stable streaming talking-head generation at up to 200 FPS.

  7. ViDS: Video Diffusion Shader using 3D Face Tracking

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Fine-tuning a video diffusion model on dense 3DMM normal maps from Pixel3DMM lets a single portrait photo be animated with a driving video's expressions and pose, surpassing landmark- and latent-based portrait animati...

  8. Towards Flexible, Natural, Efficient Interaction for Conversational Talking Face Generation

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    InterTalk is a motion-based real-time framework for flexible multi-round multi-person conversational talking face generation using motion feedback, iterative strategies, and facial component disentanglement, supported...

  9. ARGen: Affect-Reinforced Generative Augmentation towards Vision-based Dynamic Emotion Perception

    cs.CV 2026-04 unverdicted novelty 6.0 of 10

    ARGen uses AU-guided prompts and a reinforcement-learned diffusion strategy to synthesize scarce-class facial expression videos that improve dynamic emotion recognition.

  10. AHOY! Animatable Humans under Occlusion from YouTube Videos with Gaussian Splatting and Video Diffusion Priors

    cs.CV 2026-03 conditional novelty 6.0 of 10

    Identity-finetuned video diffusion plus RF-Inversion can supply multi-view body supervision that lets 3D Gaussian avatars be completed and animated from heavily occluded monocular video.

  11. Avatar Forcing: Real-Time Interactive Head Avatar Generation for Natural Conversation

    cs.LG 2026-01 conditional novelty 6.0 of 10

    A causal diffusion-forcing model generates interactive head-avatar motion with 500ms motion-generation latency and learns expressive reactions via DPO with synthetic negative samples.

  12. Bitrate-Controlled Diffusion for Disentangling Motion and Content in Video

    cs.CV 2025-09 conditional novelty 6.0 of 10

    A self-supervised diffusion framework with a low-bitrate vector-quantization bottleneck learns disentangled motion and content latents supporting motion transfer and auto-regressive generation.

  13. FantasyTalking2: Timestep-Layer Adaptive Preference Optimization for Audio-Driven Portrait Animation

    cs.CV 2025-08 unverdicted novelty 6.0 of 10

    A three-part system, Talking-Critic, Talking-NSQ, and TLPO, aligns diffusion portrait animation models to human preferences and improves lip-sync, motion naturalness, and visual quality.

  14. LaVieID: Local Autoregressive Diffusion Transformers for Identity-Preserving Video Creation

    cs.CV 2025-08 conditional novelty 6.0 of 10

    LaVieID improves identity-preserving text-to-video by routing local facial parts into early DiT blocks and autoregressively refining denoised video tokens in temporal chunks.

  15. EvoMakeup: High-Fidelity and Controllable Makeup Editing with MakeupQuad

    cs.CV 2025-08 unverdicted novelty 6.0 of 10

    A new synthetic paired dataset and a distillation-aware framework enable a single model to do reference-based and text-guided facial makeup editing that transfers to real photos.

  16. X-NeMo: Expressive Neural Motion Reenactment via Disentangled Latent Attention

    cs.CV 2025-07 conditional novelty 6.0 of 10

    X-NeMo trains a 1D identity-agnostic motion descriptor end-to-end with a diffusion model, enabling zero-shot portrait animation with improved identity and expression fidelity.

  17. JOLT3D: Joint Learning of Talking Heads and 3DMM Parameters with Application to Lip-Sync

    cs.CV 2025-07 conditional novelty 6.0 of 10

    JOLT3D jointly trains a 3DMM reconstruction network with a talking head generator, then uses FACS mouth blendshapes from a diffusion model to lip-sync videos while preserving the original chin contour.

  18. MagicAnime: A Hierarchically Annotated, Multimodal and Multitasking Dataset with Benchmarks for Cartoon Animation Generation

    cs.CV 2025-07 conditional novelty 6.0 of 10

    MagicAnime is a 400k-clip multimodal cartoon dataset with hierarchical annotations and benchmarks for image-to-video, pose-driven, face reenactment, and audio-driven animation generation.

  19. HairShifter: Consistent and High-Fidelity Video Hair Transfer via Anchor-Guided Animation

    cs.CV 2025-07 conditional novelty 6.0 of 10

    HairShifter transfers a reference hairstyle onto a person throughout a video by animating a high-quality anchor frame and using a gated decoder that preserves non-hair regions.

  20. ARIG: Autoregressive Interactive Head Generation for Real-time Conversations

    cs.CV 2025-07 conditional novelty 6.0 of 10

    ARIG introduces a real-time, frame-wise autoregressive head generation framework with diffusion-based continuous motion prediction, improving interactive realism over clip-wise methods.

  21. Long-Term TalkingFace Generation via Motion-Prior Conditional Diffusion Model

    cs.CV 2025-02 conditional novelty 6.0 of 10

    A motion-prior diffusion model with archived-frame memory improves identity, lip-sync, and head-motion consistency in long talking-face videos.

  22. Split and Drive: Dual-Axis Disentanglement for Real-Time Gaussian Head Avatars

    cs.CV 2026-07 conditional novelty 5.0 of 10

    A single-image 3DGS head avatar with internalized motion encoding and three region-specialized Gaussian branches runs real-time end-to-end and matches or beats recent baselines on reenactment metrics.

  23. TaoMate: Anchor-Guided Memory Bridging Evolving and Reference States for Real-Time Audio-Video Digital Human Generation

    cs.CV 2026-07 conditional novelty 5.0 of 10

    Anchor-guided fixed-capacity multimodal memory plus stage-parallel few-step denoising stabilizes long-form causal audio-video digital humans at real-time speed.

  24. Real-Time Human-Centric World Modeling for Upper-Body Human-Object Interaction

    cs.CV 2026-07 conditional novelty 5.0 of 10

    A distilled video model jointly controls multi-scale upper-body motion latents and two discrete contact states to generate real-time human–object interaction at 25 FPS.

  25. FA-LAM: Focus-Aware Large Avatar Model for One-Shot 4D Animatable Gaussian Head

    cs.CV 2026-07 conditional novelty 5.0 of 10

    A one-shot model for animatable 3D/4D Gaussian head reconstruction that adds attention regularization, decoupled reconstruction-animation training, and autoregressive visibility-gated fusion, reporting consistent metr...

  26. MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation

    cs.CV 2025-08 reject novelty 5.0 of 10

    A new autoregressive video-generation framework for interactive digital humans, with a 64x compression autoencoder and a diffusion renderer, claims real-time multimodal control but is only demonstrated for audio input.

  27. Celeb-DF++: A Large-scale Challenging Video DeepFake Benchmark for Generalizable Forensics

    cs.CV 2025-07 conditional novelty 5.0 of 10

    Celeb-DF++ is a more diverse deepfake video benchmark, spanning 22 generation methods and three manipulation scenarios, on which 24 published detectors drop to about 70 to 72 percent average AUC in cross-method tests.

  28. Navigating Large-Pose Challenge for High-Fidelity Face Reenactment with Video Diffusion Model

    cs.CV 2025-07 conditional novelty 5.0 of 10

    FRVD warps a source face toward driving poses with implicit keypoints, then repairs lost details inside Stable Video Diffusion's latent space, reporting gains over seven baselines on large-pose reenactment benchmarks.

  29. StableAnimator++: Overcoming Pose Misalignment and Face Distortion for Human Image Animation

    cs.CV 2025-07 conditional novelty 5.0 of 10

    StableAnimator++ combines learnable SVD-guided pose alignment, a distribution-aware ID Adapter, and an HJB-based inference-time face optimizer to preserve identity in human image animation under severe pose misalignment.

  30. MoDA: Multi-modal Diffusion Architecture for Talking Head Generation

    cs.GR 2025-07 conditional novelty 5.0 of 10

    MoDA uses flow matching in a compact face-motion space with a progressively fused multi-modal transformer to generate expressive, lip-synced talking-head videos from a single image and audio.

  31. FixTalk: Taming Identity Leakage for High-Quality Talking Head Generation in Extreme Cases

    cs.CV 2025-07 conditional novelty 5.0 of 10

    FixTalk adds two modules to a real-time GAN talking-head model, decoupling identity from motion to stop identity leakage while using a memory to recover details and reduce artifacts.

  32. MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation

    cs.CV 2025-06 conditional novelty 5.0 of 10

    MirrorMe adapts the LTX video diffusion transformer to generate real-time, high-fidelity audio-driven halfbody animations with identity preservation and hand pose control.

  33. SyncTalk++: High-Fidelity and Efficient Synchronized Talking Heads Synthesis Using Gaussian Splatting

    cs.CV 2025-06 conditional novelty 5.0 of 10

    SyncTalk++ synthesizes speech-driven talking-head videos via 3D Gaussian Splatting and reports state-of-the-art synchronization and quality at up to 101 FPS.

  34. Speaking images. A novel framework for the automated self-description of artworks

    cs.CV 2025-05 conditional novelty 5.0 of 10

    A four-stage open-source AI pipeline turns a digitized artwork into a short video where a depicted person animates and narrates the scene.

  35. OpenWorldLib: A Unified Codebase and Definition of Advanced World Models

    cs.CV 2026-04 unverdicted novelty 4.0 of 10

    OpenWorldLib offers a standardized codebase and definition for world models that combine perception, interaction, and memory to understand and predict the world.

  36. Human Motion Video Generation: A Survey

    cs.CV 2025-09 conditional novelty 4.0 of 10

    A comprehensive survey with a five-phase pipeline model for human motion video generation, covering over 200 papers and adding a new benchmark comparison of nine pose-guided methods.

  37. LIA-X: Interpretable Latent Portrait Animator

    cs.CV 2025-08 conditional novelty 4.0 of 10

    Adding an L1 sparsity penalty to the motion dictionary of the LIA portrait animator produces disentangled, human-interpretable motion vectors that support controllable image and video editing and scale to roughly one ...

  38. CartoonAlive: Towards Expressive Live2D Modeling from Single Portraits

    cs.CV 2025-07 conditional novelty 4.0 of 10

    CartoonAlive automatically generates an animatable Live2D cartoon character from a single portrait image by combining 3DMM-inspired blendshapes with landmark-guided parameter regression.

  39. KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model

    cs.CV 2025-07 conditional novelty 4.0 of 10

    KptLLM++ unifies keypoint semantic understanding, visual-prompt detection, and text-prompt detection in a single multimodal LLM, reporting SOTA accuracy on COCO, AP-10K, Human-Art, and other benchmarks.

  40. LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models

    cs.CV 2025-06 conditional novelty 4.0 of 10

    Using consistency distillation, INT8 quantization, and pipeline parallelism, the LLIA system generates portrait video from audio at 78 FPS, with 140 ms initial latency.

  41. Controllable and Expressive One-Shot Video Head Swapping

    cs.CV 2025-06

Pith tools