Pith. sign in

REVIEW 10 cited by

MagicAnimate: Temporally Consistent Human Image Animation using Diffusion Model

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.16498 v1 pith:LJT7KIYF submitted 2023-11-27 cs.CV cs.GR

classification cs.CVcs.GR
keywords animationimagereferencevideotemporalmodelaimsappearance
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper studies the human image animation task, which aims to generate a video of a certain reference identity following a particular motion sequence. Existing animation works typically employ the frame-warping technique to animate the reference image towards the target motion. Despite achieving reasonable results, these approaches face challenges in maintaining temporal consistency throughout the animation due to the lack of temporal modeling and poor preservation of reference identity. In this work, we introduce MagicAnimate, a diffusion-based framework that aims at enhancing temporal consistency, preserving reference image faithfully, and improving animation fidelity. To achieve this, we first develop a video diffusion model to encode temporal information. Second, to maintain the appearance coherence across frames, we introduce a novel appearance encoder to retain the intricate details of the reference image. Leveraging these two innovations, we further employ a simple video fusion technique to encourage smooth transitions for long video animation. Empirical results demonstrate the superiority of our method over baseline approaches on two benchmarks. Notably, our approach outperforms the strongest baseline by over 38% in terms of video fidelity on the challenging TikTok dancing dataset. Code and model will be made available.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. TruePose: Human-Parsing-guided Attention Diffusion for Full-ID Preserving Pose Transfer

    cs.CV 2025-02 conditional novelty 6.0 of 10

    TruePose uses human-parsing-guided attention and CLIP-based regional alignment in a dual-UNet diffusion framework to preserve facial identity and clothing texture during pose transfer.

  2. P3S-Diffusion:A Selective Subject-driven Generation Framework via Point Supervision

    cs.CV 2024-12 conditional novelty 6.0 of 10

    P3S-Diffusion generates personalized images by using one or two user clicks on a reference image to select and preserve the target subject.

  3. CFSynthesis: Controllable and Free-view 3D Human Video Synthesis

    cs.CV 2024-12 conditional novelty 6.0 of 10

    CFSynthesis generates free-view human videos from one reference image by conditioning a diffusion model on a textured SMPL body model and separately encoded foreground and background.

  4. SwiftTry: Fast and Consistent Video Virtual Try-On with Diffusion Models

    cs.CV 2024-12 conditional novelty 6.0 of 10

    SwiftTry makes diffusion-based video virtual try-on faster and more consistent by shifting non-overlapping video chunks during sampling and caching features across denoising steps.

  5. MoTrans: Customized Motion Transfer with Text-driven Video Diffusion Models

    cs.CV 2024-12 conditional novelty 6.0 of 10

    MoTrans transfers specific motions from reference videos to new subjects using a two-stage fine-tuning scheme with recaptioned prompts, appearance injection, and a motion-specific verb embedding.

  6. TaoMate: Anchor-Guided Memory Bridging Evolving and Reference States for Real-Time Audio-Video Digital Human Generation

    cs.CV 2026-07 conditional novelty 5.0 of 10

    Anchor-guided fixed-capacity multimodal memory plus stage-parallel few-step denoising stabilizes long-form causal audio-video digital humans at real-time speed.

  7. Animate-X++: Universal Character Image Animation with Dynamic Backgrounds

    cs.CV 2025-08 conditional novelty 5.0 of 10

    Animate-X++ turns cartoon images into pose-driven animations with text-controlled moving backgrounds, claiming state-of-the-art results on a new synthetic anthropomorphic benchmark.

  8. FixTalk: Taming Identity Leakage for High-Quality Talking Head Generation in Extreme Cases

    cs.CV 2025-07 conditional novelty 5.0 of 10

    FixTalk adds two modules to a real-time GAN talking-head model, decoupling identity from motion to stop identity leakage while using a memory to recover details and reduce artifacts.

  9. Whole-Body Conditioned Egocentric Video Prediction

    cs.CV 2025-06 conditional novelty 5.0 of 10

    An autoregressive conditional diffusion transformer predicts future egocentric video from whole-body 3D pose sequences, trained on Nymeria, with atomic action and long-horizon evaluations.

  10. Consistent Human Image and Video Generation with Spatially Conditioned Diffusion

    cs.CV 2024-12 conditional novelty 5.0 of 10

    Spatially conditioning a diffusion model by concatenating a reference human image with the noisy target, and adding causal self-attention, improves appearance consistency in human image and video animation.

Pith tools