REVIEW 10 cited by
MagicAnimate: Temporally Consistent Human Image Animation using Diffusion Model
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
This paper studies the human image animation task, which aims to generate a video of a certain reference identity following a particular motion sequence. Existing animation works typically employ the frame-warping technique to animate the reference image towards the target motion. Despite achieving reasonable results, these approaches face challenges in maintaining temporal consistency throughout the animation due to the lack of temporal modeling and poor preservation of reference identity. In this work, we introduce MagicAnimate, a diffusion-based framework that aims at enhancing temporal consistency, preserving reference image faithfully, and improving animation fidelity. To achieve this, we first develop a video diffusion model to encode temporal information. Second, to maintain the appearance coherence across frames, we introduce a novel appearance encoder to retain the intricate details of the reference image. Leveraging these two innovations, we further employ a simple video fusion technique to encourage smooth transitions for long video animation. Empirical results demonstrate the superiority of our method over baseline approaches on two benchmarks. Notably, our approach outperforms the strongest baseline by over 38% in terms of video fidelity on the challenging TikTok dancing dataset. Code and model will be made available.
Forward citations
Cited by 10 Pith papers
-
TruePose: Human-Parsing-guided Attention Diffusion for Full-ID Preserving Pose Transfer
TruePose uses human-parsing-guided attention and CLIP-based regional alignment in a dual-UNet diffusion framework to preserve facial identity and clothing texture during pose transfer.
-
P3S-Diffusion:A Selective Subject-driven Generation Framework via Point Supervision
P3S-Diffusion generates personalized images by using one or two user clicks on a reference image to select and preserve the target subject.
-
CFSynthesis: Controllable and Free-view 3D Human Video Synthesis
CFSynthesis generates free-view human videos from one reference image by conditioning a diffusion model on a textured SMPL body model and separately encoded foreground and background.
-
SwiftTry: Fast and Consistent Video Virtual Try-On with Diffusion Models
SwiftTry makes diffusion-based video virtual try-on faster and more consistent by shifting non-overlapping video chunks during sampling and caching features across denoising steps.
-
MoTrans: Customized Motion Transfer with Text-driven Video Diffusion Models
MoTrans transfers specific motions from reference videos to new subjects using a two-stage fine-tuning scheme with recaptioned prompts, appearance injection, and a motion-specific verb embedding.
-
TaoMate: Anchor-Guided Memory Bridging Evolving and Reference States for Real-Time Audio-Video Digital Human Generation
Anchor-guided fixed-capacity multimodal memory plus stage-parallel few-step denoising stabilizes long-form causal audio-video digital humans at real-time speed.
-
Animate-X++: Universal Character Image Animation with Dynamic Backgrounds
Animate-X++ turns cartoon images into pose-driven animations with text-controlled moving backgrounds, claiming state-of-the-art results on a new synthetic anthropomorphic benchmark.
-
FixTalk: Taming Identity Leakage for High-Quality Talking Head Generation in Extreme Cases
FixTalk adds two modules to a real-time GAN talking-head model, decoupling identity from motion to stop identity leakage while using a memory to recover details and reduce artifacts.
-
Whole-Body Conditioned Egocentric Video Prediction
An autoregressive conditional diffusion transformer predicts future egocentric video from whole-body 3D pose sequences, trained on Nymeria, with atomic action and long-horizon evaluations.
-
Consistent Human Image and Video Generation with Spatially Conditioned Diffusion
Spatially conditioning a diffusion model by concatenating a reference human image with the noisy target, and adding causal self-attention, improves appearance consistency in human image and video animation.
Discussion (0). Continue with ORCID to comment.