Pith. sign in

REVIEW 21 cited by

ID-Animator: Zero-Shot Identity-Preserving Human Video Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.15275 v3 pith:YSKK5QE4 submitted 2024-04-23 cs.CV

classification cs.CV
keywords generationvideoid-animatorhumanidentityfacialmodelstraining
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Generating high-fidelity human video with specified identities has attracted significant attention in the content generation community. However, existing techniques struggle to strike a balance between training efficiency and identity preservation, either requiring tedious case-by-case fine-tuning or usually missing identity details in the video generation process. In this study, we present \textbf{ID-Animator}, a zero-shot human-video generation approach that can perform personalized video generation given a single reference facial image without further training. ID-Animator inherits existing diffusion-based video generation backbones with a face adapter to encode the ID-relevant embeddings from learnable facial latent queries. To facilitate the extraction of identity information in video generation, we introduce an ID-oriented dataset construction pipeline that incorporates unified human attributes and action captioning techniques from a constructed facial image pool. Based on this pipeline, a random reference training strategy is further devised to precisely capture the ID-relevant embeddings with an ID-preserving loss, thus improving the fidelity and generalization capacity of our model for ID-specific video generation. Extensive experiments demonstrate the superiority of ID-Animator to generate personalized human videos over previous models. Moreover, our method is highly compatible with popular pre-trained T2V models like animatediff and various community backbone models, showing high extendability in real-world applications for video generation where identity preservation is highly desired. Our codes and checkpoints are released at https://github.com/ID-Animator/ID-Animator.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 21 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ID-V2V: Identity-Preserving Video Restylization

    cs.CV 2026-07 conditional novelty 6.0 of 10

    ID-V2V restyles video by conditioning a diffusion model on edited keyframes, depth, relit faces, and face normals, so scene edits propagate while facial identity and performance are preserved.

  2. GroupVideo: Multi-Identity Customized Text-to-Video Generation

    cs.CV 2026-07 conditional novelty 6.0 of 10

    GroupVideo generates multi-person videos from reference photos plus text, using multimodal identity alignment and ID localization to keep each person's identity consistent.

  3. HOMIE: Human-object Centric Video Personalization via Multimodal Intelligent Enhancement

    cs.CV 2026-07 conditional novelty 6.0 of 10

    HOMIE unifies inter- and intra-subject video personalization by injecting MLLM-derived relational features into DiT self-attention (GMG) and tagging tokens with modality/reference embeddings (MRE), reporting SOTA on a...

  4. Delving into the Temporal Challenges of Unified Video Protection Against Image-to-Video and Fine-Tuning-based Customization

    cs.CV 2026-07 conditional novelty 6.0 of 10

    TC-UAP learns a shared multi-frame adversarial perturbation that protects videos of the same identity from both fine-tuning-based and reference-based video customization, remaining effective on unseen clips and under ...

  5. Aura: Consistent Multi-Subject Video Generation via VLM-Grounded Semantic Alignment

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Aura combines VLM meta-queries, T5-teacher alignment, subject-aware RoPE shifts, memory tokens, and a large AIGC-curated dataset to claim SOTA multi-element subject-to-video generation under OpenS2V-Eval Total score.

  6. OmniCustom: Sync Audio-Video Customization Via Joint Audio-Video Generation Model

    cs.SD 2026-02 conditional novelty 6.0 of 10

    A zero-shot model that generates a video of a reference face speaking user-chosen text with a reference voice timbre.

  7. HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning

    cs.CV 2025-09 conditional novelty 6.0 of 10

    HuMo uses a two-stage training scheme and a face-focus trick to generate human videos that follow text, keep a reference person's identity, and sync speech to audio, beating several single-task systems on benchmarks.

  8. LaVieID: Local Autoregressive Diffusion Transformers for Identity-Preserving Video Creation

    cs.CV 2025-08 conditional novelty 6.0 of 10

    LaVieID improves identity-preserving text-to-video by routing local facial parts into early DiT blocks and autoregressively refining denoised video tokens in temporal chunks.

  9. DreamActor-H1: High-Fidelity Human-Product Demonstration Video Generation via Motion-designed Diffusion Transformers

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A diffusion transformer model generates human-product demonstration videos from paired human and product images while preserving both identities through masked cross-attention and motion template guidance.

  10. PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement

    cs.CV 2025-06 conditional novelty 6.0 of 10

    PolyVivid combines VLLM-based grounding, 3D-RoPE positional encoding, and attention-inherited identity injection to generate customized videos with multiple consistent subjects and text-specified interactions.

  11. UNIC: Unified In-Context Video Editing

    cs.CV 2025-06 conditional novelty 6.0 of 10

    One diffusion transformer handles ID insert, swap, delete, stylization, propagation, and re-camera control in a single model using in-context token concatenation with task-aware positional encoding and bias.

  12. AnimeShooter: A Multi-Shot Animation Dataset for Reference-Guided Video Generation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    AnimeShooter provides hierarchical story and shot annotations plus reference images for 148K one-minute animation stories, and AnimeShooterGen trained on it shows improved cross-shot consistency.

  13. FlowMo: Variance-Based Flow Guidance for Coherent Motion in Video Generation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    FlowMo reduces temporal artifacts in video generation by guiding the denoising process to lower the maximum patch-wise variance of consecutive-frame differences in the latent space.

  14. OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A benchmark and five-million-clip dataset for evaluating and training subject-to-video generation models, with three new metrics for subject consistency, naturalness, and text alignment.

  15. Vera: Identity-Faithful Human Subject-to-Video Generation

    cs.CV 2026-07 conditional novelty 5.0 of 10

    Vera improves identity consistency in human subject-to-video generation using cross-clip identity-aligned data, face-weighted masked loss, and layer-aware reference attention.

  16. Keyframe-Anchored Identity Preservation for Sequential-Action Video Generation

    cs.CV 2026-07 conditional novelty 5.0 of 10

    A keyframe-anchored, training-free pipeline—terminal-state prompts, chained keyframe generation, and identity-aware sampling—ranks third on the IPVG26 Track 2 leaderboard.

  17. Identity-Preserving Text-to-Video Generation via Training-Free Prompt, Image, and Guidance Enhancement

    cs.CV 2025-09 conditional novelty 5.0 of 10

    A training-free prompt, image, and guidance enhancement framework improves face consistency and video quality for identity-preserving text-to-video generation, winning the ACM Multimedia 2025 IPVG challenge.

  18. Tora2: Motion and Appearance Customized Diffusion Transformer for Multi-Entity Video Generation

    cs.CV 2025-07 conditional novelty 5.0 of 10

    Tora2 adds decoupled personalization embeddings, gated self-attention binding, and contrastive learning to Tora, enabling simultaneous appearance and trajectory customization for multiple entities in generated video.

  19. FullDiT2: Efficient In-Context Conditioning for Video Diffusion Transformers

    cs.CV 2025-06 conditional novelty 5.0 of 10

    FullDiT2 accelerates FullDiT-style in-context conditioning for video by dynamic token selection and selective context caching, cutting per-step time by 2-3x with minimal quality loss.

  20. Human Motion Video Generation: A Survey

    cs.CV 2025-09 conditional novelty 4.0 of 10

    A comprehensive survey with a five-phase pipeline model for human motion video generation, covering over 200 papers and adding a new benchmark comparison of nine pose-guided methods.

  21. A Summer Meridional Subsurface Temperature Dipole Mode in the South China Sea

    physics.ao-ph 2025-08 unverdicted novelty 4.0 of 10

    The manuscript body does not match the abstract, so the ocean dipole claim is unsupported by any presented evidence.

Pith tools