Pith. sign in

REVIEW 18 cited by

HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2505.04512 v2 pith:SCPKYW2D submitted 2025-05-07 cs.CV

classification cs.CV
keywords generationvideocustomizedhunyuancustommoduleconsistencymulti-modalacross
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Customized video generation aims to produce videos featuring specific subjects under flexible user-defined conditions, yet existing methods often struggle with identity consistency and limited input modalities. In this paper, we propose HunyuanCustom, a multi-modal customized video generation framework that emphasizes subject consistency while supporting image, audio, video, and text conditions. Built upon HunyuanVideo, our model first addresses the image-text conditioned generation task by introducing a text-image fusion module based on LLaVA for enhanced multi-modal understanding, along with an image ID enhancement module that leverages temporal concatenation to reinforce identity features across frames. To enable audio- and video-conditioned generation, we further propose modality-specific condition injection mechanisms: an AudioNet module that achieves hierarchical alignment via spatial cross-attention, and a video-driven injection module that integrates latent-compressed conditional video through a patchify-based feature-alignment network. Extensive experiments on single- and multi-subject scenarios demonstrate that HunyuanCustom significantly outperforms state-of-the-art open- and closed-source methods in terms of ID consistency, realism, and text-video alignment. Moreover, we validate its robustness across downstream tasks, including audio and video-driven customized video generation. Our results highlight the effectiveness of multi-modal conditioning and identity-preserving strategies in advancing controllable video generation. All the code and models are available at https://hunyuancustom.github.io.

Discussion (0). Sign in to comment.

Forward citations

Cited by 18 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. UniMoCa: Unifying Motion and Camera Controls as Visual Proxies for Faithful Human Video Generation

    cs.CV 2026-08 conditional novelty 6.0 of 10

    A visual proxy that renders human motion under the driving camera and overlays camera trajectory markers lets a video diffusion model control both body motion and camera movement from a single visual conditioning space.

  2. HOMIE: Human-object Centric Video Personalization via Multimodal Intelligent Enhancement

    cs.CV 2026-07 conditional novelty 6.0 of 10

    HOMIE unifies inter- and intra-subject video personalization by injecting MLLM-derived relational features into DiT self-attention (GMG) and tagging tokens with modality/reference embeddings (MRE), reporting SOTA on a...

  3. Aura: Consistent Multi-Subject Video Generation via VLM-Grounded Semantic Alignment

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Aura combines VLM meta-queries, T5-teacher alignment, subject-aware RoPE shifts, memory tokens, and a large AIGC-curated dataset to claim SOTA multi-element subject-to-video generation under OpenS2V-Eval Total score.

  4. Rethinking Position Embedding as a Context Controller for Multi-Reference and Multi-Shot Video Generation

    cs.CV 2026-04 conditional novelty 6.0 of 10

    SideInfo-RoPE encodes reference-identity agreement as an extra rotary axis, disambiguating similar characters in multi-reference multi-shot video generation while keeping full semantic attention.

  5. RefAlign: Representation Alignment for Reference-to-Video Generation

    cs.CV 2026-03 conditional novelty 6.0 of 10

    Explicit training-time alignment of DiT reference features to a VFM (with pull/push loss) raises OpenS2V-Eval TotalScore over prior R2V methods with no inference cost.

  6. MVHOI: Bridge Multi-view Condition to Complex Human-Object Interaction Video Reenactment via 3D Foundation Model

    cs.CV 2026-03 conditional novelty 6.0 of 10

    Using a 3D foundation model to produce viewpoint-aware anchors plus multi-view reference textures enables realistic human-object-interaction reenactment with large out-of-plane rotations.

  7. OmniCustom: Sync Audio-Video Customization Via Joint Audio-Video Generation Model

    cs.SD 2026-02 conditional novelty 6.0 of 10

    A zero-shot model that generates a video of a reference face speaking user-chosen text with a reference voice timbre.

  8. InsertAnywhere: Geometrically Grounded and Optics-Aware Video Object Insertion

    cs.CV 2025-12 conditional novelty 6.0 of 10

    InsertAnywhere inserts a reference object into arbitrary videos by reconstructing 4D geometry to propagate a user-given placement across frames and fine-tuning video diffusion on ROSE++, a removal-to-insertion dataset...

  9. CustomX: Unified Character, Action, and Scene Customization in Video World Models

    cs.CV 2025-12 conditional novelty 6.0 of 10

    AniX generates controllable videos of a user-supplied character performing typed actions inside a user-supplied 3D scene by fine-tuning a pre-trained video generator on small locomotion datasets.

  10. Phantom-Data : Towards a General Subject-Consistent Video Generation Dataset

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Phantom-Data provides around one million cross-context, identity-consistent reference-video pairs for subject-to-video generation, and training on it improves prompt following and visual quality.

  11. DreamActor-H1: High-Fidelity Human-Product Demonstration Video Generation via Motion-designed Diffusion Transformers

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A diffusion transformer model generates human-product demonstration videos from paired human and product images while preserving both identities through masked cross-attention and motion template guidance.

  12. PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement

    cs.CV 2025-06 conditional novelty 6.0 of 10

    PolyVivid combines VLLM-based grounding, 3D-RoPE positional encoding, and attention-inherited identity injection to generate customized videos with multiple consistent subjects and text-specified interactions.

  13. OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A benchmark and five-million-clip dataset for evaluating and training subject-to-video generation models, with three new metrics for subject consistency, naturalness, and text alignment.

  14. Keyframe-Anchored Identity Preservation for Sequential-Action Video Generation

    cs.CV 2026-07 conditional novelty 5.0 of 10

    A keyframe-anchored, training-free pipeline—terminal-state prompts, chained keyframe generation, and identity-aware sampling—ranks third on the IPVG26 Track 2 leaderboard.

  15. DreamSwapV: Mask-guided Subject Swapping for Any Customized Video Editing

    cs.CV 2025-08 unverdicted novelty 5.0 of 10

    DreamSwapV performs mask-guided, subject-agnostic subject swapping in videos using multiple conditions, an adaptive mask strategy, and a new benchmark.

  16. OmniV2V: Versatile Video Generation and Editing via Dynamic Content Manipulation

    cs.CV 2025-06 conditional novelty 5.0 of 10

    OmniV2V is one diffusion-transformer model that performs eight video generation and editing tasks by combining mask, pose, image, and text-instruction conditions.

  17. Hunyuan-Game: Industrial-grade Intelligent Game Creation Model

    cs.CV 2025-05 reject novelty 4.0 of 10

    Tencent's Hunyuan-Game applies diffusion transformers to game asset creation across nine image and video generation tasks, with self-reported gains that are partly contradicted by its own evaluation table.

  18. A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality

    cs.CV 2025-07 conditional novelty 3.0 of 10

    A survey of 32 long-video generation papers, presenting a taxonomy and component recommendations for backbones, text encoders, objectives, and positional encodings.

Pith tools