REVIEW 18 cited by
HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Customized video generation aims to produce videos featuring specific subjects under flexible user-defined conditions, yet existing methods often struggle with identity consistency and limited input modalities. In this paper, we propose HunyuanCustom, a multi-modal customized video generation framework that emphasizes subject consistency while supporting image, audio, video, and text conditions. Built upon HunyuanVideo, our model first addresses the image-text conditioned generation task by introducing a text-image fusion module based on LLaVA for enhanced multi-modal understanding, along with an image ID enhancement module that leverages temporal concatenation to reinforce identity features across frames. To enable audio- and video-conditioned generation, we further propose modality-specific condition injection mechanisms: an AudioNet module that achieves hierarchical alignment via spatial cross-attention, and a video-driven injection module that integrates latent-compressed conditional video through a patchify-based feature-alignment network. Extensive experiments on single- and multi-subject scenarios demonstrate that HunyuanCustom significantly outperforms state-of-the-art open- and closed-source methods in terms of ID consistency, realism, and text-video alignment. Moreover, we validate its robustness across downstream tasks, including audio and video-driven customized video generation. Our results highlight the effectiveness of multi-modal conditioning and identity-preserving strategies in advancing controllable video generation. All the code and models are available at https://hunyuancustom.github.io.
Forward citations
Cited by 18 Pith papers
-
UniMoCa: Unifying Motion and Camera Controls as Visual Proxies for Faithful Human Video Generation
A visual proxy that renders human motion under the driving camera and overlays camera trajectory markers lets a video diffusion model control both body motion and camera movement from a single visual conditioning space.
-
HOMIE: Human-object Centric Video Personalization via Multimodal Intelligent Enhancement
HOMIE unifies inter- and intra-subject video personalization by injecting MLLM-derived relational features into DiT self-attention (GMG) and tagging tokens with modality/reference embeddings (MRE), reporting SOTA on a...
-
Aura: Consistent Multi-Subject Video Generation via VLM-Grounded Semantic Alignment
Aura combines VLM meta-queries, T5-teacher alignment, subject-aware RoPE shifts, memory tokens, and a large AIGC-curated dataset to claim SOTA multi-element subject-to-video generation under OpenS2V-Eval Total score.
-
Rethinking Position Embedding as a Context Controller for Multi-Reference and Multi-Shot Video Generation
SideInfo-RoPE encodes reference-identity agreement as an extra rotary axis, disambiguating similar characters in multi-reference multi-shot video generation while keeping full semantic attention.
-
RefAlign: Representation Alignment for Reference-to-Video Generation
Explicit training-time alignment of DiT reference features to a VFM (with pull/push loss) raises OpenS2V-Eval TotalScore over prior R2V methods with no inference cost.
-
MVHOI: Bridge Multi-view Condition to Complex Human-Object Interaction Video Reenactment via 3D Foundation Model
Using a 3D foundation model to produce viewpoint-aware anchors plus multi-view reference textures enables realistic human-object-interaction reenactment with large out-of-plane rotations.
-
OmniCustom: Sync Audio-Video Customization Via Joint Audio-Video Generation Model
A zero-shot model that generates a video of a reference face speaking user-chosen text with a reference voice timbre.
-
InsertAnywhere: Geometrically Grounded and Optics-Aware Video Object Insertion
InsertAnywhere inserts a reference object into arbitrary videos by reconstructing 4D geometry to propagate a user-given placement across frames and fine-tuning video diffusion on ROSE++, a removal-to-insertion dataset...
-
CustomX: Unified Character, Action, and Scene Customization in Video World Models
AniX generates controllable videos of a user-supplied character performing typed actions inside a user-supplied 3D scene by fine-tuning a pre-trained video generator on small locomotion datasets.
-
Phantom-Data : Towards a General Subject-Consistent Video Generation Dataset
Phantom-Data provides around one million cross-context, identity-consistent reference-video pairs for subject-to-video generation, and training on it improves prompt following and visual quality.
-
DreamActor-H1: High-Fidelity Human-Product Demonstration Video Generation via Motion-designed Diffusion Transformers
A diffusion transformer model generates human-product demonstration videos from paired human and product images while preserving both identities through masked cross-attention and motion template guidance.
-
PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement
PolyVivid combines VLLM-based grounding, 3D-RoPE positional encoding, and attention-inherited identity injection to generate customized videos with multiple consistent subjects and text-specified interactions.
-
OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation
A benchmark and five-million-clip dataset for evaluating and training subject-to-video generation models, with three new metrics for subject consistency, naturalness, and text alignment.
-
Keyframe-Anchored Identity Preservation for Sequential-Action Video Generation
A keyframe-anchored, training-free pipeline—terminal-state prompts, chained keyframe generation, and identity-aware sampling—ranks third on the IPVG26 Track 2 leaderboard.
-
DreamSwapV: Mask-guided Subject Swapping for Any Customized Video Editing
DreamSwapV performs mask-guided, subject-agnostic subject swapping in videos using multiple conditions, an adaptive mask strategy, and a new benchmark.
-
OmniV2V: Versatile Video Generation and Editing via Dynamic Content Manipulation
OmniV2V is one diffusion-transformer model that performs eight video generation and editing tasks by combining mask, pose, image, and text-instruction conditions.
-
Hunyuan-Game: Industrial-grade Intelligent Game Creation Model
Tencent's Hunyuan-Game applies diffusion transformers to game asset creation across nine image and video generation tasks, with self-reported gains that are partly contradicted by its own evaluation table.
-
A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality
A survey of 32 long-video generation papers, presenting a taxonomy and component recommendations for backbones, text encoders, objectives, and positional encodings.
Discussion (0). Sign in to comment.