Pith. sign in

Paper Citation Record · LEDGER

HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 38 inbound Pith citation observations for arXiv:2505.04512.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.04512 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 38 of 38 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 38 of 38 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T12:16:44.665393Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T21:10:09.669744Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 8dd27455-0272-48ec-a382-50a76bb3bf6f · inbound

Hunyuan-Game: Industrial-grade Intelligent Game Creation Model cites this paper.

Hunyuan-Game: Industrial-grade Intelligent Game Creation Model HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:30.549695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:30.549695Z digest=sha256:22157fd8b69ee6bce7f5f15850beba622291ff1910463f23ba455145d35abd21

Observation 63e9d5a2-120e-4bf2-8687-2cb8e04ecf70 · inbound

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation cites this paper.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:30.485928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:30.485928Z digest=sha256:92294c0569f0892d1f3d1cfe5529a2beafafc50f575dc7c5d5efb7714b68df60

Observation abfbb6ad-57e0-46c4-8e59-66bedea81e69 · inbound

OmniV2V: Versatile Video Generation and Editing via Dynamic Content Manipulation cites this paper.

OmniV2V: Versatile Video Generation and Editing via Dynamic Content Manipulation HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T11:39:51.666724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:39:51.666724Z digest=sha256:a882ec14b4351506fc43fd5511b9641a48d36a3c6791c2272c3f98616554a723

Observation 788063b8-c336-4922-8764-25d8cbfc7f70 · inbound

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement cites this paper.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T05:30:32.934408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:30:32.934408Z digest=sha256:b26e3d508432947534037ce9f1f505f57911984c6b7e88a3ffbed0bb43e11e18

Observation 074be566-ef03-4cca-86e2-8188ae16870a · inbound

DreamActor-H1: High-Fidelity Human-Product Demonstration Video Generation via Motion-designed Diffusion Transformers cites this paper.

DreamActor-H1: High-Fidelity Human-Product Demonstration Video Generation via Motion-designed Diffusion Transformers HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T04:27:37.363932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:27:37.363932Z digest=sha256:4dd3989f31c76fa00cd5bd0a5c194b26ec3ce37e20ede80e872a4a40bf7b3978

Observation 4fe3aad9-1022-4d98-a760-24ab29b51e31 · inbound

Phantom-Data : Towards a General Subject-Consistent Video Generation Dataset cites this paper.

Phantom-Data : Towards a General Subject-Consistent Video Generation Dataset HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:09.055319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:09.055319Z digest=sha256:31bc3edf604365b0872a97c5dd4886b03a9708e75a7bba91db5306084f23c2e0

Observation b641c452-348b-4b38-88ec-e82f018c0546 · inbound

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality cites this paper.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:26.356204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:26.356204Z digest=sha256:2ec39f1e3c14d377d7744c61b0bacefde26c66ef1d38847308814b13941b3d91

Observation 0a61444a-cf7f-48d6-94ed-a824c5c1e72d · inbound

DreamSwapV: Mask-guided Subject Swapping for Any Customized Video Editing cites this paper.

DreamSwapV: Mask-guided Subject Swapping for Any Customized Video Editing HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T18:34:25.012542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T18:34:25.012542Z digest=sha256:acb8095d94668a153c76b5da24b0897fc9c3c0f69d8bb74494fc55f7a48e9c3a

Observation 812b033c-b4c9-46d7-a44b-6d9dd7a06258 · inbound

InsertAnywhere: Geometrically Grounded and Optics-Aware Video Object Insertion cites this paper.

InsertAnywhere: Geometrically Grounded and Optics-Aware Video Object Insertion HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T15:20:18.764867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T15:20:18.764867Z digest=sha256:2b3148df90d0c39a96110c0de88e17eb80ce8564ec24b82b12ae83f2330ff901

Observation e5bf8773-d8bb-4ee5-9162-fd11c56f9371 · inbound

CustomX: Unified Character, Action, and Scene Customization in Video World Models cites this paper.

CustomX: Unified Character, Action, and Scene Customization in Video World Models HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T15:28:52.239638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T15:28:52.239638Z digest=sha256:ab3630c75d24ef9bbd9a71a39c3f44825dfb18bbd5aa8ccc7f1fe11f6b4070d5

Observation 6ee46033-c95f-4400-813a-8b67dfa1477b · inbound

OmniCustom: Sync Audio-Video Customization Via Joint Audio-Video Generation Model cites this paper.

OmniCustom: Sync Audio-Video Customization Via Joint Audio-Video Generation Model HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T00:11:12.422942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:11:12.422942Z digest=sha256:8d8888b4b43d2effaf2a9c22901dbe5c625362de0ab4db6a830fdb4374d8e08a

Observation c97ce011-c67f-4cbf-a1d6-0403d5d9fecc · inbound

MVHOI: Bridge Multi-view Condition to Complex Human-Object Interaction Video Reenactment via 3D Foundation Model cites this paper.

MVHOI: Bridge Multi-view Condition to Complex Human-Object Interaction Video Reenactment via 3D Foundation Model HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T05:51:18.169044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T05:51:18.169044Z digest=sha256:895258a1fb15ba0e47959c5c3813bb3c7a591ef2153fdb49c3a0e057193a149c

Observation d24af3bf-84ac-43b0-ae30-3b1dd448ad90 · inbound

RefAlign: Representation Alignment for Reference-to-Video Generation cites this paper.

RefAlign: Representation Alignment for Reference-to-Video Generation HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-13T18:01:19.034578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T18:01:19.034578Z digest=sha256:1e3cc793770b2ba3996b4ce316b723e7981893f15952b5078ec178f26ba96576

Observation e4ce79ab-ccc8-420b-b23d-c466bd122994 · inbound

Rethinking Position Embedding as a Context Controller for Multi-Reference and Multi-Shot Video Generation cites this paper.

Rethinking Position Embedding as a Context Controller for Multi-Reference and Multi-Shot Video Generation HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-13T12:23:05.876881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T12:23:05.876881Z digest=sha256:924101649492ed98f1e82088fafa701c1768fa2b9cf10bf364764368e9152eb5

Observation 0163d288-73d9-4b31-97ba-08b701df773b · inbound

Evolution of Video Generative Foundations cites this paper.

Evolution of Video Generative Foundations HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation

Reference 204

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:05:51.753975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T18:41:38.616611Z digest=sha256:2d45876f5da42a2f010007530d5dbca9f26fb9948943765c928959fc0bc2a857

Observation 04fe287e-dfd8-4469-bcbb-8a9e6bc57b04 · inbound

Prompt Relay: Inference-Time Temporal Control for Multi-Event Video Generation cites this paper.

Prompt Relay: Inference-Time Temporal Control for Multi-Event Video Generation HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:01:00.395392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T15:42:35.693519Z digest=sha256:7a18fc0c3596d80883078d4fc7e19c76d7cc2bd2dcb65e2453b3e2c3e4319170

Observation f08526f3-b1aa-49ac-a63c-5947baa19432 · inbound

Script-a-Video: Deep Structured Audio-visual Captions via Factorized Streams and Relational Grounding cites this paper.

Script-a-Video: Deep Structured Audio-visual Captions via Factorized Streams and Relational Grounding HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:11:03.662032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T15:07:45.595260Z digest=sha256:918fc1f18f539a6a4e5f8232d4364ad07760a6c583f7f76ba6db4e8430f2eaa5

Observation c1d8d2a0-a7d1-48f9-9bb7-bff428615352 · inbound

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation cites this paper.

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:11:00.806138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T15:09:02.727887Z digest=sha256:86fcf1caa8fa7b4357c124c722e99a9831b15e9912e5b950f7717dfbf668911a

Observation 952199a1-ebad-4157-a3ce-83405553fdcc · inbound

Controllable Video Object Insertion via Multiview Priors cites this paper.

Controllable Video Object Insertion via Multiview Priors HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:45:20.706056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T11:44:17.033051Z digest=sha256:cf17fb5f476a9a07407dc123e87ecca3aca3cc526c4b9077c1cc74118ea4c499

Observation 65333d83-84de-4bec-87d3-da71d12a5d45 · inbound

TS-Attn: Temporal-wise Separable Attention for Multi-Event Video Generation cites this paper.

TS-Attn: Temporal-wise Separable Attention for Multi-Event Video Generation HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T02:53:29.903427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T02:45:10.577070Z digest=sha256:04046966b09a6bdb7aec006e8f5df63879343f58fe4bd2051f33534afe0a5937

Observation eacb0a72-0353-47e2-931e-c83cd4fbc94b · inbound

MMControl: Unified Multi-Modal Control for Joint Audio-Video Generation cites this paper.

MMControl: Unified Multi-Modal Control for Joint Audio-Video Generation HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:46:28.018405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T02:55:09.008954Z digest=sha256:d37ae4a798e82efb07f12d165ff2bb3b13ca753ada9e460394dca618185af27b

Observation f3ac5160-9ad1-4dfb-b028-b606168c8028 · inbound

FaithfulFaces: Pose-Faithful Facial Identity Preservation for Text-to-Video Generation cites this paper.

FaithfulFaces: Pose-Faithful Facial Identity Preservation for Text-to-Video Generation HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:21:09.751609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T17:40:00.224358Z digest=sha256:de1b4e9599b4879002cc1ece814e6b47fadca7f18e1cee4491b5b77cad9d8af3

Observation d89586ac-4a70-4d6b-8094-c6a6994a90a8 · inbound

UniCustom: Unified Visual Conditioning for Multi-Reference Image Generation cites this paper.

UniCustom: Unified Visual Conditioning for Multi-Reference Image Generation HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:02:23.937729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T05:53:21.851578Z digest=sha256:15e7b7e8dff44ea691c2ab4db8473282c7304e3346cd90d07962bbf517e32f36

Observation 586ba538-ab1b-4b26-8fe9-2a6b78fff3c1 · inbound

UniCustom: Unified Visual Conditioning for Multi-Reference Image Generation cites this paper.

UniCustom: Unified Visual Conditioning for Multi-Reference Image Generation HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:03:03.463171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-14T22:00:01.349754Z digest=sha256:a32e4096e0e23f131ffc5af760dd6d7abed6f0d8a801d2dc519cae9426570bb2

Observation 0f54d741-72e8-4725-929e-dbc3f06fb308 · inbound

Omni-Customizer: End-to-End MultiModal Customization for Joint Audio-Video Generation cites this paper.

Omni-Customizer: End-to-End MultiModal Customization for Joint Audio-Video Generation HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-20T14:13:21.373265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T14:08:30.802619Z digest=sha256:e461cd88a75e27342372a23232d7ded781aee0d3606b18b83bba02a2d352bad5

Observation 2af97305-875a-456f-9aab-b10078e1f81f · inbound

Spatial-Temporal Decoupled Reference Conditioning for Identity-Preserving Text-to-Video Generation cites this paper.

Spatial-Temporal Decoupled Reference Conditioning for Identity-Preserving Text-to-Video Generation HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:26:17.501082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T15:25:22.778550Z digest=sha256:73fffabdf936af86d03427224c9b61769f82d0d57b9a9f650e996833766ac367

Observation a795f952-7525-4266-a9d2-855d4aa99514 · inbound

MetaWorld: Scaling Multi-Agent Video World Model from Single-view Video Data cites this paper.

MetaWorld: Scaling Multi-Agent Video World Model from Single-view Video Data HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:56:20.256912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T14:52:30.406683Z digest=sha256:5d7b939c6225c785e28c187cc85bf26733fc14461447d33c08375e897f3adc04

Observation e690e724-3c34-452d-9f64-4b51f8002a07 · inbound

CineDance: Towards Next-Generation Multi-Shot Long-Form Cinematic Audio-Video Generation cites this paper.

CineDance: Towards Next-Generation Multi-Shot Long-Form Cinematic Audio-Video Generation HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T00:07:28.122494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T17:30:25.371658Z digest=sha256:f6ad44d66b9f3a614581011bbb209f2f4527da24778ede46bcbd11f5bc3c6a41

Observation 8d18100d-2885-40e1-892b-3bdca5f5cedb · inbound

HarmoView: Harmonizing Multi-View Constraints for Identity-Consistent Video Generation cites this paper.

HarmoView: Harmonizing Multi-View Constraints for Identity-Consistent Video Generation HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-03T04:37:36.656864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T13:49:50.272650Z digest=sha256:b06bc5e5b0ecb98457a61ca5d7c4cadb0e50b4cfa4f3871f5feeb93aad76d810

Observation 0c8b23d7-dc6d-456c-9fb3-b8ba38c205fc · inbound

ARGUS: Stacked Multi-View Identity Mosaic Injection for Subject-Preserving Video Generation cites this paper.

ARGUS: Stacked Multi-View Identity Mosaic Injection for Subject-Preserving Video Generation HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-07-03T08:17:45.961436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T10:46:56.871174Z digest=sha256:32c6396ce906f72286439328f2798831961d51a7857aadf8641b5db4b45c45a8

Observation 7f39caec-ed05-4f26-aee3-ac525f169cd6 · inbound

DomainShuttle: Freeform Open Domain Subject-driven Text-to-video Generation cites this paper.

DomainShuttle: Freeform Open Domain Subject-driven Text-to-video Generation HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-04T21:10:09.671686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-25T19:00:23.260939Z digest=sha256:38cbe562f9050decf2bbe10765b86d1b0b2dc2a80691f81d6f44df802bd8d745

Observation bf6267c4-094e-4ab3-9066-b21b8598b93f · inbound

Ink3D: Sculpting 3D Assets with Extremely Complex Textures via Video Generative Models cites this paper.

Ink3D: Sculpting 3D Assets with Extremely Complex Textures via Video Generative Models HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T13:16:58.123698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-02T13:16:16.676244Z digest=sha256:dc29411095aafc88b8d225f0ac0a91cc0e56476c038e718ba9fb46d3dac59032

Observation fd6a0b46-6194-4af3-baa6-6848e97a5470 · inbound

Aura: Consistent Multi-Subject Video Generation via VLM-Grounded Semantic Alignment cites this paper.

Aura: Consistent Multi-Subject Video Generation via VLM-Grounded Semantic Alignment HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-11T20:11:31.576642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T20:11:31.576642Z digest=sha256:c4420bc33b23558acba701521bb8a657534217260593985c7564c94e9e4d2ca2

Observation 63374138-3231-45c3-9427-fbcf3959b7d1 · inbound

Keyframe-Anchored Identity Preservation for Sequential-Action Video Generation cites this paper.

Keyframe-Anchored Identity Preservation for Sequential-Action Video Generation HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T16:30:01.157173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:30:01.157173Z digest=sha256:973765734931645e9e68e4bc46bb8ba6dcb82a50f88b4afb81d4490db93c9095

Observation 18cd2add-2d16-4bc1-8d33-4e7a7b66130b · inbound

HOMIE: Human-object Centric Video Personalization via Multimodal Intelligent Enhancement cites this paper.

HOMIE: Human-object Centric Video Personalization via Multimodal Intelligent Enhancement HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T15:42:20.695859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:42:20.695859Z digest=sha256:3b6b75f52c1aecfd3c91f942c757d241c9d37332bdd26ed9d74ced566d212a64

Observation f34af647-a454-47b8-99ea-6fd7f628fb95 · inbound

UniMoCa: Unifying Motion and Camera Controls as Visual Proxies for Faithful Human Video Generation cites this paper.

UniMoCa: Unifying Motion and Camera Controls as Visual Proxies for Faithful Human Video Generation HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-04T17:55:19.705417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:55:19.705417Z digest=sha256:27013676eddff393506601d1c22a7dbe222af8b1407797add413e5748a2c169e

Observation f7d9a41d-ccf7-4740-93f2-eabb73556743 · inbound

VideoArgus: Agentic Rubric-Grounded Unified Evaluation for Video Generation and Editing cites this paper.

VideoArgus: Agentic Rubric-Grounded Unified Evaluation for Video Generation and Editing HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T12:16:44.665393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:16:44.665393Z digest=sha256:2fc3e4fb3c3380f58a755cfe660b32e072c6419d6021ac8df8dd46f077f9ca58

Observation 7b68b13c-0de6-45ee-807e-dc1472fd2363 · inbound

Vorch-IR: Long-Form Unified Multimodal Identity Replacement Video Generation cites this paper.

Vorch-IR: Long-Form Unified Multimodal Identity Replacement Video Generation HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-08T04:54:58.654243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:54:58.654243Z digest=sha256:d5fa5435fdc112205ebec60b1ad6a30d93dcca8475446dd9850054f16299647f