Pith. sign in

Paper Citation Record · LEDGER

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement

As of 18 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 4 inbound Pith citation observations for arXiv:2506.07848.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.07848 v1

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:30:33.055697Z

measured 52 of 52 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-11T20:11:31.576642Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T22:56:20.169424Z

Reference resolution

48 of 48 outbound references displayed

  • verified exact0
  • verified fuzzy7
  • unresolved41
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c4dcf0d4-b5af-472c-bf9a-4064cd2a576f · outbound

This paper cites Qwen2.5-VL Technical Report.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement Qwen2.5-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T05:30:32.874829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:30:32.874829Z digest=sha256:ec0f1099bf173c5d0a0197eae907a3511a119b67cbcced78ce7ef52b49380230

Observation 65232c13-4a34-4297-b1a9-9fa062e9ebd3 · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T05:30:32.880099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:30:32.880099Z digest=sha256:9c4a70de32997fa8e24437faa8545e436dcd348ee35d9b09be4bfba1d7184e75

Observation 26b2b07c-bdf0-48a2-a14d-bf3ccb8b3f0a · outbound

This paper cites Carreira and A.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement Carreira and A

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:30:33.612763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:30:32.884690Z digest=sha256:e026bb1f472e7545e325f7def6fea8fc12e8ea2a17add069569646425777d9bb

Observation 9e15de1b-3f68-4af7-9c4c-21e0d3498141 · outbound

This paper cites Chefer, S.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement Chefer, S

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:30:33.602556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:30:32.888262Z digest=sha256:e887b687fe7acc9ff66e5cd4289d58b176fff9182529f6afd98b8e8f043e1178

Observation acd89131-1c62-4c1a-9edb-ce308a4e98e7 · outbound

This paper cites an unresolved cited work.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:30:33.592113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:30:32.892748Z digest=sha256:e68727d9a91949461c41bcf2ff3fff4df84aa75580505af094f427f6df6b0fe0

Observation 014d3878-8673-46d3-a57e-5a234eb77074 · outbound

This paper cites an unresolved cited work.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:30:33.581859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:30:32.896304Z digest=sha256:e18433e805c18e49f9f2f6c75c5bcbdb034da905400f1bde70f2ad31691129d6

Observation e7ee9805-6d3c-4c63-8ac3-f8620a328848 · outbound

This paper cites Multi-subject Open-set Personalization in Video Generation.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement Multi-subject Open-set Personalization in Video Generation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T05:30:32.900997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:30:32.900997Z digest=sha256:c1468c6c514ebd7d5a3b01a157525130d4294dd890567efba61212b76d27c34a

Observation 5b0db90e-4a28-4768-967f-fc6571f623e4 · outbound

This paper cites an unresolved cited work.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:30:33.570754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:30:32.904426Z digest=sha256:8eb24f82b959ee7553cf61958e7e42489fec8c1698914821fd36881f4c367558

Observation 9fd35efb-5bd5-4fc8-b149-aa5b49f6e6b2 · outbound

This paper cites Esser, S.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement Esser, S

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:30:33.559434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:30:32.907803Z digest=sha256:df9a8320d41e83c0b3f00cd8c9718a8116172f72d1aaef0b38340abf94b51090

Observation d13d751c-5cc6-4351-86f9-1e9a92164b10 · outbound

This paper cites SkyReels-A2: Compose Anything in Video Diffusion Transformers.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement SkyReels-A2: Compose Anything in Video Diffusion Transformers

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T05:30:32.911278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:30:32.911278Z digest=sha256:85c38f8c3b9de75a41bbc6ec35b17b36fd887cbc51dc7dd6496c4dce093dc824

Observation 641b8e7e-ff02-4cd9-8491-e5529d7b9da5 · outbound

This paper cites AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T05:30:32.915736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:30:32.915736Z digest=sha256:05b308ba27b0cbf78f6bcbab19aa5fc23955447a917d812f2763ba6b371deb75

Observation 71e00960-50b1-449a-aae8-59de3dc5b93f · outbound

This paper cites Hailuo.https://hailuoai.video/, 2025.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement Hailuo.https://hailuoai.video/, 2025

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:30:33.549238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:30:32.919468Z digest=sha256:75d2bcd23d95c695d8c40aa3c2eb2f5f19d4f12b6d32f6a9a19b60e56300dbd4

Observation 09860b72-42ef-44ba-9626-a86e246c6e5c · outbound

This paper cites ID-Animator: Zero-Shot Identity-Preserving Human Video Generation.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement ID-Animator: Zero-Shot Identity-Preserving Human Video Generation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T05:30:32.923584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:30:32.923584Z digest=sha256:e6f6a86203fa86625aa22571ded85fb6451ba350871288fe181d82ed864af129

Observation 51327993-868a-4a4a-8765-b91f4ff8d639 · outbound

This paper cites an unresolved cited work.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T05:30:32.927378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:30:32.927378Z digest=sha256:3cd08055cadaa85121db3b73a5c86465e176037f86de4900272c4b959535ebf5

Observation 6ee2d3f2-0ecd-49bb-9c5e-6daf89019b76 · outbound

This paper cites MotionMaster: Training-free Camera Motion Transfer For Video Generation.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement MotionMaster: Training-free Camera Motion Transfer For Video Generation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T05:30:32.930745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:30:32.930745Z digest=sha256:72f2264af414683a323e8cd42c8efbcf209b025aeed52626b9bc54f38cbf41ee

Observation 788063b8-c336-4922-8764-25d8cbfc7f70 · outbound

This paper cites HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T05:30:32.934408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:30:32.934408Z digest=sha256:d34d4dbcaf28f6853bc3174be2a97d534036f7ec6a07eef228a6496eb38ebe7e

Observation a2ebc5f8-5042-4246-882d-8f37098a098e · outbound

This paper cites ConceptMaster: Multi-Concept Video Customization on Diffusion Transformer Models Without Test-Time Tuning.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement ConceptMaster: Multi-Concept Video Customization on Diffusion Transformer Models Without Test-Time Tuning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T05:30:32.938000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:30:32.938000Z digest=sha256:94ffbdf40e9518d39fe8ff3b724dacf3e485bf62c754e88c158f4a3366fbcc44

Observation 540271da-c943-4812-9335-e520231aefe1 · outbound

This paper cites Huang, Y.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement Huang, Y

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T05:30:32.942747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:30:32.942747Z digest=sha256:7b3b3184db7930bb862318f8a39d129cd10f78a0f15bbf4438d3cd188d38ef63

Observation 52ecb108-716a-4ba1-91b0-b5532fb7517b · outbound

This paper cites Jiang, T.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement Jiang, T

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:30:33.525452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:30:32.946940Z digest=sha256:d74810783a7bbf3705ae7ffef3d6caf9420e2cdc35d4dde0c9eae4bee5dd7728

Observation adf7c961-802f-4aa7-b3aa-19e0f32162d4 · outbound

This paper cites VACE: All-in-One Video Creation and Editing.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement VACE: All-in-One Video Creation and Editing

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T05:30:32.950352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:30:32.950352Z digest=sha256:9bf4055a7609734b5b28985931c60553dddf22d7673f610158afb9706acb2724

Observation f88c9a3f-7b87-4f0e-b7b8-ba0fa0638e13 · outbound

This paper cites Keling.https://klingai.com/cn/, 2025.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement Keling.https://klingai.com/cn/, 2025

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:30:33.515383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:30:32.953610Z digest=sha256:c9221e54b2cf31dd1398b36ef4917619697d0d981d9c02cfa5105c144e0770f0

Observation 7c5a358e-d494-44a7-bdc6-fdf727b93f40 · outbound

This paper cites YOLOv11: An Overview of the Key Architectural Enhancements.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement YOLOv11: An Overview of the Key Architectural Enhancements

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T05:30:32.956666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:30:32.956666Z digest=sha256:8a6beb90d5ecf716f650720d417845bf1fbf0d85638d63f9e85a23f40ff22e7e

Observation 852fb42d-a9c3-4703-b06c-181505fcf0b5 · outbound

This paper cites HunyuanVideo: A Systematic Framework For Large Video Generative Models.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T05:30:32.960064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:30:32.960064Z digest=sha256:3d485f908629ceff6b61cde87914d8c81fd15d62351574ba95b5db73e5bebb06

Observation b117c48f-1a0b-4a48-bd43-d0cf9555ecd3 · outbound

This paper cites an unresolved cited work.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:30:33.505384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:30:32.963301Z digest=sha256:85019c2aae557ff6b7e7cb6d2e712457bfa0941f27266d5a51649ef3dce60949

Observation 20c8f669-c750-4943-8128-6c68e642efac · outbound

This paper cites OmniV2V: Versatile Video Generation and Editing via Dynamic Content Manipulation.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement OmniV2V: Versatile Video Generation and Editing via Dynamic Content Manipulation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T05:30:32.966807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:30:32.966807Z digest=sha256:dc5d754b8c8df5cd18560a48f43f8bf6bd5d4dc4314a495d97d1c6155acb74a2

Observation 7ad2a0c6-d9a0-4088-b3d5-d113549528fa · outbound

This paper cites an unresolved cited work.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T05:30:32.970705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:30:32.970705Z digest=sha256:5ba4aa732c6e08d7263a4e5c71bda27685039aff3aabc0f5a776e6919ba6e73b

Observation e62295d4-8fa2-4396-8deb-fac1b7d52bf6 · outbound

This paper cites Phantom: Subject-consistent video generation via cross-modal alignment.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement Phantom: Subject-consistent video generation via cross-modal alignment

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T05:30:32.974027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:30:32.974027Z digest=sha256:b496b3638acd8ecf6a627d35d3e0280c61a2472163136b9fdd6058412f6f9649

Observation 98a3f92e-9485-4ece-8bf1-d5fcfb7d2279 · outbound

This paper cites an unresolved cited work.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:30:33.489553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:30:32.978124Z digest=sha256:3f8be73e6ded177bf7c57586de08fdab76b49284c72c8b0ac4ae629c2abf600e

Observation da3df5d8-0e4d-47c9-b56b-87c5fdc313fc · outbound

This paper cites Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T05:30:32.981781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:30:32.981781Z digest=sha256:1afffeb29012587708e26fb39b0923b3fde13fde1c9a6c17606790d92dec3b38

Observation 9a65e3a7-a8e5-431d-b192-dfaedde6aae6 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement DINOv2: Learning Robust Visual Features without Supervision

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T05:30:32.985757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:30:32.985757Z digest=sha256:77b87d8004fe6e5889ace93508cab0739c745de970a7b752e01996a9064841b2

Observation 20e6f371-1ccb-40d5-bbc1-4607a7dcaf99 · outbound

This paper cites Peebles and S.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement Peebles and S

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T05:30:32.989540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:30:32.989540Z digest=sha256:fc5d10e6865c8a67389c74e86378bb6ba03b2b0b8a82de2064e219206290dd00

Observation eba37099-d398-4880-b291-d32f32d5b804 · outbound

This paper cites Pika.https://pika.art/, 2025.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement Pika.https://pika.art/, 2025

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T05:30:32.994167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:30:32.994167Z digest=sha256:43e32f6fecfa7bdcfc04f6293352e0e501ebd4db87daf63f62aec18ddacd4357

Observation fe2140d3-0410-4633-b890-35f32edab8e9 · outbound

This paper cites Movie Gen: A Cast of Media Foundation Models.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement Movie Gen: A Cast of Media Foundation Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T05:30:32.997934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:30:32.997934Z digest=sha256:149251417a59a5244d4a1b3d7c29902604217f57802478823fce342352177311

Observation 1193e3a8-e1ee-497a-bbb3-d7bfede70c0f · outbound

This paper cites Radford, J.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement Radford, J

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T05:30:33.001566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:30:33.001566Z digest=sha256:f23bc9080b515e90f19d987a7a7adb2c214c6ac1f5e0df5b3e13dd2cd61c0f4c

Observation d58ba38d-d620-4cb4-b579-57496e270ca5 · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement SAM 2: Segment Anything in Images and Videos

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T05:30:33.004775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:30:33.004775Z digest=sha256:17f8ec25d5e48f0b746075e8b3545ab1c97789885b2b794d7057bd7d479e1012

Observation 6bd4dd8c-3ca5-418b-9e70-6d987e5a7d91 · outbound

This paper cites Rombach, A.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement Rombach, A

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T05:30:33.008178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:30:33.008178Z digest=sha256:d8b0b2cf2213373db279373f196148a5aad7141fc88f8852a0440334f177f90c

Observation 22903861-99de-4589-9c12-032ff825573f · outbound

This paper cites an unresolved cited work.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:30:33.456056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:30:33.012354Z digest=sha256:18211db938531ffd4bd0c8c81cc792d51b9775aded6854f3425d50c550fc058c

Observation 2170f9ea-fbf8-402f-9956-73cff330ae2a · outbound

This paper cites OminiControl: Minimal and Universal Control for Diffusion Transformer.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement OminiControl: Minimal and Universal Control for Diffusion Transformer

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T05:30:33.016398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:30:33.016398Z digest=sha256:bd7ada0bd671214df2ae5b8800e3b9ba9f89060baad9ce7a850518aff8a62041

Observation 8ab8caa8-94fe-4a1f-bd69-fe8ea910bf47 · outbound

This paper cites Vidu.https://www.vidu.cn/, 2025.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement Vidu.https://www.vidu.cn/, 2025

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:30:33.444854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:30:33.020971Z digest=sha256:1b74997349f1b315b84b220d87ed29af00fa2147733168e2f7a276d69aae1a1c

Observation aaa5cc40-d3d4-4670-90f3-c313426ea273 · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement Wan: Open and Advanced Large-Scale Video Generative Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T05:30:33.024592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:30:33.024592Z digest=sha256:27f48f9d991b28e6b4736a02d6de07f155e13f6974b032ed9ea0c0078a8cc5c1

Observation b968ddfb-b492-40fb-9464-ff3a9d280d70 · outbound

This paper cites Koala-36M: A Large-scale Video Dataset Improving Consistency between Fine-grained Conditions and Video Content.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement Koala-36M: A Large-scale Video Dataset Improving Consistency between Fine-grained Conditions and Video Content

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T05:30:33.029140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:30:33.029140Z digest=sha256:98db1ba5ff478bc206c72710aee0e262a9ed979ee2d731e93b123bc71718c836

Observation 05864c36-0f1c-4b0e-8657-9709df16deae · outbound

This paper cites an unresolved cited work.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement Unresolved cited work

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T05:30:33.033126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:30:33.033126Z digest=sha256:e6674328891ff2fce9a2529cb81fa05d0f4dab494d0af258a8b50db40337c27e

Observation 71b94bfd-edf0-4df3-b9c6-891b5e00de7a · outbound

This paper cites an unresolved cited work.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:30:33.434457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:30:33.037361Z digest=sha256:04473e2f8acbd481f47062942bd0267b93a7d049d446150971124cf9f70b6004

Observation 09a7381a-dea6-4daf-adfb-680a20eb4248 · outbound

This paper cites an unresolved cited work.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:30:33.423512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:30:33.041175Z digest=sha256:bc93097a8b6d57b137e6ba0b949f5c3741a1f7bb88b947afaac2f4b2f158367c

Observation f8307a88-ae46-4052-bf27-b46c3ef04f15 · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T05:30:33.044724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:30:33.044724Z digest=sha256:70009ba20e73c3e9a170bc97d9230d1b8f53526e7f5a3932e7b4177407a0d87f

Observation 97b49b59-a33d-416b-bff2-fb5d884efcde · outbound

This paper cites IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T05:30:33.048874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:30:33.048874Z digest=sha256:67708510fbb021f8d3f92316bb7f94afde4344c1ee9785f62392c857345cc3aa

Observation c3761732-b039-47c0-a26f-1e59f3b149d6 · outbound

This paper cites Identity-Preserving Text-to-Video Generation by Frequency Decomposition.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement Identity-Preserving Text-to-Video Generation by Frequency Decomposition

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T05:30:33.052194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:30:33.052194Z digest=sha256:045cc74f9cf3a635097349b4c2e10c1fa185d647f1ccdd14c3e9d37414237628

Observation ef30a557-886f-49d0-8ac8-70202a503d7d · outbound

This paper cites Allegro: Open the Black Box of Commercial-Level Video Generation Model.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement Allegro: Open the Black Box of Commercial-Level Video Generation Model

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T05:30:33.055697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:30:33.055697Z digest=sha256:dbcbe1a5f30f928a3842dc6d77fb2e0dd5f748d4792b749d4d951d14fe03eb40

Pith citing papers

Observation 3d494d21-a065-448a-af60-769d4e10c20c · inbound

Evolution of Video Generative Foundations cites this paper.

Evolution of Video Generative Foundations PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement

Reference 206

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:05:51.123093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T18:41:38.616611Z digest=sha256:a4f1f39d76d63e710cf05d1aa97b04b9ea3c65db0bcb0da0f033fe0bc203ec38

Observation 3d7b7d1c-4e97-49d1-9e0f-205daa78cebf · inbound

PresentAgent-2: Towards Generalist Multimodal Presentation Agents cites this paper.

PresentAgent-2: Towards Generalist Multimodal Presentation Agents PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:32:06.483612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-13T02:29:42.157339Z digest=sha256:e760a73f97c147dd04859f60f4c0d1633d784bd87d92201cd4d88b6deca998bf

Observation 49a8ca39-c798-42c2-9a7a-58f88f3ce637 · inbound

MetaWorld: Scaling Multi-Agent Video World Model from Single-view Video Data cites this paper.

MetaWorld: Scaling Multi-Agent Video World Model from Single-view Video Data PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:56:20.171795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-28T14:52:30.406683Z digest=sha256:417bea78686ca12857a31e00a87adaf4617075db44a635c38950cec70ee8674a

Observation f1f519fa-6a6d-4a2d-baa5-688f1e0aaf91 · inbound

Aura: Consistent Multi-Subject Video Generation via VLM-Grounded Semantic Alignment cites this paper.

Aura: Consistent Multi-Subject Video Generation via VLM-Grounded Semantic Alignment PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-11T20:11:31.576642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T20:11:31.576642Z digest=sha256:7af3114f30fa1ef77023d01375fd08daa69a8ac54649c67fdb4c080ad69ec3a9