Pith. sign in

Paper Citation Record · LEDGER

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement

As of 9 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 4 inbound Pith citation observations for arXiv:2506.07848.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.07848 v1

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:30:33.055697Z

measured 52 of 52 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-11T20:11:31.576642Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T22:56:20.169424Z

Reference resolution

48 of 48 outbound references displayed

  • verified exact0
  • verified fuzzy7
  • unresolved41
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c4dcf0d4-b5af-472c-bf9a-4064cd2a576f · outbound

This paper cites Qwen2.5-VL Technical Report.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement Qwen2.5-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T05:30:32.874829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:30:32.874829Z digest=sha256:605a5f0a8afaa923e3b5c038649e0d288bfeef6b03304d12f7541c34775a9b99

Observation 65232c13-4a34-4297-b1a9-9fa062e9ebd3 · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T05:30:32.880099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:30:32.880099Z digest=sha256:0ce90642eafc41fe812e8a3a92204baccd4f4fe211960d938a9922949b38c2b5

Observation 26b2b07c-bdf0-48a2-a14d-bf3ccb8b3f0a · outbound

This paper cites Carreira and A.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement Carreira and A

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:30:33.612763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:30:32.884690Z digest=sha256:dc0cd8ab012916a75ed4e4b02b8af2e6646fcc0f007f77ff1667885644af3291

Observation 9e15de1b-3f68-4af7-9c4c-21e0d3498141 · outbound

This paper cites Chefer, S.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement Chefer, S

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:30:33.602556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:30:32.888262Z digest=sha256:909ada1a4b05a3642c2dc4649576501407d592bd5224c2acf669a3c0c38e411d

Observation acd89131-1c62-4c1a-9edb-ce308a4e98e7 · outbound

This paper cites an unresolved cited work.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:30:33.592113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:30:32.892748Z digest=sha256:2a6cd789237a9c274811d7337b20297aac84c650584c51bc9c1573e4f14e7160

Observation 014d3878-8673-46d3-a57e-5a234eb77074 · outbound

This paper cites an unresolved cited work.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:30:33.581859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:30:32.896304Z digest=sha256:6a564be4d74025678c6a622ff14b0351d4d90d522dd2e8f06aa26284c017b350

Observation e7ee9805-6d3c-4c63-8ac3-f8620a328848 · outbound

This paper cites Multi-subject Open-set Personalization in Video Generation.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement Multi-subject Open-set Personalization in Video Generation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T05:30:32.900997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:30:32.900997Z digest=sha256:13e2762ea05d8a47d423b22e33daf8204df1d181cefe2e9355cf82a42d1e34f5

Observation 5b0db90e-4a28-4768-967f-fc6571f623e4 · outbound

This paper cites an unresolved cited work.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:30:33.570754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:30:32.904426Z digest=sha256:9ed00dd3d6ac25e0d7df98bfc733dea70fa83ac03fc87ca2adeacedbc9e6c9ce

Observation 9fd35efb-5bd5-4fc8-b149-aa5b49f6e6b2 · outbound

This paper cites Esser, S.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement Esser, S

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:30:33.559434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:30:32.907803Z digest=sha256:d6e94bc80e3a346b5f900a7e3b9e250656971e719a11a2a9124dfcb8c99f6db7

Observation d13d751c-5cc6-4351-86f9-1e9a92164b10 · outbound

This paper cites SkyReels-A2: Compose Anything in Video Diffusion Transformers.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement SkyReels-A2: Compose Anything in Video Diffusion Transformers

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T05:30:32.911278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:30:32.911278Z digest=sha256:7971fda39e8d26224385a0c82d2f5f34c2dd645ef809ee6584f5680e18c729b5

Observation 641b8e7e-ff02-4cd9-8491-e5529d7b9da5 · outbound

This paper cites AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T05:30:32.915736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:30:32.915736Z digest=sha256:88e8208c72d664b1f01faba3636ba87ff8cc9f826c6a3be4fd6204c303244cad

Observation 71e00960-50b1-449a-aae8-59de3dc5b93f · outbound

This paper cites Hailuo.https://hailuoai.video/, 2025.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement Hailuo.https://hailuoai.video/, 2025

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:30:33.549238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:30:32.919468Z digest=sha256:57333b18ebdf5c85420a89a5405bc1919e9385e64e13c7e13a719452f8f955da

Observation 09860b72-42ef-44ba-9626-a86e246c6e5c · outbound

This paper cites ID-Animator: Zero-Shot Identity-Preserving Human Video Generation.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement ID-Animator: Zero-Shot Identity-Preserving Human Video Generation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T05:30:32.923584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:30:32.923584Z digest=sha256:1dbf5416ddb1244c89c95854870ba2a008feb3c4037dc3d8e2814c0542e1382a

Observation 51327993-868a-4a4a-8765-b91f4ff8d639 · outbound

This paper cites an unresolved cited work.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T05:30:32.927378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:30:32.927378Z digest=sha256:2650a79574f2bdfc8aa1d3f241e8086bfe168a8bd47f32b243f06cbdf5be1ec8

Observation 6ee2d3f2-0ecd-49bb-9c5e-6daf89019b76 · outbound

This paper cites MotionMaster: Training-free Camera Motion Transfer For Video Generation.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement MotionMaster: Training-free Camera Motion Transfer For Video Generation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T05:30:32.930745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:30:32.930745Z digest=sha256:ac87b7ab2c18a3b6854d6dc25bfa0747f2ca882f20bfc35a7801ba5ca04fc957

Observation 788063b8-c336-4922-8764-25d8cbfc7f70 · outbound

This paper cites HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T05:30:32.934408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:30:32.934408Z digest=sha256:496eb85992a6280f2f0b1afc0cb57586b87f35b90ae28568a9f4445aef16677a

Observation a2ebc5f8-5042-4246-882d-8f37098a098e · outbound

This paper cites ConceptMaster: Multi-Concept Video Customization on Diffusion Transformer Models Without Test-Time Tuning.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement ConceptMaster: Multi-Concept Video Customization on Diffusion Transformer Models Without Test-Time Tuning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T05:30:32.938000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:30:32.938000Z digest=sha256:d31b6e773f71a6e9ec930844c142695d18400d8dd965575737d91b576337a4d2

Observation 540271da-c943-4812-9335-e520231aefe1 · outbound

This paper cites Huang, Y.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement Huang, Y

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T05:30:32.942747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:30:32.942747Z digest=sha256:65c7970daa73b77f1b0231fa59f5ebed8db197d901d401e3f10630c2d62fd308

Observation 52ecb108-716a-4ba1-91b0-b5532fb7517b · outbound

This paper cites Jiang, T.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement Jiang, T

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:30:33.525452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:30:32.946940Z digest=sha256:0a780d9831c5c86370706c4b5352fa3c50204794692ce20a043f236a6b163151

Observation adf7c961-802f-4aa7-b3aa-19e0f32162d4 · outbound

This paper cites VACE: All-in-One Video Creation and Editing.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement VACE: All-in-One Video Creation and Editing

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T05:30:32.950352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:30:32.950352Z digest=sha256:85d35fae460f773dba689406006a42f17e4996dba49690f86f0ed5047bb82fcc

Observation f88c9a3f-7b87-4f0e-b7b8-ba0fa0638e13 · outbound

This paper cites Keling.https://klingai.com/cn/, 2025.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement Keling.https://klingai.com/cn/, 2025

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:30:33.515383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:30:32.953610Z digest=sha256:87739d5ab8183905e2e0cbfb1ef8178c3989828464afe1529d95abcbe9a4dfee

Observation 7c5a358e-d494-44a7-bdc6-fdf727b93f40 · outbound

This paper cites YOLOv11: An Overview of the Key Architectural Enhancements.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement YOLOv11: An Overview of the Key Architectural Enhancements

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T05:30:32.956666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:30:32.956666Z digest=sha256:b621a62c42c912b1a9aff5c65f326022ac865e757687ec4322915015a799c580

Observation 852fb42d-a9c3-4703-b06c-181505fcf0b5 · outbound

This paper cites HunyuanVideo: A Systematic Framework For Large Video Generative Models.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T05:30:32.960064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:30:32.960064Z digest=sha256:8fcfd631bf23560f25605d1cf074a6f6bb26e71f9d99d117ed0ca89df1cb6211

Observation b117c48f-1a0b-4a48-bd43-d0cf9555ecd3 · outbound

This paper cites an unresolved cited work.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:30:33.505384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:30:32.963301Z digest=sha256:02ec845f3780869c26901bf87211c33eeb39130877771e8d003e4295a02ec1b3

Observation 20c8f669-c750-4943-8128-6c68e642efac · outbound

This paper cites OmniV2V: Versatile Video Generation and Editing via Dynamic Content Manipulation.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement OmniV2V: Versatile Video Generation and Editing via Dynamic Content Manipulation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T05:30:32.966807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:30:32.966807Z digest=sha256:3da8fecbcf9434f93e7ab460fba7b8a432ae7dcf476ec881016965e71593c7bc

Observation 7ad2a0c6-d9a0-4088-b3d5-d113549528fa · outbound

This paper cites an unresolved cited work.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T05:30:32.970705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:30:32.970705Z digest=sha256:2f5e36312aeb0bd1f705fe60ce82d7bf856dd4839c2c807c9bcf408d647ca8a3

Observation e62295d4-8fa2-4396-8deb-fac1b7d52bf6 · outbound

This paper cites Phantom: Subject-consistent video generation via cross-modal alignment.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement Phantom: Subject-consistent video generation via cross-modal alignment

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T05:30:32.974027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:30:32.974027Z digest=sha256:df40bb4214174f1fe3d8b745999a3da90acdf05b7bedfc499586401e860d88db

Observation 98a3f92e-9485-4ece-8bf1-d5fcfb7d2279 · outbound

This paper cites an unresolved cited work.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:30:33.489553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:30:32.978124Z digest=sha256:7eca3d02a2e43d5dfaa401bcea6454ff16f6590e4c8e77be673797610157b593

Observation da3df5d8-0e4d-47c9-b56b-87c5fdc313fc · outbound

This paper cites Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T05:30:32.981781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:30:32.981781Z digest=sha256:6d5df60ba0a7f776121df1654721d97d8baeeaa305d86566edcc9a0955b299ca

Observation 9a65e3a7-a8e5-431d-b192-dfaedde6aae6 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement DINOv2: Learning Robust Visual Features without Supervision

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T05:30:32.985757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:30:32.985757Z digest=sha256:36954346f91889a0d4d60ae2205daf653d5ba8dcb9afc01472c4f5602d1b8535

Observation 20e6f371-1ccb-40d5-bbc1-4607a7dcaf99 · outbound

This paper cites Peebles and S.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement Peebles and S

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T05:30:32.989540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:30:32.989540Z digest=sha256:e34a905948b8a050216fadf76886d3b37e40ed3c3bae890ff9a780131e4d74ba

Observation eba37099-d398-4880-b291-d32f32d5b804 · outbound

This paper cites Pika.https://pika.art/, 2025.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement Pika.https://pika.art/, 2025

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T05:30:32.994167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:30:32.994167Z digest=sha256:0daa1d3f81b0a9bdd553d5e345a44adf4e7f0f38ad2c3d0f263ee6e9dfd14cd8

Observation fe2140d3-0410-4633-b890-35f32edab8e9 · outbound

This paper cites Movie Gen: A Cast of Media Foundation Models.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement Movie Gen: A Cast of Media Foundation Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T05:30:32.997934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:30:32.997934Z digest=sha256:f41a3efdda60c94d2c868206b2872f2b840cd129b9dfe5edeac5be5fdee9198f

Observation 1193e3a8-e1ee-497a-bbb3-d7bfede70c0f · outbound

This paper cites Radford, J.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement Radford, J

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T05:30:33.001566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:30:33.001566Z digest=sha256:bdc1d4cef6f50a4d0268b51d74dac6f1dbdd21faffb6346d8cbe2f52eefcdd4c

Observation d58ba38d-d620-4cb4-b579-57496e270ca5 · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement SAM 2: Segment Anything in Images and Videos

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T05:30:33.004775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:30:33.004775Z digest=sha256:c052bd9f08dcefdaef5b18ad40a100c19082e6f0415e9e3345c788fff4cd0be5

Observation 6bd4dd8c-3ca5-418b-9e70-6d987e5a7d91 · outbound

This paper cites Rombach, A.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement Rombach, A

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T05:30:33.008178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:30:33.008178Z digest=sha256:ff4d5551ed394b7d03da82c6bfc609e24d7a5b276ae17c1cd9994eafe82af4eb

Observation 22903861-99de-4589-9c12-032ff825573f · outbound

This paper cites an unresolved cited work.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:30:33.456056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:30:33.012354Z digest=sha256:37d17414ed4af26bdf8b76ff6da7e4d7dd8aeeeb6479f78940a9f4c31b72cae5

Observation 2170f9ea-fbf8-402f-9956-73cff330ae2a · outbound

This paper cites OminiControl: Minimal and Universal Control for Diffusion Transformer.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement OminiControl: Minimal and Universal Control for Diffusion Transformer

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T05:30:33.016398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:30:33.016398Z digest=sha256:8949bd8eb5c30c9be17899f67f4c4de57d8903c0f4b7331d012a5a6c952c6767

Observation 8ab8caa8-94fe-4a1f-bd69-fe8ea910bf47 · outbound

This paper cites Vidu.https://www.vidu.cn/, 2025.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement Vidu.https://www.vidu.cn/, 2025

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:30:33.444854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:30:33.020971Z digest=sha256:b843e9a2f3ead3496c6ecf2837abeacbf50dc80beabdbe7c0e4ac0a504fc3aed

Observation aaa5cc40-d3d4-4670-90f3-c313426ea273 · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement Wan: Open and Advanced Large-Scale Video Generative Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T05:30:33.024592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:30:33.024592Z digest=sha256:4a5e6c3f7244ea5d319ec2551c32337bf462a666c9e0f264b4a1fa7cccfb8af2

Observation b968ddfb-b492-40fb-9464-ff3a9d280d70 · outbound

This paper cites Koala-36M: A Large-scale Video Dataset Improving Consistency between Fine-grained Conditions and Video Content.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement Koala-36M: A Large-scale Video Dataset Improving Consistency between Fine-grained Conditions and Video Content

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T05:30:33.029140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:30:33.029140Z digest=sha256:a5063a1fb30ed729f4b286d374aca67f81fa5d4e287d31ee7b7fc6e8cf5e9cf5

Observation 05864c36-0f1c-4b0e-8657-9709df16deae · outbound

This paper cites an unresolved cited work.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement Unresolved cited work

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T05:30:33.033126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:30:33.033126Z digest=sha256:e26a09bfd87b55921268593f4dad4809cfff8516c56f5df40f2df8ffc428fdd9

Observation 71b94bfd-edf0-4df3-b9c6-891b5e00de7a · outbound

This paper cites an unresolved cited work.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:30:33.434457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:30:33.037361Z digest=sha256:160b31f6fd82799b951b9305aeed6c38c14957cf0df6be2bde98810e0c034f9b

Observation 09a7381a-dea6-4daf-adfb-680a20eb4248 · outbound

This paper cites an unresolved cited work.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:30:33.423512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:30:33.041175Z digest=sha256:fe1610931dd5b78a382719c9be5fc8af94426bcfeb1b5c940f4e85007a6d52e8

Observation f8307a88-ae46-4052-bf27-b46c3ef04f15 · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T05:30:33.044724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:30:33.044724Z digest=sha256:e2d4399d4a7617a47ae1ef7daccf7b26f3dfe2ccf8a818f63f38fe4df42493e6

Observation 97b49b59-a33d-416b-bff2-fb5d884efcde · outbound

This paper cites IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T05:30:33.048874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:30:33.048874Z digest=sha256:edda08e59c1fd1a8d9ccb5cd10146fc139d001f3781326510ea63ac517f616e5

Observation c3761732-b039-47c0-a26f-1e59f3b149d6 · outbound

This paper cites Identity-Preserving Text-to-Video Generation by Frequency Decomposition.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement Identity-Preserving Text-to-Video Generation by Frequency Decomposition

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T05:30:33.052194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:30:33.052194Z digest=sha256:1888004c9bfd4549672cb8bd0426540fb027bde310b5cc869fde5791e40533e9

Observation ef30a557-886f-49d0-8ac8-70202a503d7d · outbound

This paper cites Allegro: Open the Black Box of Commercial-Level Video Generation Model.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement Allegro: Open the Black Box of Commercial-Level Video Generation Model

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T05:30:33.055697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:30:33.055697Z digest=sha256:277cd4889b1ab5b69519f90cba0606fdb6e856933562dc978fd75fe14b82d307

Pith citing papers

Observation 3d494d21-a065-448a-af60-769d4e10c20c · inbound

Evolution of Video Generative Foundations cites this paper.

Evolution of Video Generative Foundations PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement

Reference 206

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:05:51.123093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T18:41:38.616611Z digest=sha256:917901a461761e9b01e829096638e57140797ccf6dd14a57fa760b0160aac102

Observation 3d7b7d1c-4e97-49d1-9e0f-205daa78cebf · inbound

PresentAgent-2: Towards Generalist Multimodal Presentation Agents cites this paper.

PresentAgent-2: Towards Generalist Multimodal Presentation Agents PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:32:06.483612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T02:29:42.157339Z digest=sha256:d733fc8eed899d562b3d63f2787e96998f15d9787998f14291bfa8bb89a858f4

Observation 49a8ca39-c798-42c2-9a7a-58f88f3ce637 · inbound

MetaWorld: Scaling Multi-Agent Video World Model from Single-view Video Data cites this paper.

MetaWorld: Scaling Multi-Agent Video World Model from Single-view Video Data PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:56:20.171795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T14:52:30.406683Z digest=sha256:5bd91b5fbff3c500c13d353233da995c9542ff7a046afac04db2eb8d1df3de23

Observation f1f519fa-6a6d-4a2d-baa5-688f1e0aaf91 · inbound

Aura: Consistent Multi-Subject Video Generation via VLM-Grounded Semantic Alignment cites this paper.

Aura: Consistent Multi-Subject Video Generation via VLM-Grounded Semantic Alignment PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-11T20:11:31.576642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T20:11:31.576642Z digest=sha256:2cc9786db928fcd849cad3fa14403d4e55f8fad75e847ca4d5d8fffcb1ea8d4a