Pith. sign in

Paper Citation Record · LEDGER

MOVi: Training-free Text-conditioned Multi-Object Video Generation

As of 8 August 2026, this Paper Citation Record lists 67 of 67 outbound references and 0 inbound Pith citation observations for arXiv:2505.22980.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.22980 v1

Coverage vector

measured 67 of 67 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:01:02.423170Z

measured 67 of 67 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

67 of 67 outbound references displayed

  • verified exact1
  • verified fuzzy34
  • unresolved32
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6c51cfaf-ecdf-40d7-ad8c-681e9defe6fb · outbound

This paper cites GPT-4 Technical Report.

MOVi: Training-free Text-conditioned Multi-Object Video Generation GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T13:00:56.424420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:00:56.424420Z digest=sha256:7b3b197a3dc12c03d911fefe34699ece6288c52280ecbe5bb2881382fc1b4ed2

Observation 1d82a762-a827-434e-bece-cc68cfb7d6f8 · outbound

This paper cites Frozen in time: A joint video and im- age encoder for end-to-end retrieval.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Frozen in time: A joint video and im- age encoder for end-to-end retrieval

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:11.203309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:00:56.521401Z digest=sha256:206fd6068a88b61f4e45717bd8d04be7e59f367f2acb44b14fbe903292ad33b7

Observation be240a50-c82f-4585-9e23-60e5a9698fa5 · outbound

This paper cites Lumiere: A Space-Time Diffusion Model for Video Generation.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Lumiere: A Space-Time Diffusion Model for Video Generation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T13:00:56.611108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:00:56.611108Z digest=sha256:1848c0a403875f83746d682926b2fb9b7fbbbd346e325750e16dae0ff5048885

Observation 43d0e7af-36c5-4746-99e5-dd6195be5d4a · outbound

This paper cites Multidiffusion: Fusing diffusion paths for con- trolled image generation.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Multidiffusion: Fusing diffusion paths for con- trolled image generation

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:10.845129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:00:56.724522Z digest=sha256:9214ea7957c021ce8ca7b64ac92b42b3eed058acfb33fdc5cf7fec612b9fc14b

Observation 842663ce-81b1-4b8c-8aa1-fa07edf8a6ab · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T13:00:56.872776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:00:56.872776Z digest=sha256:646db918c99d54897443e6fb15225b80234f254745a380ae080aa87aa9b3da5a

Observation 36d37ad9-256f-4803-b3a0-9f2618fed3e0 · outbound

This paper cites Align your latents: High-resolution video synthesis with latent diffusion models.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Align your latents: High-resolution video synthesis with latent diffusion models

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:10.495466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:00:56.970597Z digest=sha256:fa1e11435e95afbb2f978eefea89420ad48046f0b654ee974cb9ecf7a3d6b908

Observation a7ccfdea-b328-485b-9aa7-2f6d59470f66 · outbound

This paper cites VideoDreamer: Customized Multi-Subject Text-to-Video Generation with Disen-Mix Finetuning on Language-Video Foundation Models.

MOVi: Training-free Text-conditioned Multi-Object Video Generation VideoDreamer: Customized Multi-Subject Text-to-Video Generation with Disen-Mix Finetuning on Language-Video Foundation Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T13:00:57.039138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:00:57.039138Z digest=sha256:1316934025d78959de8d85ababb08220bf9b3f9df0c8f9534a6a68627d85088a

Observation 9b6071c2-fd71-4667-990e-1f617c379fd0 · outbound

This paper cites Videocrafter2: Overcoming data limitations for high- quality video diffusion models.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Videocrafter2: Overcoming data limitations for high- quality video diffusion models

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:10.125256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:00:57.143508Z digest=sha256:78345a517ddf2615afbbcfefbc8b66b6208714c967c02c9e8a06866f1a68814e

Observation b335be6a-f2fe-4445-abf3-c8838c6ffc80 · outbound

This paper cites Panda-70m: Captioning 70m videos with multiple cross-modality teachers.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Panda-70m: Captioning 70m videos with multiple cross-modality teachers

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:09.819009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:00:57.188341Z digest=sha256:e3191badcb57e06ca4ea68a9a27d5c56784bf38176782ebe0269a4a04adfb1e6

Observation 86692de2-767f-4608-8f4a-42cd3c20b68f · outbound

This paper cites Sora as an agi world model? a complete survey on text-to-video generation.arXiv preprint arXiv:2403.05131, 2024.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Sora as an agi world model? a complete survey on text-to-video generation.arXiv preprint arXiv:2403.05131, 2024

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T13:00:57.269517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:00:57.269517Z digest=sha256:d04fc37231809ba2c37783e1f3b13e671b195cdfa8d6ad722915b1c690490fc6

Observation 9ccaa405-75ef-4749-abaa-783b7ae51a71 · outbound

This paper cites Data-Juicer Sandbox: A Comprehensive Suite for Multimodal Data-Model Co-development.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Data-Juicer Sandbox: A Comprehensive Suite for Multimodal Data-Model Co-development

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:09.551068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:00:57.433196Z digest=sha256:1b45120a61c4d58bd64e13436edf8b37e5f0d4d3afaab6056a346ce269149bd1

Observation 9b25e274-37d8-4f14-9df1-b3ea4c4abc93 · outbound

This paper cites DiffSynth-Studio: Enjoy the magic of Diffusion models!https://github.com/ modelscope/DiffSynth-Studio, 2024.

MOVi: Training-free Text-conditioned Multi-Object Video Generation DiffSynth-Studio: Enjoy the magic of Diffusion models!https://github.com/ modelscope/DiffSynth-Studio, 2024

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:09.166354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:00:57.582602Z digest=sha256:2b742beed9e75b874ea7cb932af550eda488258e0c67b2f0e1253543a8ec1c1f

Observation 3935d16d-c838-4c6e-a3c0-37efa2f38078 · outbound

This paper cites Animatediff: Animate your personalized text-to-image diffusion models without specific tuning.International Conference on Learn- ing Representations, 2024.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Animatediff: Animate your personalized text-to-image diffusion models without specific tuning.International Conference on Learn- ing Representations, 2024

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:08.832656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:00:57.717505Z digest=sha256:745727b26a35230e50cea7a74631ec500eb829c1d7092aeba6153cde5e1c7f89

Observation 37fc4c8d-8a6f-458c-a9a0-ec81da22db58 · outbound

This paper cites Latent Video Diffusion Models for High-Fidelity Long Video Generation.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Latent Video Diffusion Models for High-Fidelity Long Video Generation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T13:00:57.842598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:00:57.842598Z digest=sha256:ab9b42a62f26f105b3efeeaa0a9c72c10ce37cdca8de34cdf81f9f07a0a7468d

Observation 6f32561d-1157-4416-8631-f864ad6099ec · outbound

This paper cites CLIPScore: a reference- free evaluation metric for image captioning.

MOVi: Training-free Text-conditioned Multi-Object Video Generation CLIPScore: a reference- free evaluation metric for image captioning

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:08.484965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:00:57.985309Z digest=sha256:5df51a5e95cfb2d893aa58d45a329c48a95a53610e38160270bbac7b0dee627e

Observation d1933484-4e79-4734-841f-0027a4878843 · outbound

This paper cites Imagen Video: High Definition Video Generation with Diffusion Models.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Imagen Video: High Definition Video Generation with Diffusion Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T13:00:58.093105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:00:58.093105Z digest=sha256:54a33597cede4c572407896065bdbbcf62b68e0a5354728a6c7a1d4962425d5b

Observation 33db6231-caa2-4980-810c-6cb240c89537 · outbound

This paper cites Denois- ing diffusion probabilistic models.Advances in neural information processing systems, 33:6840–6851, 2020.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Denois- ing diffusion probabilistic models.Advances in neural information processing systems, 33:6840–6851, 2020

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T13:00:58.197779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:00:58.197779Z digest=sha256:d1955f35cde3134781e77b3735e8615b8c9d9884e860d12c4f4fed9029ab4570

Observation 9a3b2a17-96da-4b8b-b8e0-5a45bb38fc4c · outbound

This paper cites Video Diffusion Models.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Video Diffusion Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T13:00:58.322441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:00:58.322441Z digest=sha256:b884243f608c0136631acaa63085c40163fe7822fc880ee0394f775a08e74fdb

Observation caa1bd13-a401-4165-ba1b-c5537d101575 · outbound

This paper cites CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers.

MOVi: Training-free Text-conditioned Multi-Object Video Generation CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:00:58.427430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:00:58.427430Z digest=sha256:5d7a10d07460509521e37aed23090cd577cba2a9c91bfe047bb106d8c0a33a3f

Observation 50e588c3-2169-4d3c-acd8-eae01d532af9 · outbound

This paper cites Vbench: Comprehensive benchmark suite for video generative models.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Vbench: Comprehensive benchmark suite for video generative models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:08.315535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:00:58.539695Z digest=sha256:4735d9380118f02b861fedf62184a68fa7d5522e069a82290ea02de3808238fc

Observation 6015e7c6-9ffa-4ff7-b1aa-9ed77f66ca46 · outbound

This paper cites High-quality Text-to-video Models.

MOVi: Training-free Text-conditioned Multi-Object Video Generation High-quality Text-to-video Models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:08.307585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:00:58.658289Z digest=sha256:53c8c4eaab0048ea950a2040fd7fea829f88e683d9864687b42f54962fc8f0d7

Observation b974e219-4fb0-4c1e-ae84-95e7459971bd · outbound

This paper cites KLING AI: Next-Generation AI Creative Studio.https://www.klingai.com/, 2024.

MOVi: Training-free Text-conditioned Multi-Object Video Generation KLING AI: Next-Generation AI Creative Studio.https://www.klingai.com/, 2024

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:08.180073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:00:58.748386Z digest=sha256:3f4d76e83ae0b683e4d2a00fd498f18ced25d9a152d36ef1ee06090a9d3d9066

Observation 272c985e-138a-496a-b615-57a1b4732816 · outbound

This paper cites Multi-concept cus- tomization of text-to-image diffusion.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Multi-concept cus- tomization of text-to-image diffusion

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:07.862669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:00:58.858770Z digest=sha256:5d8dec2f511607d552f07c0f7be63ba49c9ac9f959de98b1529504dfbe9da857

Observation 2233dae1-8b17-4b73-b21d-c9aca587af4f · outbound

This paper cites TrackDiffusion: Tracklet-Conditioned Video Generation via Diffusion Models.

MOVi: Training-free Text-conditioned Multi-Object Video Generation TrackDiffusion: Tracklet-Conditioned Video Generation via Diffusion Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T13:00:58.952626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:00:58.952626Z digest=sha256:d2edafa22b913da7845a6fbd1ee33745149a86e71416792d01da1ab8dcb8d3e5

Observation be3c788d-8ee8-4078-a85b-caa23787f32f · outbound

This paper cites Gligen: Open-set grounded text-to- image generation.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Gligen: Open-set grounded text-to- image generation

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:07.638758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:00:59.025274Z digest=sha256:bbd10c12397c52d11c4ff587aa7d09dcec3b7bbf6f3820235b7d1dfc8dfa6b7d

Observation 80044332-1829-4001-9a49-611707efe2e6 · outbound

This paper cites Movideo: Motion-aware video generation with diffusion model.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Movideo: Motion-aware video generation with diffusion model

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:07.496035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:00:59.091672Z digest=sha256:f961ce2f19eb4f0d53c83d535aa594466c283d2e741411730647e473290abd48

Observation ff403e82-cb7a-4ac2-a7ee-489674d20e87 · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36, 2024.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Visual instruction tuning.Advances in neural information processing systems, 36, 2024

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:07.343576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:00:59.172626Z digest=sha256:0a06230a8feb90113a301fbf628462fd0840463e381f9589945a3facc3950b79

Observation a634ebcf-1c1d-4882-9502-b3558c670a96 · outbound

This paper cites Detector Guidance for Multi-Object Text-to-Image Generation.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Detector Guidance for Multi-Object Text-to-Image Generation

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:01:03.014410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:00:59.268187Z digest=sha256:3a62001af4c2ea19c54c851298e58590e29fa1a3eca3c14917ddcccdf8a4a766

Observation be93cda6-eecb-4039-958a-99d6463f79db · outbound

This paper cites Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T13:00:59.360313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:00:59.360313Z digest=sha256:5449bf76b753ab752f976085faef9cba395ecc989a690177a3cb4b193482c218

Observation 4ab781fb-589e-4f9b-93fb-8ee640cf5507 · outbound

This paper cites Lumaai.https://lumalabs.ai/,.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Lumaai.https://lumalabs.ai/,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:07.181827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:00:59.480131Z digest=sha256:a10088dda8ccfc30f4481a837ac28c77448052440c8061cc5741ec80ac7245fa

Observation 32ae22fd-4d49-4c6d-ad5c-14872d010344 · outbound

This paper cites Vidm: Video implicit diffusion models.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Vidm: Video implicit diffusion models

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:06.675171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:00:59.677588Z digest=sha256:6d147bcf08b4620efcfbbea09c62a639f427445a394094a0ccbfee2d9b2bb6b7

Observation 37de2084-99d7-4691-bb74-cfadb942a24d · outbound

This paper cites Hailuo AI: Captivating AI Videos Gen- erated with Hailuo AI .https://hailuoai.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Hailuo AI: Captivating AI Videos Gen- erated with Hailuo AI .https://hailuoai

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:06.512260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:00:59.744083Z digest=sha256:921167bcae69393043a300912d26a0ba9c31d7f4cb4a88858b688d5078bb5da2

Observation 00d20cf3-5a94-4aab-a214-7fbe8a35b8a0 · outbound

This paper cites Dreamix: Video Diffusion Models are General Video Editors.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Dreamix: Video Diffusion Models are General Video Editors

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T13:00:59.801514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:00:59.801514Z digest=sha256:fd2d4443043b305750cb6ea814a44a829daea5f5f6c032781aee1e414fb8b051

Observation ed0700fb-40f1-46c2-b9df-8dc13412dc53 · outbound

This paper cites WorldSimBench: Towards Video Generation Models as World Simulators.

MOVi: Training-free Text-conditioned Multi-Object Video Generation WorldSimBench: Towards Video Generation Models as World Simulators

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T13:00:59.855458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:00:59.855458Z digest=sha256:7e57410b705b5bdfdfbec9fc5947841fc8169b35f39dd482c6e0de921637db8b

Observation ccd886c3-2fdb-4825-9188-ee47f941fb16 · outbound

This paper cites FreeTraj: Tuning-Free Trajectory Control in Video Diffusion Models.

MOVi: Training-free Text-conditioned Multi-Object Video Generation FreeTraj: Tuning-Free Trajectory Control in Video Diffusion Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T13:00:59.932469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:00:59.932469Z digest=sha256:7af648d9e5eabf306cb0525a7e0421595839fbae34d5dc55cc605e8ce389a3ca

Observation b61f693c-65c2-4189-8e45-07b895997203 · outbound

This paper cites High- resolution image synthesis with latent diffusion mod- els.

MOVi: Training-free Text-conditioned Multi-Object Video Generation High- resolution image synthesis with latent diffusion mod- els

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T13:00:59.991815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:00:59.991815Z digest=sha256:88c3aba80fae4e04e60a1415fc8f44a4008c320188a6576438ed98372648e2ce

Observation 077b15c3-383c-4714-9e5e-ec3e29d03024 · outbound

This paper cites Gen-2: Generate novel videos with text, im- ages or video clips, 2024.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Gen-2: Generate novel videos with text, im- ages or video clips, 2024

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:06.325296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:01:00.054526Z digest=sha256:645a8a32e5be07155c420c0f7077b8b38342697e1fd92413be9d39c07d590ffa

Observation 0d5210f5-31b9-4dd7-9af0-34f0764c78a1 · outbound

This paper cites Introducing Gen-3 Alpha: A New Frontier for Video Generation, 2024.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Introducing Gen-3 Alpha: A New Frontier for Video Generation, 2024

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:06.166910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:01:00.104985Z digest=sha256:81e668e89c0f0752ad8cc98295a844dd692ae8a3c2a7128d06a3859a7c3679ee

Observation 55510d57-fa72-422a-81f8-f760346fd4d4 · outbound

This paper cites Make-A-Video: Text-to-Video Generation without Text-Video Data.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Make-A-Video: Text-to-Video Generation without Text-Video Data

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T13:01:00.153544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:01:00.153544Z digest=sha256:c67c07c14d6f9a66c40b097a3176688c41059b55dee0f4bd37560c0f1eef180b

Observation d567d424-2a2d-472e-ab42-d20348480375 · outbound

This paper cites Denoising Diffusion Implicit Models.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Denoising Diffusion Implicit Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T13:01:00.214571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:01:00.214571Z digest=sha256:b091c7571921b81f368bd3ce4209599bb99bc4d949642d754d2c172838feca55

Observation b94524fd-15c5-4b14-a04b-404da42a760b · outbound

This paper cites UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild.

MOVi: Training-free Text-conditioned Multi-Object Video Generation UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T13:01:00.283843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:01:00.283843Z digest=sha256:05cc75236f75baffaaf2c8b6eec1c0a7db55d1837c395a15039496bd157a0798

Observation cf967497-233e-4b07-818c-d3b7f7e4bfe7 · outbound

This paper cites Raft: Recurrent all-pairs field transforms for optical flow.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Raft: Recurrent all-pairs field transforms for optical flow

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:06.011982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:01:00.331605Z digest=sha256:0855f178445d8a4d1cf6eba4dc3aed2b8b75f9f0c160e5356a5b1e9e7dfddd7f

Observation d42c452d-9d30-4735-b207-c73d02727e5c · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

MOVi: Training-free Text-conditioned Multi-Object Video Generation LLaMA: Open and Efficient Foundation Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T13:01:00.389339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:01:00.389339Z digest=sha256:9f223a1e3b6c991a4c83af1b688307e8d4431680f2bbb6985a47c5e9601f09e2

Observation 9c91d916-d50f-466b-ae54-23f87b483035 · outbound

This paper cites Fvd: A new metric for video generation.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Fvd: A new metric for video generation

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:05.828168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:01:00.436115Z digest=sha256:cb3358757ee7ae8c8d0a3aa57fd6fe24020435c3e4f2b2a4602057b6680c7fcb

Observation c9b98d0c-7c9f-4cf4-8689-727ac648b204 · outbound

This paper cites Vchitect 2.0: Embark on a Visual Fan- tasy Journey.https://vchitect.intern- ai.org.cn/, 2024.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Vchitect 2.0: Embark on a Visual Fan- tasy Journey.https://vchitect.intern- ai.org.cn/, 2024

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:05.693951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:01:00.504131Z digest=sha256:bd21f047d6eb2a4ba2d6c3034e75b841c9e28fd264fa2e85c27af645911b9f11

Observation 7728518b-2d64-4a24-aa08-5fe07fb488c1 · outbound

This paper cites Phenaki: Variable Length Video Generation From Open Domain Textual Description.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Phenaki: Variable Length Video Generation From Open Domain Textual Description

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T13:01:00.551776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:01:00.551776Z digest=sha256:30fccd32d104c3523dd625c1582867b9173da90c3a8409e4979e9820403691fa

Observation b92d328c-fdd7-4fc2-8881-e6dc1d5cc576 · outbound

This paper cites ModelScope Text-to-Video Technical Report.

MOVi: Training-free Text-conditioned Multi-Object Video Generation ModelScope Text-to-Video Technical Report

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T13:01:00.629908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:01:00.629908Z digest=sha256:be5bc2f93af93fa591def63d9880b9b9ca24477589d7f545565b025787cb525d

Observation 76bd7851-616f-47fe-a686-038e71947e91 · outbound

This paper cites Boximator: Generating Rich and Controllable Motions for Video Synthesis.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Boximator: Generating Rich and Controllable Motions for Video Synthesis

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T13:01:00.698486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:01:00.698486Z digest=sha256:e691f2e334c9d69e05d20417161f7568046ebeb3e7bdfea082682954f9305722

Observation 0bd7da98-74db-4435-9e17-d1bf73de581c · outbound

This paper cites CogVLM: Visual Expert for Pretrained Language Models.

MOVi: Training-free Text-conditioned Multi-Object Video Generation CogVLM: Visual Expert for Pretrained Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T13:01:00.755454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:01:00.755454Z digest=sha256:67f77fc7f77c1f5ad563faebd6bfcae57956a53df051e2787533f22277b274b1

Observation 77514227-2a4a-46c9-a9b4-dbef5600b2ae · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Emu3: Next-Token Prediction is All You Need

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T13:01:00.803646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:01:00.803646Z digest=sha256:a6138dc571f3dc1f0529512454cbf4ffd72b2368547e16c4172b2d7ea2b91114

Observation 5527bad4-bd6c-4d2a-a004-2ab86e607bdb · outbound

This paper cites WorldDreamer: Towards General World Models for Video Generation via Predicting Masked Tokens.

MOVi: Training-free Text-conditioned Multi-Object Video Generation WorldDreamer: Towards General World Models for Video Generation via Predicting Masked Tokens

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T13:01:00.850247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:01:00.850247Z digest=sha256:f33103ea27c9a12320af1d1e5739e8df5ad1a9de62c15d75ea243abfa69fa6a6

Observation 5b73f522-f386-48ad-8fa6-f7b34042c99f · outbound

This paper cites LaVie: High-quality video generation with cascaded latent diffusion mod- els.IJCV, 2024.

MOVi: Training-free Text-conditioned Multi-Object Video Generation LaVie: High-quality video generation with cascaded latent diffusion mod- els.IJCV, 2024

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:05.494450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:01:00.900024Z digest=sha256:ae984a38d78e5108224adc1b23ec3cd560fb40093c50774cde7fe585bdfd00ec

Observation 0969d8ed-1fb6-4f3b-99cf-9b277420cd86 · outbound

This paper cites Customvideo: Cus- tomizing text-to-video generation with multiple sub- jects.arXiv preprint arXiv:2401.09962, 2024.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Customvideo: Cus- tomizing text-to-video generation with multiple sub- jects.arXiv preprint arXiv:2401.09962, 2024

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T13:01:00.965204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:01:00.965204Z digest=sha256:64b8809ea18a5b3e947773a3d7ea7bf6c0a36863892ac6580ce5392d03ba1b2a

Observation 25ef5394-9984-4e2d-ba69-6861e8466566 · outbound

This paper cites Grit: A generative region-to-text transformer for ob- ject understanding.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Grit: A generative region-to-text transformer for ob- ject understanding

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:05.242780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:01:01.016503Z digest=sha256:89bfdbd3a4446d92093df2feccfae2f9840ce1c86850ee3fda5b649ed77f2dc7

Observation a5b9fce2-6191-40d8-9afd-ef79d6c2815a · outbound

This paper cites FreeInit: Bridging Initialization Gap in Video Diffusion Models.

MOVi: Training-free Text-conditioned Multi-Object Video Generation FreeInit: Bridging Initialization Gap in Video Diffusion Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T13:01:01.085135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:01:01.085135Z digest=sha256:f49c38d11f23c1410e61c7fd997940af1a2e400c6a2b686d8d348ab00158f971

Observation 81dc84f9-3b18-4b54-ba9f-d8405d3a5837 · outbound

This paper cites Dynamicrafter: An- imating open-domain images with video diffusion pri- ors.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Dynamicrafter: An- imating open-domain images with video diffusion pri- ors

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:04.971631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:01:01.159014Z digest=sha256:d88023534b11a0a46351303415ea15eecabd097a1d3701d534024ed9ebf6805e

Observation 591e9272-6616-4805-bcaa-020b06632107 · outbound

This paper cites MSR- VTT: A large video description dataset for bridging video and language.

MOVi: Training-free Text-conditioned Multi-Object Video Generation MSR- VTT: A large video description dataset for bridging video and language

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:04.732407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:01:01.271572Z digest=sha256:31f29a3b143fc36fa279c5dbbf119682488bb1e7d9d298d3f225a9553fb73cae

Observation 5dd678f3-210e-4785-8b6f-57f08c563b28 · outbound

This paper cites Advancing high-resolution video- language representation with large-scale video tran- scriptions.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Advancing high-resolution video- language representation with large-scale video tran- scriptions

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:04.458872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:01:01.363298Z digest=sha256:bf1a7ac63f88c9f8179397c5b34991436564e5a3705589c32de3979b71e67a4f

Observation bdfe9b52-8ebb-4196-94bb-615dfa0b7647 · outbound

This paper cites Video in- stance segmentation.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Video in- stance segmentation

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:04.185001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:01:01.481638Z digest=sha256:d275f5aed4ce0b8b6087d1bed7211e0bcc4bb7b7f452fbe2f680a5b1bfbd4ecb

Observation dc07aed5-accc-4f00-9c2c-b2af7ecc7e2a · outbound

This paper cites EVA: Zero-shot Accurate Attributes and Multi-Object Video Editing.

MOVi: Training-free Text-conditioned Multi-Object Video Generation EVA: Zero-shot Accurate Attributes and Multi-Object Video Editing

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T13:01:01.607138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:01:01.607138Z digest=sha256:77285dd3c503d8cd29fa0913036bf4d33847c70043a484de32495f791373570d

Observation b7e031dc-3a43-4f2c-95a1-ce678daf7a23 · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

MOVi: Training-free Text-conditioned Multi-Object Video Generation CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T13:01:01.724034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:01:01.724034Z digest=sha256:784cd1566a3f556c9d63f96d0c71af6e05142e08af232b19cac8d33aea4820cd

Observation 1095eefb-eac3-49e4-a793-9812637b7ad5 · outbound

This paper cites Show-1: Marrying pixel and latent diffusion models for text-to-video generation.Interna- tional Journal of Computer Vision, pages 1–15, 2024.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Show-1: Marrying pixel and latent diffusion models for text-to-video generation.Interna- tional Journal of Computer Vision, pages 1–15, 2024

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:03.849630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:01:01.876054Z digest=sha256:2d5cb0e9b44148d7ffafb18998661abd3710b2c2bb7bd508cda53ec23e7f5373

Observation 274b38f8-d867-4814-bdab-22d674d6b567 · outbound

This paper cites I2VGen-XL: High-Quality Image-to-Video Synthesis via Cascaded Diffusion Models.

MOVi: Training-free Text-conditioned Multi-Object Video Generation I2VGen-XL: High-Quality Image-to-Video Synthesis via Cascaded Diffusion Models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T13:01:02.040690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:01:02.040690Z digest=sha256:4bc60106667cdba3fed70a887073a60f4caed9be8b82fe89aa7ed56d20bc1fe5

Observation 4a86a191-9976-4ebc-9f37-0abae1793df8 · outbound

This paper cites Real-time vehicle detection based on improved yolo v5.Sustainability, 14(19):12274, 2022.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Real-time vehicle detection based on improved yolo v5.Sustainability, 14(19):12274, 2022

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:03.624415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:01:02.166347Z digest=sha256:7a63f95f1a7c9b848d18b2c3b7580469bd74aa5d9aff4ba21aad4692f4a5b1da

Observation 2111eff1-6b07-4239-8e8b-1b32f1cb52c7 · outbound

This paper cites Open-sora: Democratizing efficient video production for all, March 2024.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Open-sora: Democratizing efficient video production for all, March 2024

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:03.398534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:01:02.307369Z digest=sha256:3458c1005504b0ad270257d0f66d9362fb9def55618b7d0f82d244472833e41d

Observation 808ab046-08ce-426e-8557-ccebdeab53d0 · outbound

This paper cites MagicVideo: Efficient Video Generation With Latent Diffusion Models.

MOVi: Training-free Text-conditioned Multi-Object Video Generation MagicVideo: Efficient Video Generation With Latent Diffusion Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T13:01:02.423170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:01:02.423170Z digest=sha256:a6675383771b3be69a483394a80846951652678e9387474e5767155cfc2babdc

Observation 53331444-f74f-4ea2-b685-554cd1e63306 · outbound

This paper cites 2, 5, 7, 8.

MOVi: Training-free Text-conditioned Multi-Object Video Generation 2, 5, 7, 8

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:01:06.966840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:00:59.590268Z digest=sha256:49e3f0f2bbe78a833db7317bd741859359680e86f1f595d63959f8765fe0b74d

Pith citing papers

No inbound Pith citation observations are available.