Pith. sign in

Paper Citation Record · LEDGER

Dynamic-I2V: Exploring Image-to-Video Generation Models via Multimodal LLM

As of 9 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 4 inbound Pith citation observations for arXiv:2505.19901.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.19901 v3

Coverage vector

measured 44 of 44 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:09:02.074963Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T17:45:17.010883Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T08:53:15.680626Z

Reference resolution

44 of 44 outbound references displayed

  • verified exact0
  • verified fuzzy31
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 72aa2235-ed7d-4a1e-b956-d344de8f3c81 · outbound

This paper cites Qwen Technical Report.

Dynamic-I2V: Exploring Image-to-Video Generation Models via Multimodal LLM Qwen Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:08:21.686085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:08:21.686085Z digest=sha256:6771b27a75d8bb84bd77926f0c15ba5b664c6aee59ae6d76ede24693c98d8370

Observation 8ad673e4-4a94-4317-9f21-cd09aa1de307 · outbound

This paper cites Frozen in time: A joint video and image encoder for end-to- end retrieval.

Dynamic-I2V: Exploring Image-to-Video Generation Models via Multimodal LLM Frozen in time: A joint video and image encoder for end-to- end retrieval

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:12.544623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:08:21.760629Z digest=sha256:4ad8366c037d24caae6690d335a14f53379bff66e700af38b02009c04f26b4a4

Observation 5625801f-3387-4d65-abba-e3727c721a87 · outbound

This paper cites Stable video diffusion: Scaling latent video diffusion models to large datasets, 2023.

Dynamic-I2V: Exploring Image-to-Video Generation Models via Multimodal LLM Stable video diffusion: Scaling latent video diffusion models to large datasets, 2023

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:08:21.883855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:08:21.883855Z digest=sha256:5e5727d9899cb5c81ad7ab750062b9ce17dc2647acaf30472b546d33b4ec2feb

Observation 83c440f4-422f-42d8-8c32-4838e8fda54f · outbound

This paper cites an unresolved cited work.

Dynamic-I2V: Exploring Image-to-Video Generation Models via Multimodal LLM Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:09:12.320714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:08:21.967524Z digest=sha256:fb5b4cd497a685060a1f4e0d8d041fd34da3df4136eda77c0141afd86840dfc2

Observation dcc775db-a1f0-4b45-ac3d-874a0371499f · outbound

This paper cites Panda-70m: Captioning 70m videos with multiple cross-modality teachers.

Dynamic-I2V: Exploring Image-to-Video Generation Models via Multimodal LLM Panda-70m: Captioning 70m videos with multiple cross-modality teachers

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:11.983452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:08:22.067285Z digest=sha256:556bf49bc07be8ac4a0daad944ee340b5a83fc607ce403e64e0d9cba3e749b37

Observation 1d864c9a-0500-4747-9984-74679bbc0d47 · outbound

This paper cites Seine: Short-to-long video diffusion model for generative transition and prediction, 2023.

Dynamic-I2V: Exploring Image-to-Video Generation Models via Multimodal LLM Seine: Short-to-long video diffusion model for generative transition and prediction, 2023

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:11.686020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:08:22.145040Z digest=sha256:5dc4475f57f4ba126984e25ef9488ddc7fb65cb7f9f3deea3bafd989ab385338

Observation 4e6ab43e-a61d-4b1b-9c5e-3edd53e09905 · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks, 2024.

Dynamic-I2V: Exploring Image-to-Video Generation Models via Multimodal LLM Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks, 2024

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:08:22.230147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:08:22.230147Z digest=sha256:4e272e416165558fd9987a1ca9d6f8c1b43bf4255e4898464f3d5dd4d14d0b5d

Observation 0576af17-93b7-454c-953d-79fa927e7b7a · outbound

This paper cites Glm: General language model pretraining with autoregressive blank infilling, 2022.

Dynamic-I2V: Exploring Image-to-Video Generation Models via Multimodal LLM Glm: General language model pretraining with autoregressive blank infilling, 2022

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:11.482059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:08:22.332373Z digest=sha256:ac0c9774974028b461f72329986a4b31580bbd200426d44249a6d9bf30ddb5a7

Observation f74e934f-cd88-4374-9ec0-3a4106fdb662 · outbound

This paper cites Guiding instruction-based image editing via multimodal large language models, 2024.

Dynamic-I2V: Exploring Image-to-Video Generation Models via Multimodal LLM Guiding instruction-based image editing via multimodal large language models, 2024

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:11.246322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:08:22.450490Z digest=sha256:dd95e094d8e6dd9c6cb7d2c2786178d43782d6ffd299367bb1992e9e0ed5fa3e

Observation ec9e167b-c6ef-45e7-8b56-bd58cd3585f3 · outbound

This paper cites Seed-x: Multi- modal models with unified multi-granularity comprehension and generation, 2025.

Dynamic-I2V: Exploring Image-to-Video Generation Models via Multimodal LLM Seed-x: Multi- modal models with unified multi-granularity comprehension and generation, 2025

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:11.085598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:08:22.553800Z digest=sha256:bf8419c57af1bd6e93f5d6b1047a1a4b8825a780abaf559993359918101ee50a

Observation cfa1c5df-408b-4b15-a2ef-dff08987f7bd · outbound

This paper cites Denoising diffu- sion probabilistic models, 2020.

Dynamic-I2V: Exploring Image-to-Video Generation Models via Multimodal LLM Denoising diffu- sion probabilistic models, 2020

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:08:22.632062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:08:22.632062Z digest=sha256:2910bd315090fbf19d6debd58308b6000448451a1ed3e064a190c6c5f3acd34e

Observation 0e3d4581-83d7-41de-835c-79bcdd1d5839 · outbound

This paper cites Smartedit: Exploring complex instruction-based image editing with multimodal large language models, 2023.

Dynamic-I2V: Exploring Image-to-Video Generation Models via Multimodal LLM Smartedit: Exploring complex instruction-based image editing with multimodal large language models, 2023

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:10.804747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:08:22.749027Z digest=sha256:aac2f4870dba4f44370d64fabbdcfa3ad769161e934ef1d66e1ee4562d45ef10

Observation d4496099-00c6-4eea-bced-54d3b2ae21e0 · outbound

This paper cites Smartedit: Exploring complex instruction-based image editing with multimodal large lan- guage models.

Dynamic-I2V: Exploring Image-to-Video Generation Models via Multimodal LLM Smartedit: Exploring complex instruction-based image editing with multimodal large lan- guage models

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:10.604479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:08:22.809997Z digest=sha256:7bdb852e0e929616a3a13a4973bf3373ca426a2ae0f1ef531150f6f15d8bdd11

Observation fc616fa2-5fd4-437c-a51d-1e71083f5844 · outbound

This paper cites VBench: Com- prehensive benchmark suite for video generative models.

Dynamic-I2V: Exploring Image-to-Video Generation Models via Multimodal LLM VBench: Com- prehensive benchmark suite for video generative models

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:10.369470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:08:22.880124Z digest=sha256:68b177fb6f9d5b299b75c63888817fa3db328deb44f2703cb22ee44e477dad7c

Observation 3cf50900-c011-46d9-99cb-c6d648a3569a · outbound

This paper cites Vbench++: Comprehensive and versatile benchmark suite for video generative models, 2024.

Dynamic-I2V: Exploring Image-to-Video Generation Models via Multimodal LLM Vbench++: Comprehensive and versatile benchmark suite for video generative models, 2024

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:10.148639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:08:22.987071Z digest=sha256:4526568f1379e39d375b413b421f4828e0847e82759bd465b04abf48db1463f1

Observation b6cd7b14-1bcf-45bd-b352-c42ba2864528 · outbound

This paper cites T2vbench: Benchmarking temporal dynamics for text-to- video generation.

Dynamic-I2V: Exploring Image-to-Video Generation Models via Multimodal LLM T2vbench: Benchmarking temporal dynamics for text-to- video generation

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:09.840250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:08:23.122120Z digest=sha256:432678ada9df410fd78c96decc7a36df049d12851aebf461824ec27543b9836e

Observation 1ed9d684-8fa4-4a97-ade1-cc1f88d81b87 · outbound

This paper cites Miradata: A large-scale video dataset with long durations and structured captions, 2024.

Dynamic-I2V: Exploring Image-to-Video Generation Models via Multimodal LLM Miradata: A large-scale video dataset with long durations and structured captions, 2024

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:09.663855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:08:23.220189Z digest=sha256:dacbeee21e2eaa8c86c7f2366d2262ff5294e5fad4af912ad1995bd9e1875f96

Observation aa339d00-f0ab-4f39-a214-1bc1aa13278b · outbound

This paper cites Co- tracker: It is better to track together, 2024.

Dynamic-I2V: Exploring Image-to-Video Generation Models via Multimodal LLM Co- tracker: It is better to track together, 2024

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:09.528308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:08:23.310192Z digest=sha256:b175f5f61dfe0cd8768a61aae6da23a2070c83bd731661c8f8a6b54a18f439a7

Observation 6492a1b0-062b-4f98-bf61-4d9ca71a7a3a · outbound

This paper cites Hunyuanvideo: A systematic framework for large video generative models, 2025.

Dynamic-I2V: Exploring Image-to-Video Generation Models via Multimodal LLM Hunyuanvideo: A systematic framework for large video generative models, 2025

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:09.378606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:08:59.387247Z digest=sha256:33afc891f8e16fd26600dcb37c4af9bd2f6f719d98cfc2abdb1c4ab25a62e759

Observation f02c080d-1631-4442-978d-a4ceffc245af · outbound

This paper cites Animateanything: Consistent and control- lable animation for video generation, 2024.

Dynamic-I2V: Exploring Image-to-Video Generation Models via Multimodal LLM Animateanything: Consistent and control- lable animation for video generation, 2024

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:09.146564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:08:59.484759Z digest=sha256:e446c98b457406f2ceadf44b3af51b4243d72d9d9ddeb9f388ef0bb80aae32ac

Observation 9894e930-f8e9-4f01-a7a3-45612583e2a6 · outbound

This paper cites Evaluation of Text-to-Video Generation Models: A Dynamics Perspective.

Dynamic-I2V: Exploring Image-to-Video Generation Models via Multimodal LLM Evaluation of Text-to-Video Generation Models: A Dynamics Perspective

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:08:59.645506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:08:59.645506Z digest=sha256:bdb3d59b03a7d158a96128a28bd518da56aba3507bf41320787336793fdd45ab

Observation 118e7302-b84f-4d47-a6c4-71f943984499 · outbound

This paper cites Video-llava: Learning united visual representa- tion by alignment before projection, 2024.

Dynamic-I2V: Exploring Image-to-Video Generation Models via Multimodal LLM Video-llava: Learning united visual representa- tion by alignment before projection, 2024

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:08.794907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:08:59.767550Z digest=sha256:d5adeaafbb8fd332157a307e9ae6786b1f192c2da9f11a4bcbd216ee478b0459

Observation 68ec7cf1-dcec-4e83-89dd-d13b1165fbdd · outbound

This paper cites Stiv: Scalable text and image conditioned video generation, 2024.

Dynamic-I2V: Exploring Image-to-Video Generation Models via Multimodal LLM Stiv: Scalable text and image conditioned video generation, 2024

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:08.292123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:08:59.834819Z digest=sha256:567033d3ede9ab913de90ce630ecbf61d4c64ea07cf2b7a3e2d7efcb06bbef43

Observation 7c4adeb1-41e7-413a-92d7-7852e8ce260b · outbound

This paper cites Evalcrafter: Benchmarking and evalu- ating large video generation models.

Dynamic-I2V: Exploring Image-to-Video Generation Models via Multimodal LLM Evalcrafter: Benchmarking and evalu- ating large video generation models

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:08.006335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:08:59.955072Z digest=sha256:5384c273679d26661083532c8df36e749d63361f721ee34316ca1c423a539bae

Observation 7bbe86ff-38a8-40e1-b972-e31d7e7a6155 · outbound

This paper cites Qwen2vl-flux: Unifying image and text guidance for controllable image generation, 2024.

Dynamic-I2V: Exploring Image-to-Video Generation Models via Multimodal LLM Qwen2vl-flux: Unifying image and text guidance for controllable image generation, 2024

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:07.764847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:09:00.095203Z digest=sha256:63b04e29953830e3203dbb451e869167dccb26c8742cb91ccdc0c2d9929c35c3

Observation 1806a057-43f3-4377-9110-5ac706fe34dd · outbound

This paper cites Openvid-1m: A large-scale high-quality dataset for text-to- video generation, 2024.

Dynamic-I2V: Exploring Image-to-Video Generation Models via Multimodal LLM Openvid-1m: A large-scale high-quality dataset for text-to- video generation, 2024

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:07.464961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:09:00.210997Z digest=sha256:8a8b8a90376d7735f16d4d7aae37f6a195b973c4205159610ccc2c7f08df9848

Observation 996c4bfb-196b-4917-96b6-78e9ee86cc08 · outbound

This paper cites Kosmos-G: Generating Images in Context with Multimodal Large Language Models.

Dynamic-I2V: Exploring Image-to-Video Generation Models via Multimodal LLM Kosmos-G: Generating Images in Context with Multimodal Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:09:00.316876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:09:00.316876Z digest=sha256:d930fdb8eeef200cd0fb5ad26c1aa3837b772367000fb9fd9a66604335e251f8

Observation 39c33e2c-7768-43f3-9af6-eb5146e9763d · outbound

This paper cites Kosmos-g: Generating images in context with multimodal large language models, 2024.

Dynamic-I2V: Exploring Image-to-Video Generation Models via Multimodal LLM Kosmos-g: Generating images in context with multimodal large language models, 2024

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:09:00.442187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:09:00.442187Z digest=sha256:c6ba33508adbad5d9f740b69d51fdb9275a5273b5b4fdcbbd51fb2ec0310f662

Observation d56617e2-834d-40d7-a800-f12d40feb9f5 · outbound

This paper cites Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Askell Amanda, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever.

Dynamic-I2V: Exploring Image-to-Video Generation Models via Multimodal LLM Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Askell Amanda, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:07.004397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:09:00.714761Z digest=sha256:22b1c839bbbbfae926fbad3dd35af1ec0438c8fac4d3edce6374cb083019b5c2

Observation 00f5ed96-84e4-4fe4-b58a-e6aff330096b · outbound

This paper cites an unresolved cited work.

Dynamic-I2V: Exploring Image-to-Video Generation Models via Multimodal LLM Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:09:06.740500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:09:00.874887Z digest=sha256:459fd658e750f66f4f245ef653b4947df6ddef417e2bfe299837105bddb4fffa

Observation bfe830b5-088e-4fdb-91af-422c53f1b5e6 · outbound

This paper cites ViBe: A Text-to-Video Benchmark for Evaluating Hallucination in Large Multimodal Models.

Dynamic-I2V: Exploring Image-to-Video Generation Models via Multimodal LLM ViBe: A Text-to-Video Benchmark for Evaluating Hallucination in Large Multimodal Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:09:00.950330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:09:00.950330Z digest=sha256:d662291a6774ad23f37e8ebf2adf165c7c1381b92e99af9196d421ed362192aa

Observation 1f3e9a63-1144-4afd-9dd2-759a833fec09 · outbound

This paper cites Consisti2v: Enhancing visual consistency for image-to-video generation, 2024.

Dynamic-I2V: Exploring Image-to-Video Generation Models via Multimodal LLM Consisti2v: Enhancing visual consistency for image-to-video generation, 2024

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:06.454768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:09:01.030751Z digest=sha256:5f1dd382deea61f2b6fd5029387c577302141855ce140a0aec4c130f2ff75c1d

Observation 1a46b82e-7a5d-458e-a224-fada33098c07 · outbound

This paper cites T2v-compbench: A comprehen- sive benchmark for compositional text-to-video generation,.

Dynamic-I2V: Exploring Image-to-Video Generation Models via Multimodal LLM T2v-compbench: A comprehen- sive benchmark for compositional text-to-video generation,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:06.184742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:09:01.084975Z digest=sha256:e190fd38fc827cdef6219f5bf09115ede0b82d2aec0add865e4b7b0c328c1562

Observation de6b89a0-6ec9-46a8-9d6e-1f89c5f7747a · outbound

This paper cites Generative multi- modal models are in-context learners, 2024.

Dynamic-I2V: Exploring Image-to-Video Generation Models via Multimodal LLM Generative multi- modal models are in-context learners, 2024

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:05.961472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:09:01.215729Z digest=sha256:0c6a4c19725565eb78bdd2c134c92c1ea9db0648c2b3e6776474e8aba41e93bf

Observation 6ea87874-43f1-4d34-a47b-ecc0262efe57 · outbound

This paper cites Mimir: Improving video diffusion models for precise text understanding, 2024.

Dynamic-I2V: Exploring Image-to-Video Generation Models via Multimodal LLM Mimir: Improving video diffusion models for precise text understanding, 2024

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:05.690139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:09:01.319520Z digest=sha256:f6fcb842f8d496e7e43f9df56610d4848192185416450c819e33bef5c9ead838

Observation 45f36936-4ee2-4d7f-8ce8-6f79c40c8365 · outbound

This paper cites Mige: A unified framework for multimodal instruction-based image generation and editing,.

Dynamic-I2V: Exploring Image-to-Video Generation Models via Multimodal LLM Mige: A unified framework for multimodal instruction-based image generation and editing,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:05.199254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:09:01.393704Z digest=sha256:0d0db2690b1f3f8e7de540d09c864aae7cc0b9ac0d6837cf016178057f4c8e43

Observation 730d09a2-221d-4bc3-8125-82d5864d176a · outbound

This paper cites Llama: Open and efficient foundation language mod- els, 2023.

Dynamic-I2V: Exploring Image-to-Video Generation Models via Multimodal LLM Llama: Open and efficient foundation language mod- els, 2023

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:04.835253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:09:01.489642Z digest=sha256:abcbbabaea781ed5b6db334cc4bcf71fae2c2f0108c6126d7cf6ef6fc5299a76

Observation a6edc8cb-6c1f-44b8-8a93-bc17d5bf146f · outbound

This paper cites Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution, 2024.

Dynamic-I2V: Exploring Image-to-Video Generation Models via Multimodal LLM Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution, 2024

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T14:09:01.600084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:09:01.600084Z digest=sha256:92d3799b136fbebc4bfc352af74084cea1d86584dd13b786dc74ba4f835d4402

Observation a0b20784-0f36-48a6-addc-bdc49c96f5f1 · outbound

This paper cites Koala-36m: A large-scale video dataset improving consistency between fine-grained conditions and video content, 2024.

Dynamic-I2V: Exploring Image-to-Video Generation Models via Multimodal LLM Koala-36m: A large-scale video dataset improving consistency between fine-grained conditions and video content, 2024

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:04.464912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:09:01.702109Z digest=sha256:5c45ee437289b49c7c7e85cd3831a9fb1eac4e181194a0cb2a8038ce4a05f440

Observation 64ec8ce9-8681-4706-b58a-95507a8155e6 · outbound

This paper cites Cogvlm: Visual expert for pretrained language models, 2024.

Dynamic-I2V: Exploring Image-to-Video Generation Models via Multimodal LLM Cogvlm: Visual expert for pretrained language models, 2024

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T14:09:01.795413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:09:01.795413Z digest=sha256:44fce8e6798b527ff23a11561a9ddcff1e54df4b5a352f12568d49db5ad90713

Observation eaa55a40-d8c5-4b28-8f62-263b46c10119 · outbound

This paper cites Dynamicrafter: Animating open-domain im- ages with video diffusion priors, 2023.

Dynamic-I2V: Exploring Image-to-Video Generation Models via Multimodal LLM Dynamicrafter: Animating open-domain im- ages with video diffusion priors, 2023

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:04.015122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:09:01.859480Z digest=sha256:fc7ebc5d5ca160225208e42cf1a4a880e4ff89efd95b50abc4762fd56d15f399

Observation c963fe9e-be22-4ef9-84d6-98c511a2faff · outbound

This paper cites Easyanimate: A high-performance long video generation method based on transformer architecture, 2024.

Dynamic-I2V: Exploring Image-to-Video Generation Models via Multimodal LLM Easyanimate: A high-performance long video generation method based on transformer architecture, 2024

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:03.655071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:09:01.928371Z digest=sha256:f968a6068163c078e0c60692fdd5c9602e9865b55d65881b2fd178b8d547e156

Observation 21470ee1-bf72-496c-b900-b03e1445f430 · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

Dynamic-I2V: Exploring Image-to-Video Generation Models via Multimodal LLM CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T14:09:01.994943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:09:01.994943Z digest=sha256:909c5afc7aea7d6256962f8f36bb1f881a8cba1b1dc67849db4c04c7964a4a42

Observation 0ee7e785-f87e-4cd5-b017-4364e2d1485d · outbound

This paper cites I2vgen-xl: High-quality image-to-video synthesis via cascaded diffusion models, 2023.

Dynamic-I2V: Exploring Image-to-Video Generation Models via Multimodal LLM I2vgen-xl: High-quality image-to-video synthesis via cascaded diffusion models, 2023

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:09:03.305165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:09:02.074963Z digest=sha256:181e527802e6c15ed1eb664e508360c4e99412dfa3c69bfcfdafdd7ef4d2eba1

Pith citing papers

Observation e30efa21-81b8-40a6-877a-7caca3d92cd3 · inbound

Waver: Wave Your Way to Lifelike Video Generation cites this paper.

Waver: Wave Your Way to Lifelike Video Generation Dynamic-I2V: Exploring Image-to-Video Generation Models via Multimodal LLM

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T17:45:17.010883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:45:17.010883Z digest=sha256:28ee34e88aac4b1cb5d29cc327e0ea87836e0b6e84bf40ca6d0e9233fdf0a83a

Observation e67b3fd2-7807-4c03-8bda-11054505215b · inbound

Image-to-Video Diffusion: From Foundations to Open Frontiers cites this paper.

Image-to-Video Diffusion: From Foundations to Open Frontiers Dynamic-I2V: Exploring Image-to-Video Generation Models via Multimodal LLM

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-20T15:08:25.017943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:7782335e0e368d341c376d7772ad7f4add7973b89a46423d2b3921215cd20519

Observation 080444ab-0976-4f05-9c09-eaea31c4d8c4 · inbound

SuperVoxelGPT: Adaptive and Ordered 3D Tokenization for Autoregressive Shape Generation cites this paper.

SuperVoxelGPT: Adaptive and Ordered 3D Tokenization for Autoregressive Shape Generation Dynamic-I2V: Exploring Image-to-Video Generation Models via Multimodal LLM

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T08:53:15.681935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T08:52:09.461853Z digest=sha256:9f1094e6065969a183c7eb0c474aa6698dc61b96d124e1351b5e03803d6beeb6

Observation e4f19425-5f3b-4bda-94f3-5cf54547d0c4 · inbound

EventOD: Event-Aware OD Flow Generation via LLM-Guided Semantic Modulation cites this paper.

EventOD: Event-Aware OD Flow Generation via LLM-Guided Semantic Modulation Dynamic-I2V: Exploring Image-to-Video Generation Models via Multimodal LLM

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-02T09:52:01.777787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:52:01.777787Z digest=sha256:97ca5fd9030d0a8436822ef1315b37f196bbb03de8322be5fd726d9ddecf9488