Pith. sign in

Paper Citation Record · LEDGER

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation

As of 18 August 2026, this Paper Citation Record lists 100 of 187 outbound references and 1 inbound Pith citation observation for arXiv:2412.18688.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.18688 v2

Coverage vector

measured 100 of 187 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T04:37:14.801702Z

measured 101 of 101 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-20T10:56:16.054496Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T10:58:13.835764Z

Reference resolution

100 of 187 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved100
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3f5c52f7-f554-4423-98b2-4ccc57f1ccd9 · outbound

This paper cites Video generation models as world simulators by open a.i, 2024.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Video generation models as world simulators by open a.i, 2024

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.383491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.383491Z digest=sha256:fabe869ed9ccbed4f26861e4ceb7dabae6e6aca96ccc9fc43921c28ce1c59461

Observation 1d6830a5-4ef6-4a7a-8123-2b99072bcdf8 · outbound

This paper cites Introducing chatgpt by open a.i, 2022.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Introducing chatgpt by open a.i, 2022

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.387900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.387900Z digest=sha256:71fdd8d0ab8b8d2377ef551ed156dd484b11bd9d368279fd65b6caf46805b6e5

Observation 73e61bf6-ecc7-4a84-b4f1-7cd2999a00ad · outbound

This paper cites Meta llama models, 2023.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Meta llama models, 2023

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.391439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.391439Z digest=sha256:2a20918aad5393c8a79926e2686574653eadf4321fdc73c5180341d94df4742d

Observation 6dba3691-0a87-47da-8617-83bc04eecbe7 · outbound

This paper cites Google gemini series, 2023.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Google gemini series, 2023

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.395190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.395190Z digest=sha256:1a8a6607532d55b2957b69da32324abceeef5b0bf11e4673585129998fd4e39c

Observation a2576c1b-ec00-41d2-9df2-f3a8f6501a1f · outbound

This paper cites Anthropic by claude, 2023.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Anthropic by claude, 2023

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.400143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.400143Z digest=sha256:b8a4890071c47bc7d20da1445a39b8c77c147684ac740f4b49750cb060b8da19

Observation 49f9dc04-d8d1-4b41-b5d7-42ff15665d72 · outbound

This paper cites Mistral large model by mistral, 2024.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Mistral large model by mistral, 2024

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.404176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.404176Z digest=sha256:75ff1a5b867a53c4959d5b8bc97cb56f542aaa3338720ad7435b1c8093642cc5

Observation f6b7943f-0f31-4d0e-9c8c-90ed161f509e · outbound

This paper cites Hierarchical text-conditional image generation with clip latents, 2022.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Hierarchical text-conditional image generation with clip latents, 2022

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.408021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.408021Z digest=sha256:1eaa77b8b27b59217f52281101e8c7571a1cd8d206374bc38cf7ccddd6fda725

Observation a5b2b837-f056-439f-ac48-4f0c739eafb5 · outbound

This paper cites Scaling rectified flow transformers for high-resolution image synthesis, 2024.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Scaling rectified flow transformers for high-resolution image synthesis, 2024

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.411624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.411624Z digest=sha256:c2a70510d64861d2072cc3766b735b6c4515a9a0e400052d93f6ed8f257beddf

Observation 44c30872-f7a0-42fc-a98a-ffa15714742b · outbound

This paper cites The midjourney v5.2 model for image generation, 2024.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation The midjourney v5.2 model for image generation, 2024

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.416060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.416060Z digest=sha256:cbc0128119552c66bbf07259766751f0d2e392839db88cdf6702abedd69dad34

Observation 6db003ce-e59e-4e92-9358-1448c9b2bf96 · outbound

This paper cites VideodirectorGPT: Consistent multi-scene video generation via LLM-guided planning, 2024.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation VideodirectorGPT: Consistent multi-scene video generation via LLM-guided planning, 2024

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.423947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.423947Z digest=sha256:c4db9ac2ecb109fa0c5e1aa0239162b6e96a3e4cd6dce526b1cc0a6d6fc25d9b

Observation eb451ad4-0f24-47cb-a8c4-f194b88e4719 · outbound

This paper cites Make-a-video: Text-to-video generation without text-video data.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Make-a-video: Text-to-video generation without text-video data

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.428291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.428291Z digest=sha256:07d2ef3fd24776da48ebf0e424aaeec6d0be08a0e9a22ff35e9524ef5744cd77

Observation caefaae3-3e22-46ba-998c-01e17f0cf005 · outbound

This paper cites Gen2 by runway ml, 2023.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Gen2 by runway ml, 2023

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.432170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.432170Z digest=sha256:d678530eb9b8842e113df9a7e3053c9b9930f64cb21f08e06218f8761c486e4b

Observation 548c3fb9-bd96-4d36-b82b-7ae71c8772a5 · outbound

This paper cites Cogvideo: Large-scale pretraining for text-to-video generation via transformers.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Cogvideo: Large-scale pretraining for text-to-video generation via transformers

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.435916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.435916Z digest=sha256:90031bdb2ed2d38167168f05db2da781451def71c8223ff2df8497629f6b0bda

Observation 87325dd2-a846-47d6-9444-4dc5139caccc · outbound

This paper cites Gen-l-video: Multi-text to long video generation via temporal co-denoising, 2023.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Gen-l-video: Multi-text to long video generation via temporal co-denoising, 2023

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.443960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.443960Z digest=sha256:ce3bdb3a50c9ece8803eda0c4052ded48b0a9b8055b89279287c9e0fabb1344c

Observation 019d4e69-d36e-4506-88b0-161f410528ed · outbound

This paper cites Gen-4 alpha by midjourney, 2024.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Gen-4 alpha by midjourney, 2024

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.447977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.447977Z digest=sha256:44bbbe29395072f62b1cb5afa6ba308bb0bb79ed6b6907b1e7417a1469d636af

Observation 27302c44-f4a3-4abf-8d6b-d0e0c7b08586 · outbound

This paper cites A survey on long video generation: Challenges, methods, and prospects, 2024.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation A survey on long video generation: Challenges, methods, and prospects, 2024

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.451981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.451981Z digest=sha256:2dc6d480b4439479201e4756df65defc7b88b5391c1d4e41f9cf9e6e85a4e37d

Observation 7a0018e5-fb00-4a80-85a6-0e33bf002ca1 · outbound

This paper cites A survey on generative ai and llm for video generation, understanding, and streaming, 2024.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation A survey on generative ai and llm for video generation, understanding, and streaming, 2024

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.455327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.455327Z digest=sha256:67283563391658fa699d18b26f028d4ec7ba95647167835ce5a6b9ae2c6ff6ec

Observation 8cad888a-61a3-4a28-ab96-7b46cbe14209 · outbound

This paper cites Sora: A review on background, technology, limitations, and opportunities of large vision models, 2024.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Sora: A review on background, technology, limitations, and opportunities of large vision models, 2024

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.458733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.458733Z digest=sha256:098f678cda77d4fcc3b0ca09b4cf6244db502f24e04c8bc49471272bc8f27332

Observation 8ce84fef-7f8d-445f-85ab-717bc83509e8 · outbound

This paper cites Free-bloom: Zero-shot text-to-video generator with LLM director and LDM animator.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Free-bloom: Zero-shot text-to-video generator with LLM director and LDM animator

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.462880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.462880Z digest=sha256:347a7906c52d91fa392a549c22e89f41e5b0894fe1a7053aae7f4aab9ac51730

Observation b33cc5f5-e159-4f70-b46d-6983ea86a88e · outbound

This paper cites Flowzero: Zero-shot text-to-video synthesis with llm-driven dynamic scene syntax, 2023.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Flowzero: Zero-shot text-to-video synthesis with llm-driven dynamic scene syntax, 2023

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.466710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.466710Z digest=sha256:e5802d8ba6e12efc9b05dd4c63f95e79f17f441ea51219aea100504e5980bead

Observation 53516192-30a6-49d7-adbd-aa0feaa1087a · outbound

This paper cites Align your latents: High- resolution video synthesis with latent diffusion models.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Align your latents: High- resolution video synthesis with latent diffusion models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.469740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.469740Z digest=sha256:c9611fb396a02a564ab256c11782e533b7b5d133bb487347916b4730ff1185b8

Observation 1db71ba0-b625-424c-a9fd-83f4767161af · outbound

This paper cites Mavin: Multi-action video generation with diffusion models via transition video infilling, 2024.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Mavin: Multi-action video generation with diffusion models via transition video infilling, 2024

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.473186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.473186Z digest=sha256:da62946493835d3e4082dc462444324bb66a0b5f2beb0832c7c5fff0c3c251ff

Observation 29eaa73f-033d-47ea-9e4c-183dffd434db · outbound

This paper cites SEINE: Short-to-long video diffusion model for generative transition and prediction.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation SEINE: Short-to-long video diffusion model for generative transition and prediction

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.476649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.476649Z digest=sha256:4a15a858f1e185326617259281eca55634eab1b0c75cc83aef5ea823983e08b5

Observation e7d90b62-6189-4c5e-ba2c-30b251999f9e · outbound

This paper cites Learning transferable visual models from natural language supervision, 2021.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Learning transferable visual models from natural language supervision, 2021

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.480144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.480144Z digest=sha256:397ca89c6e08a4f7bb36b5ef0d14296ad91ff7f4a88b53683b08b21858e870d5

Observation a3e3613a-024c-450f-bee0-a403509c066d · outbound

This paper cites Generative adversarial networks.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Generative adversarial networks

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.483541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.483541Z digest=sha256:7d86a3ed959e6569ff3c0c5382a076d9c3188b64beea02bbfa3e25b6b014cb5e

Observation 25f2d8a0-6316-4d33-a13c-9e33b9c973fc · outbound

This paper cites Unsupervised representation learning with deep convolutional generative adversarial networks, 2016.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Unsupervised representation learning with deep convolutional generative adversarial networks, 2016

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.488168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.488168Z digest=sha256:1f478746929bd67dde74affa9436b06c51d5b615a7de6bd7227004bd7c34703f

Observation a6101203-873c-4616-a164-fea0b5d4ccbe · outbound

This paper cites Deep generative image models using a laplacian pyramid of adversarial networks.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Deep generative image models using a laplacian pyramid of adversarial networks

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.492876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.492876Z digest=sha256:1f69e6e42936aecd34be8fb36ec854a1ef84115a66df859fe90754a5d2714dce

Observation 6bcb1f2f-56be-45d9-b22c-2b7ff2309283 · outbound

This paper cites Stackgan: Text to photo-realistic image synthesis with stacked generative adversarial networks.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Stackgan: Text to photo-realistic image synthesis with stacked generative adversarial networks

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.496953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.496953Z digest=sha256:ba3921b27fba41194744b17c99d8e6d8e037c4991903c8fa9bd6b37cc99368c2

Observation 3b43eca4-4348-4faa-998a-04ab263bc7bc · outbound

This paper cites Attngan: Fine-grained text to image generation with attentional generative adversarial networks.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Attngan: Fine-grained text to image generation with attentional generative adversarial networks

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.501400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.501400Z digest=sha256:f52436eeb8567a8bc9390f84970a1797037b1e5d20ab4db0faa93af8f46f7ff1

Observation 471641e0-f60c-46e4-be61-2485a4e9bad9 · outbound

This paper cites Analyzing and Improving the Image Quality of StyleGAN.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Analyzing and Improving the Image Quality of StyleGAN

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.506500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.506500Z digest=sha256:478d3e41a9b0bd321fc5aff630377300395a29cd0a10dac8d621e0c402b824a1

Observation 747ac8e1-caa9-45ce-96ee-418a2f1364df · outbound

This paper cites an unresolved cited work.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.510690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.510690Z digest=sha256:6e6f8278db2b7ad87f5110a1b2f0a1549d2f299c6a12ed5421b123be22e25cd7

Observation a8ccc019-4d0c-4d9a-ba45-8ae6980ec084 · outbound

This paper cites Stylegan2 distillation for feed-forward image manipulation.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Stylegan2 distillation for feed-forward image manipulation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.515413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.515413Z digest=sha256:6a8696db951ce5a0aa2e5d95b7d22ee35e505eabf16acadc6fe06f2abd6d880b

Observation f964f7cf-5ee7-417d-a84f-417245c0816f · outbound

This paper cites an unresolved cited work.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.520232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.520232Z digest=sha256:a9e016539610bb70255112eb1a13e1b507845743284b1c74bfa227cd08ee72a1

Observation a5deffb8-65f6-4ee0-b9af-fe946c7c2278 · outbound

This paper cites Deep multi-scale video prediction beyond mean square error, 2016.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Deep multi-scale video prediction beyond mean square error, 2016

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.523881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.523881Z digest=sha256:d34152c64cb56210778179a8b30e11ddc35df0367858d65639211098c0cbb260

Observation cb3c7ede-6527-44d8-90ee-7ea998f07ca6 · outbound

This paper cites Generating videos with scene dynamics, 2016.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Generating videos with scene dynamics, 2016

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.527985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.527985Z digest=sha256:9a96251d6b08d492f7b8a26b22898778a96310f2e2bbc30ad23230bc8bbd8756

Observation 585fb8f7-0c81-43d4-a131-f0bcbccb5de0 · outbound

This paper cites Learning to generate time-lapse videos using multi-stage dynamic generative adversarial networks.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Learning to generate time-lapse videos using multi-stage dynamic generative adversarial networks

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.531670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.531670Z digest=sha256:9bfd66aca0964d0f464bc7211ba2c550f034c71be54aef209eb86a81da8f2943

Observation 9956f5ac-8c22-4648-bafe-41cebffc43d5 · outbound

This paper cites To create what you tell: Generating videos from captions, 2018.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation To create what you tell: Generating videos from captions, 2018

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.534876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.534876Z digest=sha256:cc9d1bd31466540bc5ffff442ed66af7a904e3817d92885e27cd33e7c971a047

Observation 771ed588-f2af-443e-b25b-ce4f98d2c5b6 · outbound

This paper cites Generating videos with dynamics-aware implicit generative adversarial networks.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Generating videos with dynamics-aware implicit generative adversarial networks

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.538591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.538591Z digest=sha256:077154821b0a015b3849f9721df238ff64a16920418f02622b4364be1ad77463

Observation 3a56a940-90e5-4593-b4f3-ba8c55d5907d · outbound

This paper cites Stylegan-v: A continuous video generator with the price, image quality and perks of stylegan2.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Stylegan-v: A continuous video generator with the price, image quality and perks of stylegan2

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.542963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.542963Z digest=sha256:19f4fbd01dc649c1d2abebc7a17f5a1075c7164302ffad0546077384854f2603

Observation 8164accd-5bcc-43e5-9479-a94d0cddb4b3 · outbound

This paper cites Auto-encoders.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Auto-encoders

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.547400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.547400Z digest=sha256:33ee75c727cd94f66ef7f0c9ec0fae88faa4e99aa3576b3441a7d14298151587

Observation 50890519-6ac8-48af-b85e-3762cfbd14b1 · outbound

This paper cites Auto-encoding variational bayes, 2022.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Auto-encoding variational bayes, 2022

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.551351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.551351Z digest=sha256:09a72ba51240a9319d4f2d1a61e761dde9024e6d3dd1bb328115c2d5bdd9be87

Observation 5f70f62d-b848-41af-bbdb-adcf31b3f3c1 · outbound

This paper cites Masked autoencoders are scalable vision learners.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Masked autoencoders are scalable vision learners

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.555242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.555242Z digest=sha256:749df84952dcb9b474cdd1b1dc50175e4bcb091a58231f41a50839e1010eec41

Observation 61f3e908-ac56-4dc1-8a69-f718088e875d · outbound

This paper cites High-Resolution Image Synthesis with Latent Diffusion Models.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation High-Resolution Image Synthesis with Latent Diffusion Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.558857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.558857Z digest=sha256:2f095e5850508a0ad3163ef422cb52eeb427ba4e4347feb413efa5a36d0fbac7

Observation 1a2a6069-d35a-4b7f-88a0-124b30532e7a · outbound

This paper cites Neural discrete representation learning, 2018.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Neural discrete representation learning, 2018

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.562789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.562789Z digest=sha256:a7e37b6999b2705620ffca94e5eb159383b22590d279780715190775d55a1064

Observation b61f57a4-e37a-4dcf-9e92-2af188daac12 · outbound

This paper cites Videogpt: Video generation using vq-vae and transformers, 2021.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Videogpt: Video generation using vq-vae and transformers, 2021

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.567054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.567054Z digest=sha256:5910fbad8489ac6102146adceac0a71c6870a04d5bc04e0cc089c0d9ba0e0538

Observation a5ee5ed3-4b19-462f-b91c-372ee992677b · outbound

This paper cites Taming Transformers for High-Resolution Image Synthesis.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Taming Transformers for High-Resolution Image Synthesis

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.571955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.571955Z digest=sha256:cd3875019d0e1b1cb9a6250e58ed13f98ca92df5e389b8df135f09ead48d68f2

Observation 297c76c4-49f7-43c9-a868-b2a61026568c · outbound

This paper cites Clip: Connecting vision and language with contrastive learning.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Clip: Connecting vision and language with contrastive learning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.575632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.575632Z digest=sha256:d7a67c8ad2eba4ffa984622c226abccc837929b4df52ca7b42f2dc0fd92ffd02

Observation 7886cfb5-4676-4cf1-ba4a-debe256d3d34 · outbound

This paper cites Hierarchical patch vae-gan: Generating diverse videos from a single sample, 2020.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Hierarchical patch vae-gan: Generating diverse videos from a single sample, 2020

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.580081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.580081Z digest=sha256:022ea1378fb91dfdb80fe21fcfdbe00b114268102e161e699cc7df724081e7ab

Observation 4ba1198f-ce4d-4e72-ab5c-387897c1f47f · outbound

This paper cites Visual anomaly detection in video by variational autoencoder, 2022.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Visual anomaly detection in video by variational autoencoder, 2022

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.584402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.584402Z digest=sha256:5e77dc4487e6ccb1417e51b7f33b2f19edcd858323256750f97ce395c4ce9451

Observation 58e45002-cae4-4d24-a161-718163aa8fcc · outbound

This paper cites VideoMAC: Video Masked Autoencoders Meet ConvNets.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation VideoMAC: Video Masked Autoencoders Meet ConvNets

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.588528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.588528Z digest=sha256:f459435c32dc69e9b8b30eaefbd290337a4651d7d174274081b4f4a08fd2de74

Observation d9e113b1-ea31-4de7-990f-c0216fbf308e · outbound

This paper cites Hauptmann, Ming-Hsuan Yang, Yuan Hao, Irfan Essa, and Lu Jiang.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Hauptmann, Ming-Hsuan Yang, Yuan Hao, Irfan Essa, and Lu Jiang

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.592709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.592709Z digest=sha256:cd3a088b600c8c7900b837399e8322606172c97477f3d4893412b9b62f45668d

Observation e711213b-20c4-45f5-9782-039ff7bb9352 · outbound

This paper cites MAGVLT: Masked Generative Vision-and-Language Transformer.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation MAGVLT: Masked Generative Vision-and-Language Transformer

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.596657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.596657Z digest=sha256:31871ddba954a641f6b593816f7cf7dbb7de6a1758f9ed1fb69927d4706c35d0

Observation f54feb84-e2f7-4198-a32d-a5ccf71850a6 · outbound

This paper cites Multi-generator generative adversarial nets, 2017.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Multi-generator generative adversarial nets, 2017

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.601123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.601123Z digest=sha256:bec8aed56c3778d2387904c3be6df0a2368c51ef1d874907d9e505424f77b1f9

Observation 672f3283-3b57-4842-b154-c837708ea08f · outbound

This paper cites Gomez, Łukasz Kaiser, and Illia Polosukhin.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Gomez, Łukasz Kaiser, and Illia Polosukhin

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.605272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.605272Z digest=sha256:164a9dba9997448bdd16ed98558624536704a12663776e5374906c8e0b651034

Observation ab67dac8-77ec-41ef-8cba-0fd9cccc342f · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation An image is worth 16x16 words: Transformers for image recognition at scale

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.609789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.609789Z digest=sha256:2ba3c33903ce2b407317c27c66a7e4fd1f0def2df5fe6202480d0ff7458423fd

Observation 674be2a5-a75d-4ef7-97e0-36899ebd4a5e · outbound

This paper cites Zero-shot text-to-image generation.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Zero-shot text-to-image generation

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.614593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.614593Z digest=sha256:625a71768fed2c6575aaa866cd30209db2524587b6de2641fae9d877ac9fef20

Observation 24476f66-b907-4166-8f91-e195258a65c5 · outbound

This paper cites Discrete variational autoencoders, 2017.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Discrete variational autoencoders, 2017

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.618550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.618550Z digest=sha256:582d8151d2574d8d1fa3a3980ba3ccbc3c50322f7e37b39a617171706d08d634

Observation 8f4e5b6e-8a70-4976-90e8-aebd69aac744 · outbound

This paper cites Cogview: mastering text-to-image generation via transformers.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Cogview: mastering text-to-image generation via transformers

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.622499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.622499Z digest=sha256:41ec77202ff0bcc18eb08d306ac77a4cd904c136fe8f1180b45163592865f304

Observation 0759d869-67bf-4a33-b9bb-78a8a21537b0 · outbound

This paper cites Cogview2: faster and better text-to-image generation via hierarchical transformers.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Cogview2: faster and better text-to-image generation via hierarchical transformers

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.627068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.627068Z digest=sha256:3ed6a3d9a7a30109dae014d396ea9872fe30f174f3ccccb954c85565d53f7575

Observation 1cd61a27-7f57-4bc3-9689-2c26826f016f · outbound

This paper cites ViViT: A Video Vision Transformer.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation ViViT: A Video Vision Transformer

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.630981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.630981Z digest=sha256:046f1d89e57a69d62c1bcfdafa66a526c457a11f88c86f81a9f5e87677a20b10

Observation bf455fbe-2f4b-43ad-bb6e-d674fed3e654 · outbound

This paper cites Phenaki: Variable length video generation from open domain textual descriptions.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Phenaki: Variable length video generation from open domain textual descriptions

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.634625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.634625Z digest=sha256:d50e2399535ba5d654c1c620a4657549e90fa04e4df6b5e1a1a33e26b2fd11a5

Observation 98e6907e-fa47-424b-8263-76092dc5c8ce · outbound

This paper cites an unresolved cited work.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Unresolved cited work

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.639033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.639033Z digest=sha256:518ecce8bf58a1abb290e4d5db7141f60e85afa8c717d6f3150b2b425766876c

Observation d51c6616-a625-465e-a36b-9ee835221f0b · outbound

This paper cites Compositional 3d-aware video generation with llm director.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Compositional 3d-aware video generation with llm director

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.646492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.646492Z digest=sha256:22f97b9ca069b162fe79ab3c7bee0a5fa087cefc61f9aaba5795f93848afb5e8

Observation ce7d04d6-70c4-4cfe-b6b6-fba8dafb62be · outbound

This paper cites an unresolved cited work.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Unresolved cited work

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.651064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.651064Z digest=sha256:7e61e311b4eed52044220c3e2f1e66e1cb3fc07462c622c650a8d376c21c9c3e

Observation c0624193-859c-441a-9360-7b327783af10 · outbound

This paper cites Vlogger: Make your dream a vlog.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Vlogger: Make your dream a vlog

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.659432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.659432Z digest=sha256:12a0eef2b9bb623a5e5257b0bf44b9263a2f96708d557dcee7facf87f2e1027b

Observation 28464bab-cacd-42a0-8b8e-6a36d6b8dfba · outbound

This paper cites Weiss, Niru Maheswaranathan, and Surya Ganguli.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Weiss, Niru Maheswaranathan, and Surya Ganguli

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.663326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.663326Z digest=sha256:3476acf21e7ca44d1365c420ab5ec57bb4ed83b1a3b865a4e38d97c6b6277a53

Observation 809e94af-516a-42df-8fa9-9cda0b60434d · outbound

This paper cites Denoising diffusion probabilistic models.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Denoising diffusion probabilistic models

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.667295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.667295Z digest=sha256:5e50fbf80cb8e1ab5c2721c7709e327ccfa17d491df79f294689d21d6bccb2aa

Observation 9c6abda2-81d8-4615-9f4a-5b466b82d892 · outbound

This paper cites Generative modeling by estimating gradients of the data distribution, 2020.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Generative modeling by estimating gradients of the data distribution, 2020

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.671215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.671215Z digest=sha256:11a69541a253c42b9486379a05f06b31a8f48623a504df6232c6ef15df16c646

Observation c0612117-642a-4c76-892f-5cf5655b624e · outbound

This paper cites Bert: Pre-training of deep bidirectional transformers for language understanding, 2019.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Bert: Pre-training of deep bidirectional transformers for language understanding, 2019

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.674802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.674802Z digest=sha256:d7e1ffeb3e8f6ec255b0a83d9a7b27a6cac31f53c07ea98f5cf28fcf978fb3bb

Observation 980af4e7-9f53-4960-96ba-144781e41d43 · outbound

This paper cites an unresolved cited work.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Unresolved cited work

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.678399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.678399Z digest=sha256:6c658c84ee858a6d03426b58c9650afa75e7bc6af29ec4e7a2501c89d20fdc59

Observation 26429073-5858-480c-b75f-2363e69b35e5 · outbound

This paper cites Video generation models as world simulators.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Video generation models as world simulators

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.682202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.682202Z digest=sha256:182ac4b888765a010268c2516085b886e9f01a549a113251fb786faa797768d0

Observation d54f6a1c-577c-4c6b-9ab5-9036e18a0c1c · outbound

This paper cites Scalable diffusion models with transformers.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Scalable diffusion models with transformers

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.685861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.685861Z digest=sha256:a0125a0e45b666d533234ae90265283984c982a44cd6bb6d3daefc1826632743

Observation 228990ab-044b-4fa8-948d-bde650ce43ae · outbound

This paper cites Cogview: Mastering text-to-image generation via transformers.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Cogview: Mastering text-to-image generation via transformers

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.689203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.689203Z digest=sha256:b047f4ff3cff2a7d38137455ecc131a9b579fea756cfd7b67394933180dc5b32

Observation d1716863-5645-48f7-9882-90e13cd091ff · outbound

This paper cites Videotetris: Towards compositional text-to-video generation.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Videotetris: Towards compositional text-to-video generation

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.692629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.692629Z digest=sha256:b375b552d47ee6f1ad18541cb337b4b0e142ef8606e8a8847f67319947a701b9

Observation b4aa16af-7992-453a-8999-5dff9ab53953 · outbound

This paper cites Cogview2: Faster and better text-to-image generation via hierarchical transformers.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Cogview2: Faster and better text-to-image generation via hierarchical transformers

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.696025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.696025Z digest=sha256:aa48cc0244f6663d6bebefee4c292eb302e8ec5b8598a4b5601b11d781bcbee4

Observation a992bdac-de6e-4506-a688-280629395fdf · outbound

This paper cites NUWA-infinity: Autoregressive over autoregressive generation for infinite visual synthesis.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation NUWA-infinity: Autoregressive over autoregressive generation for infinite visual synthesis

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.699925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.699925Z digest=sha256:6df5aade753fae3d3b28c7d7e91170506cff638703c6cb13a379e4c1461cae6d

Observation e0b12db0-ee9f-4511-b94f-0b7e592b278c · outbound

This paper cites Ross, Bryan Seybold, and Lu Jiang.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Ross, Bryan Seybold, and Lu Jiang

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.703594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.703594Z digest=sha256:57b4d79de0ec55a5446275e36e1936fbc7e0512cd5ff8f53708db3b4e113485a

Observation 78620d1d-b30a-4889-8c72-12a49b0d737f · outbound

This paper cites ART •V: Auto-Regressive Text-to-Video Generation with Diffusion Models.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation ART •V: Auto-Regressive Text-to-Video Generation with Diffusion Models

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.707116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.707116Z digest=sha256:34440be76c859df7f41a73b08ae853513bc14296a70f5c4c85917471f4f5b3a6

Observation 6571763f-b74d-4894-8b49-14ac584d540d · outbound

This paper cites Grid diffusion models for text-to-video generation.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Grid diffusion models for text-to-video generation

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.711430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.711430Z digest=sha256:308c8983f1355cf1057c1db37ab388c5e12614ebe8dc9a08282055b54bfb2d9d

Observation b86fa8d6-a11d-4e74-9d81-a7bb7737db67 · outbound

This paper cites Arlon: Boosting diffusion transformers with autoregressive models for long video generation, 2025.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Arlon: Boosting diffusion transformers with autoregressive models for long video generation, 2025

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.720203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.720203Z digest=sha256:0f4a645680da6e915aa54f63a9c5e91134ce92a641e54cbd45180cc77fca02ea

Observation bd6c911e-4094-40f0-9bb3-02dfd32ce945 · outbound

This paper cites Stylegan-v: A continuous video generator with the price, image quality and perks of stylegan2.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Stylegan-v: A continuous video generator with the price, image quality and perks of stylegan2

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.724027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.724027Z digest=sha256:45a00eee73ac9fec54f43964288d52a5b1c1b0c064cd634ed1308e4779bb4187

Observation 5f8b9b6d-ee6a-44b6-9fb2-c37d3c6a3108 · outbound

This paper cites an unresolved cited work.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Unresolved cited work

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.728231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.728231Z digest=sha256:74f28229b49f7d42c3aeee017a17f1444b953f9ff1fcfd5e81ee87dcdb774941

Observation 26cdc14d-8f9b-4186-8e3d-eca183b36d4a · outbound

This paper cites Srinivasan, Matthew Tancik, Jonathan T.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Srinivasan, Matthew Tancik, Jonathan T

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.731871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.731871Z digest=sha256:9d335497838ba51303cff76f126cc605f45ffb677c7acd7ac188b9c8fcfdcbe1

Observation 1cfdf55d-3754-499c-a59d-a0dcff3e2ac7 · outbound

This paper cites Video probabilistic diffusion models in projected latent space.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Video probabilistic diffusion models in projected latent space

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.735767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.735767Z digest=sha256:aeec350b4a0a553ef35384a290b5a4ec512186079ffd882ed10d5da7c834811e

Observation ed4713e1-e138-4a54-a0d6-a55c607b859a · outbound

This paper cites Towards end-to-end generative modeling of long videos with memory-efficient bidirectional transformers.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Towards end-to-end generative modeling of long videos with memory-efficient bidirectional transformers

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.739757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.739757Z digest=sha256:ec8a0e9bb1e4456ca95b4619fd0049d8175c62d7d150792ef637019692d796f1

Observation 42503dbb-4772-4626-87d6-e3c6cdc03858 · outbound

This paper cites Streamingt2v: Consistent, dynamic, and extendable long video generation from text.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Streamingt2v: Consistent, dynamic, and extendable long video generation from text

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.744775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.744775Z digest=sha256:09d553387bbf742634adda5bdf933f381edcf5485c46a7750231129f3f1f8e72

Observation ce5f2492-ebb5-420a-8b07-136f55fc2c05 · outbound

This paper cites Vid-gpt: Introducing gpt-style autoregressive generation in video diffusion models, 2024.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Vid-gpt: Introducing gpt-style autoregressive generation in video diffusion models, 2024

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.748917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.748917Z digest=sha256:37dbbae13b00283776d1ece3a2d1a75ac51ce1886f668277daafb42131e52aaa

Observation faec5de8-142b-4706-9525-e9461a8ae33e · outbound

This paper cites Flexifilm: Long video generation with flexible conditions, 2024.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Flexifilm: Long video generation with flexible conditions, 2024

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.753119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.753119Z digest=sha256:702feeca4f02146791df53d14ef47f9b65ca8271e6be6a46a80a853b0410f48b

Observation 0a691872-05e8-4712-b187-4c5bde993731 · outbound

This paper cites Mora: Enabling generalist video generation via a multi-agent framework, 2024.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Mora: Enabling generalist video generation via a multi-agent framework, 2024

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.756938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.756938Z digest=sha256:c982dfcba30ec0ae215bf44cd9e6c25b82cccb1d7b0725435878f045e8f05e57

Observation abd7f13e-d49e-4ebf-8995-fbaa27bad587 · outbound

This paper cites Vidgen: Long-form text-to-video generation with temporal, narrative and visual consistency for high quality story-visualisation tasks.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Vidgen: Long-form text-to-video generation with temporal, narrative and visual consistency for high quality story-visualisation tasks

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.760816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.760816Z digest=sha256:ca8573cfc7ab95d83929e06d5d5eadf1c257437c346acf21751cc16ba1585664

Observation 3fdc9612-8db6-496e-8e11-3ea6f1af96d3 · outbound

This paper cites Bissyand, and Saad Ezzini.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Bissyand, and Saad Ezzini

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.764885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.764885Z digest=sha256:8081c5eb23fdd353ed1378a3fda17d41b2d9041a2bc9b68b65cc0b1cec6bc0a6

Observation 3df02390-71d4-41d1-8fb6-fe2e817cad53 · outbound

This paper cites Kubrick: Multimodal agent collaborations for synthetic video generation, 2024.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Kubrick: Multimodal agent collaborations for synthetic video generation, 2024

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.768889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.768889Z digest=sha256:c3cd2b55e7341cdd527478859b82a9f41a8b89edcb0f2c0f008f4713b3ab477a

Observation 6efd1475-3e4f-4d52-9704-cea6801df19d · outbound

This paper cites Modelscope text-to-video technical report, 2023.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Modelscope text-to-video technical report, 2023

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.772557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.772557Z digest=sha256:eb3224b8fd0fbc8e007874b1f09dd9c14d4d0d466219cc7b8f074f0c4fddb5ad

Observation d1255750-54c6-49cf-92e6-f11883e31db6 · outbound

This paper cites Free-bloom: zero-shot text-to-video generator with llm director and ldm animator.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Free-bloom: zero-shot text-to-video generator with llm director and ldm animator

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.776147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.776147Z digest=sha256:c0174999c655e05d48897b7d199077286064943e16cee046a58dc67078aa0368

Observation e012fc7e-558b-4ad8-9fe4-3f393f872eaf · outbound

This paper cites Denoising diffusion implicit models.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Denoising diffusion implicit models

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.779257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.779257Z digest=sha256:b4458521b5ba3bceec9ead9b317d841b366de2267865aa5ec97a670839ef199c

Observation 62041886-b668-4390-bd61-c099e3b420b8 · outbound

This paper cites LLM-grounded video diffusion models.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation LLM-grounded video diffusion models

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.783124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.783124Z digest=sha256:b14b725e74f9353bcb3cd82c58a15fc10bedc17fb44898fdd8998ea2b17f87ad

Observation d40aef87-2cf7-4a7b-bf98-e75a1218d6a6 · outbound

This paper cites Mevg: Multi-event video generation with text-to-video models.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Mevg: Multi-event video generation with text-to-video models

Reference 101

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.787359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.787359Z digest=sha256:f48cb967495f6035ed4c995df6d7cf6f4a7df09f114aa624742d4280b7972f3b

Observation 91f84484-2120-48d6-888c-7bca4ee84596 · outbound

This paper cites an unresolved cited work.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Unresolved cited work

Reference 102

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.791354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.791354Z digest=sha256:11712620983afac09f5d43bc4292d8b02202d48195b683bf72fe9003cd09724b

Observation e539ca34-483d-4f48-b056-be507e070be9 · outbound

This paper cites Videomerge: Towards training-free long video generation, 2025.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Videomerge: Towards training-free long video generation, 2025

Reference 103

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.796952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.796952Z digest=sha256:dea2d126c215e0f88fc474627fe553fe2c8599111c46bc01f66a98ed90ac39c6

Observation ede348e8-a4a3-42a1-8b65-f391622a7597 · outbound

This paper cites Mevg: Multi-event video generation with text-to-video models, 2024.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Mevg: Multi-event video generation with text-to-video models, 2024

Reference 104

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.801702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.801702Z digest=sha256:93fab88b8efae618435d5ac376373a9e11575c97a901be12e83f396be36b3832

Pith citing papers

Observation ed620966-a4cf-4a2b-a333-065cb5f1806b · inbound

Enhancing Train-Free Infinite-Frame Generation for Consistent Long Videos cites this paper.

Enhancing Train-Free Infinite-Frame Generation for Consistent Long Videos Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:58:13.837416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T10:56:16.054496Z digest=sha256:9b924751b3571ac5a90192ad58900983ff1b7431e7a41ac8e19862346b9c11f8