Pith. sign in

Paper Citation Record · LEDGER

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation

As of 18 August 2026, this Paper Citation Record lists 100 of 187 outbound references and 1 inbound Pith citation observation for arXiv:2412.18688.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.18688 v2

Coverage vector

measured 100 of 187 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T04:37:14.801702Z

measured 101 of 101 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-20T10:56:16.054496Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T10:58:13.835764Z

Reference resolution

100 of 187 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved100
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3f5c52f7-f554-4423-98b2-4ccc57f1ccd9 · outbound

This paper cites Video generation models as world simulators by open a.i, 2024.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Video generation models as world simulators by open a.i, 2024

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.383491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.383491Z digest=sha256:32efc5007a944dfd4dc76829e6d3a2b3b6d7f1b44799ea3c630cf276e26d52e7

Observation 1d6830a5-4ef6-4a7a-8123-2b99072bcdf8 · outbound

This paper cites Introducing chatgpt by open a.i, 2022.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Introducing chatgpt by open a.i, 2022

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.387900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.387900Z digest=sha256:6c006a4a65b39f32214b9af36c992588a333b2f5615de2750708343618c5734c

Observation 73e61bf6-ecc7-4a84-b4f1-7cd2999a00ad · outbound

This paper cites Meta llama models, 2023.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Meta llama models, 2023

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.391439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.391439Z digest=sha256:b10f879bb3cde044890ae06e7461faa4c93a6dfd1f849172580e5464f919b326

Observation 6dba3691-0a87-47da-8617-83bc04eecbe7 · outbound

This paper cites Google gemini series, 2023.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Google gemini series, 2023

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.395190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.395190Z digest=sha256:1a512c56aada5d05c835a98378c56be70187206b38a82b7056d4666f1e61ab70

Observation a2576c1b-ec00-41d2-9df2-f3a8f6501a1f · outbound

This paper cites Anthropic by claude, 2023.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Anthropic by claude, 2023

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.400143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.400143Z digest=sha256:5ec1314c772971884fcec6f9d839f2c55e09adc74e37e1af4acb9d3e17214e87

Observation 49f9dc04-d8d1-4b41-b5d7-42ff15665d72 · outbound

This paper cites Mistral large model by mistral, 2024.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Mistral large model by mistral, 2024

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.404176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.404176Z digest=sha256:81ec4594324ce238139659a1d1e4b1fb61587722529aa1666ac277ae34190800

Observation f6b7943f-0f31-4d0e-9c8c-90ed161f509e · outbound

This paper cites Hierarchical text-conditional image generation with clip latents, 2022.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Hierarchical text-conditional image generation with clip latents, 2022

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.408021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.408021Z digest=sha256:9fbc73ce589b998d61f5fd0f7081b5e91909220e0f9b44124601302787f9fec1

Observation a5b2b837-f056-439f-ac48-4f0c739eafb5 · outbound

This paper cites Scaling rectified flow transformers for high-resolution image synthesis, 2024.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Scaling rectified flow transformers for high-resolution image synthesis, 2024

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.411624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.411624Z digest=sha256:8b7a7d3fc435f2a35cd213c06698a0f0070b60a2bc439bd99271e93736eaf541

Observation 44c30872-f7a0-42fc-a98a-ffa15714742b · outbound

This paper cites The midjourney v5.2 model for image generation, 2024.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation The midjourney v5.2 model for image generation, 2024

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.416060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.416060Z digest=sha256:cef65d60295198839bc5cf1a552a1046e9b6037a25aeeac0364cd41c51b8bbe6

Observation 6db003ce-e59e-4e92-9358-1448c9b2bf96 · outbound

This paper cites VideodirectorGPT: Consistent multi-scene video generation via LLM-guided planning, 2024.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation VideodirectorGPT: Consistent multi-scene video generation via LLM-guided planning, 2024

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.423947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.423947Z digest=sha256:4b4a1895c7eef0b1fc725ef3046c68a9258f2db826f65a5f0fb32ecd95fa711b

Observation eb451ad4-0f24-47cb-a8c4-f194b88e4719 · outbound

This paper cites Make-a-video: Text-to-video generation without text-video data.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Make-a-video: Text-to-video generation without text-video data

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.428291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.428291Z digest=sha256:6fdf2089101dba104c58a42a9b70d3badd6b5bbc99fe9bbfe30c669edf25f0f9

Observation caefaae3-3e22-46ba-998c-01e17f0cf005 · outbound

This paper cites Gen2 by runway ml, 2023.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Gen2 by runway ml, 2023

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.432170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.432170Z digest=sha256:2606fec0c5ad96da2e895b410b8e7f89e036f060950eaaa0ed532396cf74bc7c

Observation 548c3fb9-bd96-4d36-b82b-7ae71c8772a5 · outbound

This paper cites Cogvideo: Large-scale pretraining for text-to-video generation via transformers.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Cogvideo: Large-scale pretraining for text-to-video generation via transformers

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.435916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.435916Z digest=sha256:4da5a341e18dc44efbb482c19a455d5c05dbda9fe440933f804eb0d7112d092c

Observation 87325dd2-a846-47d6-9444-4dc5139caccc · outbound

This paper cites Gen-l-video: Multi-text to long video generation via temporal co-denoising, 2023.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Gen-l-video: Multi-text to long video generation via temporal co-denoising, 2023

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.443960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.443960Z digest=sha256:70f31857ba14cf45eced862bae4a7007098cccc433e8980c2822fe204e0aa124

Observation 019d4e69-d36e-4506-88b0-161f410528ed · outbound

This paper cites Gen-4 alpha by midjourney, 2024.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Gen-4 alpha by midjourney, 2024

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.447977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.447977Z digest=sha256:1c2d94939a71cb1d19751fd9d1cac9a7da502381e6a3430b01699d10b7b06026

Observation 27302c44-f4a3-4abf-8d6b-d0e0c7b08586 · outbound

This paper cites A survey on long video generation: Challenges, methods, and prospects, 2024.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation A survey on long video generation: Challenges, methods, and prospects, 2024

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.451981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.451981Z digest=sha256:d5c31c6528849e2451e8333e7a527b6c570ab186621959b6bcb716ea52b2b1a0

Observation 7a0018e5-fb00-4a80-85a6-0e33bf002ca1 · outbound

This paper cites A survey on generative ai and llm for video generation, understanding, and streaming, 2024.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation A survey on generative ai and llm for video generation, understanding, and streaming, 2024

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.455327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.455327Z digest=sha256:adbfd32897512049188d7842a5b5add34b67088fd918be2208522fe9d745df8d

Observation 8cad888a-61a3-4a28-ab96-7b46cbe14209 · outbound

This paper cites Sora: A review on background, technology, limitations, and opportunities of large vision models, 2024.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Sora: A review on background, technology, limitations, and opportunities of large vision models, 2024

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.458733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.458733Z digest=sha256:1c3dd7c7abeb852a060719f179b7d8b79cc7d6b19f5a9f9a90a2af8ffe2ad61b

Observation 8ce84fef-7f8d-445f-85ab-717bc83509e8 · outbound

This paper cites Free-bloom: Zero-shot text-to-video generator with LLM director and LDM animator.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Free-bloom: Zero-shot text-to-video generator with LLM director and LDM animator

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.462880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.462880Z digest=sha256:74456d0456b2279ecb457d1ab504547d7da4217515692e9531e19aeb70592f2c

Observation b33cc5f5-e159-4f70-b46d-6983ea86a88e · outbound

This paper cites Flowzero: Zero-shot text-to-video synthesis with llm-driven dynamic scene syntax, 2023.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Flowzero: Zero-shot text-to-video synthesis with llm-driven dynamic scene syntax, 2023

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.466710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.466710Z digest=sha256:88a3afc395fd2842759078a173dba57c8787596432d6b73e73c9cbb7ace9120e

Observation 53516192-30a6-49d7-adbd-aa0feaa1087a · outbound

This paper cites Align your latents: High- resolution video synthesis with latent diffusion models.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Align your latents: High- resolution video synthesis with latent diffusion models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.469740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.469740Z digest=sha256:b561257319564e51a648b54e81f963b82e74b66fef0ab0f9364a938f9400bf5c

Observation 1db71ba0-b625-424c-a9fd-83f4767161af · outbound

This paper cites Mavin: Multi-action video generation with diffusion models via transition video infilling, 2024.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Mavin: Multi-action video generation with diffusion models via transition video infilling, 2024

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.473186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.473186Z digest=sha256:eb0b31b14e082a88dc0e693fa115459a2cb07e65cd02db8c0e9f1fb995b6729b

Observation 29eaa73f-033d-47ea-9e4c-183dffd434db · outbound

This paper cites SEINE: Short-to-long video diffusion model for generative transition and prediction.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation SEINE: Short-to-long video diffusion model for generative transition and prediction

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.476649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.476649Z digest=sha256:88e7adf309f7b27792138eebd750949c29beb749913e4700c0b24294d42b4c3b

Observation e7d90b62-6189-4c5e-ba2c-30b251999f9e · outbound

This paper cites Learning transferable visual models from natural language supervision, 2021.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Learning transferable visual models from natural language supervision, 2021

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.480144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.480144Z digest=sha256:6c57c61de631d6225accb5cef25cc6dfd6370ec61864fba2b4ced230e88d6dcd

Observation a3e3613a-024c-450f-bee0-a403509c066d · outbound

This paper cites Generative adversarial networks.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Generative adversarial networks

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.483541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.483541Z digest=sha256:d9888c7749e059ab083f6318d816c78b7a79aa855bd2f658eca4797b330d19fb

Observation 25f2d8a0-6316-4d33-a13c-9e33b9c973fc · outbound

This paper cites Unsupervised representation learning with deep convolutional generative adversarial networks, 2016.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Unsupervised representation learning with deep convolutional generative adversarial networks, 2016

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.488168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.488168Z digest=sha256:7135ebe458941cc773cead6c5bdd00d9dbd6f910ad2b9ba6215242b3bd2d27b0

Observation a6101203-873c-4616-a164-fea0b5d4ccbe · outbound

This paper cites Deep generative image models using a laplacian pyramid of adversarial networks.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Deep generative image models using a laplacian pyramid of adversarial networks

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.492876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.492876Z digest=sha256:148732e0e4eb5cfd66eb63a3e40600440f741927c868972c3517c8ba1f277ab6

Observation 6bcb1f2f-56be-45d9-b22c-2b7ff2309283 · outbound

This paper cites Stackgan: Text to photo-realistic image synthesis with stacked generative adversarial networks.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Stackgan: Text to photo-realistic image synthesis with stacked generative adversarial networks

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.496953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.496953Z digest=sha256:aa0c3e26cc5bf3358dfb457e472c1e43f0e804db7960a5ea81725991c6acceec

Observation 3b43eca4-4348-4faa-998a-04ab263bc7bc · outbound

This paper cites Attngan: Fine-grained text to image generation with attentional generative adversarial networks.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Attngan: Fine-grained text to image generation with attentional generative adversarial networks

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.501400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.501400Z digest=sha256:922c44f290f8dcb91b72b8c7a45d475a625ea7e13286e16dd184a344a380490f

Observation 471641e0-f60c-46e4-be61-2485a4e9bad9 · outbound

This paper cites Analyzing and Improving the Image Quality of StyleGAN.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Analyzing and Improving the Image Quality of StyleGAN

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.506500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.506500Z digest=sha256:d7caf2384c5cea54aace26f1c4efad53980eaf4fc88c227dc0ca81d677ff258e

Observation 747ac8e1-caa9-45ce-96ee-418a2f1364df · outbound

This paper cites an unresolved cited work.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.510690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.510690Z digest=sha256:13d8c86dcc0a8ceb603ea315ec599ff367d77e0d334ca5e48f68a39567ae761a

Observation a8ccc019-4d0c-4d9a-ba45-8ae6980ec084 · outbound

This paper cites Stylegan2 distillation for feed-forward image manipulation.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Stylegan2 distillation for feed-forward image manipulation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.515413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.515413Z digest=sha256:0966491aaa82c511d036df4b293dd7f53b223d20a570c2e0bebaa478b59319d6

Observation f964f7cf-5ee7-417d-a84f-417245c0816f · outbound

This paper cites an unresolved cited work.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.520232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.520232Z digest=sha256:8c5a6f630c2dd632c1f9ae9538fb23811a47bf842563025c12e6984e5d72422a

Observation a5deffb8-65f6-4ee0-b9af-fe946c7c2278 · outbound

This paper cites Deep multi-scale video prediction beyond mean square error, 2016.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Deep multi-scale video prediction beyond mean square error, 2016

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.523881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.523881Z digest=sha256:94daaf802528092f17915aeb37ac38ef6ac6e5aa18945cb50a8d5528c437509d

Observation cb3c7ede-6527-44d8-90ee-7ea998f07ca6 · outbound

This paper cites Generating videos with scene dynamics, 2016.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Generating videos with scene dynamics, 2016

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.527985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.527985Z digest=sha256:2ea976ee85f96cb408e72e4b79ca7dda390d4cec05e5e028ac906f30884480a0

Observation 585fb8f7-0c81-43d4-a131-f0bcbccb5de0 · outbound

This paper cites Learning to generate time-lapse videos using multi-stage dynamic generative adversarial networks.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Learning to generate time-lapse videos using multi-stage dynamic generative adversarial networks

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.531670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.531670Z digest=sha256:632f40e4f42d8721bd68d4e2319e58abc7dc6b47032e011b714d9ddb993811e1

Observation 9956f5ac-8c22-4648-bafe-41cebffc43d5 · outbound

This paper cites To create what you tell: Generating videos from captions, 2018.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation To create what you tell: Generating videos from captions, 2018

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.534876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.534876Z digest=sha256:cb9dee56d84a2aeabee098cb1333d2ad11bba02237154ad1aeed1d601ba61c9c

Observation 771ed588-f2af-443e-b25b-ce4f98d2c5b6 · outbound

This paper cites Generating videos with dynamics-aware implicit generative adversarial networks.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Generating videos with dynamics-aware implicit generative adversarial networks

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.538591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.538591Z digest=sha256:4b06c638a20934d4b366fad9e70e56f287be88e534ad6e14077e1e20a804ba20

Observation 3a56a940-90e5-4593-b4f3-ba8c55d5907d · outbound

This paper cites Stylegan-v: A continuous video generator with the price, image quality and perks of stylegan2.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Stylegan-v: A continuous video generator with the price, image quality and perks of stylegan2

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.542963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.542963Z digest=sha256:6ec7d8008c774d87a9bf78d71036bd41c2eb1cf423c4f7177e16611e11cde874

Observation 8164accd-5bcc-43e5-9479-a94d0cddb4b3 · outbound

This paper cites Auto-encoders.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Auto-encoders

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.547400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.547400Z digest=sha256:af9cc8a24dcfbd0ee40e657d7852ccc380e59054c4c034662f0b4b23d0fab868

Observation 50890519-6ac8-48af-b85e-3762cfbd14b1 · outbound

This paper cites Auto-encoding variational bayes, 2022.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Auto-encoding variational bayes, 2022

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.551351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.551351Z digest=sha256:118dbd2c0d72c4542576d727266d9f7715a53817aad59f312585d6f140e271ee

Observation 5f70f62d-b848-41af-bbdb-adcf31b3f3c1 · outbound

This paper cites Masked autoencoders are scalable vision learners.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Masked autoencoders are scalable vision learners

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.555242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.555242Z digest=sha256:9f991117a3eb3b122bbfc2f2d492fb5abcdd5f66f6476aa4fffb8aa433968dc6

Observation 61f3e908-ac56-4dc1-8a69-f718088e875d · outbound

This paper cites High-Resolution Image Synthesis with Latent Diffusion Models.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation High-Resolution Image Synthesis with Latent Diffusion Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.558857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.558857Z digest=sha256:938ec9d8037a72ab84512bc29c4f8088fde1e350763a533ec0f48b0f729bfbc1

Observation 1a2a6069-d35a-4b7f-88a0-124b30532e7a · outbound

This paper cites Neural discrete representation learning, 2018.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Neural discrete representation learning, 2018

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.562789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.562789Z digest=sha256:60bf50b4963c368c5ba736e233d1d170a351f934a2b03958f681c99187839fac

Observation b61f57a4-e37a-4dcf-9e92-2af188daac12 · outbound

This paper cites Videogpt: Video generation using vq-vae and transformers, 2021.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Videogpt: Video generation using vq-vae and transformers, 2021

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.567054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.567054Z digest=sha256:da4ee465dfdca2387852457856bd71c29f3fd3c25eabb59512b846054fe370a5

Observation a5ee5ed3-4b19-462f-b91c-372ee992677b · outbound

This paper cites Taming Transformers for High-Resolution Image Synthesis.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Taming Transformers for High-Resolution Image Synthesis

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.571955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.571955Z digest=sha256:1f35ee1e1c9f4f6cdc73fc824ff42c8ba809de743d24003145c6cd4b302d1196

Observation 297c76c4-49f7-43c9-a868-b2a61026568c · outbound

This paper cites Clip: Connecting vision and language with contrastive learning.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Clip: Connecting vision and language with contrastive learning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.575632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.575632Z digest=sha256:865323a9b634a9e18293b38deca1c369dbc9f2fa1c2615abaa2f30b4b80df1d3

Observation 7886cfb5-4676-4cf1-ba4a-debe256d3d34 · outbound

This paper cites Hierarchical patch vae-gan: Generating diverse videos from a single sample, 2020.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Hierarchical patch vae-gan: Generating diverse videos from a single sample, 2020

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.580081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.580081Z digest=sha256:c635795a73554cee6f9deaee1fef611ea2af0d2c202751f9901b45ca01f249c7

Observation 4ba1198f-ce4d-4e72-ab5c-387897c1f47f · outbound

This paper cites Visual anomaly detection in video by variational autoencoder, 2022.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Visual anomaly detection in video by variational autoencoder, 2022

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.584402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.584402Z digest=sha256:5c351eccc6a22edfe43f75e09f599194a1f7a45b53d6ad3c03104a928441270a

Observation 58e45002-cae4-4d24-a161-718163aa8fcc · outbound

This paper cites VideoMAC: Video Masked Autoencoders Meet ConvNets.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation VideoMAC: Video Masked Autoencoders Meet ConvNets

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.588528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.588528Z digest=sha256:857195ce72c71547f3427670136830fed202490e329f5d2f659008119e002eb5

Observation d9e113b1-ea31-4de7-990f-c0216fbf308e · outbound

This paper cites Hauptmann, Ming-Hsuan Yang, Yuan Hao, Irfan Essa, and Lu Jiang.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Hauptmann, Ming-Hsuan Yang, Yuan Hao, Irfan Essa, and Lu Jiang

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.592709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.592709Z digest=sha256:b807e333305493f6a330aee40458d7fa75667ba6981293c74c5523d3747661c9

Observation e711213b-20c4-45f5-9782-039ff7bb9352 · outbound

This paper cites MAGVLT: Masked Generative Vision-and-Language Transformer.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation MAGVLT: Masked Generative Vision-and-Language Transformer

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.596657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.596657Z digest=sha256:546419debd3313763d131f3d84411273e7b04d0684825a3bcbbb043dc0b5ace3

Observation f54feb84-e2f7-4198-a32d-a5ccf71850a6 · outbound

This paper cites Multi-generator generative adversarial nets, 2017.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Multi-generator generative adversarial nets, 2017

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.601123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.601123Z digest=sha256:a8ac44639161cc11a1a7a32307f8c4a23d0846c32424f9f56de86cfa89ab53fa

Observation 672f3283-3b57-4842-b154-c837708ea08f · outbound

This paper cites Gomez, Łukasz Kaiser, and Illia Polosukhin.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Gomez, Łukasz Kaiser, and Illia Polosukhin

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.605272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.605272Z digest=sha256:ca53b922ec1414384a31644a07b3141801c051737696a50a36de9dc2b4e6de46

Observation ab67dac8-77ec-41ef-8cba-0fd9cccc342f · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation An image is worth 16x16 words: Transformers for image recognition at scale

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.609789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.609789Z digest=sha256:e2b9f81c5a478035a7bb669abfb5ee4529a89c05ad3a07847b617b538da7d3d4

Observation 674be2a5-a75d-4ef7-97e0-36899ebd4a5e · outbound

This paper cites Zero-shot text-to-image generation.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Zero-shot text-to-image generation

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.614593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.614593Z digest=sha256:4ba424a339df8f6b9386ae156849609926caf237fdeddb39c1c65d860811546d

Observation 24476f66-b907-4166-8f91-e195258a65c5 · outbound

This paper cites Discrete variational autoencoders, 2017.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Discrete variational autoencoders, 2017

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.618550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.618550Z digest=sha256:f2321fd652647cca61109a0ec86bd5fc656f4863f5cd75c8b59506a1f7378059

Observation 8f4e5b6e-8a70-4976-90e8-aebd69aac744 · outbound

This paper cites Cogview: mastering text-to-image generation via transformers.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Cogview: mastering text-to-image generation via transformers

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.622499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.622499Z digest=sha256:30dfab2e3e448b177aa9ac4235860206f6fcba4fcf65194b341e71611985fa82

Observation 0759d869-67bf-4a33-b9bb-78a8a21537b0 · outbound

This paper cites Cogview2: faster and better text-to-image generation via hierarchical transformers.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Cogview2: faster and better text-to-image generation via hierarchical transformers

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.627068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.627068Z digest=sha256:54d9982b57384e4ea8b12e900be0824791c0f552e443083d08c3bb414cced46a

Observation 1cd61a27-7f57-4bc3-9689-2c26826f016f · outbound

This paper cites ViViT: A Video Vision Transformer.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation ViViT: A Video Vision Transformer

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.630981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.630981Z digest=sha256:b15bbfece51a339f9fc3b27c7f6d97cabc112ba3674820b1247eaab912e5a9d4

Observation bf455fbe-2f4b-43ad-bb6e-d674fed3e654 · outbound

This paper cites Phenaki: Variable length video generation from open domain textual descriptions.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Phenaki: Variable length video generation from open domain textual descriptions

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.634625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.634625Z digest=sha256:15d8dfbbe94abbcc9f639c61b7805f567eb77498ed883e6ce93d66b4f6acb6a1

Observation 98e6907e-fa47-424b-8263-76092dc5c8ce · outbound

This paper cites an unresolved cited work.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Unresolved cited work

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.639033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.639033Z digest=sha256:a9b279c8446c2b7a4b588e43f6832379e31ce62f83ed0c5576962b8ed1be3a5e

Observation d51c6616-a625-465e-a36b-9ee835221f0b · outbound

This paper cites Compositional 3d-aware video generation with llm director.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Compositional 3d-aware video generation with llm director

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.646492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.646492Z digest=sha256:e3b6c73024b4810f7539b20bfc5c189f532b5bb476030abd335eda38ce394773

Observation ce7d04d6-70c4-4cfe-b6b6-fba8dafb62be · outbound

This paper cites an unresolved cited work.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Unresolved cited work

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.651064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.651064Z digest=sha256:c8830974f308fcceb62de51ce94a9e18455484d5a991008ef405c4474323bd54

Observation c0624193-859c-441a-9360-7b327783af10 · outbound

This paper cites Vlogger: Make your dream a vlog.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Vlogger: Make your dream a vlog

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.659432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.659432Z digest=sha256:c1874639ef058165c98ff5f84aa0ebb625c303eed6c988b0fb9a971f52b6b3f8

Observation 28464bab-cacd-42a0-8b8e-6a36d6b8dfba · outbound

This paper cites Weiss, Niru Maheswaranathan, and Surya Ganguli.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Weiss, Niru Maheswaranathan, and Surya Ganguli

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.663326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.663326Z digest=sha256:8a66bd1fdf1e0a95abcdecba38c840c937c0c926aa91471c6d8284b97c64481b

Observation 809e94af-516a-42df-8fa9-9cda0b60434d · outbound

This paper cites Denoising diffusion probabilistic models.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Denoising diffusion probabilistic models

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.667295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.667295Z digest=sha256:58a64fe5dc94819bbf3f6d95d32e834b8ea1ecf0213e190dee806463862bb6ad

Observation 9c6abda2-81d8-4615-9f4a-5b466b82d892 · outbound

This paper cites Generative modeling by estimating gradients of the data distribution, 2020.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Generative modeling by estimating gradients of the data distribution, 2020

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.671215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.671215Z digest=sha256:216b8dccd0ad973918fd7cc154386176952ef08d579992be0ef5e8dacd067867

Observation c0612117-642a-4c76-892f-5cf5655b624e · outbound

This paper cites Bert: Pre-training of deep bidirectional transformers for language understanding, 2019.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Bert: Pre-training of deep bidirectional transformers for language understanding, 2019

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.674802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.674802Z digest=sha256:8d73c34a8031e0d2ce263c16301610b92cf514bd6f22fdbca401ba76d02e3b79

Observation 980af4e7-9f53-4960-96ba-144781e41d43 · outbound

This paper cites an unresolved cited work.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Unresolved cited work

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.678399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.678399Z digest=sha256:b0ceabd26366ca3d913065950cf482322e47adb8d4add3b0ecd8a5909dfb37bc

Observation 26429073-5858-480c-b75f-2363e69b35e5 · outbound

This paper cites Video generation models as world simulators.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Video generation models as world simulators

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.682202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.682202Z digest=sha256:c155f1b6c2ccd46e021866fa3c6f8aeb673b507d07c5b55cad21d3c24f3a8cdc

Observation d54f6a1c-577c-4c6b-9ab5-9036e18a0c1c · outbound

This paper cites Scalable diffusion models with transformers.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Scalable diffusion models with transformers

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.685861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.685861Z digest=sha256:a1c8f07c624116bfb4cc85a28fec50489e7b5b270d2afb058010a2adaa80e3fa

Observation 228990ab-044b-4fa8-948d-bde650ce43ae · outbound

This paper cites Cogview: Mastering text-to-image generation via transformers.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Cogview: Mastering text-to-image generation via transformers

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.689203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.689203Z digest=sha256:6253b525f5b9c57da1203c0d3b9488976e5e6c0a08ac7deac0f656125360b413

Observation d1716863-5645-48f7-9882-90e13cd091ff · outbound

This paper cites Videotetris: Towards compositional text-to-video generation.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Videotetris: Towards compositional text-to-video generation

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.692629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.692629Z digest=sha256:1ac002445bc57693daa9fa4092593b708abba9dea50e493d01e5065247cc7f07

Observation b4aa16af-7992-453a-8999-5dff9ab53953 · outbound

This paper cites Cogview2: Faster and better text-to-image generation via hierarchical transformers.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Cogview2: Faster and better text-to-image generation via hierarchical transformers

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.696025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.696025Z digest=sha256:8bb0c5db48af9506e6e7307b76c8ba465f1189009e1944c27ed6d8f7685894fb

Observation a992bdac-de6e-4506-a688-280629395fdf · outbound

This paper cites NUWA-infinity: Autoregressive over autoregressive generation for infinite visual synthesis.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation NUWA-infinity: Autoregressive over autoregressive generation for infinite visual synthesis

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.699925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.699925Z digest=sha256:69b87893b83cff8084e9cde60cb366981515ccb180435b0825c2f6d1560cf3f5

Observation e0b12db0-ee9f-4511-b94f-0b7e592b278c · outbound

This paper cites Ross, Bryan Seybold, and Lu Jiang.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Ross, Bryan Seybold, and Lu Jiang

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.703594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.703594Z digest=sha256:beb2027697983fbdcb829584d3d87f07c1a355a3a20cf234384b4e7c9e2d194a

Observation 78620d1d-b30a-4889-8c72-12a49b0d737f · outbound

This paper cites ART •V: Auto-Regressive Text-to-Video Generation with Diffusion Models.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation ART •V: Auto-Regressive Text-to-Video Generation with Diffusion Models

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.707116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.707116Z digest=sha256:9f2097238a120fc48529a9f1380c648d95c3db87f7b9b33948db63b7a755fd06

Observation 6571763f-b74d-4894-8b49-14ac584d540d · outbound

This paper cites Grid diffusion models for text-to-video generation.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Grid diffusion models for text-to-video generation

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.711430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.711430Z digest=sha256:21500b6edf91f1a8266e7d90dbc8a975a4c1dfd0415172182bd4fe997450479a

Observation b86fa8d6-a11d-4e74-9d81-a7bb7737db67 · outbound

This paper cites Arlon: Boosting diffusion transformers with autoregressive models for long video generation, 2025.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Arlon: Boosting diffusion transformers with autoregressive models for long video generation, 2025

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.720203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.720203Z digest=sha256:acc64bfe8a53a60282fb9d23b492474767789aad44e96305db099bd205cd588c

Observation bd6c911e-4094-40f0-9bb3-02dfd32ce945 · outbound

This paper cites Stylegan-v: A continuous video generator with the price, image quality and perks of stylegan2.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Stylegan-v: A continuous video generator with the price, image quality and perks of stylegan2

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.724027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.724027Z digest=sha256:545b3d660cc82725e87ccdb69a5a1f42206b7c5aebac5d94ec8678154c52f9da

Observation 5f8b9b6d-ee6a-44b6-9fb2-c37d3c6a3108 · outbound

This paper cites an unresolved cited work.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Unresolved cited work

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.728231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.728231Z digest=sha256:8703d8f668c59093dfbfeae3f68983c872df6b5060971cef98ae83654d435beb

Observation 26cdc14d-8f9b-4186-8e3d-eca183b36d4a · outbound

This paper cites Srinivasan, Matthew Tancik, Jonathan T.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Srinivasan, Matthew Tancik, Jonathan T

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.731871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.731871Z digest=sha256:9cfbb213190cb008fdfaccaef9e6c5843bad9bdfccd8d48f43c244907e7aa991

Observation 1cfdf55d-3754-499c-a59d-a0dcff3e2ac7 · outbound

This paper cites Video probabilistic diffusion models in projected latent space.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Video probabilistic diffusion models in projected latent space

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.735767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.735767Z digest=sha256:8bffed3d68ad902f0fd018efd59d1f53f9d6da6235dc782bfce071abc08fbedd

Observation ed4713e1-e138-4a54-a0d6-a55c607b859a · outbound

This paper cites Towards end-to-end generative modeling of long videos with memory-efficient bidirectional transformers.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Towards end-to-end generative modeling of long videos with memory-efficient bidirectional transformers

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.739757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.739757Z digest=sha256:d240e792b60852c1166d6b80b51804ec6fd2a802ab74375fb168d230d8ee0b1c

Observation 42503dbb-4772-4626-87d6-e3c6cdc03858 · outbound

This paper cites Streamingt2v: Consistent, dynamic, and extendable long video generation from text.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Streamingt2v: Consistent, dynamic, and extendable long video generation from text

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.744775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.744775Z digest=sha256:a5cd72642080f58a8e610312cdcc59085fbe333412db03ea92c0a242a72ae400

Observation ce5f2492-ebb5-420a-8b07-136f55fc2c05 · outbound

This paper cites Vid-gpt: Introducing gpt-style autoregressive generation in video diffusion models, 2024.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Vid-gpt: Introducing gpt-style autoregressive generation in video diffusion models, 2024

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.748917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.748917Z digest=sha256:421a76e8bb90c8fdb4500779615563152c2efaa2dad81be7713b8221c4aa7b92

Observation faec5de8-142b-4706-9525-e9461a8ae33e · outbound

This paper cites Flexifilm: Long video generation with flexible conditions, 2024.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Flexifilm: Long video generation with flexible conditions, 2024

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.753119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.753119Z digest=sha256:9a3eb14d3c34b0fbe2d5122c3169e60e8ed6ed899d4a5f59c07c6314fda99266

Observation 0a691872-05e8-4712-b187-4c5bde993731 · outbound

This paper cites Mora: Enabling generalist video generation via a multi-agent framework, 2024.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Mora: Enabling generalist video generation via a multi-agent framework, 2024

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.756938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.756938Z digest=sha256:2fa3f1298dffc48ab52e437ea5e425f472c9578fef22da39382518b6d75e5f7e

Observation abd7f13e-d49e-4ebf-8995-fbaa27bad587 · outbound

This paper cites Vidgen: Long-form text-to-video generation with temporal, narrative and visual consistency for high quality story-visualisation tasks.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Vidgen: Long-form text-to-video generation with temporal, narrative and visual consistency for high quality story-visualisation tasks

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.760816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.760816Z digest=sha256:d23969d29a0a15564e222f305737653a1b32eb54d3911ffe71cae89a0306116f

Observation 3fdc9612-8db6-496e-8e11-3ea6f1af96d3 · outbound

This paper cites Bissyand, and Saad Ezzini.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Bissyand, and Saad Ezzini

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.764885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.764885Z digest=sha256:8a5bdb44139d7c98c221eff5fbf2c2783d51bef41556e856f5e2c5c0cf60e41e

Observation 3df02390-71d4-41d1-8fb6-fe2e817cad53 · outbound

This paper cites Kubrick: Multimodal agent collaborations for synthetic video generation, 2024.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Kubrick: Multimodal agent collaborations for synthetic video generation, 2024

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.768889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.768889Z digest=sha256:4fe5af57d753e702ded549aa6508cd7878d1a34af4a0821105f1c42ccc3abd12

Observation 6efd1475-3e4f-4d52-9704-cea6801df19d · outbound

This paper cites Modelscope text-to-video technical report, 2023.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Modelscope text-to-video technical report, 2023

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.772557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.772557Z digest=sha256:409eba9e74c4c577177892160cd184ad55ca794c10109055dd52102dc54f63df

Observation d1255750-54c6-49cf-92e6-f11883e31db6 · outbound

This paper cites Free-bloom: zero-shot text-to-video generator with llm director and ldm animator.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Free-bloom: zero-shot text-to-video generator with llm director and ldm animator

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.776147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.776147Z digest=sha256:5967128c7aa1f1d1a4d9f0655c8129a5e9741249d2aaf7a449cd953b40479739

Observation e012fc7e-558b-4ad8-9fe4-3f393f872eaf · outbound

This paper cites Denoising diffusion implicit models.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Denoising diffusion implicit models

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.779257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.779257Z digest=sha256:323b24ec30bb8b4cbf1f34d237642a6a1472e77942fd8bfddc9da464c8a82d0e

Observation 62041886-b668-4390-bd61-c099e3b420b8 · outbound

This paper cites LLM-grounded video diffusion models.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation LLM-grounded video diffusion models

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.783124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.783124Z digest=sha256:30763e7232672bb30e2bc63b11ef763edf7eb43c9e2b2c22c315c6e578a8e0e1

Observation d40aef87-2cf7-4a7b-bf98-e75a1218d6a6 · outbound

This paper cites Mevg: Multi-event video generation with text-to-video models.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Mevg: Multi-event video generation with text-to-video models

Reference 101

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.787359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.787359Z digest=sha256:f32a8cefbe256e6e47f2b877fb496f183649754ac5531f44e67504a9d7940195

Observation 91f84484-2120-48d6-888c-7bca4ee84596 · outbound

This paper cites an unresolved cited work.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Unresolved cited work

Reference 102

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.791354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.791354Z digest=sha256:3bdb1beec053b407ba415d8b5ac14997b873e8f83c71338029e8ece5eb430ce7

Observation e539ca34-483d-4f48-b056-be507e070be9 · outbound

This paper cites Videomerge: Towards training-free long video generation, 2025.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Videomerge: Towards training-free long video generation, 2025

Reference 103

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.796952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.796952Z digest=sha256:c4de2eb20f23a52797324dd8d6ac056b6b3ddb2a0fde7c90e0250fa7e2295ae6

Observation ede348e8-a4a3-42a1-8b65-f391622a7597 · outbound

This paper cites Mevg: Multi-event video generation with text-to-video models, 2024.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Mevg: Multi-event video generation with text-to-video models, 2024

Reference 104

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:14.801702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:14.801702Z digest=sha256:23268c3122d9f3587e57cd20852565ee062aa654836c1857989cd2b7ddc92188

Pith citing papers

Observation ed620966-a4cf-4a2b-a333-065cb5f1806b · inbound

Enhancing Train-Free Infinite-Frame Generation for Consistent Long Videos cites this paper.

Enhancing Train-Free Infinite-Frame Generation for Consistent Long Videos Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:58:13.837416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T10:56:16.054496Z digest=sha256:ab83b2a0d24d1a94d1409833e363240ad226361ff0a6591a76a12281fad0ed31