Pith. sign in

Paper Citation Record · LEDGER

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality

As of 14 August 2026, this Paper Citation Record lists 100 of 110 outbound references and 1 inbound Pith citation observation for arXiv:2507.07202.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.07202 v1

Coverage vector

measured 100 of 110 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:51:31.757497Z

measured 101 of 101 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T13:41:42.111186Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 110 outbound references displayed

  • verified exact1
  • verified fuzzy31
  • unresolved68
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b94a691a-eb77-415d-84b9-0688d3d65428 · outbound

This paper cites Frozen in time: A joint video and image encoder for end-to-end retrieval.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Frozen in time: A joint video and image encoder for end-to-end retrieval

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:24.002704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:24.002704Z digest=sha256:05ed725433dd65cf2da8c6b299a43f251c3ceea381308feed0f5318983a084a0

Observation e31a678e-b35e-4cab-af6a-b1a2482c01da · outbound

This paper cites Lumiere: A Space-Time Diffusion Model for Video Generation.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Lumiere: A Space-Time Diffusion Model for Video Generation

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:24.068306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:24.068306Z digest=sha256:94de007d3dd2997be76bfa088bf5666932cd98b4e06b4078d0be4e2168f7d527

Observation a3e5038a-6c1b-4730-8a7d-dee5bdec4704 · outbound

This paper cites Is space-time attention all you need for video understanding? In Proceedings of the 38th International Conference on Ma- chine Learning (ICML), pages 813–824.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Is space-time attention all you need for video understanding? In Proceedings of the 38th International Conference on Ma- chine Learning (ICML), pages 813–824

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:24.169776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:24.169776Z digest=sha256:6ef54a245847f2c51e9c1d26306c160619339ca27369decc1816fb235fead07e

Observation e8e5648f-1b7b-4433-85bb-64a28f7f72c9 · outbound

This paper cites Generating Long Videos of Dynamic Scenes.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Generating Long Videos of Dynamic Scenes

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:24.241885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:24.241885Z digest=sha256:b84247bb1fad94ee4d9b55f932e441dff02f4614c1f86db65ebb1b4cb46797a5

Observation 157891b2-3d5e-4a19-b32e-a82f76fdbc3b · outbound

This paper cites an unresolved cited work.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:24.367980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:24.367980Z digest=sha256:a66300e69f84ceaf01082d71e651474dfd9e508c188c5f2d3e3642d08839616d

Observation a939f0a1-a0a5-4ace-ab90-69434ac45f4b · outbound

This paper cites SkyReels-V2: Infinite-length Film Generative Model.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality SkyReels-V2: Infinite-length Film Generative Model

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:24.412898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:24.412898Z digest=sha256:87abe47b76ae2f89c34742fb6ecc8eb68714f4dfc74d12db413bfe8381b42938

Observation 5b10905a-0c5e-4e86-8436-51d7b60dede4 · outbound

This paper cites Pixart- α: Fast train- ing of diffusion transformer for photorealistic text-to-image synthesis, 2023.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Pixart- α: Fast train- ing of diffusion transformer for photorealistic text-to-image synthesis, 2023

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:24.497048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:24.497048Z digest=sha256:0eedc67a5110535cd039502edcf3f76486ddb448cbf0a022ca1f4a6477bd461c

Observation 9b81b3cc-7a17-45a0-be57-63b4dca53bfd · outbound

This paper cites Panda-70m: Captioning 70m videos with multiple cross-modality teachers.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Panda-70m: Captioning 70m videos with multiple cross-modality teachers

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:24.568977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:24.568977Z digest=sha256:fe9165ed567d15c74ae528aa82f64764b076caf6ec9099ecd57d422ad4e2d5ce

Observation 4ae97d87-06df-42d8-83df-1ca90836f6c6 · outbound

This paper cites Panda-70M: Captioning 70M Videos with Multiple Cross-Modality Teachers.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Panda-70M: Captioning 70M Videos with Multiple Cross-Modality Teachers

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:24.668995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:24.668995Z digest=sha256:54b14479319581456f65a5790bd9c807645a696f92a4a7944822cadfa43379d3

Observation da6cfe17-97e9-49b9-9301-a12ff7e27cb2 · outbound

This paper cites Multi-subject open-set personalization in video generation.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Multi-subject open-set personalization in video generation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:24.735298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:24.735298Z digest=sha256:df4b7aeea20ed73eba46a971d1ab0c7111260ed56e093e09de3642438d7f3c29

Observation 1b6bbcf0-5e04-4977-adf3-2be0877510d1 · outbound

This paper cites Seine: Short-to-long video diffusion model for generative transition and prediction.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Seine: Short-to-long video diffusion model for generative transition and prediction

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:24.807647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:24.807647Z digest=sha256:ffb496bdc8501b2473ca70d561b6571b5b383e5b12170f5d448d646d6e048e29

Observation 79609944-7dee-4bd5-aeee-e6dcfefb663d · outbound

This paper cites HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:24.893203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:24.893203Z digest=sha256:56ff69226037566b79972a1e5bdc89d2bf7367a62914a58a02910e80b3ca2e45

Observation 97c88819-f091-44d0-80e6-7f03cec56924 · outbound

This paper cites Unimax: Fairer and more effective language sampling for large-scale multilingual pretraining.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Unimax: Fairer and more effective language sampling for large-scale multilingual pretraining

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:24.996882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:24.996882Z digest=sha256:93fcff3c719b25977eb146f5eccc9c8031683c040d6cf32335d636ea99226018

Observation dec23150-c339-4940-89cf-4e4bcef08eec · outbound

This paper cites Zhao, Yanping Huang, An- drew M.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Zhao, Yanping Huang, An- drew M

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:25.096771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:25.096771Z digest=sha256:749640bc026fb1521cc073ee973b33005d97d4607ff702f166e3166c049b21d3

Observation 8c3ace60-0f84-4185-97fc-77933b4cf28c · outbound

This paper cites Efficient video prediction via sparsely conditioned flow matching.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Efficient video prediction via sparsely conditioned flow matching

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:25.165135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:25.165135Z digest=sha256:7b1ff52be942ee7c45776c55f75dcddb4171006999fc3371e8d97ba3d0417bbd

Observation 09d01f88-6286-4db2-a306-09979a1da62b · outbound

This paper cites an unresolved cited work.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:25.262633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:25.262633Z digest=sha256:5e6a6b53b958be97b56a9c1d23410eb8e53551694859b7829f8990cc3511b819

Observation 467800da-d809-462b-929f-532a496a77f5 · outbound

This paper cites Scaling rectified flow transformers for high-resolution image synthesis.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Scaling rectified flow transformers for high-resolution image synthesis

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:25.343816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:25.343816Z digest=sha256:648e9e90f175a005166d409aed3a582e8d82666b203bf1ecc3d37fe18c0f67c8

Observation 01543e06-f156-483f-9450-0f16dc4bce26 · outbound

This paper cites Ca2-vdm: Efficient autoregres- sive video diffusion model with causal generation and cache sharing.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Ca2-vdm: Efficient autoregres- sive video diffusion model with causal generation and cache sharing

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:25.427361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:25.427361Z digest=sha256:d301bed43d33a8f1d02dcd4dea2ac9b6ec9d29697593fbf6fbe23a6a98e9ae8a

Observation 8c26b3bf-c206-4a97-b9b7-d19f0d5f3a80 · outbound

This paper cites Berg, Arash Vah- dat, Alexei A.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Berg, Arash Vah- dat, Alexei A

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:25.487912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:25.487912Z digest=sha256:c13860c0aff277357b40ffa13e9cad803629c4772b13ecc456038fb902ea4ee6

Observation 12a2a0a7-d6a8-42dc-9b9e-ce9c5b091005 · outbound

This paper cites Seedance 1.0: Exploring the Boundaries of Video Generation Models.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Seedance 1.0: Exploring the Boundaries of Video Generation Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:25.538623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:25.538623Z digest=sha256:812859fd48cb5082443b6351453751d66638952faa20cfcf70e11a18a92c2887

Observation 0f4120ff-7f4d-4994-96ff-2b597a561a2c · outbound

This paper cites Zico Kolter, and Kaiming He.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Zico Kolter, and Kaiming He

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:25.587961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:25.587961Z digest=sha256:a80e204db1d72dffa27b77bf78bbc936718f2697d59328455942f3fe15f846a8

Observation bd2b6afa-ceff-42da-9e82-3472f370d1d2 · outbound

This paper cites Veo 3: Neural video generation with native audio.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Veo 3: Neural video generation with native audio

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:25.668222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:25.668222Z digest=sha256:696109231e18c184f495c5e4cf20b72f916474596435f8c31bda99b2189327e3

Observation d5f63aa6-20b5-4b09-9f41-02afe4000d2e · outbound

This paper cites Animatediff: Animate your personalized text-to-image models.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Animatediff: Animate your personalized text-to-image models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:25.762753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:25.762753Z digest=sha256:4c4eee37db582018d667838e234508396b94231d78cf972ac8a4d818ac0ddd4b

Observation 17d3516f-6d2e-4467-877e-ccfcaba77b2a · outbound

This paper cites LTX-Video: Realtime Video Latent Diffusion.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality LTX-Video: Realtime Video Latent Diffusion

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:25.825780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:25.825780Z digest=sha256:03d1c6567e8e6c2cbb7e694d37d8bdb491e74ebb0b34a69e09118e5e89825fd1

Observation ab2eda46-0dcb-4226-a0ed-2acf84203285 · outbound

This paper cites Clipscore: A reference-free eval- uation metric for image captioning.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Clipscore: A reference-free eval- uation metric for image captioning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:25.897788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:25.897788Z digest=sha256:0fb051d4ba5a4b9695a7e09c25054dca50f777378901fd50f9433c47985b0bb7

Observation a8ecd40b-fb18-4ba6-89da-10c64e9d3d41 · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilib- rium.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Gans trained by a two time-scale update rule converge to a local nash equilib- rium

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:25.982396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:25.982396Z digest=sha256:36ded723e3253e265ac4755feba0ee91c6ae245700135cb4d66a56fd626819ba

Observation 0d792ce2-6701-4d6a-885e-5cace2211c75 · outbound

This paper cites Denoising dif- fusion probabilistic models.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Denoising dif- fusion probabilistic models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:26.076771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:26.076771Z digest=sha256:9b24ca694e2b029f5e4ec263337d0907bf1ea8e13a76608e9aac8914c68eefad

Observation 724d2876-d37d-46f0-b8d8-783c86ffe333 · outbound

This paper cites an unresolved cited work.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:26.143548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:26.143548Z digest=sha256:a80cfc628c9a202917414a5f57930b36a6b0ea5cca4863b6eb81aeb4bd744ea0

Observation a2cc2f49-fcca-40a1-acd1-754b6f86eb82 · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:26.211695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:26.211695Z digest=sha256:0d07ba3420fe66e3af19c4209fd03bae76acf67a32d5d865118ef786ac5a304d

Observation b0fe2f5a-e804-4281-9262-52d1888ca8e0 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality LoRA: Low-Rank Adaptation of Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:26.294142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:26.294142Z digest=sha256:3cd98eab8790457ca8d0010e8aaaf8ba0fc8dfe67feb808b39a8cc62662dcc5a

Observation b641c452-348b-4b38-88ec-e82f018c0546 · outbound

This paper cites HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:26.356204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:26.356204Z digest=sha256:e41b68c3e57e55679fb3ac104bafd7d2af09305aa5cd6d8346f9b08cd06e8099

Observation ba01330a-351c-4cf4-8100-3014152fd200 · outbound

This paper cites Step-Video-TI2V Technical Report: A State-of-the-Art Text-Driven Image-to-Video Generation Model.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Step-Video-TI2V Technical Report: A State-of-the-Art Text-Driven Image-to-Video Generation Model

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:26.438943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:26.438943Z digest=sha256:ad4befab000c2b46849a2127ce7104b6937ae6741b0bdd3064c4bb56790d2351

Observation 5b2e4943-7699-4f24-afe0-f875b9a9d6f2 · outbound

This paper cites ConceptMaster: Multi-Concept Video Customization on Diffusion Transformer Models Without Test-Time Tuning.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality ConceptMaster: Multi-Concept Video Customization on Diffusion Transformer Models Without Test-Time Tuning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:26.486990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:26.486990Z digest=sha256:299c477a6e4876991c136fc5fba450ce7ee05850a29ae385309eea535fffef00

Observation c90a055f-9636-4d81-99f3-beee7dc447ae · outbound

This paper cites Vbench: Comprehensive benchmark suite for video generative mod- els.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Vbench: Comprehensive benchmark suite for video generative mod- els

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:26.585805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:26.585805Z digest=sha256:4e9b36f4264a85a8dd882c7640231e2ae6865e3bffe6f39737c2d753316b31aa

Observation 003ec7ca-ed6a-4fdf-8a0c-61d3aa62ac68 · outbound

This paper cites Pika 1.5: Realistic scene and motion synthesis,.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Pika 1.5: Realistic scene and motion synthesis,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:26.659409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:26.659409Z digest=sha256:28e9184e725c78b2d5945b0259528d761f7a81d4b19bf05728d453da0b590bf3

Observation fc8dc39c-f6fd-4297-bc0a-4401d128302d · outbound

This paper cites Sim2real: Synthetic video data for vision-based robotic manipulation learning.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Sim2real: Synthetic video data for vision-based robotic manipulation learning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:26.858784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:26.858784Z digest=sha256:c59b85dc948b5b5a6dea580c0ec94c7d0eeb63f49205610ec30f63754654c618

Observation d0724895-c789-4cf7-9ec5-6f9c48184a28 · outbound

This paper cites VACE: All-in-One Video Creation and Editing.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality VACE: All-in-One Video Creation and Editing

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:26.935841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:26.935841Z digest=sha256:8b8ab372543fa5508a48545035b51c550d0f87c4fb5e7a1e9464d0a161fa6df6

Observation a8063218-9d9c-4d40-adb9-5d7ac640a87a · outbound

This paper cites Pyramidal flow matching for efficient video generative modeling.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Pyramidal flow matching for efficient video generative modeling

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:26.973195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:26.973195Z digest=sha256:c59c0a8890f04a08a6e1adf794e3f2464c1cb028a3fd772f3af01df81ca086ce

Observation 332aff63-e3e3-4168-9dfb-fdc6d5a7d6c2 · outbound

This paper cites Miradata: A large-scale video dataset with long du- rations and structured captions, 2024.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Miradata: A large-scale video dataset with long du- rations and structured captions, 2024

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:27.028318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:27.028318Z digest=sha256:36fd148399b6a361e1ee5070f76a9530c7e6d29545184c38a8f6cdc3953069ae

Observation f890aa14-7922-4460-ab83-fc509f308c55 · outbound

This paper cites Miradata: A large-scale video dataset with long 10 durations and structured captions.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Miradata: A large-scale video dataset with long 10 durations and structured captions

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:27.111928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:27.111928Z digest=sha256:d7650622f0047688f4b851ec8db79dd4b2fc7088503510d9df9339f77eebedba

Observation c7f7f318-975b-44cb-8d4a-372deeb3bf44 · outbound

This paper cites Cinediff: Diffusion models for cinematic video synthesis.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Cinediff: Diffusion models for cinematic video synthesis

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:27.149466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:27.149466Z digest=sha256:731a3a52a327d81cd2298793b16cd1723aa1ba5facd0504d67bf52b90d03e6bd

Observation c7410d0e-6651-492e-90a3-7e67b9c510fe · outbound

This paper cites HunyuanVideo: A Systematic Framework For Large Video Generative Models.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:27.228352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:27.228352Z digest=sha256:42917a9345fc0e1d131d0329e7e35136425940d302cbce496c14e752366adcf0

Observation 99aad170-c98e-46ec-af89-91034147206d · outbound

This paper cites Minimax hailuo: Scalable multi- subject video generation, 2025.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Minimax hailuo: Scalable multi- subject video generation, 2025

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:27.336902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:27.336902Z digest=sha256:a3f92f2919f5b9f390a04eb391adcd31d8db72cc4c39ed5e44135b43f629d1d3

Observation 017b6017-8503-4488-b8fb-9c2f7364b022 · outbound

This paper cites Openhumanvid: A large-scale high- quality dataset for enhancing human-centric video genera- tion.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Openhumanvid: A large-scale high- quality dataset for enhancing human-centric video genera- tion

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:51:33.606674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T18:51:27.417606Z digest=sha256:626ea3188bfbc5da94d249f6a4198fc6ff6102a7cad9f75a66c590649fbbdac8

Observation 57e461e5-e01b-4fa5-b908-3544ad5da534 · outbound

This paper cites Open-Sora Plan: Open-Source Large Video Generation Model.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Open-Sora Plan: Open-Source Large Video Generation Model

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:27.464912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:27.464912Z digest=sha256:ba66d90eab9fd040aae9411ca9af1e3ccba5d4b4739498d864d4414fac0e01aa

Observation 75a20dc8-5800-4d8b-9cbc-dcbd3b4c8144 · outbound

This paper cites VideoDirectorGPT: Consistent Multi-scene Video Generation via LLM-Guided Planning.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality VideoDirectorGPT: Consistent Multi-scene Video Generation via LLM-Guided Planning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:27.528264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:27.528264Z digest=sha256:374e020eed15ef1262163d4390f7f0d59be2f2caab4497750080915f2348c1b0

Observation 1be0cb62-7ebe-440e-91cb-cf0afe8f4a05 · outbound

This paper cites an unresolved cited work.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:51:33.597009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T18:51:27.591944Z digest=sha256:6e7adb600fb3d8e4eca5161d8a57645323286a26b97dce4aeb2096a09bb36969

Observation 78f0d332-7e44-4678-bd3e-d7f2777dd5a5 · outbound

This paper cites Phantom: Subject-consistent video generation via cross-modal alignment.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Phantom: Subject-consistent video generation via cross-modal alignment

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:27.649430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:27.649430Z digest=sha256:d18d143c36f33b39a4b92587ad4a13d7d9cb87ead15c21428459c14f705b96c1

Observation a395f1b4-e549-42f3-a11f-044a5ac1f887 · outbound

This paper cites AutoStory: Generating Diverse Storytelling Images with Minimal Human Effort.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality AutoStory: Generating Diverse Storytelling Images with Minimal Human Effort

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:51:33.048674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T18:51:27.714173Z digest=sha256:818a1fbebeb6af3f93f6c874b3de2a9034ee0b750c5bef06c26e779fa0aad4f8

Observation 14c0c2e1-baad-4300-a957-4c7151215b45 · outbound

This paper cites Latte: Latent diffusion transformer for video generation.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Latte: Latent diffusion transformer for video generation

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:51:33.588141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T18:51:27.753527Z digest=sha256:8987c932c59b9bc5b9fd1add44b31b58b009b862356c7db4f9442dab88710bb3

Observation 424ebe18-0559-4982-bacb-f234c5cf7497 · outbound

This paper cites Gamegan: Video generation for atari games.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Gamegan: Video generation for atari games

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:51:33.578596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T18:51:27.811437Z digest=sha256:08e83dcdeb3735763eddffff72001fa654bba6439306de6d069a5e0e23ec93e3

Observation 780a0571-1608-4adf-a561-3341a31ab6e5 · outbound

This paper cites Sora: Openai’s text-to-video generator, 2024.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Sora: Openai’s text-to-video generator, 2024

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:51:33.568719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T18:51:27.857440Z digest=sha256:7016f8d6a9b302a88bfd2956265ca53b093075d78220c3a11511825a6df5e0f2

Observation a03c9693-0c04-4c64-af21-97ddf8c8aa87 · outbound

This paper cites Scalable diffusion mod- els with transformers.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Scalable diffusion mod- els with transformers

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:51:33.558545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T18:51:27.901975Z digest=sha256:86b8184975709ed51ca664f384b3160af573a35ddb98f50d093157ada41582e7

Observation d0a17abb-1372-4d11-91c2-955192748180 · outbound

This paper cites Open-Sora 2.0: Training a Commercial-Level Video Generation Model in $200k.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Open-Sora 2.0: Training a Commercial-Level Video Generation Model in $200k

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:27.953286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:27.953286Z digest=sha256:273fa4748217e80e4481ff7af7d0d0d7155c8ba9b673eeab925a6ce7e3a1cc0a

Observation 4c1f6efc-ac82-4ad9-b8ea-a7da7c53f5e8 · outbound

This paper cites Movie Gen: A Cast of Media Foundation Models.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Movie Gen: A Cast of Media Foundation Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:27.992227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:27.992227Z digest=sha256:6ff4c7d3d689ffe1554e60cef4c43f5c6716d8c95854563901c5a4894de7b96b

Observation c73196b0-e227-4152-aff1-2f0008a20497 · outbound

This paper cites Learning transferable visual models from natural language supervision.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Learning transferable visual models from natural language supervision

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:51:33.548601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T18:51:28.043194Z digest=sha256:cfd576a05493cbbfa8d2f35ee070095d0610c7a409e0fd84ee75be393351e3fd

Observation 43dd8388-8eef-4038-bb7c-43218e56275e · outbound

This paper cites an unresolved cited work.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:51:33.538990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T18:51:28.093926Z digest=sha256:33cffd6bf3d05c20aac7dd38451365ba596d0d1ea6cba90cdb812b34ead2df3b

Observation 193a288d-a51e-48db-9925-16bcbf086d0b · outbound

This paper cites Gener- ating diverse high-fidelity images with vq-vae-2.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Gener- ating diverse high-fidelity images with vq-vae-2

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:51:33.529569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T18:51:28.127742Z digest=sha256:5745d011a6636ac4d2856041ba91df061bcec2734cf62f215fa71ddd33a90cab

Observation 1923a6b5-f8d4-4bd4-b227-1f59a54c8e57 · outbound

This paper cites Runway gen-3: Advanced video syn- thesis platform, 2024.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Runway gen-3: Advanced video syn- thesis platform, 2024

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:51:33.520108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T18:51:28.154624Z digest=sha256:b68936e2846e4eefc711aa32f74dffbdb5732e0abdba3ddfd56f4cd357d5c936

Observation 629336ec-88d4-4118-bcce-b4a194e5ae67 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality High-resolution image synthesis with latent diffusion models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:28.210651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:28.210651Z digest=sha256:37129ba6ae498670dbaf3e0e3efa966b0e862c663c2c83cc6059c38a609697f2

Observation 2bdb081f-9b27-4e7c-a486-11961415d555 · outbound

This paper cites U- net: Convolutional networks for biomedical image segmen- tation.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality U- net: Convolutional networks for biomedical image segmen- tation

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:51:33.505284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T18:51:28.272230Z digest=sha256:dd34ac95b27ff7224c74ddea56bb041cad49bc900b1c8d8f66a17d3988329d14

Observation 2986ff15-fdb4-4378-a3d2-44752c23a84b · outbound

This paper cites Goodfellow, Wojciech Zaremba, Vicki Cheung, Alec Radford, and Xi Chen.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Goodfellow, Wojciech Zaremba, Vicki Cheung, Alec Radford, and Xi Chen

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:51:33.496424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T18:51:28.359682Z digest=sha256:fdacc09df09bdb567260548dce00b7249e15668b11e20a6f7cfced476a5ab92b

Observation 25d87521-01e9-43a9-b2c0-2a304a3eb5db · outbound

This paper cites Flashattention-3: Fast and accurate attention with asynchrony and low-precision,.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Flashattention-3: Fast and accurate attention with asynchrony and low-precision,

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:51:33.487809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T18:51:28.399654Z digest=sha256:a7fdf43acabd4ffdc777b1812b3a36e112e50a86a16e367620f55847c7aecfe8

Observation 837e33c7-e193-4212-9f6d-db048e6219c0 · outbound

This paper cites Videopoet: Large language models are zero-shot video generators.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Videopoet: Large language models are zero-shot video generators

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:51:33.478980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T18:51:28.461098Z digest=sha256:677618b3869cb4992471d182083407c2e39151b68090772ab632e038f35a9333

Observation fc06091d-be55-4325-92c2-5669ad12e32c · outbound

This paper cites Denois- ing diffusion implicit models.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Denois- ing diffusion implicit models

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:51:33.469672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T18:51:28.496431Z digest=sha256:eed2f164e02f731cf0928fac11cf99512ba350d4220aeb4f03ccfff541165ca0

Observation 594bba73-2e54-457b-8ad6-55eca9fd08dd · outbound

This paper cites Kling2.0: Proprietary high-fidelity video generation, 2024.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Kling2.0: Proprietary high-fidelity video generation, 2024

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:51:33.460582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T18:51:28.532366Z digest=sha256:d1720d7e07b2bf404dd6ab159a1d25a99b4fc255314cefd7227114a021bf845c

Observation 27325cd4-c1a1-4879-9699-5c1953e11ec8 · outbound

This paper cites MAGI-1: Autoregressive Video Generation at Scale.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality MAGI-1: Autoregressive Video Generation at Scale

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:28.578351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:28.578351Z digest=sha256:3f9e6aaf9da96a68ab7867e456e5539de09a604586db4d4085f168b302d8a40e

Observation afbce4be-0622-49fc-88fa-eb122facfc40 · outbound

This paper cites Mocogan: Decomposing motion and content for video generation.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Mocogan: Decomposing motion and content for video generation

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:51:33.451254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T18:51:28.649535Z digest=sha256:0063927852438244f478f8e093ce2c252f5c6f21a5cdbaba58563e0b146f82f7

Observation df8baf8e-f220-41ab-ab0f-bb296487f7a1 · outbound

This paper cites Towards Accurate Generative Models of Video: A New Metric & Challenges.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Towards Accurate Generative Models of Video: A New Metric & Challenges

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:28.755370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:28.755370Z digest=sha256:99e6f5c0fcbc45d3a586abd48460b9ca862061343597055254387ac14a3f58e4

Observation 345d4139-8d48-45b2-8d12-b8ff27894019 · outbound

This paper cites Neural discrete representation learn- ing.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Neural discrete representation learn- ing

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:51:33.441921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T18:51:28.845014Z digest=sha256:4c63f636cdbff97f7c15b4dca3ab00c6004fa0c65034d5a38b955ddcc412bcdc

Observation 37e02c48-704a-483f-ab1a-ed5581b25862 · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Wan: Open and Advanced Large-Scale Video Generative Models

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:28.893268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:28.893268Z digest=sha256:887c54a736193c83931a28c27c9c1148e1302a5a18df71f7e06f9ba0db375054

Observation c9402324-cd38-4aab-8a99-fb2ca61e9150 · outbound

This paper cites ModelScope Text-to-Video Technical Report.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality ModelScope Text-to-Video Technical Report

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:28.967399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:28.967399Z digest=sha256:87efe2723489098276ee081748e966d09fd9478a9248e42016540ce00ffe775c

Observation 55664aa2-5f58-4555-9c35-0de6339c5977 · outbound

This paper cites Text2video: Gener- ating educational videos from text.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Text2video: Gener- ating educational videos from text

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:51:33.433200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T18:51:29.028121Z digest=sha256:bda53346c4a99a32540de9d2ae9c89f81d993312057674e48a205a37ea490991

Observation 079cc167-5b3a-4d9f-89ed-0c0b57dbb5ff · outbound

This paper cites Koala- 36m: A large-scale video dataset improving consistency between fine-grained conditions and video content, 2024.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Koala- 36m: A large-scale video dataset improving consistency between fine-grained conditions and video content, 2024

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:51:33.424383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T18:51:29.134036Z digest=sha256:43dd8f73118b30367daa3af7e7c9aed99f97405943950a530c022c6984f33c50

Observation 3db225ab-3e21-479c-999b-2be81a2f815a · outbound

This paper cites Koala-36M: A Large-scale Video Dataset Improving Consistency between Fine-grained Conditions and Video Content.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Koala-36M: A Large-scale Video Dataset Improving Consistency between Fine-grained Conditions and Video Content

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:29.180801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:29.180801Z digest=sha256:58509bd24622efc168dbed4c3ca4db89f356d8d48196ca3ddd700ab3c18fcca8

Observation 2c21f850-8a82-478d-9044-85fec28bfd67 · outbound

This paper cites Diffuse and Disperse: Image Generation with Representation Regularization.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Diffuse and Disperse: Image Generation with Representation Regularization

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:29.240432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:29.240432Z digest=sha256:9e650ed0071b503f6db14269b7c457ccd634982d027b880f0da9a7fd0a67016b

Observation cfc808d4-8e26-44b4-83a7-5c2aba9e5c35 · outbound

This paper cites Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:29.322880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:29.322880Z digest=sha256:3abeb881e8821eba8de6f12759369211ae5725667c210f426efefebcc1616ff9

Observation 8196cb90-41fc-4e2d-af9f-03ceb81ba722 · outbound

This paper cites Videofactory: Swap attention in spatiotemporal diffusions for text-to-video gen- eration.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Videofactory: Swap attention in spatiotemporal diffusions for text-to-video gen- eration

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:51:33.415822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T18:51:29.387266Z digest=sha256:97ece18439b7bca33a0456cea006af6cc5fd6a5660047c637a60ac48acd3a88d

Observation f6648244-5260-4af2-b69d-1baa0b764ec2 · outbound

This paper cites Non-local neural networks.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Non-local neural networks

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:51:33.406460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T18:51:29.473010Z digest=sha256:fe713e998759c04228485835bfd1850a197dad89ea64c1a642a9abd7c573afab

Observation b1bca7bc-db88-4221-8c66-7be59d4e6f92 · outbound

This paper cites Yuille, Zicheng Liu, and Emad Barsoum.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Yuille, Zicheng Liu, and Emad Barsoum

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:29.564360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:29.564360Z digest=sha256:42a29bc90f5cb19ea23b6e701cd2304ce0ae6a942ed093e24b913c4aac1e57a0

Observation 08562ac5-5180-4707-9435-0fbf8578180c · outbound

This paper cites Loong: Generating Minute-level Long Videos with Autoregressive Language Models.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Loong: Generating Minute-level Long Videos with Autoregressive Language Models

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:29.715057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:29.715057Z digest=sha256:7d03385a779c2907f6f216344b1198c7ba3960f61ecb8249c455e6a586102fb8

Observation ec0cb88a-42e8-4d6a-bc3e-c70189248a6a · outbound

This paper cites Bovik, Hamid R.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Bovik, Hamid R

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:51:33.396397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T18:51:29.761047Z digest=sha256:aae9bf5089a6352b2914fa92c373b1989ae1489472fa79bb389bb0b381c090fe

Observation 39361e02-e3e2-40d7-a1ca-bb2daa1e6d9a · outbound

This paper cites Humanvid: Demystifying train- ing data for camera-controllable human image animation.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Humanvid: Demystifying train- ing data for camera-controllable human image animation

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:51:33.388182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T18:51:29.921360Z digest=sha256:83d7b8807f9ac65dd5c1414a0747de1ed293a16e683754bfa93e16ee55f9f386

Observation 9d980628-5c98-412d-a19f-c778cdc8fdca · outbound

This paper cites FreeFlux: Understanding and Exploiting Layer-Specific Roles in RoPE-Based MMDiT for Versatile Image Editing.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality FreeFlux: Understanding and Exploiting Layer-Specific Roles in RoPE-Based MMDiT for Versatile Image Editing

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:29.977844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:29.977844Z digest=sha256:562e2bfd20e7f1224aae4a1414c243dc449f570bf5f6d3ce13766199b2c3af9d

Observation d7efc319-0dfa-4f4a-b26b-a9d530746040 · outbound

This paper cites Panacea: Panoramic and controllable video generation for au- tonomous driving.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Panacea: Panoramic and controllable video generation for au- tonomous driving

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:51:33.379861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T18:51:30.077913Z digest=sha256:473d7d62961388ba727b6d467fcaa71b1a2ba3b02b87b89bfc13850b2fd5c374

Observation e1a52704-b453-4aa2-b5c1-63e31b1b411d · outbound

This paper cites MovieBench: A Hierarchical Movie Level Dataset for Long Video Generation.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality MovieBench: A Hierarchical Movie Level Dataset for Long Video Generation

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:30.200027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:30.200027Z digest=sha256:058c028d7d5bb87d604788f96ca0e0e7a23bf6c1c41494b4ff077a82fb873160

Observation 9cf082b0-8e2d-4d20-b1e9-8b58e686189b · outbound

This paper cites Automated Movie Generation via Multi-Agent CoT Planning.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Automated Movie Generation via Multi-Agent CoT Planning

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:30.310486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:30.310486Z digest=sha256:a00302d4d72490ba6e105e4f398006aa13b5271c3f26b890acb11cd5ea3c1c14

Observation 282a7c20-516f-4f2c-9ae3-0d1d9d22b5a3 · outbound

This paper cites Mind the time: Temporally- controlled multi-event video generation.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Mind the time: Temporally- controlled multi-event video generation

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:51:33.371071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T18:51:30.361446Z digest=sha256:e75acf2412d4165070b9989b1666442ce6e9a92b16fe2e370e93ec5f79347323

Observation bce22bbb-dd8e-43bd-b66a-05e822aad854 · outbound

This paper cites A survey on video diffusion models.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality A survey on video diffusion models

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:51:33.362057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T18:51:30.462880Z digest=sha256:7745af765905f2d38d5cdf0169027c1a503b450936659cf71c64af61fe519bc8

Observation 285c0893-c957-41e1-822a-9b8cbd99868b · outbound

This paper cites Surgi- cal video synthesis using generative models for procedure training.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Surgi- cal video synthesis using generative models for procedure training

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:51:33.353216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T18:51:30.537696Z digest=sha256:259ba38317a9170e974757013aea31036b5d95ba37f2873c9c0887aed13d9347

Observation fd631bcc-e4b8-4e1e-98e4-b71d250ac301 · outbound

This paper cites Advancing high-resolution video-language representation with large-scale video transcriptions.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Advancing high-resolution video-language representation with large-scale video transcriptions

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:51:33.344569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T18:51:30.659490Z digest=sha256:4a572b91df581728df9cddfc978a7a5c68a19c863fc0a9fa36e9487017e744a7

Observation d67194a9-c23e-4ad0-b33f-b54bc01b11b7 · outbound

This paper cites Temporally Consistent Transformers for Video Generation.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Temporally Consistent Transformers for Video Generation

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:30.809287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:30.809287Z digest=sha256:af11087db9cc27e005602615ab39d93a7bf7e21d3b039b28b1496701c2e509e1

Observation 899138b4-ed67-4150-a186-aab007597afb · outbound

This paper cites Vript: A video is worth thousands of words.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Vript: A video is worth thousands of words

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:51:33.336035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T18:51:30.935301Z digest=sha256:c764c226ba170f49f7de0e8f38ba7cd000b71be6f035ee009723d282786b27f9

Observation c90b442e-095d-4bab-a9e6-47c48015729f · outbound

This paper cites 360° vr video genera- tion with generative adversarial networks.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality 360° vr video genera- tion with generative adversarial networks

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:51:33.326333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T18:51:31.112264Z digest=sha256:7ec31e13478c8a724effe79e87c96253d6c694c7b85805397a833736e29a8a59

Observation 2a742e0c-8837-4f4f-8bb2-41edbf20e9c1 · outbound

This paper cites VLIPP: Towards Physically Plausible Video Generation with Vision and Language Informed Physical Prior.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality VLIPP: Towards Physically Plausible Video Generation with Vision and Language Informed Physical Prior

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:31.233075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:31.233075Z digest=sha256:5df425df59cbb49c847fc63728756062c9de34be43912a247b55f1f34d22e8a8

Observation fcde2073-940e-4af3-ac8b-5c23df6659fc · outbound

This paper cites Cogvideox: Text-to-video diffusion models with an expert transformer.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Cogvideox: Text-to-video diffusion models with an expert transformer

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:51:33.315924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T18:51:31.377341Z digest=sha256:0dbd3a726aa27fada9002ebdda4602bd38196459042df85eb8cbe9bb45e9a033

Observation 83421e8d-4098-4575-b067-bba14b442e4a · outbound

This paper cites Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:31.488724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:31.488724Z digest=sha256:1c2132121918c171f5e43fd597c4f68d643628ba9eabd24e42f167777e1bc0c7

Observation 0f6cd5e5-47de-4748-bb35-16492f705483 · outbound

This paper cites Framepack: Pack- ing input frame context for next-frame prediction models.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Framepack: Pack- ing input frame context for next-frame prediction models

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:31.599998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:31.599998Z digest=sha256:a869e44b32ec418a199dde3b6c66bbb40b057334cf0fa341946642732bdf9721

Observation 36715bad-77f5-4ede-bd11-a840d56fa124 · outbound

This paper cites Efros, Eli Shecht- man, and Oliver Wang.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Efros, Eli Shecht- man, and Oliver Wang

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:51:33.305757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T18:51:31.661328Z digest=sha256:97cc2bf96c01371f1de749a5972a4d48947d3df3db96767b83b04174db33c62b

Observation e1cda6db-b13e-47c4-b547-d3ec492efd5e · outbound

This paper cites I2VGen-XL: High-Quality Image-to-Video Synthesis via Cascaded Diffusion Models.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality I2VGen-XL: High-Quality Image-to-Video Synthesis via Cascaded Diffusion Models

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:31.757497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:31.757497Z digest=sha256:9535ca9abe06b294380a25b192aab7ad4a7f37d406ec1b6dc4d77146f6bd2b5e

Pith citing papers

Observation 01be779b-48f3-4e5c-be87-d677397da751 · inbound

FilmWorld: Agentic Novel-to-Film Generation through Dynamic Cinematic World Modeling cites this paper.

FilmWorld: Agentic Novel-to-Film Generation through Dynamic Cinematic World Modeling A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T13:41:42.111186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:41:42.111186Z digest=sha256:64d0c55a1fdc0b9d53e890ef72a159cef96b581b5b01330f5b1608c805ae1a17