Pith. sign in

Paper Citation Record · LEDGER

Frame-Level Captions for Long Video Generation with Complex Multi Scenes

As of 18 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 1 inbound Pith citation observation for arXiv:2505.20827.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.20827 v1

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:52:24.986017Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T16:39:33.443014Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T16:39:34.724653Z

Reference resolution

47 of 47 outbound references displayed

  • verified exact0
  • verified fuzzy16
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 326a447c-60ba-46d3-a187-c57f939e92e8 · outbound

This paper cites Lumiere: A space-time diffusion model for video generation.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Lumiere: A space-time diffusion model for video generation

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:52:28.475980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:52:20.619862Z digest=sha256:b0b7434187cd1f8caad5707ca4a2e832f76f019bae29552a83665d48cf3b769c

Observation d58735dd-48a3-4813-89ea-9807fa66f72b · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:20.736392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:20.736392Z digest=sha256:d47a3b3faf7e7aa630534abea0db6bdf62ca4bd29773c1ca73c1d88b88830ef7

Observation a57090df-d86f-46bc-bfce-6382860c7b8d · outbound

This paper cites Align your latents: High-resolution video synthesis with latent diffusion models.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Align your latents: High-resolution video synthesis with latent diffusion models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:20.841396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:20.841396Z digest=sha256:2d4419927a3af934bbde673e6af5868b3a462cfa707e4e63835cfe5db1e9a6ff

Observation 1d51cc5a-4ff0-4606-a8f0-7645f5a55768 · outbound

This paper cites Video generation models as world simulators.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Video generation models as world simulators

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:20.907024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:20.907024Z digest=sha256:3d61176b139f11cf71530dc989008d831183f324d51b84a462cf58232e20b130

Observation 44bba26c-5ef4-4aac-a2ab-af2857f75a54 · outbound

This paper cites Activitynet: A large-scale video benchmark for human activity understanding.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Activitynet: A large-scale video benchmark for human activity understanding

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:52:28.241900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:52:20.983081Z digest=sha256:2c8695d37df9c0700a25911ac2aeeea30a34a1f5399e2d1be8169168af8f9326

Observation c3721de0-7be4-461d-a563-a61dbd409232 · outbound

This paper cites Diffusion forcing: Next-token prediction meets full-sequence diffusion.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Diffusion forcing: Next-token prediction meets full-sequence diffusion

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:21.064096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:21.064096Z digest=sha256:57a9053f2321bd5ba0aca7e3fb0e2a713a64867f0c1304e39b9852d63a662a1d

Observation 3038fb96-a49c-4bb0-a81f-681711bf0412 · outbound

This paper cites SkyReels-V2: Infinite-length Film Generative Model.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes SkyReels-V2: Infinite-length Film Generative Model

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:21.230585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:21.230585Z digest=sha256:d1cf79ebdc5ac5b39cb85c2005f404bf163fd8361360421ace13b365ec3ac324

Observation f2cd86fe-fcb3-4001-a2cc-2ac94e8811e4 · outbound

This paper cites Panda-70m: Captioning 70m videos with multiple cross-modality teachers.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Panda-70m: Captioning 70m videos with multiple cross-modality teachers

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:52:28.033266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:52:21.313961Z digest=sha256:7996d30b2adfba63d48618a33cc8e13c117db47f0dec1e2812a5e49e51bea1fb

Observation b2f7a23e-1389-4327-b6fa-4601eda84a77 · outbound

This paper cites Learning temporal coherence via self-supervision for gan-based video generation.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Learning temporal coherence via self-supervision for gan-based video generation

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:52:27.875131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:52:21.429053Z digest=sha256:5d3d38478991a7fb7824117734bce673a9d566760e78b3a022ad9fa034572800

Observation 15b9771c-e22f-4231-adeb-d36464455b75 · outbound

This paper cites Factorizing text-to-video generation by explicit image conditioning.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Factorizing text-to-video generation by explicit image conditioning

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:52:27.757449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:52:21.543209Z digest=sha256:8d1d06fabfee35ccf5d2eaff9370847a1f4ae1d075dd4b35518ceab3ee92c29b

Observation fa7f08fb-4ff9-4ff0-b7f1-7d20d9a4f1f1 · outbound

This paper cites AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:21.675607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:21.675607Z digest=sha256:7ad21a13bdf1fc05c2adfe75df2bf78986c76162dd6b28de92dd0ca5e4ee0ce0

Observation b625c430-1f3c-43d7-ab62-2dd8f00864fd · outbound

This paper cites Long Context Tuning for Video Generation.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Long Context Tuning for Video Generation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:21.800462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:21.800462Z digest=sha256:8c44887823a62c138bf3bd91c1b1e608cde56ff6d202bb70b1a41e0286e59fd8

Observation 80296d13-3267-4b87-99b6-ede1ccec4070 · outbound

This paper cites StreamingT2V: Consistent, Dynamic, and Extendable Long Video Generation from Text.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes StreamingT2V: Consistent, Dynamic, and Extendable Long Video Generation from Text

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:21.866461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:21.866461Z digest=sha256:16b355b46e0f86d219cbd6659d63319de31098cd3a7aee246216c251005d77aa

Observation 438697a2-2e0d-43c2-b8ab-abc715c43277 · outbound

This paper cites Autoregressive diffusion models.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Autoregressive diffusion models

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:52:27.602519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:52:21.937109Z digest=sha256:ae7264ea6ca4db006a9c3ba7d0f1a5f7f27684f1eb7e2042926ff28c506240fe

Observation 8b9fdd70-f981-4764-a3ae-3e19a022e48e · outbound

This paper cites Step-Video-TI2V Technical Report: A State-of-the-Art Text-Driven Image-to-Video Generation Model.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Step-Video-TI2V Technical Report: A State-of-the-Art Text-Driven Image-to-Video Generation Model

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:22.021206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:22.021206Z digest=sha256:0dd075376ce6cee44cc22449f4fe64b32bd48787cdde1d0fdde8e63d1d65bf5e

Observation b13f9e56-6b9f-4ea2-a976-c9b6e453d8ee · outbound

This paper cites FIFO-Diffusion: Generating Infinite Videos from Text without Training.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes FIFO-Diffusion: Generating Infinite Videos from Text without Training

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:22.117866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:22.117866Z digest=sha256:c953d670204c9d5d4b2580739547c3068628522ffff655ccfcad56f4728be43f

Observation d12f45af-4f3b-4d6f-9679-3cce42ef9b65 · outbound

This paper cites HunyuanVideo: A Systematic Framework For Large Video Generative Models.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:22.237972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:22.237972Z digest=sha256:42cf681975ccbd82263a5f1d3a596defd5371126de67f734471ed53c17612451

Observation a307adf3-e2ac-4d93-8f34-e9cdbdb6e240 · outbound

This paper cites A Survey on Long Video Generation: Challenges, Methods, and Prospects.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes A Survey on Long Video Generation: Challenges, Methods, and Prospects

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:22.315530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:22.315530Z digest=sha256:01c358cac515b46a0ce3d8ed35d6536336fc37629e32f949ca2f5ee825d35dd5

Observation c4852276-2f62-44ff-88f6-1bece02499fe · outbound

This paper cites Unified Video Action Model.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Unified Video Action Model

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:22.397770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:22.397770Z digest=sha256:ce9779151525d4f175a8bb17d98a750415d3a3b1b194d9e7f76be8f63ac2f84d

Observation 4ecca3c2-367f-4ce3-a333-0d009eac451e · outbound

This paper cites Open-Sora Plan: Open-Source Large Video Generation Model.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Open-Sora Plan: Open-Source Large Video Generation Model

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:22.472441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:22.472441Z digest=sha256:d95bedeca401f00b27013c68fbf9be77fa0926cbd35e0c9c908031fe8d064370

Observation a076bb32-7a26-45d0-a597-80329400b0b5 · outbound

This paper cites Videostudio: Generating consistent-content and multi-scene videos.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Videostudio: Generating consistent-content and multi-scene videos

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:52:27.452934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:52:22.546815Z digest=sha256:244a751f345b529d470eab1fd863a935c58d85484549b91e8a420164b5e13364

Observation e8eed2f6-a220-4570-a159-41bfa82eb6c3 · outbound

This paper cites Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:22.696814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:22.696814Z digest=sha256:1b7c49455e61cc2565b17642d58c1cc3923755fc76d6378c9ba21fe212d4703f

Observation e77a915f-003e-4247-86e2-ad20030aba1a · outbound

This paper cites Mevg: Multi-event video generation with text-to-video models.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Mevg: Multi-event video generation with text-to-video models

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:52:27.206936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:52:22.790070Z digest=sha256:7f325a01a288e63bae90f0a43c340d752f0c0f438294b03b9c6edf6a72f9b496

Observation c9b86993-63dc-4c26-ba95-210aa7a8b470 · outbound

This paper cites Scalable diffusion models with transformers.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Scalable diffusion models with transformers

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:22.881341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:22.881341Z digest=sha256:b63f838850100befa7bbf1dbfc7bfb17f447a109b7f0ed67c02ab09259b202c7

Observation f227f7fc-c4cc-4f16-ad55-e6e0e7674011 · outbound

This paper cites Mask$^2$DiT: Dual Mask-based Diffusion Transformer for Multi-Scene Long Video Generation.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Mask$^2$DiT: Dual Mask-based Diffusion Transformer for Multi-Scene Long Video Generation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:22.951099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:22.951099Z digest=sha256:f88fad57308e9f6ac88f4677fda5cf4fc63db3eda3e0c08b0042fc816d934fde

Observation 2cb4abcb-3de5-4d37-a93b-5e1a93019b13 · outbound

This paper cites Freenoise: Tuning-free longer video diffusion via noise rescheduling.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Freenoise: Tuning-free longer video diffusion via noise rescheduling

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:52:26.911330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:52:22.997892Z digest=sha256:9f52937e1ac5da5e6f97774374fffe61c2d273312e6a64d752d3182755af87be

Observation c69b16a1-53d6-411e-920f-de6d91f91e68 · outbound

This paper cites Seaweed-7B: Cost-Effective Training of Video Generation Foundation Model.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Seaweed-7B: Cost-Effective Training of Video Generation Foundation Model

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:23.076920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:23.076920Z digest=sha256:b93be863c87f3f14cb6a3dc24d53d3e6056edd5d9ef49aee3cc40645b7dd26d5

Observation 70b1d677-ea3a-4e73-802a-d1c568ef8139 · outbound

This paper cites Stylegan-v: A continuous video generator with the price, image quality and perks of stylegan2.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Stylegan-v: A continuous video generator with the price, image quality and perks of stylegan2

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:52:26.761609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:52:23.143892Z digest=sha256:a44e7dd82db91b62a1fa50f2911a87b68c866de8155269fcb999c4bbf3a8a9d6

Observation 212a9cf1-50d9-48f8-bbad-7a0f8af8487d · outbound

This paper cites Lightweight, Pre-trained Transformers for Remote Sensing Timeseries.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Lightweight, Pre-trained Transformers for Remote Sensing Timeseries

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:23.239488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:23.239488Z digest=sha256:b3940f86f2d4867731d0662128c9b4b61481c109ffc677eb0ab0a7dfd321002d

Observation a1d42ff8-58c3-4633-ba3c-39b2f7a4c21f · outbound

This paper cites Mocogan: Decomposing motion and content for video generation.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Mocogan: Decomposing motion and content for video generation

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:52:26.532001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:52:23.335228Z digest=sha256:573cf23346f69fde86a5e03d704be221a84adeb4b6a40675f5543ccd5ac784f3

Observation 55bd74b1-3bf0-40f5-9594-e1a9d0f0218d · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Wan: Open and Advanced Large-Scale Video Generative Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:23.434779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:23.434779Z digest=sha256:defb7d81329b9dace82cdd0f4203b7d25ea9dade14eadf7f8ebbbf33cc3cec4f

Observation f999d2fb-6d40-4027-9760-581b287740b2 · outbound

This paper cites STORYANCHORS: Generating Consistent Multi-Scene Story Frames for Long-Form Narratives.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes STORYANCHORS: Generating Consistent Multi-Scene Story Frames for Long-Form Narratives

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:23.548609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:23.548609Z digest=sha256:c595350e138928de2dc9cc5f3965a031f9b108441a8997ccac5b1c2285a40a65

Observation d5bea886-b2d0-455b-af90-1c8903bcb576 · outbound

This paper cites Gen-L-Video: Multi-Text to Long Video Generation via Temporal Co-Denoising.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Gen-L-Video: Multi-Text to Long Video Generation via Temporal Co-Denoising

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:23.659573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:23.659573Z digest=sha256:f70485043187adff650d7e66cb27327aba451c32c8af73e60e04fd2319018103

Observation a5e26b97-9690-40db-85af-400913c64a0e · outbound

This paper cites Koala-36M: A Large-scale Video Dataset Improving Consistency between Fine-grained Conditions and Video Content.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Koala-36M: A Large-scale Video Dataset Improving Consistency between Fine-grained Conditions and Video Content

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:23.736263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:23.736263Z digest=sha256:cc12e0345872edd2012e44994e82e5296738ac3c1a7a7b4035001c6b53e62ba7

Observation 8ffd8aa8-5ad0-40cc-a4ae-1faf228dae00 · outbound

This paper cites Lvbench: An extreme long video understanding benchmark, 2024.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Lvbench: An extreme long video understanding benchmark, 2024

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:23.833995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:23.833995Z digest=sha256:95929c8effeeb2eb7230fbb4f3dd63613d1fcf40948eb935ee977388499c2684

Observation 7246b089-de6e-49b3-9bd5-6dd77cde997e · outbound

This paper cites VideoAgent: Long-form Video Understanding with Large Language Model as Agent.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes VideoAgent: Long-form Video Understanding with Large Language Model as Agent

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:23.917757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:23.917757Z digest=sha256:0c42cf171963888f10636c454d4599e9dfc970699504665a6f018822aa49f559

Observation 0b4f6d54-734f-4648-98e3-9798e16b796c · outbound

This paper cites Vatex: A large-scale, high-quality multilingual dataset for video-and-language research.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Vatex: A large-scale, high-quality multilingual dataset for video-and-language research

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:52:26.296919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:52:24.045232Z digest=sha256:d4416ed0c3955b59cd41e069ae101011b5c267a0b2b1df874cb1613494461a5e

Observation 83bf8499-8c7c-4f69-83f8-60277cbbbae4 · outbound

This paper cites Imaginator: Condi- tional spatio-temporal gan for video generation.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Imaginator: Condi- tional spatio-temporal gan for video generation

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:52:26.128349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:52:24.153790Z digest=sha256:6e88b898d64c3806b7930956778bb140b16a15a5b44a8983f42e2c51595e030e

Observation 8e9afcf7-8ae4-48b7-97b1-71253170f820 · outbound

This paper cites Advancing high-resolution video-language representation with large-scale video transcriptions.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Advancing high-resolution video-language representation with large-scale video transcriptions

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:24.252204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:24.252204Z digest=sha256:578fdc2f5d524869b39b0bfd81a4586d3c1e5539cbeeb1f25a30a987b626cabe

Observation 40835e76-76e2-4297-bbc0-8aeaf401f988 · outbound

This paper cites Long Video Diffusion Generation with Segmented Cross-Attention and Content-Rich Video Data Curation.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Long Video Diffusion Generation with Segmented Cross-Attention and Content-Rich Video Data Curation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:24.370109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:24.370109Z digest=sha256:76325e8069a8c3432a3624d886f7b5d1b12780bbe83a12280e33568fbcdcaeed

Observation 1861851a-92b6-47e7-94ad-99d54fc6287e · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:24.467700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:24.467700Z digest=sha256:4c86823ff78297568dad7ffb96d7a141b4eed5894c460392dee41eac3cbdce9f

Observation f68c7dbf-cef9-4122-9acc-69273d4938b4 · outbound

This paper cites Merlot: Multimodal neural script knowledge models.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Merlot: Multimodal neural script knowledge models

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:52:25.994105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:52:24.570316Z digest=sha256:8f0098fbf2678c2c4da9aff12c436c6c2846ce3cd5ef4b2a61fb5024641ef543

Observation 4cd00ff3-e07a-4ac9-b454-5e03a94706d1 · outbound

This paper cites Moviedreamer: Hierarchical generation for coherent long visual sequence.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Moviedreamer: Hierarchical generation for coherent long visual sequence

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:24.643654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:24.643654Z digest=sha256:41892b065401a311d0a2cc4d2a21731afda780a63bfd709ddebbb083aceb0386

Observation bfaca529-075e-4c00-81e6-caaaf624d3f9 · outbound

This paper cites VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:24.723287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:24.723287Z digest=sha256:31baa14bd3e635e382157e56081301d918e13bf5f8254d64ae13b604c3f8dbc7

Observation 1a69e4fa-55ef-4a37-bc3e-75fd0c59ff2a · outbound

This paper cites Videogen-of-thought: A collaborative framework for multi-shot video generation.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Videogen-of-thought: A collaborative framework for multi-shot video generation

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:24.804117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:24.804117Z digest=sha256:ed7b354f94c79c9f3e08ddb56659c37aad8b82a193ae0f55cefe2014887469b3

Observation 63da9f05-e6a4-45aa-a677-27469349e31c · outbound

This paper cites Towards automatic learning of procedures from web instructional videos.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Towards automatic learning of procedures from web instructional videos

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:52:25.793585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:52:24.885503Z digest=sha256:d81b96e187b17346feae469dc19d1390836b3014e7a01fe1a7d0deb33ade6640

Observation 614f1169-0a76-4a97-82d3-70953de36695 · outbound

This paper cites Storydiffusion: Consistent self-attention for long-range image and video generation.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Storydiffusion: Consistent self-attention for long-range image and video generation

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:52:25.576965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:52:24.986017Z digest=sha256:1dd653d69d159b00c36f6aef869d63fb98734ec6d9fdc3e01a1a4cab8cf27506

Pith citing papers

Observation 672da83a-bf59-448e-b9c2-ffe3a00488e4 · inbound

LoViC: Efficient Long Video Generation with Context Compression cites this paper.

LoViC: Efficient Long Video Generation with Context Compression Frame-Level Captions for Long Video Generation with Complex Multi Scenes

Reference 91

Resolution
verified exact
local_arxiv, observed 2026-08-06T16:39:34.798842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T16:39:33.443014Z digest=sha256:1ade4970aeefd87c7872722febae5dbc33067fa9f28094b36c8dbddd064fe169