Pith. sign in

Paper Citation Record · LEDGER

LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity

As of 11 August 2026, this Paper Citation Record lists 76 of 76 outbound references and 6 inbound Pith citation observations for arXiv:2412.09856.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.09856 v2

Coverage vector

measured 76 of 76 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T16:43:08.574966Z

measured 82 of 82 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:03:20.460990Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-19T07:27:08.875128Z

Reference resolution

76 of 76 outbound references displayed

  • verified exact0
  • verified fuzzy25
  • unresolved51
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 67be81be-11f0-4b4e-b2e0-c5f26f9a2a65 · outbound

This paper cites Video generation models as world simu- lators.

LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity Video generation models as world simu- lators

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:43:09.974969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T16:43:08.180513Z digest=sha256:2bc46c1d86efc3bd0a24892ba17224971d93ecdaee602878ae339b6119ef4ca4

Observation 4546eec7-2c32-4ebc-8dde-fd71c16bff2a · outbound

This paper cites VideoCrafter1: Open Diffusion Models for High-Quality Video Generation.

LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity VideoCrafter1: Open Diffusion Models for High-Quality Video Generation

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:08.187386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:43:08.187386Z digest=sha256:fddef72a35769f46efb9b167b48b7c833c33049a4a2e3215ee980c32dc03cb35

Observation df97383c-259c-4653-9fba-4ce95c81c15d · outbound

This paper cites VideoCrafter2: Overcoming data limitations for high-quality video diffu- sion models.

LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity VideoCrafter2: Overcoming data limitations for high-quality video diffu- sion models

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:43:09.959193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T16:43:08.192920Z digest=sha256:109027748ea7dbca7cb563865677f1ed0fe8b5ad0e01b29bc148ddf436c092be

Observation 8f8f8991-fcd0-4764-9516-13c1a0eae939 · outbound

This paper cites DiffEdit: Diffusion-based semantic image editing with mask guidance.

LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity DiffEdit: Diffusion-based semantic image editing with mask guidance

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:08.198731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:43:08.198731Z digest=sha256:fe235bbaaad596bb23da278b8092de763c11246208d549d91b59fdf6b8a7384c

Observation 3fa3e2dc-c0a3-451d-ac0c-e511652471a0 · outbound

This paper cites Emu: Enhancing Image Generation Models Using Photogenic Needles in a Haystack.

LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity Emu: Enhancing Image Generation Models Using Photogenic Needles in a Haystack

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:08.204201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:43:08.204201Z digest=sha256:c6775f690249ec26efd42357bee51fd3ceccab224eb8fba0d00d35abe00da2fc

Observation 1a6c9d8d-052b-4f2b-82a6-6806e2bd1ac3 · outbound

This paper cites Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality.

LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:08.210442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:43:08.210442Z digest=sha256:2cc1613fc4a679bc38281afbe9547364e347c0f65c969656b54851c4648c253a

Observation 434e6069-d101-43fc-a7de-3ca178f17454 · outbound

This paper cites FlashAttention: Fast and memory-efficient exact attention with IO-awareness.

LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity FlashAttention: Fast and memory-efficient exact attention with IO-awareness

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:43:09.941900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T16:43:08.216087Z digest=sha256:55033bf2bf52043b354164d95fd9610d1a200385b68da2c1cd348d087d12d742

Observation f707be84-4ea5-4a18-8d5c-5e57d64d9060 · outbound

This paper cites Alabdul- mohsin, et al.

LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity Alabdul- mohsin, et al

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:43:09.926042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T16:43:08.221900Z digest=sha256:dd3d4f125166c809e18bf424fa4096f6df4ed4b871bce8942778473b2e4677b5

Observation 98cbeb64-83b5-48aa-82dc-8039ff9a7c47 · outbound

This paper cites The Llama 3 Herd of Models.

LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity The Llama 3 Herd of Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:08.228074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:43:08.228074Z digest=sha256:4dee6ebb0a97fa7d742c7ed8bd52593eeae77b83ec03cfd6f9472044468aad54

Observation f72e4609-bef2-4b9f-a8f0-2769afa8be51 · outbound

This paper cites Matten: Video Generation with Mamba-Attention.

LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity Matten: Video Generation with Mamba-Attention

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:08.233267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:43:08.233267Z digest=sha256:7019f399bb0b808503522dd0042cb4e220aa693b36448eef13da44e3c36ebfae

Observation e1c7108d-86e9-429d-a487-cdda0b005945 · outbound

This paper cites Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning.

LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:08.238185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:43:08.238185Z digest=sha256:d46d110fbd2618dbd486a4e6654b0b64433fc0f4be01fca662551541a9a37ae2

Observation a39fe9b4-09d7-466f-aa97-d98a4220b74b · outbound

This paper cites Mamba: Linear-Time Sequence Modeling with Selective State Spaces.

LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity Mamba: Linear-Time Sequence Modeling with Selective State Spaces

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:08.245160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:43:08.245160Z digest=sha256:e1ae35e9ae72f685919cbf0bb00bbc4fd6676025b6d310986b7ceb0a0cf6388d

Observation 87fb6a4b-1dda-4eb5-b2e8-90a140483ca2 · outbound

This paper cites HiPPO: Recurrent memory with optimal polyno- mial projections.

LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity HiPPO: Recurrent memory with optimal polyno- mial projections

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:43:09.910237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T16:43:08.250778Z digest=sha256:c716b2b8e8443784d8165f8cb6df25e301d686e6f6abd48185168505b20312df

Observation 1f31fdff-1860-4e54-b596-3d9f364712c7 · outbound

This paper cites Efficiently Modeling Long Sequences with Structured State Spaces.

LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity Efficiently Modeling Long Sequences with Structured State Spaces

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:08.255538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:43:08.255538Z digest=sha256:b7b6169943b1f4ca5b0943ac91ce704daa5f934877508bd59658333e161ddca9

Observation 8016844c-8074-43e6-ac6c-97465382bbf7 · outbound

This paper cites MambaAD: Exploring State Space Models for Multi-class Unsupervised Anomaly Detection.

LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity MambaAD: Exploring State Space Models for Multi-class Unsupervised Anomaly Detection

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:08.260634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:43:08.260634Z digest=sha256:d89054220c41aafb2935c109f4dafbf714ac5765833d53c093b2437c9cc3e5cc

Observation 69c39173-1b98-43c4-bfcd-3fc74f86e966 · outbound

This paper cites Denoising diffu- sion probabilistic models.

LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity Denoising diffu- sion probabilistic models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:08.265594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:43:08.265594Z digest=sha256:93d6fca3e792c84e937992b8e8a4290a7c08746edd9bfe0073e5248daa773c3d

Observation 73af2afb-1f10-42fb-abe2-0fb5db852272 · outbound

This paper cites Imagen Video: High Definition Video Generation with Diffusion Models.

LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity Imagen Video: High Definition Video Generation with Diffusion Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:08.270287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:43:08.270287Z digest=sha256:92309594a489649e80fb080c04b33f34e4ccb1378bc6082da079cdd80dcb4d19

Observation 5584c0cd-a842-4134-932b-2bfe21d80aea · outbound

This paper cites CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers.

LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:08.276151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:43:08.276151Z digest=sha256:eb4a1f8799b404eb18583bf0ff619f27eaff6a66cb0da8f3c968d35c8ab238aa

Observation b87f8b7d-304b-4486-a080-f9a33cad1989 · outbound

This paper cites ZigMa: A DiT-style Zigzag Mamba Diffusion Model.

LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity ZigMa: A DiT-style Zigzag Mamba Diffusion Model

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:08.281382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:43:08.281382Z digest=sha256:6ce96f48a6834d587a4d4aef668cf976d1e71319d9dd8557c97b9cb3f0f5a7e6

Observation 93a6088d-2350-4891-84b1-d660d273be6d · outbound

This paper cites VBench: Com- prehensive benchmark suite for video generative models.

LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity VBench: Com- prehensive benchmark suite for video generative models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:08.286787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:43:08.286787Z digest=sha256:f83b24c8b5b633a575b04e27a1f2d14622fa933b4f5cc3069d6380a222f91ef7

Observation ebc306c6-34c7-4484-8efc-8ab750767b79 · outbound

This paper cites MiraData: A Large-Scale Video Dataset with Long Durations and Structured Captions.

LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity MiraData: A Large-Scale Video Dataset with Long Durations and Structured Captions

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:08.291243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:43:08.291243Z digest=sha256:f109cd58749e9b763d0cf500b1e4720aa9c0495e57087d12f51d0a80c8bf3b21

Observation 304a5e6a-d94e-4091-ac81-b411f442c56b · outbound

This paper cites Imagic: Text-based real image editing with diffusion models.

LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity Imagic: Text-based real image editing with diffusion models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:08.295497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:43:08.295497Z digest=sha256:841676fa03201d52bc499757ee02cfe45760d3bc7de7e7f29d832d55121221c1

Observation c8adbf2d-ef64-428e-be91-d8aabb0f3530 · outbound

This paper cites BK-SDM: A Lightweight, Fast, and Cheap Version of Stable Diffusion.

LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity BK-SDM: A Lightweight, Fast, and Cheap Version of Stable Diffusion

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:08.300020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:43:08.300020Z digest=sha256:82f19944e5562d65769e78253b3cf19d3babf1f627d6ae68c58839018ca2cc8f

Observation 4338dae4-dca8-43f4-9558-30ae7b5529c1 · outbound

This paper cites Kling AI: Next-generation AI creative studio.

LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity Kling AI: Next-generation AI creative studio

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:43:09.861129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T16:43:08.305666Z digest=sha256:db074c6d879b94a2c3d09369c3221117e575b2095e8a851caa658a3076428109

Observation 096d883a-35b0-4f2e-ae7c-ebe3d2b8644d · outbound

This paper cites VideoPoet: A Large Language Model for Zero-Shot Video Generation.

LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:08.311136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:43:08.311136Z digest=sha256:8bb6baefbcd7950c6c4b45b75e425c687fedbc33857c45b3ae6751d462b596aa

Observation 2d3acca8-e53b-4e2b-8c1e-65fb9a39e256 · outbound

This paper cites Pika labs.

LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity Pika labs

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:43:09.843977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T16:43:08.317114Z digest=sha256:ea511c240c875920990aceddcbf1ec3b926437cb4c6350ddd90f2982b2473e53

Observation dced3c44-002a-4e03-9d1e-b8d38aaf19b1 · outbound

This paper cites xFormers: A modular and hackable trans- former modelling library.

LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity xFormers: A modular and hackable trans- former modelling library

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:43:09.826627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T16:43:08.322165Z digest=sha256:cee57595ec2880aa6ee369ca92849369acc812e1ff4c8ac50b7100b3fb27d0f5

Observation 49a8df5e-1cf7-4675-a3cd-d7887bad7f70 · outbound

This paper cites T2V-Turbo: Breaking the Quality Bottleneck of Video Consistency Model with Mixed Reward Feedback.

LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity T2V-Turbo: Breaking the Quality Bottleneck of Video Consistency Model with Mixed Reward Feedback

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:08.327092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:43:08.327092Z digest=sha256:5699748cc678e8be73c92bc340df55d042b5454eff432ddd66a455bf562508b5

Observation e5c55d33-5deb-4072-9510-1050b5d2165c · outbound

This paper cites T2V- Turbo-v2: Enhancing video generation model post-training through data, reward, and conditional guidance design.arXiv preprint arXiv:2410.05677, 2024.

LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity T2V- Turbo-v2: Enhancing video generation model post-training through data, reward, and conditional guidance design.arXiv preprint arXiv:2410.05677, 2024

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:08.332179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:43:08.332179Z digest=sha256:91abbf66581527791b15d428efd025f9e89378957ec5f9732cbfcc8ba9d56705

Observation 9d4794f3-5e65-4028-a7b2-3285f26d365e · outbound

This paper cites Flow Matching for Generative Modeling.

LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity Flow Matching for Generative Modeling

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:08.337755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:43:08.337755Z digest=sha256:892ec6bf3c127cdb2af09281535b00e9051b55fc5b8155753179a2b4a020e78b

Observation 6c9fc96a-32fe-40db-b94b-5173b6635299 · outbound

This paper cites Swin Transformer: Hierarchical vision transformer using shifted windows.

LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity Swin Transformer: Hierarchical vision transformer using shifted windows

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:43:09.811957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T16:43:08.342684Z digest=sha256:2376221db5250b9551d4f27674391141d221b7357e0a590c4c80cc53dbb4673c

Observation aba8da3d-3592-41ea-862f-9e30436b6564 · outbound

This paper cites VDT: General-purpose Video Diffusion Transformers via Mask Modeling.

LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity VDT: General-purpose Video Diffusion Transformers via Mask Modeling

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:08.347341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:43:08.347341Z digest=sha256:b67e4e3b590cf930a2960e29903eec2c5f3bf399d33e196b8ec3e619053511bc

Observation 52824075-8794-4904-a44b-b57723da92ab · outbound

This paper cites Dream machine.

LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity Dream machine

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:43:09.796124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T16:43:08.351882Z digest=sha256:bcc8c794624ea049ecda95132f334776559c6f69980e1a9f1ae9767c4f68e402

Observation 93f206e9-c541-4573-9808-e797e531f9dc · outbound

This paper cites Diffusion probabilistic models for 3D point cloud generation.

LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity Diffusion probabilistic models for 3D point cloud generation

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:43:09.780129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T16:43:08.356402Z digest=sha256:78bf8a9a2a9e35f7a56e95f97b52b76cd23267cc7ad0c1ad77a5b69f06b2a89d

Observation 45ec4db2-4e09-4354-9913-ae343f0d2ee1 · outbound

This paper cites Latent Consistency Models: Synthesizing High-Resolution Images with Few-Step Inference.

LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity Latent Consistency Models: Synthesizing High-Resolution Images with Few-Step Inference

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:08.362974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:43:08.362974Z digest=sha256:8b7fe4ecc2c97c169f6b730f3f499ee4d8f2dc01a3fe4cb8354c7b2fc851c650

Observation 9f108b34-3684-43b0-9e5b-1acf437d6dd6 · outbound

This paper cites On distillation of guided diffusion models.

LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity On distillation of guided diffusion models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:08.368317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:43:08.368317Z digest=sha256:7a7aaabd007789d31050fefd1f8ca4ce7b4e5ef7f484658f72566b41ec380a3d

Observation 3271f270-fff0-492f-893e-ee838cc80e68 · outbound

This paper cites Scaling Diffusion Mamba with Bidirectional SSMs for Efficient Image and Video Generation.

LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity Scaling Diffusion Mamba with Bidirectional SSMs for Efficient Image and Video Generation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:08.374170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:43:08.374170Z digest=sha256:10820354ab6443a5b3ee5a521287e5d496e219abffc844c40523bf45b05df96e

Observation 1d66c585-fee5-45db-bbc1-d996606ef9bb · outbound

This paper cites Transframer: Arbitrary Frame Prediction with Generative Models.

LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity Transframer: Arbitrary Frame Prediction with Generative Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:08.379341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:43:08.379341Z digest=sha256:3490baba856f6ba46d92edc6b3d93979217490120222a842416998e96ac76266

Observation 1fd1fe16-4527-444e-9c70-efd6d6929ca6 · outbound

This paper cites Scalable diffusion models with transformers.

LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity Scalable diffusion models with transformers

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:08.384349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:43:08.384349Z digest=sha256:7bceb93ff6265c6ffa4d2a4999da16f569abfb6683ed39cf0e55c4636242106c

Observation 93b99cae-553e-4607-b021-cf266fc24959 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:08.389315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:43:08.389315Z digest=sha256:854b2e0c19343fb2f3a1372dfaeebcf690bb07f09200d8c37098da4ddbbc8223

Observation 9a67080f-8ff0-4101-b294-704c6b458a16 · outbound

This paper cites Movie Gen: A Cast of Media Foundation Models.

LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity Movie Gen: A Cast of Media Foundation Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:08.395562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:43:08.395562Z digest=sha256:4c50e36f1cfb34173b5a3fa56af695d67e2df01557ed1fd963e0278083548726

Observation 94240f2a-ecc1-4110-b093-b3a8473e482f · outbound

This paper cites RawFilm: 8k cinematic royalty-free stock footage.

LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity RawFilm: 8k cinematic royalty-free stock footage

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:43:09.739126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T16:43:08.401196Z digest=sha256:1bc313424a8909b4ad3e2887d0f5e50a83b0128291488935390964f24e3da1e7

Observation 394572c4-2cf3-4c0e-b9c2-57989059a4f5 · outbound

This paper cites Samba: Simple Hybrid State Space Models for Efficient Unlimited Context Language Modeling.

LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity Samba: Simple Hybrid State Space Models for Efficient Unlimited Context Language Modeling

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:08.406817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:43:08.406817Z digest=sha256:b953ba4e9f6f5cbcd975248b142b68282a132b1bbeebe067fe1b4169f1f37cb5

Observation a9140747-2199-4b26-803c-511bd120043f · outbound

This paper cites MambaCSR: Dual-Interleaved Scanning for Compressed Image Super-Resolution With SSMs.

LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity MambaCSR: Dual-Interleaved Scanning for Compressed Image Super-Resolution With SSMs

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:08.411772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:43:08.411772Z digest=sha256:208f1140b30ff5ffbd6a346357517dfeffb1c30d229148974ae49e86ee7c4bc0

Observation 2761b435-eb27-4d91-a71d-b347209851f5 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity High-resolution image synthesis with latent diffusion models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:08.417340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:43:08.417340Z digest=sha256:1a8c62604afbe17cf75454fe040d9873413198cea60ccae0e6bc65c372bea4f6

Observation 6c18eb9f-9ee1-45aa-bc60-eea27b3554e3 · outbound

This paper cites Introducing Gen-3 alpha.

LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity Introducing Gen-3 alpha

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:43:09.713492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T16:43:08.421498Z digest=sha256:948e5df820c1b4f14ac90a07d7ebaa7e61d0d1e1108006f9debb9923d809b9de

Observation 79c12c70-83fa-4901-885e-a794f0a3b58a · outbound

This paper cites Denton, Kamyar Ghasemipour, Raphael Gontijo Lopes, Burcu Karagol Ayan, Tim Salimans, et al.

LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity Denton, Kamyar Ghasemipour, Raphael Gontijo Lopes, Burcu Karagol Ayan, Tim Salimans, et al

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:43:09.697332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T16:43:08.426432Z digest=sha256:a69deb9ca1100f08079fb1a1c8e9aa14579023010f791d4bdbeef996a5e49959

Observation 3a31a639-25df-401b-bf69-137bad6fdfad · outbound

This paper cites GLU Variants Improve Transformer.

LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity GLU Variants Improve Transformer

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:08.430805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:43:08.430805Z digest=sha256:89dc068a6587b0ff3873818374ade35c5883fa1dce9ade38c4e30df59b33187b

Observation aa30885d-27f4-41ed-8eb5-11bb5c6c12ad · outbound

This paper cites Emu Edit: Precise image editing via recognition and gen- eration tasks.

LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity Emu Edit: Precise image editing via recognition and gen- eration tasks

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:43:09.680306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T16:43:08.435207Z digest=sha256:7043a83436801e03e7a880cc0ff230001e1be017b44e08df2d3195418cb23272

Observation ce1f816d-4e84-412c-ba00-2b1ab475e770 · outbound

This paper cites NormFormer: Improved Transformer Pretraining with Extra Normalization.

LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity NormFormer: Improved Transformer Pretraining with Extra Normalization

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:08.440239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:43:08.440239Z digest=sha256:8ec045d5d8f31e4f198832105edee9d4770d4cbff68440e111dd756d046c5d49

Observation 64de6b84-dbc5-4660-b28d-ec1d84c83d1e · outbound

This paper cites Shutterstock: Stock photos, royalty-free images, graphics, vectors, videos, and music.

LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity Shutterstock: Stock photos, royalty-free images, graphics, vectors, videos, and music

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:43:09.664288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T16:43:08.445562Z digest=sha256:e73d635fa2c31cfddac69c3e00928013c0819ed7f388cc01e6133163c0b321f8

Observation cfdc7c31-2fd1-45d8-9e8f-c630d5920afc · outbound

This paper cites Make-A-Video: Text-to-Video Generation without Text-Video Data.

LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity Make-A-Video: Text-to-Video Generation without Text-Video Data

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:08.450278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:43:08.450278Z digest=sha256:424740827a507b0be2bb234e0cbf629595419f240adddcca53ead910e794ae6e

Observation 5b0c444a-68c3-4be6-b776-e0ef90e36fed · outbound

This paper cites Deep unsupervised learning using nonequilibrium thermodynamics.

LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity Deep unsupervised learning using nonequilibrium thermodynamics

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:43:09.647984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T16:43:08.455265Z digest=sha256:541b70c50e840b2d18214b4f301b0d3d83a74074e6627048065b3bced379950f

Observation 66110414-57a4-40c8-851e-b00797916aa4 · outbound

This paper cites Consistency Models.

LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity Consistency Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:08.461195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:43:08.461195Z digest=sha256:e641976638599dbbb662373a89772a51bc5465ed2cc48b0f82dbf4cc120b9f9a

Observation 9b74f8bf-1982-4a81-ab30-b09496863bc0 · outbound

This paper cites UL2: Unifying Language Learning Paradigms.

LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity UL2: Unifying Language Learning Paradigms

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:08.467829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:43:08.467829Z digest=sha256:83c53ab1b57760ab40c78d28c6304b46a060bc3a157bd9e79b338a448e09955c

Observation 4ccbca24-6cc3-449e-9468-ba4b5f26a920 · outbound

This paper cites an unresolved cited work.

LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:43:09.630832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T16:43:08.473388Z digest=sha256:142614a6d2973905fcc864394ec810b51807673a3386b169d23c32c987126981

Observation 6edfc395-e805-48a7-9f56-300d4160948b · outbound

This paper cites DiM: Diffusion Mamba for Efficient High-Resolution Image Synthesis.

LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity DiM: Diffusion Mamba for Efficient High-Resolution Image Synthesis

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:08.478362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:43:08.478362Z digest=sha256:38e739555da0af9410c8ebfbc1f77ba0dc2b1d380a53bd1c826b6767db5284f1

Observation 93cf7088-e7da-40bf-9ff8-692f012afd68 · outbound

This paper cites LION: Latent point dif- fusion models for 3D shape generation.

LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity LION: Latent point dif- fusion models for 3D shape generation

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:43:09.613422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T16:43:08.484763Z digest=sha256:b9c9e0af6148a05fd25383178447e0b728f588f65afaec267c5cad7396ecff72

Observation ff7a6544-0c04-43ad-a68e-435c6c98d6f4 · outbound

This paper cites Phenaki: Variable length video generation from open domain textual descriptions.

LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity Phenaki: Variable length video generation from open domain textual descriptions

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:08.489615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:43:08.489615Z digest=sha256:f2fc6052da2c0e1180f049b8ad7098d7d668189d788af5e0c779e7ec1f922268

Observation 5e2ed4d1-b2d8-4131-ad92-1a0a21c701cc · outbound

This paper cites An Empirical Study of Mamba-based Language Models.

LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity An Empirical Study of Mamba-based Language Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:08.494872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:43:08.494872Z digest=sha256:43d9f8e28096629549fe892a4de60bbc9ccc6791d459429e962adb123ece1a73

Observation 50e6d2a5-ff15-4a22-9ce9-9b479f9317e5 · outbound

This paper cites Jha, and Yuchen Liu.

LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity Jha, and Yuchen Liu

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:43:09.587749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T16:43:08.499943Z digest=sha256:89ca83774abc2aac486d742b903a71528b32a654fc7bad0ae0276bcefdce8478

Observation 1f44977b-4c10-4806-8cf6-d2f2b9a797bd · outbound

This paper cites ModelScope Text-to-Video Technical Report.

LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity ModelScope Text-to-Video Technical Report

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:08.504673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:43:08.504673Z digest=sha256:90e33432bdcb224f7c1d34928bf8d52140eafa4a4b99579a32f2d84bbffe8da6

Observation b87730bf-0482-421d-9086-c9aec0aed188 · outbound

This paper cites VideoLCM: Video Latent Consistency Model.

LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity VideoLCM: Video Latent Consistency Model

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:08.509845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:43:08.509845Z digest=sha256:a1761fcff70524cc02a8f76b239afa72022edc460a0f05a49a36cbbf1ac0954d

Observation fc670c36-7bbe-4b30-b704-235d62799298 · outbound

This paper cites LAVIE: High-Quality Video Generation with Cascaded Latent Diffusion Models.

LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity LAVIE: High-Quality Video Generation with Cascaded Latent Diffusion Models

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:08.514912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:43:08.514912Z digest=sha256:0daf773b52f5ac94ba0df040dcb0983153428ef5ae49975043d9cca193c4d719

Observation 13f00a1f-ca3e-4c96-9bb6-99320135d3d4 · outbound

This paper cites Loong: Generating Minute-level Long Videos with Autoregressive Language Models.

LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity Loong: Generating Minute-level Long Videos with Autoregressive Language Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:08.519752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:43:08.519752Z digest=sha256:a3dd1c716b49a46a4dcfe93d22b43885b4b3c6b3d657d4c7776f9495e0fed21b

Observation 986da5f9-b3bf-4441-a306-b32c54140b37 · outbound

This paper cites Progressive Autoregressive Video Diffusion Models.

LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity Progressive Autoregressive Video Diffusion Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:08.525261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:43:08.525261Z digest=sha256:c76d190d873bf94f5c268a1f58640ef0f11967f258e380a00d2acbb77b2e7ed2

Observation 30536006-04cb-4964-9e3c-614def0e3e23 · outbound

This paper cites Demystifying CLIP Data.

LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity Demystifying CLIP Data

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:08.531267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:43:08.531267Z digest=sha256:ace0e93a76d46298e1ea1eccafdfc7499b2c309f5e2b993c023a0be662c56722

Observation c4ec7859-8615-4d98-a3d8-131b31574633 · outbound

This paper cites ByT5: Towards a token-free future with pre-trained byte- to-byte models.

LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity ByT5: Towards a token-free future with pre-trained byte- to-byte models

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:43:09.573235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T16:43:08.536061Z digest=sha256:f709d8d5fd5d297d01d921913fe4504cfa31f357b6ad0403828e3bdc5d4d1b29

Observation b5b75df8-c35b-4007-8684-bcf332f3247a · outbound

This paper cites Dif- fusion models without attention.

LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity Dif- fusion models without attention

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:43:09.558055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T16:43:08.541049Z digest=sha256:0510163b955c77f7d471a9aae9cfc470b480b2f368b4e1c7650860feb453c5d5

Observation a3c68c58-0f51-48e8-9131-ed880c767a6d · outbound

This paper cites VideoGPT: Video Generation using VQ-VAE and Transformers.

LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity VideoGPT: Video Generation using VQ-VAE and Transformers

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:08.545220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:43:08.545220Z digest=sha256:98a33cfd6488c9a936cdfe2f1e1920b7c19718a89db3c90aaef3b3b6923551ac

Observation bfa55274-aaff-4222-8c16-acdcc4db6dd9 · outbound

This paper cites Paint by example: Exemplar-based image editing with diffusion mod- els.

LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity Paint by example: Exemplar-based image editing with diffusion mod- els

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:08.550153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:43:08.550153Z digest=sha256:4824ae1cdceb6dacbc1c2fdad94825c8e35ad27bea0f6f0eed740217ae4fb953

Observation f16bd07c-6c7e-4ddc-89cb-47d5a430be8b · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:08.554893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:43:08.554893Z digest=sha256:94ec34c65037cd19b8c61f10f01876e4605f13de5e53209df0d43cc87cfe5cf9

Observation bbb8c895-b74a-43b8-a0dd-8d11a7d08921 · outbound

This paper cites Hauptmann, Ming- Hsuan Yang, Yuan Hao, Irfan Essa, et al.

LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity Hauptmann, Ming- Hsuan Yang, Yuan Hao, Irfan Essa, et al

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:43:09.531284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T16:43:08.559795Z digest=sha256:498c034b1bb2cfecf65b472ef55585f91089b4bf1f2ee2225b55791f70ef9334

Observation e90ea079-7759-4fd6-9fc4-a3250ced6a98 · outbound

This paper cites Root mean square layer nor- malization.

LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity Root mean square layer nor- malization

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:43:09.515407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T16:43:08.565377Z digest=sha256:28797458bcdf4ee07210367db655c1888a89738e493f6ba7a7f57b54ce7e96b0

Observation d4003444-edac-41d9-9900-71d939834724 · outbound

This paper cites Show-1: Marrying pixel and latent diffusion models for text-to-video generation.

LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity Show-1: Marrying pixel and latent diffusion models for text-to-video generation

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:43:09.499522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T16:43:08.570224Z digest=sha256:f34083e0b673b75787d1de25d8c0ab1ccb6f824132d20cb9cea724aaa03b90c9

Observation 5bb90295-a349-43af-bedc-e0295fb25444 · outbound

This paper cites Open-Sora: Democratizing efficient video production for all.

LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity Open-Sora: Democratizing efficient video production for all

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:43:09.483317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T16:43:08.574966Z digest=sha256:9dea372d5f29dd7920974aee3f0e6ea69ee1af0b5d61ecd2c9d532d17a43e993

Pith citing papers

Observation 0b237622-2fa2-45f5-8386-99bd91f8cf11 · inbound

Long-Context State-Space Video World Models cites this paper.

Long-Context State-Space Video World Models LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:20.460990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:03:20.460990Z digest=sha256:225d6c6afc9b5330ee1287e08cdcfe211e8033bc501ef7e3e1b676e42deb63f9

Observation 0c076b86-9971-4220-859b-86de36be2480 · inbound

Video World Models with Long-term Spatial Memory cites this paper.

Video World Models with Long-term Spatial Memory LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:41.279115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:41.279115Z digest=sha256:5e3c16d364cfab73356f4defead1a54fca0eaf8e59763710ad3615e7d6d9f3d6

Observation de71e09f-eb27-4bd3-bf72-0227ef3f457e · inbound

Exploring Diffusion Transformer Designs via Grafting cites this paper.

Exploring Diffusion Transformer Designs via Grafting LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:04.156409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:29:04.156409Z digest=sha256:71fe2331f3880b47bbf8c2d6f4126d61266730077408372254b57bcfb1f7a341

Observation a8fe1376-17c3-4580-a606-0c671997fc7a · inbound

M4V: Multimodal Mamba for Efficient Text-to-Video Generation cites this paper.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T04:21:19.044277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:21:19.044277Z digest=sha256:44ae082eac736fd48e569824dff8b16ed04343189094bf069a0e6dfa8ec31861

Observation d2d1e529-6b84-4fb0-92c2-28dede2602f8 · inbound

GenHSI: Controllable Generation of Human-Scene Interaction Videos cites this paper.

GenHSI: Controllable Generation of Human-Scene Interaction Videos LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity

Reference 80

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:27:08.877210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:8a151bc4b60d38b75c6197148930cb90e41af1c60555326cf7b84fa02bf90a4e

Observation 17f7f2af-7ba3-415a-84f0-f6c164e11758 · inbound

VMoBA: Mixture-of-Block Attention for Video Diffusion Models cites this paper.

VMoBA: Mixture-of-Block Attention for Video Diffusion Models LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T21:34:04.914075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:34:04.914075Z digest=sha256:ada91fdded799662e1b6482251b160cfa9a638cfd1cb8d83098e399c5a778063