Pith. sign in

REVIEW 4 major objections 7 minor 187 references

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation

T0 review · 4 major / 7 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read This survey claims that long video generation is best understood as three paradigms—autoregressive, divide-and-conquer, and implicit latent-space synthesis—and argues that divide-and-conquer, especially LLM-guided planning, is the key to…

desk verdict A broad, useful survey that is let down by unreliable benchmark tables and attribution errors; worth revising, not rejecting. read the letter →

arxiv 2412.18688 v2 pith:6WK6XSZH submitted 2024-12-24 cs.CV cs.AI

classification cs.CVcs.AI
keywords longvideogenerationtext-to-videosynthesisdivide-and-conquerparadigmLLM-guideddiffusionmodelsautoregressivesurveyevaluationmetrics
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper is a survey aiming to serve as the comprehensive foundation for long video generation research. It organizes the field into three generation paradigms and dives deepest into divide-and-conquer, where an LLM plans scenes and a separate generator fills frames. The authors argue this approach addresses scalability and narrative control better than pure autoregressive generation. The survey also catalogs datasets, metrics, and open problems, positioning itself as the missing focused review of a fast-growing field.

What carries the argument

The central organizing mechanism is the divide-and-conquer paradigm: generate keyframes or short clips from prompts, often planned by an LLM, then interpolate or stitch them into a continuous long video. The LLM-as-director pattern is the load-bearing instance, in which an LLM produces a narrative blueprint (scene descriptions, layouts, bounding boxes, actions) that a separate video diffusion module executes. This is contrasted with autoregressive prediction, which conditions each frame on previous frames, and implicit generation, which synthesizes the whole video from a compressed spatiotemporal latent representation.

What would settle it

A systematic literature search for long video generation papers published between 2021 and mid-2025 that are absent from this survey and do not fit any of the three paradigms (for example, direct world-simulator models that generate video without planning, sequential prediction, or compressed latent-space synthesis) would refute the claim that the taxonomy is exhaustive. A reproducible annotation study measuring the fraction of relevant papers that independent annotators cleanly assign to one of the three paradigms would likewise test whether the partition is well-defined.

Watch

Extended reading notes

Core claim

On its own terms, the paper establishes that long video generation can be systematically mapped into three paradigms: autoregressive frame prediction, divide-and-conquer keyframe-and-interpolation, and implicit synthesis from a compressed latent space. Within divide-and-conquer, it identifies three sub-patterns: LLM-as-director, multi-stage or agent-based frameworks, and transition or compositional stitching. The paper further claims that this taxonomy, together with its catalog of datasets and evaluation metrics, provides a comprehensive and current foundation that earlier surveys lacked, particularly in its detailed treatment of divide-and-conquer.

Load-bearing premise

The survey's comprehensiveness rests on the assumption that snowball sampling of 190+ articles from a manually selected list of venues captures a representative and complete picture of the field, and that the three-way partition of generation paradigms is exhaustive.

Editorial extensions

If this is right

  • If the survey's map is correct, future long video models will increasingly separate a planning stage from a frame-generation stage, since divide-and-conquer supports parallel keyframe generation and finer narrative control.
  • The identified scarcity of large-scale, richly captioned video datasets becomes a concrete bottleneck; building datasets that combine scale with spatial and temporal detail would directly accelerate progress.
  • Metrics that rely on manual feedback, such as FETV, VBench, and MiraBench, are a scalability bottleneck, making fully automated semantic and temporal evaluation a clear research target.
  • The divide-and-conquer sub-taxonomy defines a modular design space—planner, generator, transition module—that is testable by comparing how well different combinations reconstruct the same storylines.
  • The survey's open challenges, including audio alignment and physical dynamics, point to next steps that are largely orthogonal to the core generation paradigm and can be pursued on top of any of the three approaches.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A consequence the paper leaves implicit: if the three-paradigm taxonomy is accepted, the implicit category is likely to absorb most future foundation models, while divide-and-conquer will dominate controllable and story-driven applications; the two may eventually converge.
  • The paper's emphasis on divide-and-conquer is a bet on the value of explicit planning; one could test whether planning-based methods actually beat end-to-end latent synthesis on long-range coherence when compute budgets are held equal.
  • A practical extension: use the survey's taxonomy to build a benchmark that samples methods from each paradigm and evaluates them on identical prompts and durations; the results would either validate the taxonomy's utility or reveal overlaps that blur its boundaries.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 7 minor

Summary. This survey reviews the emerging area of long video generation, organizing the literature into three generation paradigms (autoregressive, divide-and-conquer, and implicit latent-space generation) and covering backbone architectures, tokenization strategies, input control mechanisms, datasets, and evaluation metrics. The authors state in the abstract and in Section 1.1 that the survey 'would serve as a comprehensive foundation' for the field, and they position their contribution against two earlier surveys [18,19] by focusing on divide-and-conquer methods, agent-based frameworks, and short-to-long video transitions. The paper catalogs over 190 papers from 2021 to 2025, with timeline figures and several comparison tables.

Significance. If the cataloging is accurate, the survey would be a useful entry point for researchers: it gathers recent work (2023-2025) that earlier surveys do not cover in depth, and its three-paradigm taxonomy, particularly the divide-and-conquer subsection, is a reasonable high-level organization of the field. The paper also draws attention to and compares datasets and metrics, which are often treated separately in the literature. However, the central value of a survey of this kind is reliability of its factual claims; the inconsistencies in the taxonomy and in the benchmark tables described below mean that the 'comprehensive foundation' claim is not currently supported by the manuscript as written. The paper contains no derivations or experiments, so its correctness hangs entirely on the accuracy of its reporting.

major comments (4)
  1. [Section 3.1, Table 1] Table 1, which is labeled 'Auto Regressive Approaches,' lists StyleGAN-V [85] and DIGAN [40] as autoregressive methods. This contradicts the paper's own description in Section 2.1.2, where both are presented as continuous/implicit GAN generators, and it also conflicts with the definition of implicit video generation in Section 3.3, which explicitly excludes extrapolation (autoregressive) and interpolation (divide-and-conquer). As a result, the three-paradigm partition in Section 3 is not clean or exhaustive, and the reader cannot determine which papers belong to which paradigm. Please reclassify these models or justify their placement in the autoregressive category.
  2. [Table 7] Table 7, titled 'Comprehensive benchmark comparison,' mixes FVD numbers from different evaluation protocols (UCF-101, BAIR) with FID and CLIPSIM numbers from MSR-VTT, and the footnote does not resolve which cell comes from which benchmark. The table also contains duplicate entries: Phenaki appears as both [15] and [63], ModelScope and ModelScopeT2V both cite [177], VideoDirectorGPT appears as both [11] and [166], and Video LDM [179] is the same work as Video Diffusion [173]. Without per-row benchmark provenance and deduplication, these numbers cannot be verified or compared. Please provide a source for every reported value and remove the duplicate rows.
  3. [Table 8] Table 8 reports per-dimension VBench scores that are largely in the 80-99 range (e.g., LaVie: Subject Consistency 91.41, Background Consistency 97.47, Motion 96.38) but lists Overall scores around 26-28 for the same models (e.g., LaVie Overall 26.41). Since the caption states that the table evaluates models across 12 VBench dimensions, the Overall score should be derivable from those dimensions; the reported values appear to be from a different, unlabeled benchmark. In addition, Gen-2 is cited to [188], which is the RunwayML Gen-3 Alpha page, and references [13] and [17] attribute Gen-2 and Gen-4 Alpha to 'Midjourney Team,' contradicting the running text in Section 1, which identifies both as RunwayML models. These provenance errors prevent readers from trusting the table.
  4. [Section 3.1, ARLON description] The text states that ARLON [84] achieves '128× compression (8× spatial + 6× temporal downsampling),' but 8 multiplied by 6 is 48, not 128. Either the compression factor or the downsampling factors are incorrect. Please verify this claim against the ARLON paper and correct the numbers.
minor comments (7)
  1. [Section 1.4] The survey organization paragraph contains placeholder cross-references 'Section datasets' and 'Section metrics' that do not correspond to any section, and Section 1.2 uses bracket-style references such as '[3.2]' and '[5.1]' that do not match the manuscript's section numbering.
  2. [References] The reference list contains duplicate entries for the same works, including Phenaki (appearing as [15] and [63]), MEVG (as [101] and [104]), VideoDirectorGPT (as [11] and [166]), and Video Diffusion/Video LDM (as [173] and [179]). These should be consolidated.
  3. [Section 6.2] The sentence 'Here are examples from the MSR-VTT dataset...' appears twice verbatim, once in the main text and once immediately after the figure caption for Figure 17.
  4. [Section 3.3] The sentence about Cosmos cites reference [84], which is the ARLON paper; the Cosmos World Foundation Model is separately cited as [113] later in the same section. The earlier citation appears to be a mismatch.
  5. [Section 2.2.1] Reference [51], an anomaly-detection paper by one of the authors, is cited in the autoencoder background section. The connection to video generation is not explained, and it is not used to support any of the survey's central claims; consider removing it or providing a clearer rationale.
  6. [Section 7] The introduction to Section 7 contains stray LaTeX commands: 'textit Video Quality Metrics' and 'textit Semantics Quality Metrics'. Additionally, Table 5 lists 'Pandas 70m' while the reference [141] is titled 'Panda-70M'; please make the names consistent.
  7. [Figure 3] The caption of Figure 3 claims that most papers on long video generation were published in 2023-2025, but no data source or counting methodology is provided for the histogram, so the claim cannot be verified.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the survey makes no fitted predictions or derivations, and its only self-citation is peripheral and non-load-bearing.

full rationale

This is a literature survey, not a derivation or prediction paper. The abstract's claim that the survey 'would serve as a comprehensive foundation' is a scope statement about cataloging external work, not a mathematical or empirical result derived from the paper's own inputs. The survey's content consists of summaries of third-party methods, datasets, and benchmarks, with citations to the primary sources (e.g., StyleGAN-V, DIGAN, Phenaki, VideoDirectorGPT, Sora, VBench). No parameter is fitted and no quantity is predicted from another quantity within the paper, so none of the enumerated circularity patterns applies. The one self-citation, reference [51] in Section 2.2.1, is an earlier anomaly-detection paper by the authors cited only as an example of LSTM-Convolutional VAEs; it does not support the survey's central taxonomy, benchmark tables, or any comparative claim, and removing it would not change any conclusion. Concerns about Table 1 misclassifications, Table 7 duplicated/unverifiable entries, and Table 8 provenance errors are factual-accuracy and verification issues for a survey, not circular reasoning: the survey is not defining its categories in terms of its own conclusions or citing itself to force a result. The paper is self-contained as a review of external literature, and the honest finding is no significant circularity.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The survey introduces no free parameters or invented entities. Its conclusions depend on literature-coverage assumptions: representative sampling, exhaustive taxonomy, accurate transcription of prior benchmark numbers, and the claim that only two earlier surveys exist. These are domain assumptions about the completeness and correctness of the secondary reporting.

assumptions (4)
  • domain assumption Snowball sampling of over 190 articles from a selected set of venues yields a representative coverage of long video generation.
    Section 1.3 describes the search but gives no inclusion and exclusion criteria or validation against a systematic query; the survey's comprehensiveness rests on this.
  • domain assumption The three-way partition (auto-regressive, divide-and-conquer, implicit) is exhaustive and non-overlapping.
    Section 3 introduces the paradigms as 'three core paradigms' without proving mutual exclusivity; several models, such as ARLON and Grid Diffusion, fit multiple categories.
  • domain assumption There are only two related long-video surveys ([18], [19]).
    Section 1.1 asserts 'to our knowledge' but does not describe how this was verified; the positioning of the contribution depends on this.
  • domain assumption Metric values in Tables 7 and 8 are correctly transcribed from the cited papers.
    The survey provides no code or raw data for these tables; the inconsistent overall scores in Table 8 indicate this assumption is violated.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation." pith.science (2026). https://pith.science/paper/6WK6XSZH

@misc{pith2026241218688,
  author       = {Pith},
  title        = {Pith review of: Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/6WK6XSZH}},
  note         = {Machine review of arXiv:2412.18688}
}
read the original abstract

An image may convey a thousand words, but a video composed of hundreds or thousands of image frames tells a more intricate story. Despite significant progress in multimodal large language models (MLLMs), generating extended videos remains a formidable challenge. As of this writing, OpenAI's Sora, the current state-of-the-art system, is still limited to producing videos that are up to one minute in length. This limitation stems from the complexity of long video generation, which requires more than generative AI techniques for approximating density functions essential aspects such as planning, story development, and maintaining spatial and temporal consistency present additional hurdles. Integrating generative AI with a divide-and-conquer approach could improve scalability for longer videos while offering greater control. In this survey, we examine the current landscape of long video generation, covering foundational techniques like GANs and diffusion models, video generation strategies, large-scale training datasets, quality metrics for evaluating long videos, and future research areas to address the limitations of the existing video generation capabilities. We believe it would serve as a comprehensive foundation, offering extensive information to guide future advancements and research in the field of long video generation.

Figures

Figures reproduced from arXiv: 2412.18688 by the authors.

Figure 1
Figure 1. Example of semantic content not changing with the progress of frames [ [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Example Of semantic content changing with the progress of frames [ [PITH_FULL_IMAGE:figures/full_fig_p002_2.png] view at source ↗
Figure 3
Figure 3. Most papers focusing on long video generation were published in 2023–2025. [PITH_FULL_IMAGE:figures/full_fig_p003_3.png] view at source ↗
Figures from the paper (18 more)
Figure 4
Figure 4. Figure 4: Evolution of long video generation models. Later models, such as SORA [ [PITH_FULL_IMAGE:figures/full_fig_p003_4.png]
Figure 5
Figure 5. Figure 5: The basic theme of the auto-regressive approach is that it generates new frames, given the initial anchor frame, previous [PITH_FULL_IMAGE:figures/full_fig_p009_5.png]
Figure 6
Figure 6. Figure 6: Grid Diffusion Model. It first generates a grid image and then learns a spatial auto-regressive model by learning to predict [PITH_FULL_IMAGE:figures/full_fig_p010_6.png]
Figure 7
Figure 7. Figure 7: Divide-and-conquer timeline: We used the dates these papers were published in online resources, such as arXiv or [PITH_FULL_IMAGE:figures/full_fig_p012_7.png]
Figure 8
Figure 8. Figure 8: The LLM as director approach utilizes LLM as the spatiotemporal director of the script, along with a separate video generation [PITH_FULL_IMAGE:figures/full_fig_p012_8.png]
Figure 9
Figure 9. Figure 9: VideoDirectorGPT: GPT-4 generates a blueprint for video generation, including scene and entity description. Separate module [PITH_FULL_IMAGE:figures/full_fig_p013_9.png]
Figure 10
Figure 10. Figure 10: Mora utilizes a multi-agent framework. The prompt selection agent enhances prompts with detailed instructions, the [PITH_FULL_IMAGE:figures/full_fig_p013_10.png]
Figure 11
Figure 11. Figure 11: Using a Causal 3D VAE, Hunyuan Video compresses data into latent space. LLM-encoded text conditions Gaussian noise [PITH_FULL_IMAGE:figures/full_fig_p015_11.png]
Figure 13
Figure 13. Figure 13: An image is split into fixed-size patches, linearly embedded, augmented with position embeddings, and then passed through [PITH_FULL_IMAGE:figures/full_fig_p016_13.png]
Figure 12
Figure 12. Figure 12: Transformer-Based Diffusion Model Sora compressed video of variable length into fixed space-time latent compressed [PITH_FULL_IMAGE:figures/full_fig_p016_12.png]
Figure 14
Figure 14. Figure 14: The proposed 3D-VQ architecture extends the 2D VQGAN [ [PITH_FULL_IMAGE:figures/full_fig_p017_14.png]
Figure 15
Figure 15. Figure 15: A latent diffusion model with input conditioning generates data by applying a reverse diffusion process on latent representa [PITH_FULL_IMAGE:figures/full_fig_p018_15.png]
Figure 16
Figure 16. Figure 16: DirectT2V Modulated self-attention for capturing interactions between frames [ [PITH_FULL_IMAGE:figures/full_fig_p019_16.png]
Figure 17
Figure 17. Figure 17: Here are examples from the MSR-VTT dataset showcasing video clips paired with labeled sentences. Each example includes [PITH_FULL_IMAGE:figures/full_fig_p021_17.png]
Figure 18
Figure 18. Figure 18: Dover score. Samples from a dataset with human labeling of aesthetics and technical aspects of images. Dover score could be [PITH_FULL_IMAGE:figures/full_fig_p023_18.png]
Figure 19
Figure 19. Figure 19: The prompt dataset is designed to evaluate the model by focusing on three key quality aspects: (1) spatial quality (frame [PITH_FULL_IMAGE:figures/full_fig_p023_19.png]
Figure 20
Figure 20. Figure 20: GRiT locates different entities in scenes with their relations and matches with dense captions [ [PITH_FULL_IMAGE:figures/full_fig_p024_20.png]
Figure 21
Figure 21. Figure 21: FETV is multi-faceted, classifying prompts into three distinct aspects: the main content, controllable attributes, and prompt [PITH_FULL_IMAGE:figures/full_fig_p026_21.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

187 extracted references · 53 canonical work pages

  1. [85]

    Stylegan-v: A continuous video generator with the price, image quality and perks of stylegan2

    Ivan Skorokhodov, Sergey Tulyakov, and Mohamed Elhoseiny. Stylegan-v: A continuous video generator with the price, image quality and perks of stylegan2. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pages 3626–3636, June 2022

  2. [40]

    Generating videos with dynamics-aware implicit generative adversarial networks

    Sihyun Yu, Jihoon Tack, Sangwoo Mo, Hyunsu Kim, Junho Kim, Jung-Woo Ha, and Jinwoo Shin. Generating videos with dynamics-aware implicit generative adversarial networks. In International Conference on Learning Representations , 2022

  3. [63]

    Phenaki: Variable length video generation from open domain textual descriptions

    Ruben Villegas, Mohammad Babaeizadeh, Pieter-Jan Kindermans, Hernan Moraldo, Han Zhang, Mohammad Taghi Saffar, Santiago Castro, Julius Kunze, and Dumitru Erhan. Phenaki: Variable length video generation from open domain textual descriptions. In International Conference on Learning Representations, 2023

  4. [11]

    VideodirectorGPT: Consistent multi-scene video generation via LLM-guided planning, 2024

    Han Lin, Abhay Zala, Jaemin Cho, and Mohit Bansal. VideodirectorGPT: Consistent multi-scene video generation via LLM-guided planning, 2024

  5. [166]

    Videodirectorgpt: Consistent multi-scene video generation via llm-guided planning

    Xi Chen, Yaohui Wang, Shangchen Zhou, Ceyuan Yang, Yinan Zhang, Yizhou He, Yu Wang, Haoxin Yang, Ziwei Huang, Xiaodong Lu, et al. Videodirectorgpt: Consistent multi-scene video generation via llm-guided planning. arXiv:2309.15091, 2023

  6. [179]

    Videoldm: High-resolution video generation with latent diffusion models

    Andreas Blattmann, Robin Rombach, Huan Ling, Tim Dockhorn, Seung Wook Kim, Sanja Fidler, and Karsten Kreis. Videoldm: High-resolution video generation with latent diffusion models. arXiv:2304.08818, 2023

  7. [173]

    Video diffusion models

    Jonathan Ho, Tim Salimans, Alexey Gritsenko, William Chan, Mohammad Norouzi, and David J Fleet. Video diffusion models. arXiv preprint arXiv:2204.03458, 2022

  8. [188]

    Accessed June 17, 2024 [Online] https://runwayml.com/research/introducing-gen-3-alpha, 2024

    Gen-3. Accessed June 17, 2024 [Online] https://runwayml.com/research/introducing-gen-3-alpha, 2024

  9. [13]

    Gen2 by runway ml, 2023

    Midjourney Team. Gen2 by runway ml, 2023

  10. [17]

    Gen-4 alpha by midjourney, 2024

    Midjourney Team. Gen-4 alpha by midjourney, 2024

  11. [84]

    Arlon: Boosting diffusion transformers with autoregressive models for long video generation, 2025

    Zongyi Li, Shujie Hu, Shujie Liu, Long Zhou, Jeongsoo Choi, Lingwei Meng, Xun Guo, Jinyu Li, Hefei Ling, and Furu Wei. Arlon: Boosting diffusion transformers with autoregressive models for long video generation, 2025

Show all 187 references
  1. [1]

    Video generation models as world simulators by open a.i, 2024

    Sora Team. Video generation models as world simulators by open a.i, 2024

  2. [2]

    Introducing chatgpt by open a.i, 2022

    Oepn A.I Team. Introducing chatgpt by open a.i, 2022

  3. [3]

    Meta llama models, 2023

    Meta A.I Team. Meta llama models, 2023

  4. [4]

    Google gemini series, 2023

    Google A.I Team. Google gemini series, 2023

  5. [5]

    Anthropic by claude, 2023

    Claude A.I Team. Anthropic by claude, 2023

  6. [6]

    Mistral large model by mistral, 2024

    Mistral A.I Team. Mistral large model by mistral, 2024

  7. [7]

    Hierarchical text-conditional image generation with clip latents, 2022

    Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. Hierarchical text-conditional image generation with clip latents, 2022

  8. [8]

    Scaling rectified flow transformers for high-resolution image synthesis, 2024

    Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas Müller, Harry Saini, Yam Levi, Dominik Lorenz, Axel Sauer, Frederic Boesel, Dustin Podell, Tim Dockhorn, Zion English, Kyle Lacey, Alex Goodwin, Yannik Marek, and Robin Rombach. Scaling rectified flow transf...

  9. [9]

    The midjourney v5.2 model for image generation, 2024

    Midjourney A.I Team. The midjourney v5.2 model for image generation, 2024

  10. [12]

    Make-a-video: Text-to-video generation without text-video data

    Uriel Singer, Adam Polyak, Thomas Hayes, Xi Yin, Jie An, Songyang Zhang, Qiyuan Hu, Harry Yang, Oron Ashual, Oran Gafni, Devi Parikh, Sonal Gupta, and Yaniv Taigman. Make-a-video: Text-to-video generation without text-video data. In The Eleventh International Conference on Lea...

  11. [14]

    Cogvideo: Large-scale pretraining for text-to-video generation via transformers

    Wenyi Hong, Ming Ding, Wendi Zheng, Xinghan Liu, and Jie Tang. Cogvideo: Large-scale pretraining for text-to-video generation via transformers. In The Eleventh International Conference on Learning Representations , 2023

  12. [16]

    Gen-l-video: Multi-text to long video generation via temporal co-denoising, 2023

    Fu-Yun Wang, Wenshuo Chen, Guanglu Song, Han-Jia Ye, Yu Liu, and Hongsheng Li. Gen-l-video: Multi-text to long video generation via temporal co-denoising, 2023. Manuscript submitted to ACM Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation 29

  13. [18]

    A survey on long video generation: Challenges, methods, and prospects, 2024

    Chengxuan Li, Di Huang, Zeyu Lu, Yang Xiao, Qingqi Pei, and Lei Bai. A survey on long video generation: Challenges, methods, and prospects, 2024

  14. [19]

    A survey on generative ai and llm for video generation, understanding, and streaming, 2024

    Pengyuan Zhou, Lin Wang, Zhi Liu, Yanbin Hao, Pan Hui, Sasu Tarkoma, and Jussi Kangasharju. A survey on generative ai and llm for video generation, understanding, and streaming, 2024

  15. [20]

    Sora: A review on background, technology, limitations, and opportunities of large vision models, 2024

    Yixin Liu, Kai Zhang, Yuan Li, Zhiling Yan, Chujie Gao, Ruoxi Chen, Zhengqing Yuan, Yue Huang, Hanchi Sun, Jianfeng Gao, Lifang He, and Lichao Sun. Sora: A review on background, technology, limitations, and opportunities of large vision models, 2024

  16. [21]

    Free-bloom: Zero-shot text-to-video generator with LLM director and LDM animator

    Hanzhuo Huang, Yufan Feng, Cheng Shi, Lan Xu, Jingyi Yu, and Sibei Yang. Free-bloom: Zero-shot text-to-video generator with LLM director and LDM animator. In Thirty-seventh Conference on Neural Information Processing Systems , 2023

  17. [22]

    Flowzero: Zero-shot text-to-video synthesis with llm-driven dynamic scene syntax, 2023

    Yu Lu, Linchao Zhu, Hehe Fan, and Yi Yang. Flowzero: Zero-shot text-to-video synthesis with llm-driven dynamic scene syntax, 2023

  18. [23]

    Align your latents: High- resolution video synthesis with latent diffusion models

    Andreas Blattmann, Robin Rombach, Huan Ling, Tim Dockhorn, Seung Wook Kim, Sanja Fidler, and Karsten Kreis. Align your latents: High- resolution video synthesis with latent diffusion models. In 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pages ...

  19. [24]

    Mavin: Multi-action video generation with diffusion models via transition video infilling, 2024

    Bowen Zhang, Xiaofei Xie, Haotian Lu, Na Ma, Tianlin Li, and Qing Guo. Mavin: Multi-action video generation with diffusion models via transition video infilling, 2024

  20. [25]

    SEINE: Short-to-long video diffusion model for generative transition and prediction

    Xinyuan Chen, Yaohui Wang, Lingjun Zhang, Shaobin Zhuang, Xin Ma, Jiashuo Yu, Yali Wang, Dahua Lin, Yu Qiao, and Ziwei Liu. SEINE: Short-to-long video diffusion model for generative transition and prediction. In The Twelfth International Conference on Learning Representations , 2024

  21. [26]

    Learning transferable visual models from natural language supervision, 2021

    Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning transferable visual models from natural language supervision, 2021

  22. [27]

    Generative adversarial networks

    Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. Generative adversarial networks. Commun. ACM, 63(11):139–144, October 2020

  23. [28]

    Unsupervised representation learning with deep convolutional generative adversarial networks, 2016

    Alec Radford, Luke Metz, and Soumith Chintala. Unsupervised representation learning with deep convolutional generative adversarial networks, 2016

  24. [29]

    Deep generative image models using a laplacian pyramid of adversarial networks

    Emily Denton, Soumith Chintala, Arthur Szlam, and Rob Fergus. Deep generative image models using a laplacian pyramid of adversarial networks. In Proceedings of the 28th International Conference on Neural Information Processing Systems - Volume 1 , NIPS’15, page 1486–1494, Camb...

  25. [30]

    Stackgan: Text to photo-realistic image synthesis with stacked generative adversarial networks

    Han Zhang, Tao Xu, Hongsheng Li, Shaoting Zhang, Xiaogang Wang, Xiaolei Huang, and Dimitris Metaxas. Stackgan: Text to photo-realistic image synthesis with stacked generative adversarial networks. In 2017 IEEE International Conference on Computer Vision (ICCV) , pages 5908–5916, 2017

  26. [31]

    Attngan: Fine-grained text to image generation with attentional generative adversarial networks

    Tao Xu, Pengchuan Zhang, Qiuyuan Huang, Han Zhang, Zhe Gan, Xiaolei Huang, and Xiaodong He. Attngan: Fine-grained text to image generation with attentional generative adversarial networks. In 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 1316–1324, 2018

  27. [32]

    Analyzing and Improving the Image Quality of StyleGAN

    Tero Karras, Samuli Laine, Miika Aittala, Janne Hellsten, Jaakko Lehtinen, and Timo Aila. Analyzing and Improving the Image Quality of StyleGAN . In 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pages 8107–8116, Los Alamitos, CA, USA, June 2020. ...

  28. [33]

    Jun-Yan Zhu, Taesung Park, Phillip Isola, and Alexei A. Efros. Unpaired image-to-image translation using cycle-consistent adversarial networks. In 2017 IEEE International Conference on Computer Vision (ICCV) , pages 2242–2251, 2017

  29. [34]

    Stylegan2 distillation for feed-forward image manipulation

    Yuri Viazovetskyi, Vladimir Ivashkin, and Evgeny Kashin. Stylegan2 distillation for feed-forward image manipulation. In Computer Vision – ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XXII , page 170–186, Berlin, Heidelberg, 2020. Spri...

  30. [35]

    Phillip Isola, Jun-Yan Zhu, Tinghui Zhou, and Alexei A. Efros. Image-to-image translation with conditional adversarial networks. In 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , pages 5967–5976, 2017

  31. [36]

    Deep multi-scale video prediction beyond mean square error, 2016

    Michael Mathieu, Camille Couprie, and Yann LeCun. Deep multi-scale video prediction beyond mean square error, 2016

  32. [37]

    Generating videos with scene dynamics, 2016

    Carl Vondrick, Hamed Pirsiavash, and Antonio Torralba. Generating videos with scene dynamics, 2016

  33. [38]

    Learning to generate time-lapse videos using multi-stage dynamic generative adversarial networks

    Wei Xiong, Wenhan Luo, Lin Ma, Wei Liu, and Jiebo Luo. Learning to generate time-lapse videos using multi-stage dynamic generative adversarial networks. In 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 2364–2373, 2018

  34. [39]

    To create what you tell: Generating videos from captions, 2018

    Yingwei Pan, Zhaofan Qiu, Ting Yao, Houqiang Li, and Tao Mei. To create what you tell: Generating videos from captions, 2018

  35. [41]

    Stylegan-v: A continuous video generator with the price, image quality and perks of stylegan2

    Ivan Skorokhodov, Sergey Tulyakov, and Mohamed Elhoseiny. Stylegan-v: A continuous video generator with the price, image quality and perks of stylegan2. In 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pages 3616–3626, 2022

  36. [42]

    Auto-encoders

    Standford tutorial. Auto-encoders

  37. [43]

    Auto-encoding variational bayes, 2022

    Diederik P Kingma and Max Welling. Auto-encoding variational bayes, 2022

  38. [44]

    Masked autoencoders are scalable vision learners

    Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, and Ross Girshick. Masked autoencoders are scalable vision learners. In 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pages 15979–15988, 2022

  39. [45]

    High-Resolution Image Synthesis with Latent Diffusion Models

    Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bjorn Ommer. High-Resolution Image Synthesis with Latent Diffusion Models . In 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pages 10674–10685, Los Alamitos, CA, USA, June 2022....

  40. [46]

    Neural discrete representation learning, 2018

    Aaron van den Oord, Oriol Vinyals, and Koray Kavukcuoglu. Neural discrete representation learning, 2018

  41. [47]

    Videogpt: Video generation using vq-vae and transformers, 2021

    Wilson Yan, Yunzhi Zhang, Pieter Abbeel, and Aravind Srinivas. Videogpt: Video generation using vq-vae and transformers, 2021

  42. [48]

    Taming Transformers for High-Resolution Image Synthesis

    Patrick Esser, Robin Rombach, and Bjorn Ommer. Taming Transformers for High-Resolution Image Synthesis . In 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pages 12868–12878, Los Alamitos, CA, USA, June 2021. IEEE Computer Society

  43. [49]

    Clip: Connecting vision and language with contrastive learning

    Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever. Clip: Connecting vision and language with contrastive learning. In Proceedings of the 38th International Conference on Machine Learning , volume 139, pages 8821–8...

  44. [50]

    Hierarchical patch vae-gan: Generating diverse videos from a single sample, 2020

    Shir Gur, Sagie Benaim, and Lior Wolf. Hierarchical patch vae-gan: Generating diverse videos from a single sample, 2020

  45. [51]

    Visual anomaly detection in video by variational autoencoder, 2022

    Faraz Waseem, Rafael Perez Martinez, and Chris Wu. Visual anomaly detection in video by variational autoencoder, 2022

  46. [52]

    VideoMAC: Video Masked Autoencoders Meet ConvNets

    Gensheng Pei, Tao Chen, Xiruo Jiang, Huafeng Liu, Zeren Sun, and Yazhou Yao. VideoMAC: Video Masked Autoencoders Meet ConvNets . In 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pages 22733–22743, Los Alamitos, CA, USA, June 2024. IEEE Computer Society

  47. [53]

    Hauptmann, Ming-Hsuan Yang, Yuan Hao, Irfan Essa, and Lu Jiang

    Lijun Yu, Yong Cheng, Kihyuk Sohn, José Lezama, Han Zhang, Huiwen Chang, Alexander G. Hauptmann, Ming-Hsuan Yang, Yuan Hao, Irfan Essa, and Lu Jiang. Magvit: Masked generative video transformer. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogniti...

  48. [54]

    MAGVLT: Masked Generative Vision-and-Language Transformer

    Sungwoong Kim, Daejin Jo, Donghoon Lee, and Jongmin Kim. MAGVLT: Masked Generative Vision-and-Language Transformer . In 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pages 23338–23348, Los Alamitos, CA, USA, June 2023. IEEE Computer Society

  49. [55]

    Multi-generator generative adversarial nets, 2017

    Quan Hoang, Tu Dinh Nguyen, Trung Le, and Dinh Phung. Multi-generator generative adversarial nets, 2017

  50. [56]

    Gomez, Łukasz Kaiser, and Illia Polosukhin

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Proceedings of the 31st International Conference on Neural Information Processing Systems , NIPS’17, page 6000–6010, Red ...

  51. [57]

    An image is worth 16x16 words: Transformers for image recognition at scale

    Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition at...

  52. [58]

    Zero-shot text-to-image generation

    Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever. Zero-shot text-to-image generation. In Marina Meila and Tong Zhang, editors, Proceedings of the 38th International Conference on Machine Learning , volume 139 of ...

  53. [59]

    Discrete variational autoencoders, 2017

    Jason Tyler Rolfe. Discrete variational autoencoders, 2017

  54. [60]

    Cogview: mastering text-to-image generation via transformers

    Ming Ding, Zhuoyi Yang, Wenyi Hong, Wendi Zheng, Chang Zhou, Da Yin, Junyang Lin, Xu Zou, Zhou Shao, Hongxia Yang, and Jie Tang. Cogview: mastering text-to-image generation via transformers. In Proceedings of the 35th International Conference on Neural Information Processing S...

  55. [61]

    Cogview2: faster and better text-to-image generation via hierarchical transformers

    Ming Ding, Wendi Zheng, Wenyi Hong, and Jie Tang. Cogview2: faster and better text-to-image generation via hierarchical transformers. In Proceedings of the 36th International Conference on Neural Information Processing Systems , NIPS ’22, Red Hook, NY, USA, 2024. Curran Associates Inc

  56. [62]

    ViViT: A Video Vision Transformer

    Anurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun, Mario Lucic, and Cordelia Schmid. ViViT: A Video Vision Transformer . In 2021 IEEE/CVF International Conference on Computer Vision (ICCV) , pages 6816–6826, Los Alamitos, CA, USA, October 2021. IEEE Computer Society

  57. [64]

    Adam Roberts, Hyung Won Chung, Anselm Levskaya, Gaurav Mishra, James Bradbury, Daniel Andor, Sharan Narang, Brian Lester, Colin Gaffney, Afroz Mohiuddin, Curtis Hawthorne, Aitor Lewkowycz, Alex Salcianu, Marc van Zee, Jacob Austin, Sebastian Goodman, Livio Baldini Soares, Hait...

  58. [65]

    Compositional 3d-aware video generation with llm director

    Hanxin Zhu, Tianyu He, Anni Tang, Junliang Guo, Zhibo Chen, and Jiang Bian. Compositional 3d-aware video generation with llm director. In Proceedings of the 38th Annual Conference on Neural Information Processing Systems (NeurIPS 2024) , 2024. Poster presentation

  59. [66]

    Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffr...

  60. [68]

    Vlogger: Make your dream a vlog

    Shaobin Zhuang, Kunchang Li, Xinyuan Chen, Yaohui Wang, Ziwei Liu, Yu Qiao, and Yali Wang. Vlogger: Make your dream a vlog. In 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pages 8806–8817, 2024. Manuscript submitted to ACM Video Is Worth a Thous...

  61. [69]

    Weiss, Niru Maheswaranathan, and Surya Ganguli

    Jascha Sohl-Dickstein, Eric A. Weiss, Niru Maheswaranathan, and Surya Ganguli. Deep unsupervised learning using nonequilibrium thermody- namics, 2015

  62. [70]

    Denoising diffusion probabilistic models

    Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. In Proceedings of the 34th International Conference on Neural Information Processing Systems, NIPS ’20, Red Hook, NY, USA, 2020. Curran Associates Inc

  63. [71]

    Generative modeling by estimating gradients of the data distribution, 2020

    Yang Song and Stefano Ermon. Generative modeling by estimating gradients of the data distribution, 2020

  64. [72]

    Bert: Pre-training of deep bidirectional transformers for language understanding, 2019

    Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding, 2019

  65. [73]

    Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res. , 21(1), January 2020

  66. [74]

    Video generation models as world simulators

    Open A.I Team. Video generation models as world simulators

  67. [75]

    Scalable diffusion models with transformers

    William Peebles and Saining Xie. Scalable diffusion models with transformers. In 2023 IEEE/CVF International Conference on Computer Vision (ICCV), pages 4172–4182, 2023

  68. [76]

    Cogview: Mastering text-to-image generation via transformers

    Ming Ding, Zhuoyi Yang, Wenyi Hong, Wendi Zheng, Chang Zhou, Da Yin, Junyang Lin, Xu Zou, Zhou Shao, Hongxia Yang, et al. Cogview: Mastering text-to-image generation via transformers. Advances in Neural Information Processing Systems , 34:19822–19835, 2021

  69. [77]

    Videotetris: Towards compositional text-to-video generation

    Ye Tian, Ling Yang, Haotian Yang, Yuan Gao, Yufan Deng, Xintao Wang, Zhaochen Yu, Xin Tao, Pengfei Wan, Di ZHANG, and Bin CUI. Videotetris: Towards compositional text-to-video generation. In The Thirty-eighth Annual Conference on Neural Information Processing Systems , 2024

  70. [78]

    Cogview2: Faster and better text-to-image generation via hierarchical transformers

    Ming Ding, Wendi Zheng, Wenyi Hong, and Jie Tang. Cogview2: Faster and better text-to-image generation via hierarchical transformers. In S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh, editors, Advances in Neural Information Processing Systems , volume 35, p...

  71. [79]

    NUWA-infinity: Autoregressive over autoregressive generation for infinite visual synthesis

    Jian Liang, Chenfei Wu, Xiaowei Hu, Zhe Gan, Jianfeng Wang, Lijuan Wang, Zicheng Liu, Yuejian Fang, and Nan Duan. NUWA-infinity: Autoregressive over autoregressive generation for infinite visual synthesis. In Alice H. Oh, Alekh Agarwal, Danielle Belgrave, and Kyunghyun Cho, ed...

  72. [80]

    Ross, Bryan Seybold, and Lu Jiang

    Dan Kondratyuk, Lijun Yu, Xiuye Gu, José Lezama, Jonathan Huang, Grant Schindler, Rachel Hornung, Vighnesh Birodkar, Jimmy Yan, Ming-Chang Chiu, Krishna Somandepalli, Hassan Akbari, Yair Alon, Yong Cheng, Josh Dillon, Agrim Gupta, Meera Hahn, Anja Hauth, David Hendon, Alonso M...

  73. [81]

    ART •V: Auto-Regressive Text-to-Video Generation with Diffusion Models

    Wenming Weng, Ruoyu Feng, Yanhui Wang, Qi Dai, Chunyu Wang, Dacheng Yin, Zhiyuan Zhao, Kai Qiu, Jianmin Bao, Yuhui Yuan, Chong Luo, Yueyi Zhang, and Zhiwei Xiong. ART •V: Auto-Regressive Text-to-Video Generation with Diffusion Models . In 2024 IEEE/CVF Conference on Computer V...

  74. [82]

    Grid diffusion models for text-to-video generation

    Taegyeong Lee, Soyeong Kwon, and Taehwan Kim. Grid diffusion models for text-to-video generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 8734–8743, 2024

  75. [86]

    Vincent Sitzmann, Julien N. P. Martel, Alexander W. Bergman, David B. Lindell, and Gordon Wetzstein. Implicit neural representations with periodic activation functions, 2020

  76. [87]

    Srinivasan, Matthew Tancik, Jonathan T

    Ben Mildenhall, Pratul P. Srinivasan, Matthew Tancik, Jonathan T. Barron, Ravi Ramamoorthi, and Ren Ng. Nerf: Representing scenes as neural radiance fields for view synthesis, 2020

  77. [88]

    Video probabilistic diffusion models in projected latent space

    Sihyun Yu, Kihyuk Sohn, Subin Kim, and Jinwoo Shin. Video probabilistic diffusion models in projected latent space. In 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pages 18456–18466, 2023

  78. [89]

    Towards end-to-end generative modeling of long videos with memory-efficient bidirectional transformers

    Jaehoon Yoo, Semin Kim, Doyup Lee, Chiheon Kim, and Seunghoon Hong. Towards end-to-end generative modeling of long videos with memory-efficient bidirectional transformers. In 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pages 22888–22897, 2023

  79. [90]

    Streamingt2v: Consistent, dynamic, and extendable long video generation from text

    Anonymous. Streamingt2v: Consistent, dynamic, and extendable long video generation from text. In Submitted to The Thirteenth International Conference on Learning Representations , 2024. under review

  80. [91]

    Vid-gpt: Introducing gpt-style autoregressive generation in video diffusion models, 2024

    Kaifeng Gao, Jiaxin Shi, Hanwang Zhang, Chunping Wang, and Jun Xiao. Vid-gpt: Introducing gpt-style autoregressive generation in video diffusion models, 2024

  81. [92]

    Flexifilm: Long video generation with flexible conditions, 2024

    Yichen Ouyang, jianhao Yuan, Hao Zhao, Gaoang Wang, and Bo zhao. Flexifilm: Long video generation with flexible conditions, 2024

  82. [93]

    Mora: Enabling generalist video generation via a multi-agent framework, 2024

    Zhengqing Yuan, Ruoxi Chen, Zhaoxu Li, Haolong Jia, Lifang He, Chi Wang, and Lichao Sun. Mora: Enabling generalist video generation via a multi-agent framework, 2024

  83. [94]

    Vidgen: Long-form text-to-video generation with temporal, narrative and visual consistency for high quality story-visualisation tasks

    Ram Selvaraj, Ayush Singh, Shafiudeen Kameel, Rahul Samal, and Pooja Agarwal. Vidgen: Long-form text-to-video generation with temporal, narrative and visual consistency for high quality story-visualisation tasks. In 2024 IEEE 9th International Conference for Convergence in Tec...

  84. [95]

    Bissyand, and Saad Ezzini

    Zhifei Xie, Daniel Tang, Dingwei Tan, Jacques Klein, Tegawend F. Bissyand, and Saad Ezzini. Dreamfactory: Pioneering multi-scene long video generation with a multi-agent framework, 2024

  85. [96]

    Kubrick: Multimodal agent collaborations for synthetic video generation, 2024

    Liu He, Yizhi Song, Hejun Huang, Daniel Aliaga, and Xin Zhou. Kubrick: Multimodal agent collaborations for synthetic video generation, 2024

  86. [97]

    Modelscope text-to-video technical report, 2023

    Jiuniu Wang, Hangjie Yuan, Dayou Chen, Yingya Zhang, Xiang Wang, and Shiwei Zhang. Modelscope text-to-video technical report, 2023

  87. [98]

    Free-bloom: zero-shot text-to-video generator with llm director and ldm animator

    Hanzhuo Huang, Yufan Feng, Cheng Shi, Lan Xu, Jingyi Yu, and Sibei Yang. Free-bloom: zero-shot text-to-video generator with llm director and ldm animator. In Proceedings of the 37th International Conference on Neural Information Processing Systems , NIPS ’23, Red Hook, NY, USA...

  88. [99]

    Denoising diffusion implicit models

    Jiaming Song, Chenlin Meng, and Stefano Ermon. Denoising diffusion implicit models. In International Conference on Learning Representations , 2021

  89. [100]

    LLM-grounded video diffusion models

    Long Lian, Baifeng Shi, Adam Yala, Trevor Darrell, and Boyi Li. LLM-grounded video diffusion models. In The Twelfth International Conference on Learning Representations, 2024

  90. [101]

    Mevg: Multi-event video generation with text-to-video models

    Gyeongrok Oh, Jaehwan Jeong, Sieun Kim, Wonmin Byeon, Jinkyu Kim, Sungwoong Kim, and Sangpil Kim. Mevg: Multi-event video generation with text-to-video models. In Computer Vision – ECCV 2024: 18th European Conference, Milan, Italy, September 29–October 4, 2024, Proceedings, Pa...

  91. [102]

    Jingbo Yang and Adrian G. Bors. Enabling the encoder-empowered gan-based video generators for long video generation. In2023 IEEE International Conference on Image Processing (ICIP) , pages 1425–1429, 2023

  92. [103]

    Videomerge: Towards training-free long video generation, 2025

    Siyang Zhang, Harry Yang, and Ser-Nam Lim. Videomerge: Towards training-free long video generation, 2025

  93. [104]

    Mevg: Multi-event video generation with text-to-video models, 2024

    Gyeongrok Oh, Jaehwan Jeong, Sieun Kim, Wonmin Byeon, Jinkyu Kim, Sungwoong Kim, and Sangpil Kim. Mevg: Multi-event video generation with text-to-video models, 2024

  94. [105]

    Hunyuanvideo: A systematic framework for large video generative models, 2025

    Weijie Kong, Qi Tian, Zijian Zhang, Rox Min, Zuozhuo Dai, Jin Zhou, Jiangfeng Xiong, Xin Li, Bo Wu, Jianwei Zhang, Kathrina Wu, Qin Lin, Junkun Yuan, Yanxin Long, Aladdin Wang, Andong Wang, Changlin Li, Duojun Huang, Fang Yang, Hao Tan, Hongmei Wang, Jacob Song, Jiawang Bai, J...

  95. [106]

    Freenoise: Tuning-free longer video diffusion via noise rescheduling

    Haonan Qiu, Menghan Xia, Yong Zhang, Yingqing He, Xintao Wang, Ying Shan, and Ziwei Liu. Freenoise: Tuning-free longer video diffusion via noise rescheduling. In The Twelfth International Conference on Learning Representations , 2024

  96. [107]

    GLOBER: Coherent non-autoregressive video generation via GLOBal guided video decodER

    Mingzhen Sun, Weining Wang, Zihan Qin, Jiahui Sun, Sihan Chen, and Jing Liu. GLOBER: Coherent non-autoregressive video generation via GLOBal guided video decodER. In Thirty-seventh Conference on Neural Information Processing Systems , 2023

  97. [108]

    Goku: Flow based video generative foundation models, 2025

    Shoufa Chen, Chongjian Ge, Yuqi Zhang, Yida Zhang, Fengda Zhu, Hao Yang, Hongxiang Hao, Hui Wu, Zhichao Lai, Yifei Hu, Ting-Che Lin, Shilong Zhang, Fu Li, Chuan Li, Xing Wang, Yanghua Peng, Peize Sun, Ping Luo, Yi Jiang, Zehuan Yuan, Bingyue Peng, and Xiaobing Liu. Goku: Flow ...

  98. [109]

    Reducio! generating 1024×1024 video within 16 seconds using extremely compressed motion latents, 2024

    Rui Tian, Qi Dai, Jianmin Bao, Kai Qiu, Yifan Yang, Chong Luo, Zuxuan Wu, and Yu-Gang Jiang. Reducio! generating 1024×1024 video within 16 seconds using extremely compressed motion latents, 2024

  99. [110]

    Genmo Team. Mochi 1. https://github.com/genmoai/models, 2024

  100. [111]

    Moviegen: A cast of media foundation models

    Meta Research Team. Moviegen: A cast of media foundation models. arXiv preprint, October 2024. Meta AI Research Publication

  101. [112]

    From sora what we can see: A survey of text-to-video generation, 2024

    Rui Sun, Yumin Zhang, Tejal Shah, Jiahao Sun, Shuoying Zhang, Wenqi Li, Haoran Duan, Bo Wei, and Rajiv Ranjan. From sora what we can see: A survey of text-to-video generation, 2024

  102. [113]

    Cosmos world foundation model platform for physical ai, 2025

    NVIDIA, :, Niket Agarwal, Arslan Ali, Maciej Bala, Yogesh Balaji, Erik Barker, Tiffany Cai, Prithvijit Chattopadhyay, Yongxin Chen, Yin Cui, Yifan Ding, Daniel Dworakowski, Jiaojiao Fan, Michele Fenzi, Francesco Ferroni, Sanja Fidler, Dieter Fox, Songwei Ge, Yunhao Ge, Jinwei ...

  103. [114]

    Fleet, Mohammad Norouzi, and Tim Salimans

    Jonathan Ho, Chitwan Saharia, William Chan, David J. Fleet, Mohammad Norouzi, and Tim Salimans. Cascaded diffusion models for high fidelity image generation, 2021

  104. [115]

    Hauptmann, Ming-Hsuan Yang, Yuan Hao, Irfan Essa, and Lu Jiang

    Lijun Yu, Yong Cheng, Kihyuk Sohn, José Lezama, Han Zhang, Huiwen Chang, Alexander G. Hauptmann, Ming-Hsuan Yang, Yuan Hao, Irfan Essa, and Lu Jiang. Magvit: Masked generative video transformer, 2023

  105. [116]

    Hitvideo: Hierarchical tokenizers for enhancing text-to-video generation with autoregressive large language models, 2025

    Ziqin Zhou, Yifan Yang, Yuqing Yang, Tianyu He, Houwen Peng, Kai Qiu, Qi Dai, Lili Qiu, Chong Luo, and Lingqiao Liu. Hitvideo: Hierarchical tokenizers for enhancing text-to-video generation with autoregressive large language models, 2025

  106. [117]

    Large language models are frame-level directors for zero-shot text-to-video generation

    Susung Hong, Junyoung Seo, Heeseong Shin, Sunghwan Hong, and Seungryong Kim. Large language models are frame-level directors for zero-shot text-to-video generation. In First Workshop on Controllable Video Generation @ICML24 , 2024. Manuscript submitted to ACM Video Is Worth a ...

  107. [118]

    Microcinema: A divide-and-conquer approach for text-to-video generation

    Yanhui Wang, Jianmin Bao, Wenming Weng, Ruoyu Feng, Dacheng Yin, Tao Yang, Jingxu Zhang, Qi Dai, Zhiyuan Zhao, Chunyu Wang, Kai Qiu, Yuhui Yuan, Xiaoyan Sun, Chong Luo, and Baining Guo. Microcinema: A divide-and-conquer approach for text-to-video generation. In Proceedings of ...

  108. [120]

    Adding conditional control to text-to-image diffusion models, 2023

    Lvmin Zhang, Anyi Rao, and Maneesh Agrawala. Adding conditional control to text-to-image diffusion models, 2023

  109. [121]

    Videocontrolnet: A motion-guided video-to-video translation framework by using diffusion model with controlnet, 2023

    Zhihao Hu and Dong Xu. Videocontrolnet: A motion-guided video-to-video translation framework by using diffusion model with controlnet, 2023

  110. [122]

    Moonshot: Towards controllable video generation and editing with multimodal conditions, 2024

    David Junhao Zhang, Dongxu Li, Hung Le, Mike Zheng Shou, Caiming Xiong, and Doyen Sahoo. Moonshot: Towards controllable video generation and editing with multimodal conditions, 2024

  111. [123]

    Easycontrol: Transfer controlnet to video diffusion for controllable generation and interpolation, 2024

    Cong Wang, Jiaxi Gu, Panwen Hu, Haoyu Zhao, Yuanfan Guo, Jianhua Han, Hang Xu, and Xiaodan Liang. Easycontrol: Transfer controlnet to video diffusion for controllable generation and interpolation, 2024

  112. [124]

    Controlvideo: Training-free controllable text-to-video generation, 2023

    Yabo Zhang, Yuxiang Wei, Dongsheng Jiang, Xiaopeng Zhang, Wangmeng Zuo, and Qi Tian. Controlvideo: Training-free controllable text-to-video generation, 2023

  113. [125]

    Videostudio: Generating consistent-content and multi-scene videos, 2024

    Fuchen Long, Zhaofan Qiu, Ting Yao, and Tao Mei. Videostudio: Generating consistent-content and multi-scene videos, 2024

  114. [126]

    Sdxl: Improving latent diffusion models for high-resolution image synthesis, 2023

    Dustin Podell, Zion English, Kyle Lacey, Andreas Blattmann, Tim Dockhorn, Jonas Müller, Joe Penna, and Robin Rombach. Sdxl: Improving latent diffusion models for high-resolution image synthesis, 2023

  115. [127]

    Videobooth: Diffusion-based video generation with image prompts

    Yuming Jiang, Tianxing Wu, Shuai Yang, Chenyang Si, Dahua Lin, Yu Qiao, Chen Change Loy, and Ziwei Liu. Videobooth: Diffusion-based video generation with image prompts. In 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pages 6689–6700, 2024

  116. [128]

    Lavie: High-quality video generation with cascaded latent diffusion models, 2024

    Yaohui Wang, Xinyuan Chen, Xin Ma, Shangchen Zhou, Ziqi Huang, Yi Wang, Ceyuan Yang, Yinan He, Jiashuo Yu, Peiqing Yang, Yuwei Guo, Tianxing Wu, Chenyang Si, Yuming Jiang, Cunjian Chen, Chen Change Loy, Bo Dai, Dahua Lin, Yu Qiao, and Ziwei Liu. Lavie: High-quality video gener...

  117. [129]

    Ucf101: A dataset of 101 human actions classes from videos in the wild, 2012

    Khurram Soomro, Amir Roshan Zamir, and Mubarak Shah. Ucf101: A dataset of 101 human actions classes from videos in the wild, 2012

  118. [130]

    Quo vadis, action recognition? a new model and the kinetics dataset

    Joao Carreira and Andrew Zisserman. Quo vadis, action recognition? a new model and the kinetics dataset. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , July 2017

  119. [131]

    Youtube-8m: A large-scale video classification benchmark, 2016

    Sami Abu-El-Haija, Nisarg Kothari, Joonseok Lee, Paul Natsev, George Toderici, Balakrishnan Varadarajan, and Sudheendra Vijayanarasimhan. Youtube-8m: A large-scale video classification benchmark, 2016

  120. [132]

    Howto100m: Learning a text-video embedding by watching hundred million narrated video clips

    Antoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi, Ivan Laptev, and Josef Sivic. Howto100m: Learning a text-video embedding by watching hundred million narrated video clips. In 2019 IEEE/CVF International Conference on Computer Vision (ICCV) , pages 2630–2...

  121. [133]

    Msr-vtt: A large video description dataset for bridging video and language

    Jun Xu, Tao Mei, Ting Yao, and Yong Rui. Msr-vtt: A large video description dataset for bridging video and language. In 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , pages 5288–5296, 2016

  122. [134]

    Frozen in Time: A Joint Video and Image Encoder for End-to-End Retrieval

    Max Bain, Arsha Nagrani, Gul Varol, and Andrew Zisserman. Frozen in Time: A Joint Video and Image Encoder for End-to-End Retrieval . In 2021 IEEE/CVF International Conference on Computer Vision (ICCV) , pages 1708–1718, Los Alamitos, CA, USA, October 2021. IEEE Computer Society

  123. [135]

    Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning

    Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut. Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning. In Proceedings of ACL, 2018

  124. [136]

    Internvid: A large-scale video-text dataset for multimodal understanding and generation

    Yi Wang, Yinan He, Yizhuo Li, Kunchang Li, Jiashuo Yu, Xin Ma, Xinhao Li, Guo Chen, Xinyuan Chen, Yaohui Wang, Ping Luo, Ziwei Liu, Yali Wang, Limin Wang, and Yu Qiao. Internvid: A large-scale video-text dataset for multimodal understanding and generation. In The Twelfth Inter...

  125. [137]

    Video-chatgpt: Towards detailed video understanding via large vision and language models, 2024

    Muhammad Maaz, Hanoona Rasheed, Salman Khan, and Fahad Shahbaz Khan. Video-chatgpt: Towards detailed video understanding via large vision and language models, 2024

  126. [138]

    Activitynet: A large-scale video benchmark for human activity understanding

    Fabian Caba Heilbron, Victor Escorcia, Bernard Ghanem, and Juan Carlos Niebles. Activitynet: A large-scale video benchmark for human activity understanding. In 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , pages 961–970, 2015

  127. [139]

    Miradata: A large-scale video dataset with long durations and structured captions

    Xuan Ju, Yiming Gao, Zhaoyang Zhang, Ziyang Yuan, Xintao Wang, Ailing Zeng, Yu Xiong, Qiang Xu, and Ying Shan. Miradata: A large-scale video dataset with long durations and structured captions. In The Thirty-eight Conference on Neural Information Processing Systems Datasets an...

  128. [140]

    A short note about kinetics-600, 2018

    Joao Carreira, Eric Noland, Andras Banki-Horvath, Chloe Hillier, and Andrew Zisserman. A short note about kinetics-600, 2018

  129. [141]

    Panda-70M: Captioning 70M Videos with Multiple Cross-Modality Teachers

    Tsai-Shien Chen, Aliaksandr Siarohin, Willi Menapace, Ekaterina Deyneka, Hsiang-Wei Chao, Byung Eun Jeon, Yuwei Fang, Hsin-Ying Lee, Jian Ren, Ming-Hsuan Yang, and Sergey Tulyakov. Panda-70M: Captioning 70M Videos with Multiple Cross-Modality Teachers . In 2024 IEEE/CVF Confer...

  130. [142]

    Videofactory: Swap attention in spatiotemporal diffusions for text-to-video generation, 2024

    Wenjing Wang, Huan Yang, Zixi Tuo, Huiguo He, Junchen Zhu, Jianlong Fu, and Jiaying Liu. Videofactory: Swap attention in spatiotemporal diffusions for text-to-video generation, 2024

  131. [143]

    Vidprom: A million-scale real prompt-gallery dataset for text-to-video diffusion models

    Wenhao Wang and Yi Yang. Vidprom: A million-scale real prompt-gallery dataset for text-to-video diffusion models. In Proceedings of the 2024 NeurIPS Conference. NeurIPS, 2024. Poster number: 97505

  132. [144]

    Kristen Grauman, Andrew Westbury, Eugene Byrne, Vincent Cartillier, Zachary Chavis, Antonino Furnari, Rohit Girdhar, Jackson Hamburger, Hao Jiang, Devansh Kukreja, Miao Liu, Xingyu Liu, Miguel Martin, Tushar Nagarajan, Ilija Radosavovic, Santhosh Kumar Ramakrishnan, Fiona Ryan...

  133. [145]

    Synchronized video storytelling: Generating video narrations with structured storyline, 2024

    Dingyi Yang, Chunru Zhan, Ziheng Wang, Biao Wang, Tiezheng Ge, Bo Zheng, and Qin Jin. Synchronized video storytelling: Generating video narrations with structured storyline, 2024

  134. [146]

    Benchmarking aigc video quality assessment: A dataset and unified model, 2024

    Zhichao Zhang, Xinyue Li, Wei Sun, Jun Jia, Xiongkuo Min, Zicheng Zhang, Chunyi Li, Zijian Chen, Puyi Wang, Zhongpeng Ji, Fengyu Sun, Shangling Jui, and Guangtao Zhai. Benchmarking aigc video quality assessment: A dataset and unified model, 2024

  135. [147]

    Improved techniques for training gans

    Tim Salimans, Ian Goodfellow, Wojciech Zaremba, Vicki Cheung, Alec Radford, and Xi Chen. Improved techniques for training gans. InProceedings of the 30th International Conference on Neural Information Processing Systems , NIPS’16, page 2234–2242, Red Hook, NY, USA, 2016. Curra...

  136. [148]

    Rethinking the inception architecture for computer vision

    Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna. Rethinking the inception architecture for computer vision. In 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , pages 2818–2826, 2016

  137. [149]

    Gans trained by a two time-scale update rule converge to a local nash equilibrium

    Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. Gans trained by a two time-scale update rule converge to a local nash equilibrium. In Proceedings of the 31st International Conference on Neural Information Processing Systems , NIPS’17,...

  138. [150]

    FVD: A new metric for video generation, 2019

    Thomas Unterthiner, Sjoerd van Steenkiste, Karol Kurach, Raphaël Marinier, Marcin Michalski, and Sylvain Gelly. FVD: A new metric for video generation, 2019

  139. [151]

    Exploring Video Quality Assessment on User Generated Contents from Aesthetic and Technical Perspectives

    Haoning Wu, Erli Zhang, Liang Liao, Chaofeng Chen, Jingwen Hou, Annan Wang, Wenxiu Sun, Qiong Yan, and Weisi Lin. Exploring Video Quality Assessment on User Generated Contents from Aesthetic and Technical Perspectives . In 2023 IEEE/CVF International Conference on Computer Vis...

  140. [152]

    Raft: Recurrent all-pairs field transforms for optical flow

    Zachary Teed and Jia Deng. Raft: Recurrent all-pairs field transforms for optical flow. In Computer Vision – ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part II , page 402–419, Berlin, Heidelberg, 2020. Springer-Verlag

  141. [153]

    Clipscore: A reference-free evaluation metric for image captioning, 2022

    Jack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras, and Yejin Choi. Clipscore: A reference-free evaluation metric for image captioning, 2022

  142. [154]

    Godiva: Generating open-domain videos from natural descriptions, 2021

    Chenfei Wu, Lun Huang, Qianxi Zhang, Binyang Li, Lei Ji, Fan Yang, Guillermo Sapiro, and Nan Duan. Godiva: Generating open-domain videos from natural descriptions, 2021

  143. [155]

    Grit: A generative region-to-text transformer for object understanding, 2022

    Jialian Wu, Jianfeng Wang, Zhengyuan Yang, Zhe Gan, Zicheng Liu, Junsong Yuan, and Lijuan Wang. Grit: A generative region-to-text transformer for object understanding, 2022

  144. [156]

    Fetv: a benchmark for fine-grained evaluation of open-domain text-to-video generation

    Yuanxin Liu, Lei Li, Shuhuai Ren, Rundong Gao, Shicheng Li, Sishuo Chen, Xu Sun, and Lu Hou. Fetv: a benchmark for fine-grained evaluation of open-domain text-to-video generation. In Proceedings of the 37th International Conference on Neural Information Processing Systems , NI...

  145. [157]

    Vbench: Comprehensive benchmark suite for video generative models

    Ziqi Huang, Yinan He, Jiashuo Yu, Fan Zhang, Chenyang Si, Yuming Jiang, Yuanhan Zhang, Tianxing Wu, Qingyang Jin, Nattapol Chanpaisit, Yaohui Wang, Xinyuan Chen, Limin Wang, Dahua Lin, Yu Qiao, and Ziwei Liu. Vbench: Comprehensive benchmark suite for video generative models. I...

  146. [158]

    Subjective-aligned dataset and metric for text-to-video quality assessment, 2024

    Tengchuan Kou, Xiaohong Liu, Zicheng Zhang, Chunyi Li, Haoning Wu, Xiongkuo Min, Guangtao Zhai, and Ning Liu. Subjective-aligned dataset and metric for text-to-video quality assessment, 2024

  147. [159]

    Fréchet video motion distance: A metric for evaluating motion consistency in videos

    Jiahe Liu, Youran Qu, Qi Yan, Xiaohui Zeng, Lele Wang, and Renjie Liao. Fréchet video motion distance: A metric for evaluating motion consistency in videos. In First Workshop on Controllable Video Generation @ICML24 , 2024

  148. [160]

    T2v-compbench: A comprehensive benchmark for compositional text-to-video generation, 2024

    Kaiyue Sun, Kaiyi Huang, Xian Liu, Yue Wu, Zihan Xu, Zhenguo Li, and Xihui Liu. T2v-compbench: A comprehensive benchmark for compositional text-to-video generation, 2024

  149. [161]

    Videopoet: A large language model for zero-shot video generation

    Jiaxuan Guo, Yichun Li, Shangzhe Wang, Yinan Zhang, Xihui Liu, Yu Wang, Hanyang Yang, Jing Yang, and Ziwei Liu. Videopoet: A large language model for zero-shot video generation. arXiv:2312.14125, 2023

  150. [162]

    Make-a-video: Text-to-video generation without text-video data

    Uriel Singer, Adam Polyak, Thomas Hayes, Xi Yin, Junjie An, Songyang Zhang, Qiyuan Hu, Oran Yang, Omri Ashual, Oran Gafni, et al. Make-a-video: Text-to-video generation without text-video data. arXiv:2209.14792, 2022

  151. [163]

    Latent video diffusion models for high-fidelity long video generation, 2023

    Yingqing He, Tianyu Yang, Yong Zhang, Ying Shan, and Qifeng Chen. Latent video diffusion models for high-fidelity long video generation, 2023

  152. [164]

    Vlogger: Generating long videos of dynamic human activities

    Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi. Vlogger: Generating long videos of dynamic human activities. arXiv:2306.04308, 2023

  153. [165]

    Microcinema: A divide-and-conquer approach for text-to-video generation

    Yinan He, Yaohui Wang, Ceyuan Yang, Shangchen Zhou, Xiangyu Zhang, Xiaodong Yang, Yu Qiao, Dahua Lin, and Ying Shan. Microcinema: A divide-and-conquer approach for text-to-video generation. arXiv preprint arXiv:2312.04889, 2023

  154. [167]

    Jonathan Ho, Tim Salimans, Alexey Gritsenko, William Chan, Mohammad Norouzi, and David J. Fleet. Video diffusion models, 2022. Manuscript submitted to ACM Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation 35

  155. [168]

    Nuwa: Visual synthesis pre-training for neural visual world creation

    Chenfei Wu, Jian Liang, Lei Ji, Fan Yang, Yuejian Fang, Daxin Jiang, and Nan Duan. Nuwa: Visual synthesis pre-training for neural visual world creation. arXiv:2111.12417, 2021

  156. [169]

    Temporal generative adversarial nets with singular value clipping, 2017

    Masaki Saito, Eiichi Matsumoto, and Shunta Saito. Temporal generative adversarial nets with singular value clipping, 2017

  157. [170]

    Lvt: Language-vision transformer for multi-modal video understanding

    Qingqiu Huang, Wentao Yu, Yuanze Xu, Yitong Wang, and Dacheng Zhang. Lvt: Language-vision transformer for multi-modal video understanding. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) , page , 2021

  158. [171]

    Generating videos with scene dynamics

    Carl Vondrick, Hamed Pirsiavash, and Antonio Torralba. Generating videos with scene dynamics. In NeurIPS, 2016

  159. [172]

    Videofusion: Decomposed diffusion models for high-quality video generation

    Zhengxiong Luo, Dayou Chen, Yingya Zhang, Yan Huang, Liang Wang, Yujun Shen, Deli Zhao, Jingren Zhou, and Tieniu Tan. Videofusion: Decomposed diffusion models for high-quality video generation. CVPR, 2023

  160. [174]

    Magicvideo: Efficient video generation with latent diffusion models

    Rui Zhao, Yuxiang Wu, Hao Dong, Ning Zhang, Tao Yang, Wei Wei, and Xiaowei Huang. Magicvideo: Efficient video generation with latent diffusion models. arXiv:2301.11093, 2023

  161. [175]

    Pyoco: Latent diffusion priors for zero-shot video editing

    Andreas Blattmann, Robin Rombach, Huan Ling, Tim Dockhorn, Seung Wook Kim, Sanja Fidler, and Karsten Kreis. Pyoco: Latent diffusion priors for zero-shot video editing. arXiv:2303.04734, 2023

  162. [176]

    Videofactory: Synthesizing high-quality video with diffusion models

    Yaohui Wang, Xinyuan Chen, Xin Ma, Shangchen Zhou, Ziwei Huang, Yu Wang, Ceyuan Yang, Yinan He, Jiashuo Yu, Yinan Yang, et al. Videofactory: Synthesizing high-quality video with diffusion models. arXiv:2305.10874, 2023

  163. [178]

    Lavie: A layered video diffusion model with multi-modal conditioning

    Yue Ma, Yingqing He, Xiaodong Cun, Xintao Wang, Ying Shan, Xiu Li, and Qifeng Chen. Lavie: A layered video diffusion model with multi-modal conditioning. arXiv:2309.15130, 2023

  164. [180]

    Internvid: A large-scale video-text dataset for multimodal understanding and generation

    Yaohui Wang, Xinyuan Chen, Xin Ma, Shangchen Zhou, Ziwei Huang, Yu Wang, Ceyuan Yang, Yinan He, Jiashuo Yu, Yinan Yang, et al. Internvid: A large-scale video-text dataset for multimodal understanding and generation. arXiv:2307.06942, 2023

  165. [181]

    Make pixels dance: High-dynamic video generation, 2023

    Yan Zeng, Guoqiang Wei, Jiani Zheng, Jiaxin Zou, Yang Wei, Yuchen Zhang, and Hang Li. Make pixels dance: High-dynamic video generation, 2023

  166. [182]

    Emu video: Factorizing text-to-video generation by explicit image conditioning, 2024

    Rohit Girdhar, Mannat Singh, Andrew Brown, Quentin Duval, Samaneh Azadi, Sai Saketh Rambhatla, Akbar Shah, Xi Yin, Devi Parikh, and Ishan Misra. Emu video: Factorizing text-to-video generation by explicit image conditioning, 2024

  167. [183]

    Lumiere: A space-time diffusion model for video generation

    Omer Bar-Tal, Dolev Ofri-Amar, Rafail Fridman, Yoni Kasten, and Tali Dekel. Lumiere: A space-time diffusion model for video generation. arXiv preprint arXiv:2401.12945, 2024

  168. [184]

    Llmscore: Unveiling the power of large language models in text-to-image synthesis evaluation, 2023

    Yujie Lu, Xianjun Yang, Xiujun Li, Xin Eric Wang, and William Yang Wang. Llmscore: Unveiling the power of large language models in text-to-image synthesis evaluation, 2023

  169. [185]

    Lavie: High-quality video generation with cascaded latent diffusion models

    Yaohui Wang, Xinyuan Chen, Xin Ma, Shangchen Zhou, Ziqi Huang, Yi Wang, Ceyuan Yang, Yinan He, Jiashuo Yu, Peiqing Yang, et al. Lavie: High-quality video generation with cascaded latent diffusion models. arXiv preprint arXiv:2309.15103, 2023

  170. [186]

    Modelscope text-to-video technical report.arXiv preprint arXiv:2308.06571, 2023

    Jiuniu Wang, Hangjie Yuan, Dayou Chen, Yingya Zhang, Xiang Wang, and Shiwei Zhang. Modelscope text-to-video technical report.arXiv preprint arXiv:2308.06571, 2023

  171. [187]

    Videocrafter2: Overcoming data limitations for high-quality video diffusion models, 2024

    Haoxin Chen, Yong Zhang, Xiaodong Cun, Menghan Xia, Xintao Wang, Chao Weng, and Ying Shan. Videocrafter2: Overcoming data limitations for high-quality video diffusion models, 2024

  172. [189]

    Animatediff: Animate your personalized text-to-image diffusion models without specific tuning

    Yuwei Guo, Ceyuan Yang, Anyi Rao, Yaohui Wang, Yu Qiao, Dahua Lin, and Bo Dai. Animatediff: Animate your personalized text-to-image diffusion models without specific tuning. 2024

  173. [190]

    Latte: Latent diffusion transformer for video generation

    Xin Ma, Yaohui Wang, Gengyun Jia, Xinyuan Chen, Ziwei Liu, Yuan-Fang Li, Cunjian Chen, and Yu Qiao. Latte: Latent diffusion transformer for video generation. arXiv preprint arXiv:2401.03048, 2024

  174. [191]

    Accessed September 25, 2023 [Online] https://www.pika.art/, 2023

    Pika labs. Accessed September 25, 2023 [Online] https://www.pika.art/, 2023

  175. [192]

    Accessed June 6, 2024 [Online] https://klingai.kuaishou.com/, 2024

    Kling. Accessed June 6, 2024 [Online] https://klingai.kuaishou.com/, 2024

  176. [193]

    Cogvideox: Text-to-video diffusion models with an expert transformer

    Zhuoyi Yang, Jiayan Teng, Wendi Zheng, Ming Ding, Shiyu Huang, Jiazheng Xu, Yuanming Yang, Wenyi Hong, Xiaohan Zhang, Guanyu Feng, et al. Cogvideox: Text-to-video diffusion models with an expert transformer. arXiv preprint arXiv:2408.06072, 2024. Manuscript submitted to ACM

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.