REVIEW 4 major objections 7 minor 187 references
Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation
T0 review · 4 major / 7 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read This survey claims that long video generation is best understood as three paradigms—autoregressive, divide-and-conquer, and implicit latent-space synthesis—and argues that divide-and-conquer, especially LLM-guided planning, is the key to…
desk verdict A broad, useful survey that is let down by unreliable benchmark tables and attribution errors; worth revising, not rejecting. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central organizing mechanism is the divide-and-conquer paradigm: generate keyframes or short clips from prompts, often planned by an LLM, then interpolate or stitch them into a continuous long video. The LLM-as-director pattern is the load-bearing instance, in which an LLM produces a narrative blueprint (scene descriptions, layouts, bounding boxes, actions) that a separate video diffusion module executes. This is contrasted with autoregressive prediction, which conditions each frame on previous frames, and implicit generation, which synthesizes the whole video from a compressed spatiotemporal latent representation.
What would settle it
A systematic literature search for long video generation papers published between 2021 and mid-2025 that are absent from this survey and do not fit any of the three paradigms (for example, direct world-simulator models that generate video without planning, sequential prediction, or compressed latent-space synthesis) would refute the claim that the taxonomy is exhaustive. A reproducible annotation study measuring the fraction of relevant papers that independent annotators cleanly assign to one of the three paradigms would likewise test whether the partition is well-defined.
Extended reading notes
Core claim
On its own terms, the paper establishes that long video generation can be systematically mapped into three paradigms: autoregressive frame prediction, divide-and-conquer keyframe-and-interpolation, and implicit synthesis from a compressed latent space. Within divide-and-conquer, it identifies three sub-patterns: LLM-as-director, multi-stage or agent-based frameworks, and transition or compositional stitching. The paper further claims that this taxonomy, together with its catalog of datasets and evaluation metrics, provides a comprehensive and current foundation that earlier surveys lacked, particularly in its detailed treatment of divide-and-conquer.
Load-bearing premise
The survey's comprehensiveness rests on the assumption that snowball sampling of 190+ articles from a manually selected list of venues captures a representative and complete picture of the field, and that the three-way partition of generation paradigms is exhaustive.
Editorial extensions
If this is right
- If the survey's map is correct, future long video models will increasingly separate a planning stage from a frame-generation stage, since divide-and-conquer supports parallel keyframe generation and finer narrative control.
- The identified scarcity of large-scale, richly captioned video datasets becomes a concrete bottleneck; building datasets that combine scale with spatial and temporal detail would directly accelerate progress.
- Metrics that rely on manual feedback, such as FETV, VBench, and MiraBench, are a scalability bottleneck, making fully automated semantic and temporal evaluation a clear research target.
- The divide-and-conquer sub-taxonomy defines a modular design space—planner, generator, transition module—that is testable by comparing how well different combinations reconstruct the same storylines.
- The survey's open challenges, including audio alignment and physical dynamics, point to next steps that are largely orthogonal to the core generation paradigm and can be pursued on top of any of the three approaches.
Reading between the lines
- A consequence the paper leaves implicit: if the three-paradigm taxonomy is accepted, the implicit category is likely to absorb most future foundation models, while divide-and-conquer will dominate controllable and story-driven applications; the two may eventually converge.
- The paper's emphasis on divide-and-conquer is a bet on the value of explicit planning; one could test whether planning-based methods actually beat end-to-end latent synthesis on long-range coherence when compute budgets are held equal.
- A practical extension: use the survey's taxonomy to build a benchmark that samples methods from each paradigm and evaluates them on identical prompts and durations; the results would either validate the taxonomy's utility or reveal overlaps that blur its boundaries.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This survey reviews the emerging area of long video generation, organizing the literature into three generation paradigms (autoregressive, divide-and-conquer, and implicit latent-space generation) and covering backbone architectures, tokenization strategies, input control mechanisms, datasets, and evaluation metrics. The authors state in the abstract and in Section 1.1 that the survey 'would serve as a comprehensive foundation' for the field, and they position their contribution against two earlier surveys [18,19] by focusing on divide-and-conquer methods, agent-based frameworks, and short-to-long video transitions. The paper catalogs over 190 papers from 2021 to 2025, with timeline figures and several comparison tables.
Significance. If the cataloging is accurate, the survey would be a useful entry point for researchers: it gathers recent work (2023-2025) that earlier surveys do not cover in depth, and its three-paradigm taxonomy, particularly the divide-and-conquer subsection, is a reasonable high-level organization of the field. The paper also draws attention to and compares datasets and metrics, which are often treated separately in the literature. However, the central value of a survey of this kind is reliability of its factual claims; the inconsistencies in the taxonomy and in the benchmark tables described below mean that the 'comprehensive foundation' claim is not currently supported by the manuscript as written. The paper contains no derivations or experiments, so its correctness hangs entirely on the accuracy of its reporting.
major comments (4)
- [Section 3.1, Table 1] Table 1, which is labeled 'Auto Regressive Approaches,' lists StyleGAN-V [85] and DIGAN [40] as autoregressive methods. This contradicts the paper's own description in Section 2.1.2, where both are presented as continuous/implicit GAN generators, and it also conflicts with the definition of implicit video generation in Section 3.3, which explicitly excludes extrapolation (autoregressive) and interpolation (divide-and-conquer). As a result, the three-paradigm partition in Section 3 is not clean or exhaustive, and the reader cannot determine which papers belong to which paradigm. Please reclassify these models or justify their placement in the autoregressive category.
- [Table 7] Table 7, titled 'Comprehensive benchmark comparison,' mixes FVD numbers from different evaluation protocols (UCF-101, BAIR) with FID and CLIPSIM numbers from MSR-VTT, and the footnote does not resolve which cell comes from which benchmark. The table also contains duplicate entries: Phenaki appears as both [15] and [63], ModelScope and ModelScopeT2V both cite [177], VideoDirectorGPT appears as both [11] and [166], and Video LDM [179] is the same work as Video Diffusion [173]. Without per-row benchmark provenance and deduplication, these numbers cannot be verified or compared. Please provide a source for every reported value and remove the duplicate rows.
- [Table 8] Table 8 reports per-dimension VBench scores that are largely in the 80-99 range (e.g., LaVie: Subject Consistency 91.41, Background Consistency 97.47, Motion 96.38) but lists Overall scores around 26-28 for the same models (e.g., LaVie Overall 26.41). Since the caption states that the table evaluates models across 12 VBench dimensions, the Overall score should be derivable from those dimensions; the reported values appear to be from a different, unlabeled benchmark. In addition, Gen-2 is cited to [188], which is the RunwayML Gen-3 Alpha page, and references [13] and [17] attribute Gen-2 and Gen-4 Alpha to 'Midjourney Team,' contradicting the running text in Section 1, which identifies both as RunwayML models. These provenance errors prevent readers from trusting the table.
- [Section 3.1, ARLON description] The text states that ARLON [84] achieves '128× compression (8× spatial + 6× temporal downsampling),' but 8 multiplied by 6 is 48, not 128. Either the compression factor or the downsampling factors are incorrect. Please verify this claim against the ARLON paper and correct the numbers.
minor comments (7)
- [Section 1.4] The survey organization paragraph contains placeholder cross-references 'Section datasets' and 'Section metrics' that do not correspond to any section, and Section 1.2 uses bracket-style references such as '[3.2]' and '[5.1]' that do not match the manuscript's section numbering.
- [References] The reference list contains duplicate entries for the same works, including Phenaki (appearing as [15] and [63]), MEVG (as [101] and [104]), VideoDirectorGPT (as [11] and [166]), and Video Diffusion/Video LDM (as [173] and [179]). These should be consolidated.
- [Section 6.2] The sentence 'Here are examples from the MSR-VTT dataset...' appears twice verbatim, once in the main text and once immediately after the figure caption for Figure 17.
- [Section 3.3] The sentence about Cosmos cites reference [84], which is the ARLON paper; the Cosmos World Foundation Model is separately cited as [113] later in the same section. The earlier citation appears to be a mismatch.
- [Section 2.2.1] Reference [51], an anomaly-detection paper by one of the authors, is cited in the autoencoder background section. The connection to video generation is not explained, and it is not used to support any of the survey's central claims; consider removing it or providing a clearer rationale.
- [Section 7] The introduction to Section 7 contains stray LaTeX commands: 'textit Video Quality Metrics' and 'textit Semantics Quality Metrics'. Additionally, Table 5 lists 'Pandas 70m' while the reference [141] is titled 'Panda-70M'; please make the names consistent.
- [Figure 3] The caption of Figure 3 claims that most papers on long video generation were published in 2023-2025, but no data source or counting methodology is provided for the histogram, so the claim cannot be verified.
Circularity Check
No circularity: the survey makes no fitted predictions or derivations, and its only self-citation is peripheral and non-load-bearing.
full rationale
This is a literature survey, not a derivation or prediction paper. The abstract's claim that the survey 'would serve as a comprehensive foundation' is a scope statement about cataloging external work, not a mathematical or empirical result derived from the paper's own inputs. The survey's content consists of summaries of third-party methods, datasets, and benchmarks, with citations to the primary sources (e.g., StyleGAN-V, DIGAN, Phenaki, VideoDirectorGPT, Sora, VBench). No parameter is fitted and no quantity is predicted from another quantity within the paper, so none of the enumerated circularity patterns applies. The one self-citation, reference [51] in Section 2.2.1, is an earlier anomaly-detection paper by the authors cited only as an example of LSTM-Convolutional VAEs; it does not support the survey's central taxonomy, benchmark tables, or any comparative claim, and removing it would not change any conclusion. Concerns about Table 1 misclassifications, Table 7 duplicated/unverifiable entries, and Table 8 provenance errors are factual-accuracy and verification issues for a survey, not circular reasoning: the survey is not defining its categories in terms of its own conclusions or citing itself to force a result. The paper is self-contained as a review of external literature, and the honest finding is no significant circularity.
Assumptions & free parameters
assumptions (4)
- domain assumption Snowball sampling of over 190 articles from a selected set of venues yields a representative coverage of long video generation.
- domain assumption The three-way partition (auto-regressive, divide-and-conquer, implicit) is exhaustive and non-overlapping.
- domain assumption There are only two related long-video surveys ([18], [19]).
- domain assumption Metric values in Tables 7 and 8 are correctly transcribed from the cited papers.
Cite this review
Pith. "Pith review of Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation." pith.science (2026). https://pith.science/paper/6WK6XSZH
@misc{pith2026241218688,
author = {Pith},
title = {Pith review of: Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation},
year = {2026},
howpublished = {\url{https://pith.science/paper/6WK6XSZH}},
note = {Machine review of arXiv:2412.18688}
}
read the original abstract
An image may convey a thousand words, but a video composed of hundreds or thousands of image frames tells a more intricate story. Despite significant progress in multimodal large language models (MLLMs), generating extended videos remains a formidable challenge. As of this writing, OpenAI's Sora, the current state-of-the-art system, is still limited to producing videos that are up to one minute in length. This limitation stems from the complexity of long video generation, which requires more than generative AI techniques for approximating density functions essential aspects such as planning, story development, and maintaining spatial and temporal consistency present additional hurdles. Integrating generative AI with a divide-and-conquer approach could improve scalability for longer videos while offering greater control. In this survey, we examine the current landscape of long video generation, covering foundational techniques like GANs and diffusion models, video generation strategies, large-scale training datasets, quality metrics for evaluating long videos, and future research areas to address the limitations of the existing video generation capabilities. We believe it would serve as a comprehensive foundation, offering extensive information to guide future advancements and research in the field of long video generation.
Figures
Figures from the paper (18 more)
Reference graph
Works this paper leans on
-
[85]
Stylegan-v: A continuous video generator with the price, image quality and perks of stylegan2
Ivan Skorokhodov, Sergey Tulyakov, and Mohamed Elhoseiny. Stylegan-v: A continuous video generator with the price, image quality and perks of stylegan2. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pages 3626–3636, June 2022
2022
-
[40]
Generating videos with dynamics-aware implicit generative adversarial networks
Sihyun Yu, Jihoon Tack, Sangwoo Mo, Hyunsu Kim, Junho Kim, Jung-Woo Ha, and Jinwoo Shin. Generating videos with dynamics-aware implicit generative adversarial networks. In International Conference on Learning Representations , 2022
2022
-
[63]
Phenaki: Variable length video generation from open domain textual descriptions
Ruben Villegas, Mohammad Babaeizadeh, Pieter-Jan Kindermans, Hernan Moraldo, Han Zhang, Mohammad Taghi Saffar, Santiago Castro, Julius Kunze, and Dumitru Erhan. Phenaki: Variable length video generation from open domain textual descriptions. In International Conference on Learning Representations, 2023
2023
-
[11]
VideodirectorGPT: Consistent multi-scene video generation via LLM-guided planning, 2024
Han Lin, Abhay Zala, Jaemin Cho, and Mohit Bansal. VideodirectorGPT: Consistent multi-scene video generation via LLM-guided planning, 2024
2024
-
[166]
Videodirectorgpt: Consistent multi-scene video generation via llm-guided planning
Xi Chen, Yaohui Wang, Shangchen Zhou, Ceyuan Yang, Yinan Zhang, Yizhou He, Yu Wang, Haoxin Yang, Ziwei Huang, Xiaodong Lu, et al. Videodirectorgpt: Consistent multi-scene video generation via llm-guided planning. arXiv:2309.15091, 2023
arXiv 2023
-
[179]
Videoldm: High-resolution video generation with latent diffusion models
Andreas Blattmann, Robin Rombach, Huan Ling, Tim Dockhorn, Seung Wook Kim, Sanja Fidler, and Karsten Kreis. Videoldm: High-resolution video generation with latent diffusion models. arXiv:2304.08818, 2023
arXiv 2023
-
[173]
Jonathan Ho, Tim Salimans, Alexey Gritsenko, William Chan, Mohammad Norouzi, and David J Fleet. Video diffusion models. arXiv preprint arXiv:2204.03458, 2022
arXiv 2022
-
[188]
Accessed June 17, 2024 [Online] https://runwayml.com/research/introducing-gen-3-alpha, 2024
Gen-3. Accessed June 17, 2024 [Online] https://runwayml.com/research/introducing-gen-3-alpha, 2024
work page 2024
-
[13]
Gen2 by runway ml, 2023
Midjourney Team. Gen2 by runway ml, 2023
2023
-
[17]
Gen-4 alpha by midjourney, 2024
Midjourney Team. Gen-4 alpha by midjourney, 2024
2024
-
[84]
Arlon: Boosting diffusion transformers with autoregressive models for long video generation, 2025
Zongyi Li, Shujie Hu, Shujie Liu, Long Zhou, Jeongsoo Choi, Lingwei Meng, Xun Guo, Jinyu Li, Hefei Ling, and Furu Wei. Arlon: Boosting diffusion transformers with autoregressive models for long video generation, 2025
2025
Show all 187 references
-
[1]
Video generation models as world simulators by open a.i, 2024
Sora Team. Video generation models as world simulators by open a.i, 2024
2024
-
[2]
Introducing chatgpt by open a.i, 2022
Oepn A.I Team. Introducing chatgpt by open a.i, 2022
2022
-
[3]
Meta llama models, 2023
Meta A.I Team. Meta llama models, 2023
2023
-
[4]
Google gemini series, 2023
Google A.I Team. Google gemini series, 2023
2023
-
[5]
Anthropic by claude, 2023
Claude A.I Team. Anthropic by claude, 2023
2023
-
[6]
Mistral large model by mistral, 2024
Mistral A.I Team. Mistral large model by mistral, 2024
2024
-
[7]
Hierarchical text-conditional image generation with clip latents, 2022
Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. Hierarchical text-conditional image generation with clip latents, 2022
2022
-
[8]
Scaling rectified flow transformers for high-resolution image synthesis, 2024
Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas Müller, Harry Saini, Yam Levi, Dominik Lorenz, Axel Sauer, Frederic Boesel, Dustin Podell, Tim Dockhorn, Zion English, Kyle Lacey, Alex Goodwin, Yannik Marek, and Robin Rombach. Scaling rectified flow transf...
2024
-
[9]
The midjourney v5.2 model for image generation, 2024
Midjourney A.I Team. The midjourney v5.2 model for image generation, 2024
2024
-
[12]
Make-a-video: Text-to-video generation without text-video data
Uriel Singer, Adam Polyak, Thomas Hayes, Xi Yin, Jie An, Songyang Zhang, Qiyuan Hu, Harry Yang, Oron Ashual, Oran Gafni, Devi Parikh, Sonal Gupta, and Yaniv Taigman. Make-a-video: Text-to-video generation without text-video data. In The Eleventh International Conference on Lea...
2023
-
[14]
Cogvideo: Large-scale pretraining for text-to-video generation via transformers
Wenyi Hong, Ming Ding, Wendi Zheng, Xinghan Liu, and Jie Tang. Cogvideo: Large-scale pretraining for text-to-video generation via transformers. In The Eleventh International Conference on Learning Representations , 2023
2023
-
[16]
Gen-l-video: Multi-text to long video generation via temporal co-denoising, 2023
Fu-Yun Wang, Wenshuo Chen, Guanglu Song, Han-Jia Ye, Yu Liu, and Hongsheng Li. Gen-l-video: Multi-text to long video generation via temporal co-denoising, 2023. Manuscript submitted to ACM Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation 29
2023
-
[18]
A survey on long video generation: Challenges, methods, and prospects, 2024
Chengxuan Li, Di Huang, Zeyu Lu, Yang Xiao, Qingqi Pei, and Lei Bai. A survey on long video generation: Challenges, methods, and prospects, 2024
2024
-
[19]
A survey on generative ai and llm for video generation, understanding, and streaming, 2024
Pengyuan Zhou, Lin Wang, Zhi Liu, Yanbin Hao, Pan Hui, Sasu Tarkoma, and Jussi Kangasharju. A survey on generative ai and llm for video generation, understanding, and streaming, 2024
2024
-
[20]
Sora: A review on background, technology, limitations, and opportunities of large vision models, 2024
Yixin Liu, Kai Zhang, Yuan Li, Zhiling Yan, Chujie Gao, Ruoxi Chen, Zhengqing Yuan, Yue Huang, Hanchi Sun, Jianfeng Gao, Lifang He, and Lichao Sun. Sora: A review on background, technology, limitations, and opportunities of large vision models, 2024
2024
-
[21]
Free-bloom: Zero-shot text-to-video generator with LLM director and LDM animator
Hanzhuo Huang, Yufan Feng, Cheng Shi, Lan Xu, Jingyi Yu, and Sibei Yang. Free-bloom: Zero-shot text-to-video generator with LLM director and LDM animator. In Thirty-seventh Conference on Neural Information Processing Systems , 2023
2023
-
[22]
Flowzero: Zero-shot text-to-video synthesis with llm-driven dynamic scene syntax, 2023
Yu Lu, Linchao Zhu, Hehe Fan, and Yi Yang. Flowzero: Zero-shot text-to-video synthesis with llm-driven dynamic scene syntax, 2023
2023
-
[23]
Align your latents: High- resolution video synthesis with latent diffusion models
Andreas Blattmann, Robin Rombach, Huan Ling, Tim Dockhorn, Seung Wook Kim, Sanja Fidler, and Karsten Kreis. Align your latents: High- resolution video synthesis with latent diffusion models. In 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pages ...
2023
-
[24]
Mavin: Multi-action video generation with diffusion models via transition video infilling, 2024
Bowen Zhang, Xiaofei Xie, Haotian Lu, Na Ma, Tianlin Li, and Qing Guo. Mavin: Multi-action video generation with diffusion models via transition video infilling, 2024
2024
-
[25]
SEINE: Short-to-long video diffusion model for generative transition and prediction
Xinyuan Chen, Yaohui Wang, Lingjun Zhang, Shaobin Zhuang, Xin Ma, Jiashuo Yu, Yali Wang, Dahua Lin, Yu Qiao, and Ziwei Liu. SEINE: Short-to-long video diffusion model for generative transition and prediction. In The Twelfth International Conference on Learning Representations , 2024
2024
-
[26]
Learning transferable visual models from natural language supervision, 2021
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning transferable visual models from natural language supervision, 2021
2021
-
[27]
Generative adversarial networks
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. Generative adversarial networks. Commun. ACM, 63(11):139–144, October 2020
2020
-
[28]
Unsupervised representation learning with deep convolutional generative adversarial networks, 2016
Alec Radford, Luke Metz, and Soumith Chintala. Unsupervised representation learning with deep convolutional generative adversarial networks, 2016
2016
-
[29]
Deep generative image models using a laplacian pyramid of adversarial networks
Emily Denton, Soumith Chintala, Arthur Szlam, and Rob Fergus. Deep generative image models using a laplacian pyramid of adversarial networks. In Proceedings of the 28th International Conference on Neural Information Processing Systems - Volume 1 , NIPS’15, page 1486–1494, Camb...
2015
-
[30]
Stackgan: Text to photo-realistic image synthesis with stacked generative adversarial networks
Han Zhang, Tao Xu, Hongsheng Li, Shaoting Zhang, Xiaogang Wang, Xiaolei Huang, and Dimitris Metaxas. Stackgan: Text to photo-realistic image synthesis with stacked generative adversarial networks. In 2017 IEEE International Conference on Computer Vision (ICCV) , pages 5908–5916, 2017
2017
-
[31]
Attngan: Fine-grained text to image generation with attentional generative adversarial networks
Tao Xu, Pengchuan Zhang, Qiuyuan Huang, Han Zhang, Zhe Gan, Xiaolei Huang, and Xiaodong He. Attngan: Fine-grained text to image generation with attentional generative adversarial networks. In 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 1316–1324, 2018
2018
-
[32]
Analyzing and Improving the Image Quality of StyleGAN
Tero Karras, Samuli Laine, Miika Aittala, Janne Hellsten, Jaakko Lehtinen, and Timo Aila. Analyzing and Improving the Image Quality of StyleGAN . In 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pages 8107–8116, Los Alamitos, CA, USA, June 2020. ...
2020
-
[33]
Jun-Yan Zhu, Taesung Park, Phillip Isola, and Alexei A. Efros. Unpaired image-to-image translation using cycle-consistent adversarial networks. In 2017 IEEE International Conference on Computer Vision (ICCV) , pages 2242–2251, 2017
2017
-
[34]
Stylegan2 distillation for feed-forward image manipulation
Yuri Viazovetskyi, Vladimir Ivashkin, and Evgeny Kashin. Stylegan2 distillation for feed-forward image manipulation. In Computer Vision – ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XXII , page 170–186, Berlin, Heidelberg, 2020. Spri...
2020
-
[35]
Phillip Isola, Jun-Yan Zhu, Tinghui Zhou, and Alexei A. Efros. Image-to-image translation with conditional adversarial networks. In 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , pages 5967–5976, 2017
2017
-
[36]
Deep multi-scale video prediction beyond mean square error, 2016
Michael Mathieu, Camille Couprie, and Yann LeCun. Deep multi-scale video prediction beyond mean square error, 2016
2016
-
[37]
Generating videos with scene dynamics, 2016
Carl Vondrick, Hamed Pirsiavash, and Antonio Torralba. Generating videos with scene dynamics, 2016
2016
-
[38]
Learning to generate time-lapse videos using multi-stage dynamic generative adversarial networks
Wei Xiong, Wenhan Luo, Lin Ma, Wei Liu, and Jiebo Luo. Learning to generate time-lapse videos using multi-stage dynamic generative adversarial networks. In 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 2364–2373, 2018
2018
-
[39]
To create what you tell: Generating videos from captions, 2018
Yingwei Pan, Zhaofan Qiu, Ting Yao, Houqiang Li, and Tao Mei. To create what you tell: Generating videos from captions, 2018
2018
-
[41]
Stylegan-v: A continuous video generator with the price, image quality and perks of stylegan2
Ivan Skorokhodov, Sergey Tulyakov, and Mohamed Elhoseiny. Stylegan-v: A continuous video generator with the price, image quality and perks of stylegan2. In 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pages 3616–3626, 2022
2022
-
[42]
Auto-encoders
Standford tutorial. Auto-encoders
-
[43]
Auto-encoding variational bayes, 2022
Diederik P Kingma and Max Welling. Auto-encoding variational bayes, 2022
2022
-
[44]
Masked autoencoders are scalable vision learners
Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, and Ross Girshick. Masked autoencoders are scalable vision learners. In 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pages 15979–15988, 2022
2022
-
[45]
High-Resolution Image Synthesis with Latent Diffusion Models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bjorn Ommer. High-Resolution Image Synthesis with Latent Diffusion Models . In 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pages 10674–10685, Los Alamitos, CA, USA, June 2022....
2022
-
[46]
Neural discrete representation learning, 2018
Aaron van den Oord, Oriol Vinyals, and Koray Kavukcuoglu. Neural discrete representation learning, 2018
2018
-
[47]
Videogpt: Video generation using vq-vae and transformers, 2021
Wilson Yan, Yunzhi Zhang, Pieter Abbeel, and Aravind Srinivas. Videogpt: Video generation using vq-vae and transformers, 2021
2021
-
[48]
Taming Transformers for High-Resolution Image Synthesis
Patrick Esser, Robin Rombach, and Bjorn Ommer. Taming Transformers for High-Resolution Image Synthesis . In 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pages 12868–12878, Los Alamitos, CA, USA, June 2021. IEEE Computer Society
2021
-
[49]
Clip: Connecting vision and language with contrastive learning
Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever. Clip: Connecting vision and language with contrastive learning. In Proceedings of the 38th International Conference on Machine Learning , volume 139, pages 8821–8...
2021
-
[50]
Hierarchical patch vae-gan: Generating diverse videos from a single sample, 2020
Shir Gur, Sagie Benaim, and Lior Wolf. Hierarchical patch vae-gan: Generating diverse videos from a single sample, 2020
2020
-
[51]
Visual anomaly detection in video by variational autoencoder, 2022
Faraz Waseem, Rafael Perez Martinez, and Chris Wu. Visual anomaly detection in video by variational autoencoder, 2022
2022
-
[52]
VideoMAC: Video Masked Autoencoders Meet ConvNets
Gensheng Pei, Tao Chen, Xiruo Jiang, Huafeng Liu, Zeren Sun, and Yazhou Yao. VideoMAC: Video Masked Autoencoders Meet ConvNets . In 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pages 22733–22743, Los Alamitos, CA, USA, June 2024. IEEE Computer Society
2024
-
[53]
Hauptmann, Ming-Hsuan Yang, Yuan Hao, Irfan Essa, and Lu Jiang
Lijun Yu, Yong Cheng, Kihyuk Sohn, José Lezama, Han Zhang, Huiwen Chang, Alexander G. Hauptmann, Ming-Hsuan Yang, Yuan Hao, Irfan Essa, and Lu Jiang. Magvit: Masked generative video transformer. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogniti...
2023
-
[54]
MAGVLT: Masked Generative Vision-and-Language Transformer
Sungwoong Kim, Daejin Jo, Donghoon Lee, and Jongmin Kim. MAGVLT: Masked Generative Vision-and-Language Transformer . In 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pages 23338–23348, Los Alamitos, CA, USA, June 2023. IEEE Computer Society
2023
-
[55]
Multi-generator generative adversarial nets, 2017
Quan Hoang, Tu Dinh Nguyen, Trung Le, and Dinh Phung. Multi-generator generative adversarial nets, 2017
2017
-
[56]
Gomez, Łukasz Kaiser, and Illia Polosukhin
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Proceedings of the 31st International Conference on Neural Information Processing Systems , NIPS’17, page 6000–6010, Red ...
2017
-
[57]
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition at...
2021
-
[58]
Zero-shot text-to-image generation
Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever. Zero-shot text-to-image generation. In Marina Meila and Tong Zhang, editors, Proceedings of the 38th International Conference on Machine Learning , volume 139 of ...
2021
-
[59]
Discrete variational autoencoders, 2017
Jason Tyler Rolfe. Discrete variational autoencoders, 2017
2017
-
[60]
Cogview: mastering text-to-image generation via transformers
Ming Ding, Zhuoyi Yang, Wenyi Hong, Wendi Zheng, Chang Zhou, Da Yin, Junyang Lin, Xu Zou, Zhou Shao, Hongxia Yang, and Jie Tang. Cogview: mastering text-to-image generation via transformers. In Proceedings of the 35th International Conference on Neural Information Processing S...
2024
-
[61]
Cogview2: faster and better text-to-image generation via hierarchical transformers
Ming Ding, Wendi Zheng, Wenyi Hong, and Jie Tang. Cogview2: faster and better text-to-image generation via hierarchical transformers. In Proceedings of the 36th International Conference on Neural Information Processing Systems , NIPS ’22, Red Hook, NY, USA, 2024. Curran Associates Inc
2024
-
[62]
ViViT: A Video Vision Transformer
Anurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun, Mario Lucic, and Cordelia Schmid. ViViT: A Video Vision Transformer . In 2021 IEEE/CVF International Conference on Computer Vision (ICCV) , pages 6816–6826, Los Alamitos, CA, USA, October 2021. IEEE Computer Society
2021
-
[64]
Adam Roberts, Hyung Won Chung, Anselm Levskaya, Gaurav Mishra, James Bradbury, Daniel Andor, Sharan Narang, Brian Lester, Colin Gaffney, Afroz Mohiuddin, Curtis Hawthorne, Aitor Lewkowycz, Alex Salcianu, Marc van Zee, Jacob Austin, Sebastian Goodman, Livio Baldini Soares, Hait...
2022
-
[65]
Compositional 3d-aware video generation with llm director
Hanxin Zhu, Tianyu He, Anni Tang, Junliang Guo, Zhibo Chen, and Jiang Bian. Compositional 3d-aware video generation with llm director. In Proceedings of the 38th Annual Conference on Neural Information Processing Systems (NeurIPS 2024) , 2024. Poster presentation
2024
-
[66]
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffr...
2020
-
[68]
Vlogger: Make your dream a vlog
Shaobin Zhuang, Kunchang Li, Xinyuan Chen, Yaohui Wang, Ziwei Liu, Yu Qiao, and Yali Wang. Vlogger: Make your dream a vlog. In 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pages 8806–8817, 2024. Manuscript submitted to ACM Video Is Worth a Thous...
2024
-
[69]
Weiss, Niru Maheswaranathan, and Surya Ganguli
Jascha Sohl-Dickstein, Eric A. Weiss, Niru Maheswaranathan, and Surya Ganguli. Deep unsupervised learning using nonequilibrium thermody- namics, 2015
2015
-
[70]
Denoising diffusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. In Proceedings of the 34th International Conference on Neural Information Processing Systems, NIPS ’20, Red Hook, NY, USA, 2020. Curran Associates Inc
2020
-
[71]
Generative modeling by estimating gradients of the data distribution, 2020
Yang Song and Stefano Ermon. Generative modeling by estimating gradients of the data distribution, 2020
2020
-
[72]
Bert: Pre-training of deep bidirectional transformers for language understanding, 2019
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding, 2019
2019
-
[73]
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res. , 21(1), January 2020
2020
-
[74]
Video generation models as world simulators
Open A.I Team. Video generation models as world simulators
-
[75]
Scalable diffusion models with transformers
William Peebles and Saining Xie. Scalable diffusion models with transformers. In 2023 IEEE/CVF International Conference on Computer Vision (ICCV), pages 4172–4182, 2023
2023
-
[76]
Cogview: Mastering text-to-image generation via transformers
Ming Ding, Zhuoyi Yang, Wenyi Hong, Wendi Zheng, Chang Zhou, Da Yin, Junyang Lin, Xu Zou, Zhou Shao, Hongxia Yang, et al. Cogview: Mastering text-to-image generation via transformers. Advances in Neural Information Processing Systems , 34:19822–19835, 2021
2021
-
[77]
Videotetris: Towards compositional text-to-video generation
Ye Tian, Ling Yang, Haotian Yang, Yuan Gao, Yufan Deng, Xintao Wang, Zhaochen Yu, Xin Tao, Pengfei Wan, Di ZHANG, and Bin CUI. Videotetris: Towards compositional text-to-video generation. In The Thirty-eighth Annual Conference on Neural Information Processing Systems , 2024
2024
-
[78]
Cogview2: Faster and better text-to-image generation via hierarchical transformers
Ming Ding, Wendi Zheng, Wenyi Hong, and Jie Tang. Cogview2: Faster and better text-to-image generation via hierarchical transformers. In S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh, editors, Advances in Neural Information Processing Systems , volume 35, p...
2022
-
[79]
NUWA-infinity: Autoregressive over autoregressive generation for infinite visual synthesis
Jian Liang, Chenfei Wu, Xiaowei Hu, Zhe Gan, Jianfeng Wang, Lijuan Wang, Zicheng Liu, Yuejian Fang, and Nan Duan. NUWA-infinity: Autoregressive over autoregressive generation for infinite visual synthesis. In Alice H. Oh, Alekh Agarwal, Danielle Belgrave, and Kyunghyun Cho, ed...
2022
-
[80]
Ross, Bryan Seybold, and Lu Jiang
Dan Kondratyuk, Lijun Yu, Xiuye Gu, José Lezama, Jonathan Huang, Grant Schindler, Rachel Hornung, Vighnesh Birodkar, Jimmy Yan, Ming-Chang Chiu, Krishna Somandepalli, Hassan Akbari, Yair Alon, Yong Cheng, Josh Dillon, Agrim Gupta, Meera Hahn, Anja Hauth, David Hendon, Alonso M...
2024
-
[81]
ART •V: Auto-Regressive Text-to-Video Generation with Diffusion Models
Wenming Weng, Ruoyu Feng, Yanhui Wang, Qi Dai, Chunyu Wang, Dacheng Yin, Zhiyuan Zhao, Kai Qiu, Jianmin Bao, Yuhui Yuan, Chong Luo, Yueyi Zhang, and Zhiwei Xiong. ART •V: Auto-Regressive Text-to-Video Generation with Diffusion Models . In 2024 IEEE/CVF Conference on Computer V...
2024
-
[82]
Grid diffusion models for text-to-video generation
Taegyeong Lee, Soyeong Kwon, and Taehwan Kim. Grid diffusion models for text-to-video generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 8734–8743, 2024
2024
-
[86]
Vincent Sitzmann, Julien N. P. Martel, Alexander W. Bergman, David B. Lindell, and Gordon Wetzstein. Implicit neural representations with periodic activation functions, 2020
2020
-
[87]
Srinivasan, Matthew Tancik, Jonathan T
Ben Mildenhall, Pratul P. Srinivasan, Matthew Tancik, Jonathan T. Barron, Ravi Ramamoorthi, and Ren Ng. Nerf: Representing scenes as neural radiance fields for view synthesis, 2020
2020
-
[88]
Video probabilistic diffusion models in projected latent space
Sihyun Yu, Kihyuk Sohn, Subin Kim, and Jinwoo Shin. Video probabilistic diffusion models in projected latent space. In 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pages 18456–18466, 2023
2023
-
[89]
Towards end-to-end generative modeling of long videos with memory-efficient bidirectional transformers
Jaehoon Yoo, Semin Kim, Doyup Lee, Chiheon Kim, and Seunghoon Hong. Towards end-to-end generative modeling of long videos with memory-efficient bidirectional transformers. In 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pages 22888–22897, 2023
2023
-
[90]
Streamingt2v: Consistent, dynamic, and extendable long video generation from text
Anonymous. Streamingt2v: Consistent, dynamic, and extendable long video generation from text. In Submitted to The Thirteenth International Conference on Learning Representations , 2024. under review
2024
-
[91]
Vid-gpt: Introducing gpt-style autoregressive generation in video diffusion models, 2024
Kaifeng Gao, Jiaxin Shi, Hanwang Zhang, Chunping Wang, and Jun Xiao. Vid-gpt: Introducing gpt-style autoregressive generation in video diffusion models, 2024
2024
-
[92]
Flexifilm: Long video generation with flexible conditions, 2024
Yichen Ouyang, jianhao Yuan, Hao Zhao, Gaoang Wang, and Bo zhao. Flexifilm: Long video generation with flexible conditions, 2024
2024
-
[93]
Mora: Enabling generalist video generation via a multi-agent framework, 2024
Zhengqing Yuan, Ruoxi Chen, Zhaoxu Li, Haolong Jia, Lifang He, Chi Wang, and Lichao Sun. Mora: Enabling generalist video generation via a multi-agent framework, 2024
2024
-
[94]
Vidgen: Long-form text-to-video generation with temporal, narrative and visual consistency for high quality story-visualisation tasks
Ram Selvaraj, Ayush Singh, Shafiudeen Kameel, Rahul Samal, and Pooja Agarwal. Vidgen: Long-form text-to-video generation with temporal, narrative and visual consistency for high quality story-visualisation tasks. In 2024 IEEE 9th International Conference for Convergence in Tec...
2024
-
[95]
Bissyand, and Saad Ezzini
Zhifei Xie, Daniel Tang, Dingwei Tan, Jacques Klein, Tegawend F. Bissyand, and Saad Ezzini. Dreamfactory: Pioneering multi-scene long video generation with a multi-agent framework, 2024
2024
-
[96]
Kubrick: Multimodal agent collaborations for synthetic video generation, 2024
Liu He, Yizhi Song, Hejun Huang, Daniel Aliaga, and Xin Zhou. Kubrick: Multimodal agent collaborations for synthetic video generation, 2024
2024
-
[97]
Modelscope text-to-video technical report, 2023
Jiuniu Wang, Hangjie Yuan, Dayou Chen, Yingya Zhang, Xiang Wang, and Shiwei Zhang. Modelscope text-to-video technical report, 2023
2023
-
[98]
Free-bloom: zero-shot text-to-video generator with llm director and ldm animator
Hanzhuo Huang, Yufan Feng, Cheng Shi, Lan Xu, Jingyi Yu, and Sibei Yang. Free-bloom: zero-shot text-to-video generator with llm director and ldm animator. In Proceedings of the 37th International Conference on Neural Information Processing Systems , NIPS ’23, Red Hook, NY, USA...
2024
-
[99]
Denoising diffusion implicit models
Jiaming Song, Chenlin Meng, and Stefano Ermon. Denoising diffusion implicit models. In International Conference on Learning Representations , 2021
2021
-
[100]
LLM-grounded video diffusion models
Long Lian, Baifeng Shi, Adam Yala, Trevor Darrell, and Boyi Li. LLM-grounded video diffusion models. In The Twelfth International Conference on Learning Representations, 2024
2024
-
[101]
Mevg: Multi-event video generation with text-to-video models
Gyeongrok Oh, Jaehwan Jeong, Sieun Kim, Wonmin Byeon, Jinkyu Kim, Sungwoong Kim, and Sangpil Kim. Mevg: Multi-event video generation with text-to-video models. In Computer Vision – ECCV 2024: 18th European Conference, Milan, Italy, September 29–October 4, 2024, Proceedings, Pa...
2024
-
[102]
Jingbo Yang and Adrian G. Bors. Enabling the encoder-empowered gan-based video generators for long video generation. In2023 IEEE International Conference on Image Processing (ICIP) , pages 1425–1429, 2023
2023
-
[103]
Videomerge: Towards training-free long video generation, 2025
Siyang Zhang, Harry Yang, and Ser-Nam Lim. Videomerge: Towards training-free long video generation, 2025
2025
-
[104]
Mevg: Multi-event video generation with text-to-video models, 2024
Gyeongrok Oh, Jaehwan Jeong, Sieun Kim, Wonmin Byeon, Jinkyu Kim, Sungwoong Kim, and Sangpil Kim. Mevg: Multi-event video generation with text-to-video models, 2024
2024
-
[105]
Hunyuanvideo: A systematic framework for large video generative models, 2025
Weijie Kong, Qi Tian, Zijian Zhang, Rox Min, Zuozhuo Dai, Jin Zhou, Jiangfeng Xiong, Xin Li, Bo Wu, Jianwei Zhang, Kathrina Wu, Qin Lin, Junkun Yuan, Yanxin Long, Aladdin Wang, Andong Wang, Changlin Li, Duojun Huang, Fang Yang, Hao Tan, Hongmei Wang, Jacob Song, Jiawang Bai, J...
2025
-
[106]
Freenoise: Tuning-free longer video diffusion via noise rescheduling
Haonan Qiu, Menghan Xia, Yong Zhang, Yingqing He, Xintao Wang, Ying Shan, and Ziwei Liu. Freenoise: Tuning-free longer video diffusion via noise rescheduling. In The Twelfth International Conference on Learning Representations , 2024
2024
-
[107]
GLOBER: Coherent non-autoregressive video generation via GLOBal guided video decodER
Mingzhen Sun, Weining Wang, Zihan Qin, Jiahui Sun, Sihan Chen, and Jing Liu. GLOBER: Coherent non-autoregressive video generation via GLOBal guided video decodER. In Thirty-seventh Conference on Neural Information Processing Systems , 2023
2023
-
[108]
Goku: Flow based video generative foundation models, 2025
Shoufa Chen, Chongjian Ge, Yuqi Zhang, Yida Zhang, Fengda Zhu, Hao Yang, Hongxiang Hao, Hui Wu, Zhichao Lai, Yifei Hu, Ting-Che Lin, Shilong Zhang, Fu Li, Chuan Li, Xing Wang, Yanghua Peng, Peize Sun, Ping Luo, Yi Jiang, Zehuan Yuan, Bingyue Peng, and Xiaobing Liu. Goku: Flow ...
2025
-
[109]
Reducio! generating 1024×1024 video within 16 seconds using extremely compressed motion latents, 2024
Rui Tian, Qi Dai, Jianmin Bao, Kai Qiu, Yifan Yang, Chong Luo, Zuxuan Wu, and Yu-Gang Jiang. Reducio! generating 1024×1024 video within 16 seconds using extremely compressed motion latents, 2024
2024
-
[110]
Genmo Team. Mochi 1. https://github.com/genmoai/models, 2024
2024
-
[111]
Moviegen: A cast of media foundation models
Meta Research Team. Moviegen: A cast of media foundation models. arXiv preprint, October 2024. Meta AI Research Publication
2024
-
[112]
From sora what we can see: A survey of text-to-video generation, 2024
Rui Sun, Yumin Zhang, Tejal Shah, Jiahao Sun, Shuoying Zhang, Wenqi Li, Haoran Duan, Bo Wei, and Rajiv Ranjan. From sora what we can see: A survey of text-to-video generation, 2024
2024
-
[113]
Cosmos world foundation model platform for physical ai, 2025
NVIDIA, :, Niket Agarwal, Arslan Ali, Maciej Bala, Yogesh Balaji, Erik Barker, Tiffany Cai, Prithvijit Chattopadhyay, Yongxin Chen, Yin Cui, Yifan Ding, Daniel Dworakowski, Jiaojiao Fan, Michele Fenzi, Francesco Ferroni, Sanja Fidler, Dieter Fox, Songwei Ge, Yunhao Ge, Jinwei ...
2025
-
[114]
Fleet, Mohammad Norouzi, and Tim Salimans
Jonathan Ho, Chitwan Saharia, William Chan, David J. Fleet, Mohammad Norouzi, and Tim Salimans. Cascaded diffusion models for high fidelity image generation, 2021
2021
-
[115]
Hauptmann, Ming-Hsuan Yang, Yuan Hao, Irfan Essa, and Lu Jiang
Lijun Yu, Yong Cheng, Kihyuk Sohn, José Lezama, Han Zhang, Huiwen Chang, Alexander G. Hauptmann, Ming-Hsuan Yang, Yuan Hao, Irfan Essa, and Lu Jiang. Magvit: Masked generative video transformer, 2023
2023
-
[116]
Hitvideo: Hierarchical tokenizers for enhancing text-to-video generation with autoregressive large language models, 2025
Ziqin Zhou, Yifan Yang, Yuqing Yang, Tianyu He, Houwen Peng, Kai Qiu, Qi Dai, Lili Qiu, Chong Luo, and Lingqiao Liu. Hitvideo: Hierarchical tokenizers for enhancing text-to-video generation with autoregressive large language models, 2025
2025
-
[117]
Large language models are frame-level directors for zero-shot text-to-video generation
Susung Hong, Junyoung Seo, Heeseong Shin, Sunghwan Hong, and Seungryong Kim. Large language models are frame-level directors for zero-shot text-to-video generation. In First Workshop on Controllable Video Generation @ICML24 , 2024. Manuscript submitted to ACM Video Is Worth a ...
2024
-
[118]
Microcinema: A divide-and-conquer approach for text-to-video generation
Yanhui Wang, Jianmin Bao, Wenming Weng, Ruoyu Feng, Dacheng Yin, Tao Yang, Jingxu Zhang, Qi Dai, Zhiyuan Zhao, Chunyu Wang, Kai Qiu, Yuhui Yuan, Xiaoyan Sun, Chong Luo, and Baining Guo. Microcinema: A divide-and-conquer approach for text-to-video generation. In Proceedings of ...
2024
-
[120]
Adding conditional control to text-to-image diffusion models, 2023
Lvmin Zhang, Anyi Rao, and Maneesh Agrawala. Adding conditional control to text-to-image diffusion models, 2023
2023
-
[121]
Videocontrolnet: A motion-guided video-to-video translation framework by using diffusion model with controlnet, 2023
Zhihao Hu and Dong Xu. Videocontrolnet: A motion-guided video-to-video translation framework by using diffusion model with controlnet, 2023
2023
-
[122]
Moonshot: Towards controllable video generation and editing with multimodal conditions, 2024
David Junhao Zhang, Dongxu Li, Hung Le, Mike Zheng Shou, Caiming Xiong, and Doyen Sahoo. Moonshot: Towards controllable video generation and editing with multimodal conditions, 2024
2024
-
[123]
Easycontrol: Transfer controlnet to video diffusion for controllable generation and interpolation, 2024
Cong Wang, Jiaxi Gu, Panwen Hu, Haoyu Zhao, Yuanfan Guo, Jianhua Han, Hang Xu, and Xiaodan Liang. Easycontrol: Transfer controlnet to video diffusion for controllable generation and interpolation, 2024
2024
-
[124]
Controlvideo: Training-free controllable text-to-video generation, 2023
Yabo Zhang, Yuxiang Wei, Dongsheng Jiang, Xiaopeng Zhang, Wangmeng Zuo, and Qi Tian. Controlvideo: Training-free controllable text-to-video generation, 2023
2023
-
[125]
Videostudio: Generating consistent-content and multi-scene videos, 2024
Fuchen Long, Zhaofan Qiu, Ting Yao, and Tao Mei. Videostudio: Generating consistent-content and multi-scene videos, 2024
2024
-
[126]
Sdxl: Improving latent diffusion models for high-resolution image synthesis, 2023
Dustin Podell, Zion English, Kyle Lacey, Andreas Blattmann, Tim Dockhorn, Jonas Müller, Joe Penna, and Robin Rombach. Sdxl: Improving latent diffusion models for high-resolution image synthesis, 2023
2023
-
[127]
Videobooth: Diffusion-based video generation with image prompts
Yuming Jiang, Tianxing Wu, Shuai Yang, Chenyang Si, Dahua Lin, Yu Qiao, Chen Change Loy, and Ziwei Liu. Videobooth: Diffusion-based video generation with image prompts. In 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pages 6689–6700, 2024
2024
-
[128]
Lavie: High-quality video generation with cascaded latent diffusion models, 2024
Yaohui Wang, Xinyuan Chen, Xin Ma, Shangchen Zhou, Ziqi Huang, Yi Wang, Ceyuan Yang, Yinan He, Jiashuo Yu, Peiqing Yang, Yuwei Guo, Tianxing Wu, Chenyang Si, Yuming Jiang, Cunjian Chen, Chen Change Loy, Bo Dai, Dahua Lin, Yu Qiao, and Ziwei Liu. Lavie: High-quality video gener...
2024
-
[129]
Ucf101: A dataset of 101 human actions classes from videos in the wild, 2012
Khurram Soomro, Amir Roshan Zamir, and Mubarak Shah. Ucf101: A dataset of 101 human actions classes from videos in the wild, 2012
2012
-
[130]
Quo vadis, action recognition? a new model and the kinetics dataset
Joao Carreira and Andrew Zisserman. Quo vadis, action recognition? a new model and the kinetics dataset. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , July 2017
2017
-
[131]
Youtube-8m: A large-scale video classification benchmark, 2016
Sami Abu-El-Haija, Nisarg Kothari, Joonseok Lee, Paul Natsev, George Toderici, Balakrishnan Varadarajan, and Sudheendra Vijayanarasimhan. Youtube-8m: A large-scale video classification benchmark, 2016
2016
-
[132]
Howto100m: Learning a text-video embedding by watching hundred million narrated video clips
Antoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi, Ivan Laptev, and Josef Sivic. Howto100m: Learning a text-video embedding by watching hundred million narrated video clips. In 2019 IEEE/CVF International Conference on Computer Vision (ICCV) , pages 2630–2...
2019
-
[133]
Msr-vtt: A large video description dataset for bridging video and language
Jun Xu, Tao Mei, Ting Yao, and Yong Rui. Msr-vtt: A large video description dataset for bridging video and language. In 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , pages 5288–5296, 2016
2016
-
[134]
Frozen in Time: A Joint Video and Image Encoder for End-to-End Retrieval
Max Bain, Arsha Nagrani, Gul Varol, and Andrew Zisserman. Frozen in Time: A Joint Video and Image Encoder for End-to-End Retrieval . In 2021 IEEE/CVF International Conference on Computer Vision (ICCV) , pages 1708–1718, Los Alamitos, CA, USA, October 2021. IEEE Computer Society
2021
-
[135]
Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut. Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning. In Proceedings of ACL, 2018
2018
-
[136]
Internvid: A large-scale video-text dataset for multimodal understanding and generation
Yi Wang, Yinan He, Yizhuo Li, Kunchang Li, Jiashuo Yu, Xin Ma, Xinhao Li, Guo Chen, Xinyuan Chen, Yaohui Wang, Ping Luo, Ziwei Liu, Yali Wang, Limin Wang, and Yu Qiao. Internvid: A large-scale video-text dataset for multimodal understanding and generation. In The Twelfth Inter...
2024
-
[137]
Video-chatgpt: Towards detailed video understanding via large vision and language models, 2024
Muhammad Maaz, Hanoona Rasheed, Salman Khan, and Fahad Shahbaz Khan. Video-chatgpt: Towards detailed video understanding via large vision and language models, 2024
2024
-
[138]
Activitynet: A large-scale video benchmark for human activity understanding
Fabian Caba Heilbron, Victor Escorcia, Bernard Ghanem, and Juan Carlos Niebles. Activitynet: A large-scale video benchmark for human activity understanding. In 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , pages 961–970, 2015
2015
-
[139]
Miradata: A large-scale video dataset with long durations and structured captions
Xuan Ju, Yiming Gao, Zhaoyang Zhang, Ziyang Yuan, Xintao Wang, Ailing Zeng, Yu Xiong, Qiang Xu, and Ying Shan. Miradata: A large-scale video dataset with long durations and structured captions. In The Thirty-eight Conference on Neural Information Processing Systems Datasets an...
2024
-
[140]
A short note about kinetics-600, 2018
Joao Carreira, Eric Noland, Andras Banki-Horvath, Chloe Hillier, and Andrew Zisserman. A short note about kinetics-600, 2018
2018
-
[141]
Panda-70M: Captioning 70M Videos with Multiple Cross-Modality Teachers
Tsai-Shien Chen, Aliaksandr Siarohin, Willi Menapace, Ekaterina Deyneka, Hsiang-Wei Chao, Byung Eun Jeon, Yuwei Fang, Hsin-Ying Lee, Jian Ren, Ming-Hsuan Yang, and Sergey Tulyakov. Panda-70M: Captioning 70M Videos with Multiple Cross-Modality Teachers . In 2024 IEEE/CVF Confer...
2024
-
[142]
Videofactory: Swap attention in spatiotemporal diffusions for text-to-video generation, 2024
Wenjing Wang, Huan Yang, Zixi Tuo, Huiguo He, Junchen Zhu, Jianlong Fu, and Jiaying Liu. Videofactory: Swap attention in spatiotemporal diffusions for text-to-video generation, 2024
2024
-
[143]
Vidprom: A million-scale real prompt-gallery dataset for text-to-video diffusion models
Wenhao Wang and Yi Yang. Vidprom: A million-scale real prompt-gallery dataset for text-to-video diffusion models. In Proceedings of the 2024 NeurIPS Conference. NeurIPS, 2024. Poster number: 97505
2024
-
[144]
Kristen Grauman, Andrew Westbury, Eugene Byrne, Vincent Cartillier, Zachary Chavis, Antonino Furnari, Rohit Girdhar, Jackson Hamburger, Hao Jiang, Devansh Kukreja, Miao Liu, Xingyu Liu, Miguel Martin, Tushar Nagarajan, Ilija Radosavovic, Santhosh Kumar Ramakrishnan, Fiona Ryan...
-
[145]
Synchronized video storytelling: Generating video narrations with structured storyline, 2024
Dingyi Yang, Chunru Zhan, Ziheng Wang, Biao Wang, Tiezheng Ge, Bo Zheng, and Qin Jin. Synchronized video storytelling: Generating video narrations with structured storyline, 2024
2024
-
[146]
Benchmarking aigc video quality assessment: A dataset and unified model, 2024
Zhichao Zhang, Xinyue Li, Wei Sun, Jun Jia, Xiongkuo Min, Zicheng Zhang, Chunyi Li, Zijian Chen, Puyi Wang, Zhongpeng Ji, Fengyu Sun, Shangling Jui, and Guangtao Zhai. Benchmarking aigc video quality assessment: A dataset and unified model, 2024
2024
-
[147]
Improved techniques for training gans
Tim Salimans, Ian Goodfellow, Wojciech Zaremba, Vicki Cheung, Alec Radford, and Xi Chen. Improved techniques for training gans. InProceedings of the 30th International Conference on Neural Information Processing Systems , NIPS’16, page 2234–2242, Red Hook, NY, USA, 2016. Curra...
2016
-
[148]
Rethinking the inception architecture for computer vision
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna. Rethinking the inception architecture for computer vision. In 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , pages 2818–2826, 2016
2016
-
[149]
Gans trained by a two time-scale update rule converge to a local nash equilibrium
Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. Gans trained by a two time-scale update rule converge to a local nash equilibrium. In Proceedings of the 31st International Conference on Neural Information Processing Systems , NIPS’17,...
2017
-
[150]
FVD: A new metric for video generation, 2019
Thomas Unterthiner, Sjoerd van Steenkiste, Karol Kurach, Raphaël Marinier, Marcin Michalski, and Sylvain Gelly. FVD: A new metric for video generation, 2019
2019
-
[151]
Exploring Video Quality Assessment on User Generated Contents from Aesthetic and Technical Perspectives
Haoning Wu, Erli Zhang, Liang Liao, Chaofeng Chen, Jingwen Hou, Annan Wang, Wenxiu Sun, Qiong Yan, and Weisi Lin. Exploring Video Quality Assessment on User Generated Contents from Aesthetic and Technical Perspectives . In 2023 IEEE/CVF International Conference on Computer Vis...
2023
-
[152]
Raft: Recurrent all-pairs field transforms for optical flow
Zachary Teed and Jia Deng. Raft: Recurrent all-pairs field transforms for optical flow. In Computer Vision – ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part II , page 402–419, Berlin, Heidelberg, 2020. Springer-Verlag
2020
-
[153]
Clipscore: A reference-free evaluation metric for image captioning, 2022
Jack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras, and Yejin Choi. Clipscore: A reference-free evaluation metric for image captioning, 2022
2022
-
[154]
Godiva: Generating open-domain videos from natural descriptions, 2021
Chenfei Wu, Lun Huang, Qianxi Zhang, Binyang Li, Lei Ji, Fan Yang, Guillermo Sapiro, and Nan Duan. Godiva: Generating open-domain videos from natural descriptions, 2021
2021
-
[155]
Grit: A generative region-to-text transformer for object understanding, 2022
Jialian Wu, Jianfeng Wang, Zhengyuan Yang, Zhe Gan, Zicheng Liu, Junsong Yuan, and Lijuan Wang. Grit: A generative region-to-text transformer for object understanding, 2022
2022
-
[156]
Fetv: a benchmark for fine-grained evaluation of open-domain text-to-video generation
Yuanxin Liu, Lei Li, Shuhuai Ren, Rundong Gao, Shicheng Li, Sishuo Chen, Xu Sun, and Lu Hou. Fetv: a benchmark for fine-grained evaluation of open-domain text-to-video generation. In Proceedings of the 37th International Conference on Neural Information Processing Systems , NI...
2024
-
[157]
Vbench: Comprehensive benchmark suite for video generative models
Ziqi Huang, Yinan He, Jiashuo Yu, Fan Zhang, Chenyang Si, Yuming Jiang, Yuanhan Zhang, Tianxing Wu, Qingyang Jin, Nattapol Chanpaisit, Yaohui Wang, Xinyuan Chen, Limin Wang, Dahua Lin, Yu Qiao, and Ziwei Liu. Vbench: Comprehensive benchmark suite for video generative models. I...
2024
-
[158]
Subjective-aligned dataset and metric for text-to-video quality assessment, 2024
Tengchuan Kou, Xiaohong Liu, Zicheng Zhang, Chunyi Li, Haoning Wu, Xiongkuo Min, Guangtao Zhai, and Ning Liu. Subjective-aligned dataset and metric for text-to-video quality assessment, 2024
2024
-
[159]
Fréchet video motion distance: A metric for evaluating motion consistency in videos
Jiahe Liu, Youran Qu, Qi Yan, Xiaohui Zeng, Lele Wang, and Renjie Liao. Fréchet video motion distance: A metric for evaluating motion consistency in videos. In First Workshop on Controllable Video Generation @ICML24 , 2024
2024
-
[160]
T2v-compbench: A comprehensive benchmark for compositional text-to-video generation, 2024
Kaiyue Sun, Kaiyi Huang, Xian Liu, Yue Wu, Zihan Xu, Zhenguo Li, and Xihui Liu. T2v-compbench: A comprehensive benchmark for compositional text-to-video generation, 2024
2024
-
[161]
Videopoet: A large language model for zero-shot video generation
Jiaxuan Guo, Yichun Li, Shangzhe Wang, Yinan Zhang, Xihui Liu, Yu Wang, Hanyang Yang, Jing Yang, and Ziwei Liu. Videopoet: A large language model for zero-shot video generation. arXiv:2312.14125, 2023
2023 arXiv
-
[162]
Make-a-video: Text-to-video generation without text-video data
Uriel Singer, Adam Polyak, Thomas Hayes, Xi Yin, Junjie An, Songyang Zhang, Qiyuan Hu, Oran Yang, Omri Ashual, Oran Gafni, et al. Make-a-video: Text-to-video generation without text-video data. arXiv:2209.14792, 2022
2022 arXiv
-
[163]
Latent video diffusion models for high-fidelity long video generation, 2023
Yingqing He, Tianyu Yang, Yong Zhang, Ying Shan, and Qifeng Chen. Latent video diffusion models for high-fidelity long video generation, 2023
2023
-
[164]
Vlogger: Generating long videos of dynamic human activities
Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi. Vlogger: Generating long videos of dynamic human activities. arXiv:2306.04308, 2023
2023 arXiv
-
[165]
Microcinema: A divide-and-conquer approach for text-to-video generation
Yinan He, Yaohui Wang, Ceyuan Yang, Shangchen Zhou, Xiangyu Zhang, Xiaodong Yang, Yu Qiao, Dahua Lin, and Ying Shan. Microcinema: A divide-and-conquer approach for text-to-video generation. arXiv preprint arXiv:2312.04889, 2023
2023 arXiv
-
[167]
Jonathan Ho, Tim Salimans, Alexey Gritsenko, William Chan, Mohammad Norouzi, and David J. Fleet. Video diffusion models, 2022. Manuscript submitted to ACM Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation 35
2022
-
[168]
Nuwa: Visual synthesis pre-training for neural visual world creation
Chenfei Wu, Jian Liang, Lei Ji, Fan Yang, Yuejian Fang, Daxin Jiang, and Nan Duan. Nuwa: Visual synthesis pre-training for neural visual world creation. arXiv:2111.12417, 2021
2021 arXiv
-
[169]
Temporal generative adversarial nets with singular value clipping, 2017
Masaki Saito, Eiichi Matsumoto, and Shunta Saito. Temporal generative adversarial nets with singular value clipping, 2017
2017
-
[170]
Lvt: Language-vision transformer for multi-modal video understanding
Qingqiu Huang, Wentao Yu, Yuanze Xu, Yitong Wang, and Dacheng Zhang. Lvt: Language-vision transformer for multi-modal video understanding. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) , page , 2021
2021
-
[171]
Generating videos with scene dynamics
Carl Vondrick, Hamed Pirsiavash, and Antonio Torralba. Generating videos with scene dynamics. In NeurIPS, 2016
2016
-
[172]
Videofusion: Decomposed diffusion models for high-quality video generation
Zhengxiong Luo, Dayou Chen, Yingya Zhang, Yan Huang, Liang Wang, Yujun Shen, Deli Zhao, Jingren Zhou, and Tieniu Tan. Videofusion: Decomposed diffusion models for high-quality video generation. CVPR, 2023
2023
-
[174]
Magicvideo: Efficient video generation with latent diffusion models
Rui Zhao, Yuxiang Wu, Hao Dong, Ning Zhang, Tao Yang, Wei Wei, and Xiaowei Huang. Magicvideo: Efficient video generation with latent diffusion models. arXiv:2301.11093, 2023
2023 arXiv
-
[175]
Pyoco: Latent diffusion priors for zero-shot video editing
Andreas Blattmann, Robin Rombach, Huan Ling, Tim Dockhorn, Seung Wook Kim, Sanja Fidler, and Karsten Kreis. Pyoco: Latent diffusion priors for zero-shot video editing. arXiv:2303.04734, 2023
2023 arXiv
-
[176]
Videofactory: Synthesizing high-quality video with diffusion models
Yaohui Wang, Xinyuan Chen, Xin Ma, Shangchen Zhou, Ziwei Huang, Yu Wang, Ceyuan Yang, Yinan He, Jiashuo Yu, Yinan Yang, et al. Videofactory: Synthesizing high-quality video with diffusion models. arXiv:2305.10874, 2023
2023 arXiv
-
[178]
Lavie: A layered video diffusion model with multi-modal conditioning
Yue Ma, Yingqing He, Xiaodong Cun, Xintao Wang, Ying Shan, Xiu Li, and Qifeng Chen. Lavie: A layered video diffusion model with multi-modal conditioning. arXiv:2309.15130, 2023
2023 arXiv
-
[180]
Internvid: A large-scale video-text dataset for multimodal understanding and generation
Yaohui Wang, Xinyuan Chen, Xin Ma, Shangchen Zhou, Ziwei Huang, Yu Wang, Ceyuan Yang, Yinan He, Jiashuo Yu, Yinan Yang, et al. Internvid: A large-scale video-text dataset for multimodal understanding and generation. arXiv:2307.06942, 2023
2023 arXiv
-
[181]
Make pixels dance: High-dynamic video generation, 2023
Yan Zeng, Guoqiang Wei, Jiani Zheng, Jiaxin Zou, Yang Wei, Yuchen Zhang, and Hang Li. Make pixels dance: High-dynamic video generation, 2023
2023
-
[182]
Emu video: Factorizing text-to-video generation by explicit image conditioning, 2024
Rohit Girdhar, Mannat Singh, Andrew Brown, Quentin Duval, Samaneh Azadi, Sai Saketh Rambhatla, Akbar Shah, Xi Yin, Devi Parikh, and Ishan Misra. Emu video: Factorizing text-to-video generation by explicit image conditioning, 2024
2024
-
[183]
Lumiere: A space-time diffusion model for video generation
Omer Bar-Tal, Dolev Ofri-Amar, Rafail Fridman, Yoni Kasten, and Tali Dekel. Lumiere: A space-time diffusion model for video generation. arXiv preprint arXiv:2401.12945, 2024
2024 arXiv
-
[184]
Llmscore: Unveiling the power of large language models in text-to-image synthesis evaluation, 2023
Yujie Lu, Xianjun Yang, Xiujun Li, Xin Eric Wang, and William Yang Wang. Llmscore: Unveiling the power of large language models in text-to-image synthesis evaluation, 2023
2023
-
[185]
Lavie: High-quality video generation with cascaded latent diffusion models
Yaohui Wang, Xinyuan Chen, Xin Ma, Shangchen Zhou, Ziqi Huang, Yi Wang, Ceyuan Yang, Yinan He, Jiashuo Yu, Peiqing Yang, et al. Lavie: High-quality video generation with cascaded latent diffusion models. arXiv preprint arXiv:2309.15103, 2023
2023 arXiv
-
[186]
Modelscope text-to-video technical report.arXiv preprint arXiv:2308.06571, 2023
Jiuniu Wang, Hangjie Yuan, Dayou Chen, Yingya Zhang, Xiang Wang, and Shiwei Zhang. Modelscope text-to-video technical report.arXiv preprint arXiv:2308.06571, 2023
2023 arXiv
-
[187]
Videocrafter2: Overcoming data limitations for high-quality video diffusion models, 2024
Haoxin Chen, Yong Zhang, Xiaodong Cun, Menghan Xia, Xintao Wang, Chao Weng, and Ying Shan. Videocrafter2: Overcoming data limitations for high-quality video diffusion models, 2024
2024
-
[189]
Animatediff: Animate your personalized text-to-image diffusion models without specific tuning
Yuwei Guo, Ceyuan Yang, Anyi Rao, Yaohui Wang, Yu Qiao, Dahua Lin, and Bo Dai. Animatediff: Animate your personalized text-to-image diffusion models without specific tuning. 2024
2024
-
[190]
Latte: Latent diffusion transformer for video generation
Xin Ma, Yaohui Wang, Gengyun Jia, Xinyuan Chen, Ziwei Liu, Yuan-Fang Li, Cunjian Chen, and Yu Qiao. Latte: Latent diffusion transformer for video generation. arXiv preprint arXiv:2401.03048, 2024
2024 arXiv
-
[191]
Accessed September 25, 2023 [Online] https://www.pika.art/, 2023
Pika labs. Accessed September 25, 2023 [Online] https://www.pika.art/, 2023
2023
-
[192]
Accessed June 6, 2024 [Online] https://klingai.kuaishou.com/, 2024
Kling. Accessed June 6, 2024 [Online] https://klingai.kuaishou.com/, 2024
2024
-
[193]
Cogvideox: Text-to-video diffusion models with an expert transformer
Zhuoyi Yang, Jiayan Teng, Wendi Zheng, Ming Ding, Shiyu Huang, Jiazheng Xu, Yuanming Yang, Wenyi Hong, Xiaohan Zhang, Guanyu Feng, et al. Cogvideox: Text-to-video diffusion models with an expert transformer. arXiv preprint arXiv:2408.06072, 2024. Manuscript submitted to ACM
2024 arXiv
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.