REVIEW 3 major objections 5 minor 1 cited by
A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality
T0 review · 3 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read A survey of 32 video-generation papers argues that long-video storytelling is a modular problem and names the component stack—MM-DiT, MLLM encoders, dual VAEs, 3D RoPE, prompt rewriting, and MeanFlow—that most reliably solves it.
desk verdict Useful taxonomy and component tables, but the MeanFlow recommendation is supported by the wrong equations and needs correction before the survey can be trusted. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The paper's carrying device is a taxonomy tree that sorts long-video generators into six architectural styles: keyframes-to-video, discrete temporal chunks, high compression, flattened 3D space-time one-shot (with branches for foundational, single-subject, multi-subject, and multi-shot narrative planning), token-stream autoregressive, and closed source. Each branch names the dominant strategy for extending duration and managing memory, and both the taxonomy and Table 1 map every surveyed model onto these axes. The second working object is the recommended component stack—MM-DiT (a dual-stream diffusion transformer that processes text and video jointly), MLLM encoders, dual VAEs, 3D RoPE (rotary position embeddings that rotate feature pairs across time and space), MeanFlow (which learns a spatially and temporally averaged velocity field instead of instantaneous velocities), and prompt rewriting. This stack functions as the transferable recipe the survey argues holds across all taxonomy styles.
What would settle it
Take one open long-video model and run controlled swaps of each recommended component—T5 instead of an MLLM encoder, standard flow matching instead of MeanFlow, sinusoidal position embeddings instead of 3D RoPE, a single VAE instead of dual VAEs—while keeping data and compute fixed; evaluate on a multidimensional benchmark like VBench for clips longer than 16 seconds. If the recommended versions do not improve identity consistency, temporal flicker, and text-video relevance across variants and durations, the survey's central recommendation is falsified.
Extended reading notes
Core claim
The central claim is that long-video storytelling quality is not one elusive capability but a set of modular problems, each with a working solution already visible in top-performing systems. Character consistency, scene-layout stability, temporal coherence, and narrative continuity are each addressed by specific choices: dual-stream multimodal diffusion transformers (MM-DiT and Flux-MM-DiT) as the backbone; multimodal large-language-model encoders in place of T5/CLIP for text conditioning; dual VAEs that separate static appearance encoding from temporal dynamics; three-dimensional and multi-modal rotary position embeddings to keep motion coherent; LLM-based prompt rewriting and story-agent shot planning to bridge user prompts and training captions; and MeanFlow, an average-velocity training objective, to make generation faster and more stable than flow matching. The evidence offered is the recurrence of these components across the 32 surveyed models, the comparative backbone table, and benchmark figures such as MeanFlow's Fréchet Video Distance (FVD) of 128 versus flow matching's 142 on Kinetics-400.
Load-bearing premise
The load-bearing premise is that the 32 selected papers are a fair and comparable sample: if the selection is biased toward systems that already use the recommended components, or if their reported metrics come from incomparable benchmarks and durations, then the claim that these components 'consistently yield' better long videos no longer follows.
Editorial extensions
If this is right
- Systems built on MM-DiT or Flux-MM-DiT backbones should show stronger text-video alignment and lower FID than U-Net or single-stream DiT backbones of similar scale.
- Replacing T5/CLIP encoders with a multimodal LLM encoder should improve prompt adherence and keep narrative semantics consistent across longer clips.
- Separating image and video encoding into dual VAEs should cut training cost by roughly 5–10x while preserving video quality, as seen in Open-Sora 2.0.
- Using 3D RoPE or MM-RoPE should improve motion coherence and length extrapolation compared with sinusoidal position embeddings.
- Adding LLM prompt rewriting and story-agent shot planning should reduce flicker, unify style, and keep character identity stable across scene cuts.
Reading between the lines
- The paper leaves implicit a concrete ablation program: holding one base model fixed and swapping each recommended component on and off would rank the components by actual causal contribution, which the survey's recurrence-based evidence cannot do.
- Because the taxonomy includes a 'closed source' branch, the recommendations are best grounded for open-source systems; for proprietary models the component attributions are inferred from public descriptions rather than verified.
- The dataset analysis suggests a testable extension: a model trained on hierarchically annotated movie-level data (scene, shot, character attributes) should beat the same architecture trained on generic web captions specifically on narrative-coherence metrics.
- A direct way to stress-test the framework is to apply it to the 150-second regime: if the recommended stack holds, identity and layout drift should stay flat rather than grow with duration; the survey does not itself measure that slope.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This manuscript is a survey of 32 recent video-generation papers focused on long-video storytelling. It organizes methods into six architectural styles (keyframes-to-video, discrete temporal chunks, high compression, flattened 3D one-shot, token-stream autoregressive, and closed-source systems) and derives component-level recommendations: MM-DiT/Flux-MM-DiT backbones, MLLM-based text encoders, dual image/video VAEs, 3D RoPE positional encodings, LLM-driven prompt rewriting, and MeanFlow as a training objective. The paper also presents two tables labeled Table 1, an appendix of datasets and evaluation metrics, and a conclusion listing current limitations and future directions.
Significance. If accurate, the paper would provide a useful practitioner-oriented map of a fast-moving field, particularly the taxonomy and the component-level comparison across open- and closed-source systems. The paper explicitly credits the role of closed-source systems and discusses datasets and evaluation metrics, which are often omitted from shorter surveys. Its main value is as a synthesis: it gives a clear organizing frame and a concrete checklist of architectural ingredients. The paper makes no new measurements, so its soundness must be judged on internal consistency and fidelity of its literature claims. Its strongest claim—that the recommended components 'consistently yield' long-video consistency and cinematic quality—rests on a small, undocumented sample and on one internally mis-specified quantitative comparison (MeanFlow, §3.2). With those corrected, the survey could be a genuinely useful reference; in its current form, the support for the headline recommendations is weaker than the prose suggests.
major comments (3)
- [§3.2, Eqs. (7)–(9)] The equations in this section do not define the MeanFlow objective as introduced in [21]. Eq. (7) predicts optical flow u_{t→t+1} with a network F, Eq. (8) is an L1 loss against ground-truth optical flow, and Eq. (9) averages those optical-flow vectors over space and time. The cited MeanFlow method instead learns an average velocity field for an ODE that transports noise to data; it does not predict per-pixel optical flow between adjacent frames. The prose of §3.2 says the correct thing, but the quantitative claims attached to it (FVD 128 vs. 142, SSIM 0.85 vs. 0.82, LPIPS 0.12 vs. 0.15, and 4× speedup) are framed as optical-flow-estimation results. Because this is the main quantitative support for a headline component, the evidence behind the Section 4 conclusion that MeanFlow deserves recommendation is mis-specified. Please replace Eqs. (7)–(9) with the actual flow-matching/MeanFlow formalism, or clearly separate an optical-flow baseline from the MeanFlow discussion and attribute the numbers correctly.
- [Abstract; §1; §4] The central claim that the surveyed components 'consistently yield' long-video quality is not backed by a described selection protocol. The paper states that it 'comprehensively studied 32 papers' but does not state inclusion criteria, search procedure, time window, or how papers were assigned to the six taxonomy branches. Because many selected papers already use the recommended components (MM-DiT in Seedance, Phantom, and PyramidFlow; 3D RoPE in HunyuanVideo, MAGI-1, and StepVideo; MLLM encoders in HunyuanVideo), the recommendation may partly reflect a sampling bias rather than an independent comparison. The issue is not that the bias is deliberate, but that the 'consistently yield' language requires either a documented sampling frame or more cautious wording such as 'recur in recent high-performing systems.'
- [§3.2 and §4] The support for the MeanFlow and flow-matching recommendations is drawn from a single reported comparison on Kinetics-400 and UCF-101, without discussing how those results transfer to long-form storytelling with multi-subject consistency. The survey's own evaluation discussion (Appendix B) criticizes image-derived metrics such as FVD, SSIM, and LPIPS for obscuring temporal coherence and storytelling fidelity. Using those same metrics as the basis for a 'promising results' recommendation is therefore internally inconsistent; either the recommendation should be justified on the metrics the survey deems appropriate, or the mismatch should be acknowledged and the recommendation softened.
minor comments (5)
- [Section 2.4.2] AnimateDiff is described as using LoRA adapters in both U-Net and transformer blocks, but the cited source [23] does not support this description; AnimateDiff is typically a motion-module personalization method, not a LoRA-based one. Please verify the attribution and correct either the description or the reference.
- [Table 1 (both occurrences)] The paper labels three distinct tables with the same caption 'Table 1': the backbone comparison, its continuation, and the dataset/statistics table in the Appendix. Several cells in the backbone table are empty where the survey's own recommendations call for values (e.g., Params for Seedance, Sora, and Veo3; Resolution for Loong; Positional Encodings for several entries). Please renumber the tables and either fill these cells or explicitly mark them as unavailable.
- [Appendix, Table 1 legend] The legend uses the symbol 'S' for both subject count and duration ranges, and the duration ranges are not defined with explicit inequality signs (e.g., 'S Greater than 17 seconds' and 'S 5 up to 16 seconds'). Clarify the notation to avoid ambiguity between single-subject and duration classes.
- [Section 4] The list of limitations skips item (iv) and the future-work list uses '(v)' twice; please renumber the enumerated items.
- [Throughout] Names are used inconsistently, including 'W AN2.1' vs. 'Wan2.1', 'AnimatedDiff' vs. 'AnimateDiff', 'MagVit-v2' vs. 'MAGVIT-v2', and 'PyramidFlow' vs. 'Pyramid Flow'. Standardize these names for readability.
Circularity Check
No circular derivation: the survey's component recommendations rest on external benchmarks and papers; the sole self-citation is non-load-bearing, and the mis-specified MeanFlow equations are a correctness concern, not a circularity.
full rationale
This is a survey, not a derivation. Its central claim—that the recommended components (MM-DiT/Flux-MM-DiT, MLLM encoders, dual VAEs, 3D RoPE, prompt rewriting, MeanFlow-style objectives) reliably improve long-video quality—is an inductive synthesis of 32 surveyed papers and is supported by external citations and benchmarks rather than by fitting parameters to a subset of data and then predicting that same data. Section 3.2 recommends MeanFlow on the strength of FVD/SSIM/LPIPS and 4x speedup numbers attributed to the external MeanFlow paper [21]; those numbers are not produced by any equation in this manuscript. The expressions in Eqs. (7)-(9), however, define optical-flow averaging ('predicted optical flow', 'ground-truth flow', 'spatial-temporal average flow') rather than the average velocity field of the cited MeanFlow objective; this is a real technical mis-specification that weakens the evidential link, but it is not circularity because the recommendation rests on an external source, not on this paper's own formulas. The one self-citation, [46] VideoDirectorGPT (includes author Jaemin Cho), appears in the Introduction to support the general statement that 'Supporting multiple characters while preserving consistency is even more demanding [46]'; that claim is not load-bearing for any of the paper's recommendations and is consistent with standard consensus. The dataset sampling concern raised by the reader—that the 32 papers were selected without a formal protocol and evaluations use different durations and benchmarks—is a representativeness/validity limitation, which the paper itself partially acknowledges in its limitations list (e.g., 'Existing annotated datasets lack critical metadata'); it is not a derivation loop. No fitted input is relabeled as a prediction, no uniqueness theorem is imported from the authors' prior work, and no ansatz is smuggled in via citation; the taxonomy's 'Closed Source' branch is a categorization choice rather than a renamed result. Under the rubric, the paper is self-contained against external benchmarks, so the circularity score is low; the one minor self-citation, non-load-bearing, justifies the upper end of the 0-2 band.
Assumptions & free parameters
assumptions (3)
- domain assumption The 32 surveyed papers are a representative sample of the long-video generation field.
- domain assumption The architecture descriptions and performance numbers quoted from the primary papers are accurate.
- domain assumption Results reported across different papers are comparable enough to support 'consistently yield' claims.
Cite this review
Pith. "Pith review of A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality." pith.science (2026). https://pith.science/paper/MQD7WGKT
@misc{pith2026250707202,
author = {Pith},
title = {Pith review of: A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality},
year = {2026},
howpublished = {\url{https://pith.science/paper/MQD7WGKT}},
note = {Machine review of arXiv:2507.07202}
}
read the original abstract
Despite the significant progress that has been made in video generative models, existing state-of-the-art methods can only produce videos lasting 5-16 seconds, often labeled "long-form videos". Furthermore, videos exceeding 16 seconds struggle to maintain consistent character appearances and scene layouts throughout the narrative. In particular, multi-subject long videos still fail to preserve character consistency and motion coherence. While some methods can generate videos up to 150 seconds long, they often suffer from frame redundancy and low temporal diversity. Recent work has attempted to produce long-form videos featuring multiple characters, narrative coherence, and high-fidelity detail. We comprehensively studied 32 papers on video generation to identify key architectural components and training strategies that consistently yield these qualities. We also construct a comprehensive novel taxonomy of existing methods and present comparative tables that categorize papers by their architectural designs and performance characteristics.
Figures
Forward citations
Cited by 1 Pith paper
-
FilmWorld: Agentic Novel-to-Film Generation through Dynamic Cinematic World Modeling
FilmWorld generates multi-scene films from novels by materializing an explicit evolving world-state trajectory and rendering shots in parallel, beating five agents on its own FilmEval benchmark.
Reference graph
Works this paper leans on
-
[21]
Zico Kolter, and Kaiming He
Zhengyang Geng, Mingyang Deng, Xingjian Bai, J. Zico Kolter, and Kaiming He. Mean flows for one-step genera- tive modeling. CVPR, 2025. 7
2025
-
[1]
Frozen in time: A joint video and image encoder for end-to-end retrieval
Max Bain, Arsha Nagrani, G ¨ul Varol, and Andrew Zisser- man. Frozen in time: A joint video and image encoder for end-to-end retrieval. In Proceedings of the IEEE/CVF In- ternational Conference on Computer Vision, 2021. 15
2021
-
[2]
Lumiere: A space-time diffusion model for video generation
Omer Bar-Tal, Hila Chefer, Omer Tov, Charles Her- rmann, Roni Paiss, Shiran Zada, Ariel Ephrat, Junhwa Hur, Guanghui Liu, Amit Raj, Yuanzhen Li, Michael Rubinstein, Tomer Michaeli, Oliver Wang, Deqing Sun, Tali Dekel, and Inbar Mosseri. Lumiere: A space-time diffusion model for video generation. arXiv preprint arXiv:2401.12945, 2024. 2, 4, 6, 17
arXiv 2024
-
[3]
Is space-time attention all you need for video understanding? In Proceedings of the 38th International Conference on Ma- chine Learning (ICML), pages 813–824
Gedas Bertasius, Heng Wang, and Lorenzo Torresani. Is space-time attention all you need for video understanding? In Proceedings of the 38th International Conference on Ma- chine Learning (ICML), pages 813–824. PMLR, 2021. 1
2021
-
[4]
Tim Brooks, Janne Hellsten, Miika Aittala, Ting-Chun Wang, Timo Aila, Jaakko Lehtinen, Ming-Yu Liu, Alexei A. Efros, and Tero Karras. Generating long videos of dynamic scenes. arXiv preprint arXiv:2206.03429, 2022. 2
arXiv 2022
-
[5]
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Sub- biah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakan- tan, Pranav Shyam, Girish Sastry, Amanda Askell, Sand- hini Agarwal, Ariel Herbert-V oss, Gretchen Krueger, et al. Language models are few-shot learners. NeurIPS, 2020. 1
2020
-
[6]
Skyreels-v2: Infinite-length film generative model
Guibin Chen, Dixuan Lin, Jiangping Yang, Chunze Lin, Juncheng Zhu, Mingyuan Fan, Hao Zhang, Sheng Chen, Zheng Chen, Chengchen Ma, Weiming Xiong, Wei Wang, Nuo Pang, Kang Kang, Zhiheng Xu, Yuzhe Jin, Yupeng Liang, Yubing Song, Peng Zhao, Boyuan Xu, Di Qiu, Debang Li, Zhengcong Fei, Yang Li, and Yahui Zhou. Skyreels-v2: Infinite-length film generative model...
arXiv 2025
-
[7]
Pixart- α: Fast train- ing of diffusion transformer for photorealistic text-to-image synthesis, 2023
Junsong Chen, Jincheng Yu, Chongjian Ge, Lewei Yao, Enze Xie, Yue Wu, Zhongdao Wang, James Kwok, Ping Luo, Huchuan Lu, and Zhenguo Li. Pixart- α: Fast train- ing of diffusion transformer for photorealistic text-to-image synthesis, 2023. 8
2023
Show all 110 references
-
[8]
Panda-70m: Captioning 70m videos with multiple cross-modality teachers
Tsai-Shien Chen, Aliaksandr Siarohin, Willi Menapace, Ekaterina Deyneka, Hsiang-wei Chao, Byung Eun Jeon, Yuwei Fang, Hsin-Ying Lee, Jian Ren, Ming-Hsuan Yang, and Sergey Tulyakov. Panda-70m: Captioning 70m videos with multiple cross-modality teachers. In Proceedings of the IE...
2024
-
[9]
Panda-70m: Captioning 70m videos with multiple cross-modality teachers
Tsai-Shien Chen, Aliaksandr Siarohin, Willi Menapace, Ekaterina Deyneka, Hsiang-Wei Chao, Byung Eun Jeon, Yuwei Fang, Hsin-Ying Lee, Jian Ren, Ming-Hsuan Yang, and Sergey Tulyakov. Panda-70m: Captioning 70m videos with multiple cross-modality teachers. In Proceed- ings of the ...
2024 arXiv
-
[10]
Multi-subject open-set personalization in video generation
Tsai-Shien Chen, Aliaksandr Siarohin, Willi Menapace, Yuwei Fang, Kwot Sin Lee, Ivan Skorokhodov, Kfir Aberman, Jun-Yan Zhu, Ming-Hsuan Yang, and Sergey Tulyakov. Multi-subject open-set personalization in video generation. In CVPR, 2025. 2, 4, 5, 7, 8, 16
2025
-
[11]
Seine: Short-to-long video diffusion model for generative transition and prediction
Xinyuan Chen, Xin Chen, Xuan Wang, Youshan Zhuang, Xiaoxiao Li, Chaoen Xiao, Zhe Gan, and Lawrence Carin. Seine: Short-to-long video diffusion model for generative transition and prediction. ICLR, 2023. 2, 3, 6, 17
2023
-
[12]
Hunyuanvideo-avatar: High-fidelity audio-driven hu- man animation for multiple characters
Yi Chen, Sen Liang, Zixiang Zhou, Ziyao Huang, Yifeng Ma, Junshu Tang, Qin Lin, Yuan Zhou, and Qinglin Lu. Hunyuanvideo-avatar: High-fidelity audio-driven hu- man animation for multiple characters. arXiv preprint arXiv:2505.20156, 2025. 3, 5, 7, 8, 16
2025 arXiv
-
[13]
Unimax: Fairer and more effective language sampling for large-scale multilingual pretraining
Hyung Won Chung, Noah Constant, Xavier Garcia, Adam Roberts, Yi Tay, Sharan Narang, and Orhan Firat. Unimax: Fairer and more effective language sampling for large-scale multilingual pretraining. In International Conference on Learning Representations, 2023. 6
2023
-
[14]
Zhao, Yanping Huang, An- drew M
Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Yunxuan Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, Albert Webson, Shixiang Shane Gu, Zhuyun Dai, Mirac Suzgun, Xinyun Chen, Aakanksha Chowdhery, Alex Castro-Ros, Marie Pel- lat, Kevin Robin...
2024
-
[15]
Efficient video prediction via sparsely conditioned flow matching
Aram Davtyan, Sepehr Sameni, and Paolo Favaro. Efficient video prediction via sparsely conditioned flow matching. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 23263–23274, 2023. 7
2023
-
[16]
Brendan Duke, Abdalla Ahmed, Christian Wolf, Parham Aarabi, and Graham W. Taylor. Sstvos: Sparse spatiotem- poral transformers for video object segmentation. In Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 5912–5921, 2021. 1
2021
-
[17]
Scaling rectified flow transformers for high-resolution image synthesis
Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas M ¨uller, Harry Saini, Yam Levi, Dominik Lorenz, Axel Sauer, Frederic B ¨osel, Dustin Podell, Tim Dockhorn, Zion English, Kyle Lacey, Alex Goodwin, Yan- nik Marek, and Robin Rombach. Scaling rectified flow t...
2024
-
[18]
Ca2-vdm: Efficient autoregres- sive video diffusion model with causal generation and cache sharing
Kaifeng Gao, Jiaxin Shi, Hanwang Zhang, Chunping Wang, Jun Xiao, and Long Chen. Ca2-vdm: Efficient autoregres- sive video diffusion model with causal generation and cache sharing. ICML, 2025. 2, 3, 5, 17
2025
-
[19]
Berg, Arash Vah- dat, Alexei A
Tianyun Gao, Lanqing Hong, Tamara L. Berg, Arash Vah- dat, Alexei A. Efros, William T. Freeman, and Mohammad Norouzi. Dreamvideo: Composing your dream videos with customized subject and motion. CVPR, 2024. 2, 4, 5, 17
2024
-
[20]
Seedance 1.0: Exploring the boundaries of video generation models
Yu Gao, Haoyuan Guo, Tuyen Hoang, Weilin Huang, Lu Jiang, Fangyuan Kong, Huixia Li, Jiashi Li, Liang Li, Xi- aojie Li, Xunsong Li, Yifu Li, Shanchuan Lin, Zhijie Lin, Jiawei Liu, Shu Liu, Xiaonan Nie, Zhiwu Qing, Yuxi Ren, Li Sun, Zhi Tian, Rui Wang, Sen Wang, Guoqiang Wei, Gu...
2025 arXiv
-
[22]
Veo 3: Neural video generation with native audio
Google DeepMind. Veo 3: Neural video generation with native audio. Web demo, 2025. 2, 5, 16
2025
-
[23]
Animatediff: Animate your personalized text-to-image models
Shun Gu, Tianmin Shu, Yandong Guo, Zhihao Liang, Gu- oshuai Qin, and Qinghua Fei. Animatediff: Animate your personalized text-to-image models. GitHub repository,
-
[24]
Ltx-video: Realtime video latent diffusion
Yoav HaCohen, Nisan Chiprut, Benny Brazowski, Daniel Shalem, Dudu Moshe, Eitan Richardson, Eran Levin, Guy Shiran, Nir Zabari, Ori Gordon, Poriya Panet, Sapir Weiss- buch, Victor Kulikov, Yaki Bitterman, Zeev Melumian, and Ofir Bibi. Ltx-video: Realtime video latent diffusion....
2025 arXiv
-
[25]
Clipscore: A reference-free eval- uation metric for image captioning
Jack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras, and Yejin Choi. Clipscore: A reference-free eval- uation metric for image captioning. In Proceedings of the 2021 Conference on Empirical Methods in Natural Lan- guage Processing, pages 7514–7528. Association for Com- pu...
2021
-
[26]
Gans trained by a two time-scale update rule converge to a local nash equilib- rium
Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. Gans trained by a two time-scale update rule converge to a local nash equilib- rium. In Advances in Neural Information Processing Sys- tems, pages 6626–6637, 2017. 15
2017
-
[27]
Denoising dif- fusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising dif- fusion probabilistic models. In NeurIPS, 2020. 1, 6
2020
-
[28]
Jonathan Ho, Tim Salimans, Alexey Gritsenko, William Chan, Mohammad Norouzi, and David J. Fleet. Video dif- fusion models. In Advances in Neural Information Process- ing Systems (NeurIPS), 2022. 2, 3, 4, 6, 7, 17
2022
-
[29]
Jonathan Ho, Tim Salimans, Alexey Gritsenko, William Chan, Mohammad Norouzi, and David J. Fleet. Stable video diffusion. arXiv preprint arXiv:2311.15127, 2023. 2, 5, 7, 17
2023 arXiv
-
[30]
Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen- Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen
Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen- Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. Lora: Low-rank adaptation of large language models.arXiv preprint arXiv:2106.09685, 2021. 4
2021 arXiv
-
[31]
Hunyuancustom: A multimodal-driven architecture for customized video gen- eration
Teng Hu, Zhentao Yu, Zhengguang Zhou, Sen Liang, Yuan Zhou, Qin Lin, and Qinglin Lu. Hunyuancustom: A multimodal-driven architecture for customized video gen- eration. arXiv preprint arXiv:2505.04512, 2025. 2, 4, 5, 7, 8, 16
2025 arXiv
-
[32]
Step-video-ti2v: A state-of-the- art text-driven image-to-video generation model
Haoyang Huang, Guoqing Ma, Nan Duan, Xing Chen, Changyi Wan, Ranchen Ming, Tianyu Wang, Bo Wang, Zhiying Lu, Aojie Li, Xianfang Zeng, Xinhao Zhang, Gang Yu, Yuhe Yin, Qiling Wu, Wen Sun, Kang An, Xin Han, Deshan Sun, Wei Ji, Bizhu Huang, Brian Li, Chenfei Wu, Guanzhe Huang, Hu...
2025 arXiv
-
[33]
Conceptmaster: Multi-concept video customiza- tion on diffusion transformer models without test-time tun- ing
Yuzhou Huang, Ziyang Yuan, Quande Liu, Qiulin Wang, Xintao Wang, Ruimao Zhang, Pengfei Wan, Di Zhang, and Kun Gai. Conceptmaster: Multi-concept video customiza- tion on diffusion transformer models without test-time tun- ing. arXiv preprint arXiv:2501.04698, 2025. 2, 4, 5, 16
2025 arXiv
-
[34]
Vbench: Comprehensive benchmark suite for video generative mod- els
Ziqi Huang, Yinan He, Jiashuo Yu, Fan Zhang, Chenyang Si, Yuming Jiang, Yuanhan Zhang, Tianxing Wu, Qingyang Jin, Nattapol Chanpaisit, Yaohui Wang, Xinyuan Chen, Limin Wang, Dahua Lin, Yu Qiao, and Ziwei Liu. Vbench: Comprehensive benchmark suite for video generative mod- els....
-
[35]
Pika 1.5: Realistic scene and motion synthesis,
Snap Inc. Pika 1.5: Realistic scene and motion synthesis,
-
[36]
Sim2real: Synthetic video data for vision-based robotic manipulation learning
Stephen James, Ankit Gupta, and Andrew Davison. Sim2real: Synthetic video data for vision-based robotic manipulation learning. In Proceedings of the IEEE Inter- national Conference on Robotics and Automation , pages 7894–7901, 2022. 1
2022
-
[37]
Vace: All-in-one video creation and editing
Zeyinzi Jiang, Zhen Han, Chaojie Mao, Jingfeng Zhang, Yulin Pan, and Yu Liu. Vace: All-in-one video creation and editing. arXiv preprint arXiv:2503.07598, 2025. 2, 4, 5, 8, 16
2025 arXiv
-
[38]
Pyramidal flow matching for efficient video generative modeling
Yang Jin, Zhicheng Sun, Ningyuan Li, Kun Xu, Hao Jiang, Nan Zhuang, Quzhe Huang, Yang Song, Yadong Mu, and Zhouchen Lin. Pyramidal flow matching for efficient video generative modeling. ICLR, 2025. 2, 3, 5, 8
2025
-
[39]
Miradata: A large-scale video dataset with long du- rations and structured captions, 2024
Xuan Ju, Yiming Gao, Zhaoyang Zhang, Ziyang Yuan, Xin- tao Wang, Ailing Zeng, Yu Xiong, Qiang Xu, and Ying Shan. Miradata: A large-scale video dataset with long du- rations and structured captions, 2024. 1
2024
-
[40]
Miradata: A large-scale video dataset with long 10 durations and structured captions
Xuan Ju, Yiming Gao, Zhaoyang Zhang, Ziyang Yuan, Xin- tao Wang, Ailing Zeng, Yu Xiong, Qiang Xu, and Ying Shan. Miradata: A large-scale video dataset with long 10 durations and structured captions. In Proceedings of the 38th Conference on Neural Information Processing Systems...
2024
-
[41]
Cinediff: Diffusion models for cinematic video synthesis
Junhee Kim, Seong Park, and Donghyun Lee. Cinediff: Diffusion models for cinematic video synthesis. In Pro- ceedings of the IEEE International Conference on Com- puter Vision, pages 1456–1465, 2023. 1
2023
-
[42]
Hunyuanvideo: A systematic framework for large video generative models
Weijie Kong, Qi Tian, Zijian Zhang, Rox Min, Zuozhuo Dai, Jin Zhou, Jiangfeng Xiong, Xin Li, Bo Wu, Jian- wei Zhang, Kathrina Wu, Qin Lin, Junkun Yuan, Yanxin Long, Aladdin Wang, Andong Wang, Changlin Li, Duo- jun Huang, Fang Yang, Hao Tan, Hongmei Wang, Jacob Song, Jiawang Ba...
2024 arXiv
-
[43]
Minimax hailuo: Scalable multi- subject video generation, 2025
ByteDance AI Lab. Minimax hailuo: Scalable multi- subject video generation, 2025. Accessed: 17 June 2025. 2, 5
2025
-
[44]
Openhumanvid: A large-scale high- quality dataset for enhancing human-centric video genera- tion
Hui Li, Mingwang Xu, Yun Zhan, Shan Mu, Jiaye Li, Kai- hui Cheng, Yuxuan Chen, Tan Chen, Mao Ye, Jingdong Wang, and Siyu Zhu. Openhumanvid: A large-scale high- quality dataset for enhancing human-centric video genera- tion. In Proceedings of the IEEE/CVF Conference on Com- put...
2025
-
[45]
Open-sora plan: Open-source large video generation model
Bin Lin, Yunyang Ge, Xinhua Cheng, Zongjian Li, Bin Zhu, Shaodong Wang, Xianyi He, Yang Ye, Shenghai Yuan, Liuhan Chen, Tanghui Jia, Junwu Zhang, Zhenyu Tang, Yatian Pang, Bin She, Cen Yan, Zhiheng Hu, Xi- aoyi Dong, Lin Chen, Zhang Pan, Xing Zhou, Shaoling Dong, Yonghong Tian...
2024 arXiv
-
[46]
Videodirectorgpt: Consistent multi-scene video generation via llm-guided planning
Han Lin, Abhay Zala, Jaemin Cho, and Mohit Bansal. Videodirectorgpt: Consistent multi-scene video generation via llm-guided planning. arXiv preprint arXiv:2309.15091,
-
[47]
Yaron Lipman, Ricky T. Q. Chen, Heli BenHamu, Maxi- milian Nickel, and Matt Le. Flow matching for generative modeling. In Proceedings of the International Conference on Learning Representations (ICLR), 2023. 4, 7
2023
-
[48]
Phantom: Subject- consistent video generation via cross-modal alignment
Lijie Liu, Tianxiang Ma, Bingchuan Li, Zhuowei Chen, Ji- awei Liu, Qian He, and Xinglong Wu. Phantom: Subject- consistent video generation via cross-modal alignment. arXiv preprint arXiv:2502.11079, 2025. 2, 3, 4, 5, 7, 8, 16
2025 arXiv
-
[49]
Autostory: Asynchronous video gen- eration with auto-regressive diffusion
Yang Liu, Xiaolei Huang, Zhe Gan, Jian Tang, and Lawrence Carin. Autostory: Asynchronous video gen- eration with auto-regressive diffusion. arXiv preprint arXiv:2311.11243, 2023. 2, 3, 8
2023 arXiv
-
[50]
Latte: Latent diffusion transformer for video generation
Xin Ma, Yaohui Wang, Xinyuan Chen, Gengyun Jia, Zi- wei Liu, Yuan-Fang Li, Cunjian Chen, and Yu Qiao. Latte: Latent diffusion transformer for video generation. Transac- tions on Machine Learning Research, 2025. 8
2025
-
[51]
Gamegan: Video generation for atari games
Toni M ¨uller, Michael Abrash, and Ilya Sutskever. Gamegan: Video generation for atari games. InAdvances in Neural Information Processing Systems, pages 2427–2438,
-
[52]
Sora: Openai’s text-to-video generator, 2024
OpenAI. Sora: Openai’s text-to-video generator, 2024. 2, 5, 17
2024
-
[53]
Scalable diffusion mod- els with transformers
William Peebles and Saining Xie. Scalable diffusion mod- els with transformers. In Proceedings of the IEEE/CVF In- ternational Conference on Computer Vision (ICCV), pages 4195–4205, 2023. 4, 8
2023
-
[54]
Open- sora 2.0: Training a commercial-level video generation model in $200k
Xiangyu Peng, Zangwei Zheng, Chenhui Shen, Tom Young, Xinying Guo, Binluo Wang, Hang Xu, Hongxin Liu, Mingyan Jiang, Wenjun Li, Yuhui Wang, Anbang Ye, Gang Ren, Qianran Ma, Wanying Liang, Xiang Lian, Xi- wen Wu, Yuting Zhong, Zhuangyan Li, Chaoyu Gong, Guojun Lei, Leijun Cheng...
2025 arXiv
-
[55]
Sampson, Shikai Li, Si- mone Parmeggiani, Steve Fine, Tara Fowler, Vladan Petro- vic, and Yuming Du
Adam Polyak, Amit Zohar, Andrew Brown, Andros Tjan- dra, Animesh Sinha, Ann Lee, Apoorv Vyas, Bowen Shi, Chih-Yao Ma, Ching-Yao Chuang, David Yan, Dhruv Choudhary, Dingkang Wang, Geet Sethi, Guan Pang, Haoyu Ma, Ishan Misra, Ji Hou, Jialiang Wang, Kiran Ja- gadeesh, Kunpeng Li...
-
[56]
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning transferable visual models from natural language supervision. In Proceedings of th...
2021
-
[57]
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, 11 and Peter J. Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21:1–67, 2020. 6
2020
-
[58]
Gener- ating diverse high-fidelity images with vq-vae-2
Ali Razavi, Aaron van den Oord, and Oriol Vinyals. Gener- ating diverse high-fidelity images with vq-vae-2. In Ad- vances in Neural Information Processing Systems , pages 14837–14847, 2019. 7
2019
-
[59]
Runway gen-3: Advanced video syn- thesis platform, 2024
Runway Research. Runway gen-3: Advanced video syn- thesis platform, 2024. Accessed: 17 June 2025. 2, 5
2024
-
[60]
High-resolution image synthesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bj ¨orn Ommer. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 10684–10695, 2022. 6
2022
-
[61]
U- net: Convolutional networks for biomedical image segmen- tation
Olaf Ronneberger, Philipp Fischer, and Thomas Brox. U- net: Convolutional networks for biomedical image segmen- tation. In Proceedings of the International Conference on Medical Image Computing and Computer-Assisted Inter- vention (MICCAI), pages 234–241, 2015. 4, 7
2015
-
[62]
Goodfellow, Wojciech Zaremba, Vicki Cheung, Alec Radford, and Xi Chen
Tim Salimans, Ian J. Goodfellow, Wojciech Zaremba, Vicki Cheung, Alec Radford, and Xi Chen. Improved techniques for training GANs. In Advances in Neural Information Pro- cessing Systems, pages 2226–2234, 2016. 15
2016
-
[63]
Flashattention-3: Fast and accurate attention with asynchrony and low-precision,
Jay Shah, Ganesh Bikshandi, Ying Zhang, Vijay Thakkar, Pradeep Ramani, and Tri Dao. Flashattention-3: Fast and accurate attention with asynchrony and low-precision,
-
[64]
Videopoet: Large language models are zero-shot video generators
Andrew Singer, Tian Jian, Andr ´es Ma, Yifan Jiang, Linjie Yang, Daniel Khashabi, Shihan Su, Justin Johnson, Noah Snavely, Chenliang Xu, and Ming-Hsuan Yang. Videopoet: Large language models are zero-shot video generators. ICML, 2024. 2, 4, 6, 17
2024
-
[65]
Denois- ing diffusion implicit models
Jiaming Song, Chenlin Meng, and Stefano Ermon. Denois- ing diffusion implicit models. In ICLR, 2021. 1, 6
2021
-
[66]
Kling2.0: Proprietary high-fidelity video generation, 2024
Kuaishou Technology. Kling2.0: Proprietary high-fidelity video generation, 2024. Accessed: 17 June 2025. 2, 4
2024
-
[67]
Hansi Teng, Hongyu Jia, Lei Sun, Lingzhi Li, Maolin Li, Mingqiu Tang, Shuai Han, Tianning Zhang, W. ˜Q. Zhang, Weifeng Luo, Xiaoyang Kang, Yuchen Sun, Yue Cao, Yun- peng Huang, Yutong Lin, Yuxin Fang, Zewei Tao, Zheng Zhang, Zhongshu Wang, Zixun Liu, Dai Shi, Guoli Su, Hanwen ...
-
[68]
Mocogan: Decomposing motion and content for video generation
Sergey Tulyakov, Ming-Yu Liu, Xiaodong Yang, and Jan Kautz. Mocogan: Decomposing motion and content for video generation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 1526– 1535, 2018. 2
2018
-
[69]
Towards accurate generative models of video: A new met- ric & challenges
Thomas Unterthiner, Sjoerd van Steenkiste, Karol Kurach, Rapha¨el Marinier, Marcin Michalski, and Sylvain Gelly. Towards accurate generative models of video: A new met- ric & challenges. arXiv preprint arXiv:1812.01717, 2018. 15
2018 arXiv
-
[70]
Neural discrete representation learn- ing
Aaron van den Oord, Oriol Vinyals, and Koray Kavukcuoglu. Neural discrete representation learn- ing. In Advances in Neural Information Processing Systems (NeurIPS), pages 6306–6315, 2017. 7
2017
-
[71]
Wan: Open and ad- vanced large-scale video generative models
Ang Wang, Baole Ai, Bin Wen, Chaojie Mao, Chen-Wei Xie, Di Chen, Feiwu Yu, Haiming Zhao, Jianxiao Yang, Jianyuan Zeng, Jiayu Wang, Jingfeng Zhang, Jingren Zhou, Jinkai Wang, Jixuan Chen, Kai Zhu, Kang Zhao, Keyu Yan, Lianghua Huang, Mengyang Feng, Ningyi Zhang, Pandeng Li, Pin...
2025 arXiv
-
[72]
Modelscope text-to-video technical report
Jiuniu Wang, Hangjie Yuan, Dayou Chen, Yingya Zhang, Xiang Wang, and Shiwei Zhang. Modelscope text-to-video technical report. Technical Report arXiv:2308.06571, Al- ibaba DAMO Academy, 2023. 7
2023 arXiv
-
[73]
Text2video: Gener- ating educational videos from text
Lei Wang, Hui Chen, and Ming Li. Text2video: Gener- ating educational videos from text. In Proceedings of the AAAI Conference on Artificial Intelligence , pages 6543– 6551, 2023. 1
2023
-
[74]
Koala- 36m: A large-scale video dataset improving consistency between fine-grained conditions and video content, 2024
Qiuheng Wang, Yukai Shi, Jiarong Ou, Rui Chen, Ke Lin, Jiahao Wang, Boyuan Jiang, Haotian Yang, Mingwu Zheng, Xin Tao, Fei Yang, Pengfei Wan, and Di Zhang. Koala- 36m: A large-scale video dataset improving consistency between fine-grained conditions and video content, 2024. 1
2024
-
[75]
Koala- 36m: A large-scale video dataset improving consistency between fine-grained conditions and video content
Qiuheng Wang, Yukai Shi, Jiarong Ou, Rui Chen, Ke Lin, Jiahao Wang, Boyuan Jiang, Haotian Yang, Mingwu Zheng, Xin Tao, Fei Yang, Pengfei Wan, and Di Zhang. Koala- 36m: A large-scale video dataset improving consistency between fine-grained conditions and video content. In Pro- ...
2025 arXiv
-
[76]
Diffuse and disperse: Im- age generation with representation regularization
Runqian Wang and Kaiming He. Diffuse and disperse: Im- age generation with representation regularization. arXiv preprint arXiv:2506.09027, 2025. 7
2025 arXiv
-
[77]
Videofactory: Swap attention in spatiotemporal diffusions for text-to-video gen- eration
Wenjing Wang, Huan Yang, Zixi Tuo, Huiguo He, Junchen Zhu, Jianlong Fu, and Jiaying Liu. Videofactory: Swap attention in spatiotemporal diffusions for text-to-video gen- eration. arXiv preprint arXiv:2305.10874, 2023. 1
2023 arXiv
-
[78]
Videofactory: Swap attention in spatiotemporal diffusions for text-to-video gen- eration
Wenjing Wang, Huan Yang, Zixi Tuo, Huiguo He, Junchen Zhu, Jianlong Fu, and Jiaying Liu. Videofactory: Swap attention in spatiotemporal diffusions for text-to-video gen- eration. In International Conference on Learning Repre- sentations, 2024. Introduces the HD-VG-130M dataset. 15
2024
-
[79]
Non-local neural networks
Xiaolong Wang, Ross Girshick, Abhinav Gupta, and Kaim- ing He. Non-local neural networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recog- nition (CVPR), pages 7794–7803, 2018. 1 12
2018
-
[80]
Yuille, Zicheng Liu, and Emad Barsoum
Xingrui Wang, Jiang Liu, Ze Wang, Xiaodong Yu, Jialian Wu, Ximeng Sun, Yusheng Su, Alan L. Yuille, Zicheng Liu, and Emad Barsoum. Keyvid: Keyframe-aware video diffusion for audio-synchronized visual animation. arXiv preprint arXiv:2504.09656, 2025. 2, 3
2025
-
[81]
Loong: Generating minute-level long videos with autoregressive language models
Yuqing Wang, Tianwei Xiong, Daquan Zhou, Zhijie Lin, Yang Zhao, Bingyi Kang, Jiashi Feng, and Xi- hui Liu. Loong: Generating minute-level long videos with autoregressive language models. arXiv preprint arXiv:2410.02757, 2024. 2, 4, 6, 17
2024 arXiv
-
[82]
Bovik, Hamid R
Zhou Wang, Alan C. Bovik, Hamid R. Sheikh, and Eero P. Simoncelli. Image quality assessment: From error visibility to structural similarity. IEEE Transactions on Image Pro- cessing, 13(4):600–612, 2004. 7, 15
2004
-
[83]
Humanvid: Demystifying train- ing data for camera-controllable human image animation
Zhenzhi Wang, Yixuan Li, Yanhong Zeng, Youqing Fang, Yuwei Guo, Wenran Liu, Jing Tan, Kai Chen, Tianfan Xue, Bo Dai, and Dahua Lin. Humanvid: Demystifying train- ing data for camera-controllable human image animation. NeurIPS Datasets & Benchmarks, 2024. 15
2024
-
[84]
Freeflux: Understanding and exploiting layer-specific roles in rope-based mmdit for versatile image editing.arXiv preprint arXiv:2503.16153, 2025
Tianyi Wei, Yifan Zhou, Dongdong Chen, and Xingang Pan. Freeflux: Understanding and exploiting layer-specific roles in rope-based mmdit for versatile image editing.arXiv preprint arXiv:2503.16153, 2025. 8
2025 arXiv
-
[85]
Panacea: Panoramic and controllable video generation for au- tonomous driving
Jia Wen, Xiaolei Li, Kun Zhao, and Kai Lin. Panacea: Panoramic and controllable video generation for au- tonomous driving. In Proceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition, pages 6902–6912, 2024. 1
2024
-
[86]
Moviebench: A hierarchical movie- level dataset for long video generation
Weijia Wu, Mingyu Liu, Zeyu Zhu, Xi Xia, Haoen Feng, Wen Wang, Kevin Qinghong Lin, Chunhua Shen, and Mike Zheng Shou. Moviebench: A hierarchical movie- level dataset for long video generation. arXiv preprint arXiv:2411.15262, 2024. 1, 15
2024 arXiv
-
[87]
Auto- mated movie generation via multi-agent cot planning.arXiv preprint arXiv:2503.07314, 2025
Weijia Wu, Zeyu Zhu, and Mike Zheng Shou. Auto- mated movie generation via multi-agent cot planning.arXiv preprint arXiv:2503.07314, 2025. 2, 3, 8
2025 arXiv
-
[88]
Mind the time: Temporally- controlled multi-event video generation
Ziyi Wu, Aliaksandr Siarohin, Willi Menapace, Ivan Sko- rokhodov, Yuwei Fang, Varnith Chordia, Igor Gilitschen- ski, and Sergey Tulyakov. Mind the time: Temporally- controlled multi-event video generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Patter...
2025
-
[89]
A survey on video diffusion models
Zhen Xing, Qijun Feng, Haoran Chen, Qi Dai, Han Hu, Hang Xu, Zuxuan Wu, and Yu-Gang Jiang. A survey on video diffusion models. ACM Computing Surveys, 2023. 1
2023
-
[90]
Surgi- cal video synthesis using generative models for procedure training
Yijun Xu, Fei Ye, Holger Roth, and Nassir Navab. Surgi- cal video synthesis using generative models for procedure training. In Proceedings of the International Conference on Medical Image Computing and Computer-Assisted In- tervention, pages 324–332, 2020. 1
2020
-
[91]
Advancing high-resolution video-language representation with large-scale video transcriptions
Hongwei Xue, Tiankai Hang, Yanhong Zeng, Yuchong Sun, Bei Liu, Huan Yang, Jianlong Fu, and Baining Guo. Advancing high-resolution video-language representation with large-scale video transcriptions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recog...
2022
-
[92]
Temporally consistent transformers for video gen- eration
Wilson Yan, Danijar Hafner, Stephen James, and Pieter Abbeel. Temporally consistent transformers for video gen- eration. arXiv preprint arXiv:2210.02396, 2022. 2
2022 arXiv
-
[93]
Vript: A video is worth thousands of words
Dongjie Yang, Suyuan Huang, Chengqiang Lu, Xiaodong Han, Haoxin Zhang, Yan Gao, Yao Hu, and Hai Zhao. Vript: A video is worth thousands of words. In NeurIPS Datasets & Benchmarks, 2024. 15
2024
-
[94]
360° vr video genera- tion with generative adversarial networks
Li Yang, Rui Chen, and Wei Sun. 360° vr video genera- tion with generative adversarial networks. In Proceedings of ACM SIGGRAPH Asia, pages 1–10, 2021. 1
2021
-
[95]
Towards physically plau- sible video generation via vlm planning
Xindi Yang, Baolu Li, Yiming Zhang, Zhenfei Yin, Lei Bai, Liqian Ma, Zhiyong Wang, Jianfei Cai, Tien-Tsin Wong, Huchuan Lu, and Xu Jia. Towards physically plau- sible video generation via vlm planning. arXiv preprint arXiv:2503.23368, 2025. 2
2025 arXiv
-
[96]
Cogvideox: Text-to-video diffusion models with an expert transformer
Zhuoyi Yang, Jiayan Teng, Wendi Zheng, Ming Ding, Shiyu Huang, Jiazheng Xu, Yuanming Yang, Wenyi Hong, Xiaohan Zhang, Guanyu Feng, Da Yin, Yuxuan Zhang, Weihan Wang, Yean Cheng, Bin Xu, Xiaotao Gu, Yuxiao Dong, and Jie Tang. Cogvideox: Text-to-video diffusion models with an ex...
2025
-
[97]
Magvit-v2: Language model beats diffusion — tokenizer is key to visual generation
Lijun Yu, Fanghui Li, Xudong Jiang, Ming Lin, Yong Liu, and Weinan Zhang. Magvit-v2: Language model beats diffusion — tokenizer is key to visual generation. arXiv preprint arXiv:2310.05737, 2023. 2, 3, 4, 6, 17
2023 arXiv
-
[98]
Framepack: Pack- ing input frame context for next-frame prediction models
Lvmin Zhang and Maneesh Agrawala. Framepack: Pack- ing input frame context for next-frame prediction models. arXiv preprint arXiv:2504.12626, 2025. 2, 3
2025
-
[99]
Efros, Eli Shecht- man, and Oliver Wang
Richard Zhang, Phillip Isola, Alexei A. Efros, Eli Shecht- man, and Oliver Wang. The unreasonable effectiveness of deep features as a perceptual metric. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recog- nition, pages 586–595, 2018. 7, 15
2018
-
[100]
I2vgen-xl: High-quality image-to-video synthesis via cascaded diffusion models
Shiwei Zhang, Changan Chen, Hao Zhang, Jianlong Fu, Jiebo Luo, and Chuanzhi Chen. I2vgen-xl: High-quality image-to-video synthesis via cascaded diffusion models. arXiv preprint arXiv:2311.04145, 2023. 2, 3, 5, 17
2023 arXiv
-
[101]
Diffusion-driven promotional video generation for marketing
Yifan Zhang, Deepa Patel, and Arjun Roy. Diffusion-driven promotional video generation for marketing. In Proceed- ings of ACM Multimedia, pages 1122–1131, 2023. 1
2023
-
[102]
Flow-guided one-shot talking face generation with a high-resolution audio-visual dataset
Zhimeng Zhang, Lincheng Li, Yu Ding, and Changjie Fan. Flow-guided one-shot talking face generation with a high-resolution audio-visual dataset. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021. Introduces the HDTF dataset. 15
2021
-
[103]
Magicvideo: Efficient video generation with latent diffusion models
Daquan Zhou, Weimin Wang, Hanshu Yan, Weiwei Lv, Yizhe Zhu, and Jiashi Feng. Magicvideo: Efficient video generation with latent diffusion models. arXiv preprint arXiv:2211.11018, 2022. 7
2022 arXiv
-
[104]
Storydiffusion: Consistent self-attention for long-range image and video generation
Yupeng Zhou, Daquan Zhou, Ming-Ming Cheng, Ji- ashi Feng, and Qibin Hou. Storydiffusion: Consistent self-attention for long-range image and video generation. NeurIPS, 2024. 2, 3, 5, 8, 17
2024
-
[105]
Celebv- hq: A large-scale video facial attributes dataset
Hao Zhu, Wayne Wu, Wentao Zhu, Liming Jiang, Siwei Tang, Li Zhang, Ziwei Liu, and Chen Change Loy. Celebv- hq: A large-scale video facial attributes dataset. In Euro- pean Conference on Computer Vision, 2022. 15 13
2022
-
[106]
CelebV- HQ: A large-scale video facial attributes dataset
Hao Zhu, Wayne Wu, Wentao Zhu, Liming Jiang, Siwei Tang, Li Zhang, Ziwei Liu, and Chen Change Loy. CelebV- HQ: A large-scale video facial attributes dataset. In ECCV,
-
[109]
scripts
supply face tracks, skeletons and camera-motion la- bels, but most clips remain under 20 s, limiting long-form training. Vript [93] provides six-minute films with 145-word scene-level “scripts.” Large-scale video–text datasets like HD-VILA-100M [91] and Panda-70M [9] have enab...
-
[110]
(Continued) Paper’s Github Stars Date Tasks Generation Statistics Subjects Dataset Affiliation TV IV VV VE Len
1K 05’25 TV IV VV 5 26 129 M OpenHumanvid, Panda-2M Tencent Veo3 [22] - 05’25 TV IV VE 8 60 480 M - Google SkyReels-v2 [6] 2.7K 04’25 TV IV VE 30 24 720 M Koala-36M, HumanVid Skywork AI Open-Sora 2.0 [54] 26.6K 03’25 TV IV 5 24 128 M WebVid-10M, Panda-70M, HD-VG-130M, MiraData...
-
[2022]
Overview of VBench evaluation metrics [34]
1 14 Figure 2. Overview of VBench evaluation metrics [34]. VBench measures visual quality, motion smoothness, identity consistency, temporal flicker, spatial coherence, and text–video relevance to provide a fine-grained, multi-dimensional assessment of generated videos. Append...
-
[2025]
Accessed: 17 June 2025. 2, 5
2025
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.