REVIEW 4 major objections 6 minor 3 cited by
VidTwin: Video VAE with Decoupled Structure and Dynamics
T0 review · 4 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read VidTwin argues that decoupling video into Structure and Dynamics latents yields roughly 500x compression (0.20%) while preserving reconstruction quality and enabling latent diffusion training.
desk verdict A genuinely new two-stream latent geometry that compresses video ~500x with solid reconstructions, but the dynamics stream's marginal averaging structurally limits fine-grained fast motion and the headline claims outrun the evidence. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is a Q-Former, a transformer with learned query tokens that reads a sequence and outputs a few representative vectors via cross-attention. Here it is run along the temporal dimension with the spatial dimensions folded into the batch, so the Structure latent is forced to be location-independent. Dynamics extraction uses spatial downsampling followed by separate averaging over height and width, which collapses $O(h \cdot w)$ spatial information to $O(h+w)$ per frame while keeping the spatial axes. At decode time, upsampling heads align both latents to the encoder latent shape, and the video is reconstructed as $\hat{x} = D(u_S + u_D^{(h)} + u_D^{(w)})$. The additive combination is the mechanism that both compresses the representation and makes the two roles inspectable by decoding either summand alone.
What would settle it
Take a synthetic clip whose only motion is a fine checkerboard pattern shifting one pixel per frame while the average of every row and every column stays constant. If VidTwin, trained on such clips, reconstructs the moving checkerboard with high PSNR, the additive decomposition carries enough information; if the pattern blurs into a static gray field, the Dynamics latents have lost the spatial phase that the architecture assumed they could discard.
Extended reading notes
Core claim
On its own terms, the paper's central discovery is that representing a video as the sum of a Q-Former-extracted Structure latent and a per-axis-averaged Dynamics latent is enough to reconstruct high-quality video at extreme compression. The Structure latent is produced by folding spatial positions into the batch dimension and letting learned queries cross-attend along time, so it captures location-independent, low-frequency motion trends; after spatial downsampling it is a small volume. The Dynamics latent is produced by spatially downsampling the encoder output and averaging over height and width separately, shrinking each frame's motion representation from $O(h \cdot w)$ to $O(h+w)$. The two are upsampled, added elementwise, and decoded. The paper shows that this additive combination, trained with reconstruction, perceptual, adversarial, and KL losses, reaches 28.14 PSNR at 0.20% compression, outperforms uniform-latent and content-frame baselines on several quality metrics while using 2.5x to 30x smaller latents, and produces latents that support diffusion training, reported as FVD 193 on UCF-101.
Load-bearing premise
The load-bearing assumption is that adding the upsampled Structure latent to the upsampled Dynamics latents, with no interaction terms, reconstructs the video even though the Dynamics stream throws away spatial phase by averaging over height and width; everything not recoverable from per-axis averages must be smuggled through the Structure stream.
Editorial extensions
If this is right
- At 0.20% compression, a downstream DiT-style diffusion model on VidTwin latents consumes roughly 4x to 8x fewer FLOPs and 2x to 3x less training memory than on baseline latents, making video generation cheaper to train and deploy.
- The same decoder accepts latents produced by a generative model, so VidTwin can serve as the tokenizer for class-conditional or text-conditional video diffusion; the paper reports FVD 193 on UCF-101.
- Because decoding either summand alone yields interpretable content-only or motion-only video, the latent space supports cross-reenactment: structure from one video plus dynamics from another inherits object identity from the first and motion and color from the second.
- Scaling the transformer backbone from 126M to 1.3B parameters raises reconstruction PSNR from 24.83 to 27.16 at the same training steps, suggesting the design benefits from model scale.
- Compression rate and quality trade off smoothly: the same architecture at 0.11%, 0.16%, 0.20%, and 0.48% compression yields PSNRs of 24.41, 27.03, 28.14, and 30.04, so users can pick an operating point.
Reading between the lines
- As an extension, if the structure/dynamics split is as clean as the cross-reenactment examples suggest, the Structure latent could serve as a compact video-understanding representation and the Dynamics latent as a reusable motion token for controllable generation, directions the paper only gestures at.
- As a testable extension, the Dynamics stream discards spatial phase information when it averages over height and width, so any reconstruction accuracy must come from the Structure stream compensating; a video whose rapid changes cancel under both averages, such as a moving checkerboard with constant row and column means, should expose where that compensation breaks.
- Given the failure mode the paper reports for fast-moving basketball players, a natural next experiment is to give the Dynamics stream a small number of phase or position tokens instead of pure averages and measure how much compression must be sacrificed to remove the blur.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes VidTwin, a video autoencoder that encodes a video into two latent spaces: a Structure latent (extracted by a Q-Former along the temporal dimension, then spatially downsampled) said to capture global content and low-frequency motion, and a Dynamics latent (obtained by spatially downsampling and then averaging over height and width) said to capture fine-grained details and rapid motion. The two latents are separately patchified and concatenated for use as a diffusion-model training target. The authors report a 0.20% compression rate (about 500x) with 28.14 dB PSNR on MCL-JCV, claiming competitive or superior reconstruction against existing video autoencoders, favorable memory/FLOPs in a downstream DiT, and successful class-conditional generation on UCF-101. Appendix B.4 discloses a failure mode for fast-moving basketball players, who appear blurred despite a preserved background.
Significance. If the empirical claims hold, the paper makes a useful contribution: a compact video VAE that decouples structure and dynamics, potentially alleviating the latent bottleneck in video diffusion models. The reported 500x compression with reasonable reconstruction is attractive, and the explicit design of two complementary latents is conceptually novel. The resource-consumption analysis in Sec. 4.4.2 is a helpful practical benchmark. However, the paper provides no code or checkpoints, and the central interpretability claim—that the Dynamics latent captures fine-grained details and rapid motion—is not backed by an analysis of the information discarded by the spatial averaging step. The comparison to baselines is weakened by reimplementations trained on different private data. These issues mean the significance can only be provisional until the claims are qualified and the experiments are made more controlled.
major comments (4)
- [Sec. 3.2.2 and Sec. 3.3] The Dynamics latent is formed by averaging z'_D along height and width and then decoding via repetition (Fig. 2, Eq. in Sec. 3.3). Consequently, u_D^(h) is constant along width, u_D^(w) is constant along height, and their sum can only express spatial patterns of the form a_h + b_w (additive row/column effects). All joint spatial phase information—i.e., which specific locations move together—is discarded by the Dynamics branch. The claim in the Abstract and Sec. 1 that Dynamics latent 'represent[s] fine-grained details and rapid movements' is therefore structurally unsupported for spatially localized fast motion; the failure case in Appendix B.4 (blurred fast-moving players) is the expected consequence of this bottleneck, not a rare outlier. The authors should either temper the claim, provide an information-theoretic analysis of what the Dynamics branch can and cannot encode, or design a controlled experiment showing how the Structure branch and the decoder compensate for the lost phase.
- [Table 1 and Appendix C.1] The baseline comparison is not yet convincing. MAGVIT-v2, iVideoGPT, and CMD, which are used in the headline quantitative results, are reimplemented by the authors based on paper descriptions (Appendix C.1), with no official checkpoints, and are trained on datasets that are not reported. VidTwin is trained on a private 10M-pair video-text dataset (Sec. 4.1.1). The reported improvements in PSNR/SSIM/LPIPS/FVD may therefore reflect differences in training data, compute, or model scale rather than the architectural contribution. The authors should either retrain all baselines from the same data with matched compute, run official checkpoints where available (e.g., for CV-VAE and EMU-3), or at minimum report baseline training settings (data, steps, resolution) and discuss the confound explicitly.
- [Table 1, MOS scores] The subjective MOS results are reported as single averages (e.g., Sem. 4.71 vs 4.70) with no error bars, no per-evaluator variance, and no significance test. With 15 evaluators and 20 samples, differences of 0.01 are not meaningful, yet the text describes 'outperforms' based on these values. The authors should provide confidence intervals, inter-rater agreement, or a statistical test, and should avoid claiming superiority where the difference is within noise.
- [Table 2 and Sec. 4.4.3] The UCF-101 generation comparison is uncontrolled: TATS, MAGVIT-v2, and Video-LaViT are trained with different architectures, datasets (likely not the same as the authors' 10M set), and compute budgets, and no variance or confidence intervals are reported for FVD. The FVD value of 193 is far from MAGVIT-v2's 58, so the claim of being 'comparable' is not supported as stated. The paper should either provide a matched experimental protocol (same training data, steps, and compute) or restrict the claim to 'the latent is trainable in a simple diffusion setup,' which is the only claim the present experiment can actually support.
minor comments (6)
- [Sec. 3.1] The reconstruction loss is written as Lrec = ∥ˆx − x∥ without specifying the norm or whether it is summed over channels/frames; please clarify, e.g., the L2 norm.
- [Sec. 3.2.2] The phrase 'preserving spatial integrity' after the averaging operation is misleading: the operation strictly discards spatial phase. Please rephrase to describe what is actually preserved (global row/column statistics).
- [Sec. 3.5] There is a typo: 'the Structure Latent latent zS' should read 'the Structure Latent zS' (redundant word).
- [Fig. 5 and Sec. 4.4.2] The FLOPs/memory comparison uses a 'pseudo uniform DiT'; the authors should state explicitly that the resource savings are a direct consequence of the latent dimension, not of the architecture per se, and that the comparison is a resource benchmark rather than a quality benchmark.
- [Appendix C.1] The compression-rate calculation for iVideoGPT is not transparent (the formula N0d + n(T − T0)d and the numerical value 2 × 16^2 × 64 + 14 × 4^2 × 64 do not match the given expression). Please define all symbols and show the intermediate steps.
- [Appendix B.4] The failure-mode description states that the D. Latent 'captures the fast-changing players but struggles to accurately integrate them.' This is in tension with the claim in Sec. 1 and the Abstract that Dynamics Latent represents rapid movements. Please add a discussion of how this failure case relates to the averaging bottleneck and what the practical limits are.
Circularity Check
No circularity found: the paper's compression and reconstruction results come from defined dimension ratios and external-benchmark measurements, and its self-citations are confined to related work.
full rationale
The paper's headline quantitative claims are not circular. The 0.20% compression rate is an arithmetic ratio of latent dimensions to input video dimensions, explicitly defined in Sec. 4.2 and Appendix B.1, and applied uniformly to baselines; it is not a fitted parameter renamed as a prediction. Reconstruction quality (PSNR 28.14, SSIM, LPIPS, FVD) is measured against external baselines on MCL-JCV, and downstream generation is measured on UCF-101; these are standard evaluation results, not quantities equivalent to training inputs by construction. The claimed roles of Structure and Dynamics latents are supported by qualitative decoding experiments using the model's own decoder, but the paper explicitly acknowledges that isolating the latents 'inevitably introduces information loss' (Sec. 4.4.1) and does not present a formal derivation whose conclusion is assumed by its premises. The failure mode in Appendix B.4, where fast-moving basketball players appear blurred, is an admitted limitation rather than a hidden circular step. The paper's self-citations (GAIA, InstructAvatar, VidTok, Video in-context learning) appear only in related-work discussions and are not load-bearing for the method's derivation; no uniqueness theorem or ansatz is imported from the authors' prior work to force the architecture. The marginal-averaging bottleneck in the Dynamics stream is a potential information-loss concern for very fast complex motion, but that is a correctness or generalization risk, not circularity: the paper does not claim the Dynamics stream alone can reconstruct such motion, and it reports the corresponding failure. Overall, the derivation chain is self-contained against external benchmarks, so no circular step is present.
Assumptions & free parameters
free parameters (3)
- Structure latent geometry (nq, dS, hS, wS) =
nq=16, dS=4, hS=wS=7
- Dynamics latent geometry (dD, hD, wD) =
dD=8, hD=wD=7
- Loss weights (lambda_p, lambda_GAN, lambda_KL) =
0.05, 0.05, 0.001
assumptions (5)
- ad hoc to paper Video information is additively separable in latent space: the upsampled Structure and Dynamics latents sum to a faithful decoding x-hat = D(uS + uD).
- domain assumption A temporal Q-Former with nq=16 queries can summarize the low-frequency motion trends of the whole video without spatial information.
- domain assumption Spatial averaging of the dynamics stream to O(h+w) per frame retains enough information to reconstruct fine details when combined with the structure stream.
- domain assumption The reimplemented versions of MAGVIT-v2, iVideoGPT, and CMD are faithful to the original papers.
- standard math Standard VAE reparameterization, KL regularization, and DDPM/DDIM diffusion objectives yield a smooth latent space suitable for generation.
invented entities (2)
-
Structure Latent (zS)
-
Dynamics Latent (zD)
Cite this review
Pith. "Pith review of VidTwin: Video VAE with Decoupled Structure and Dynamics." pith.science (2026). https://pith.science/paper/4DAKTNNX
@misc{pith2026241217726,
author = {Pith},
title = {Pith review of: VidTwin: Video VAE with Decoupled Structure and Dynamics},
year = {2026},
howpublished = {\url{https://pith.science/paper/4DAKTNNX}},
note = {Machine review of arXiv:2412.17726}
}
read the original abstract
Recent advancements in video autoencoders (Video AEs) have significantly improved the quality and efficiency of video generation. In this paper, we propose a novel and compact video autoencoder, VidTwin, that decouples video into two distinct latent spaces: Structure latent vectors, which capture overall content and global movement, and Dynamics latent vectors, which represent fine-grained details and rapid movements. Specifically, our approach leverages an Encoder-Decoder backbone, augmented with two submodules for extracting these latent spaces, respectively. The first submodule employs a Q-Former to extract low-frequency motion trends, followed by downsampling blocks to remove redundant content details. The second averages the latent vectors along the spatial dimension to capture rapid motion. Extensive experiments show that VidTwin achieves a high compression rate of 0.20% with high reconstruction quality (PSNR of 28.14 on the MCL-JCV dataset), and performs efficiently and effectively in downstream generative tasks. Moreover, our model demonstrates explainability and scalability, paving the way for future research in video latent representation and generation. Check our project page for more details: https://vidtwin.github.io/.
Figures
Figures from the paper (5 more)
Forward citations
Cited by 3 Pith papers
-
Hi-VAE: Efficient Video Autoencoding with Global and Detailed Motion
A hierarchical motion autoencoder with a conditional diffusion decoder reconstructs 16-frame videos from latents as small as 0.07% of the input size while maintaining competitive PSNR and perceptual scores.
-
V-RAE: Rethinking Video Latent Spaces for Generation
Videos can be reconstructed and generated from temporally compressed, frozen semantic features, and the compressed space supports faster and better class-conditional video generation than conventional video VAE latents.
-
Infinite Video Understanding
The paper argues that video understanding research should aim at processing streams of arbitrary, unbounded duration and outlines the challenges, directions, and metrics needed.
Reference graph
Works this paper leans on
-
[1]
Lumiere: A space-time diffusion model for video generation, 2024
Omer Bar-Tal, Hila Chefer, Omer Tov, Charles Her- rmann, Roni Paiss, Shiran Zada, Ariel Ephrat, Junhwa Hur, Guanghui Liu, Amit Raj, Yuanzhen Li, Michael Rubinstein, Tomer Michaeli, Oliver Wang, Deqing Sun, Tali Dekel, and Inbar Mosseri. Lumiere: A space-time diffusion model for video generation, 2024. 2
work page 2024
-
[2]
Is space-time attention all you need for video understanding?,
Gedas Bertasius, Heng Wang, and Lorenzo Torresani. Is space-time attention all you need for video understanding?,
-
[3]
Stable video diffusion: Scaling latent video diffusion models to large datasets, 2023
Andreas Blattmann, Tim Dockhorn, Sumith Kulal, Daniel Mendelevitch, Maciej Kilian, Dominik Lorenz, Yam Levi, Zion English, Vikram V oleti, Adam Letts, Varun Jampani, and Robin Rombach. Stable video diffusion: Scaling latent video diffusion models to large datasets, 2023. 1
2023
-
[4]
Align your latents: High-resolution video synthesis with la- tent diffusion models, 2023
Andreas Blattmann, Robin Rombach, Huan Ling, Tim Dock- horn, Seung Wook Kim, Sanja Fidler, and Karsten Kreis. Align your latents: High-resolution video synthesis with la- tent diffusion models, 2023. 1, 2
work page 2023
-
[5]
Pixart-α: Fast training of dif- fusion transformer for photorealistic text-to-image synthesis,
Junsong Chen, Jincheng Yu, Chongjian Ge, Lewei Yao, Enze Xie, Yue Wu, Zhongdao Wang, James Kwok, Ping Luo, Huchuan Lu, and Zhenguo Li. Pixart-α: Fast training of dif- fusion transformer for photorealistic text-to-image synthesis,
-
[6]
Pixart-sigma: Weak-to-strong training of diffusion transformer for 4k text-to-image generation, 2024
Junsong Chen, Chongjian Ge, Enze Xie, Yue Wu, Lewei Yao, Xiaozhe Ren, Zhongdao Wang, Ping Luo, Huchuan Lu, and Zhenguo Li. Pixart-sigma: Weak-to-strong training of diffusion transformer for 4k text-to-image generation, 2024. 2
work page 2024
-
[7]
Od-vae: An omni-dimensional video compressor for im- proving latent video diffusion model, 2024
Liuhan Chen, Zongjian Li, Bin Lin, Bin Zhu, Qian Wang, Shenghai Yuan, Xing Zhou, Xinhua Cheng, and Li Yuan. Od-vae: An omni-dimensional video compressor for im- proving latent video diffusion model, 2024. 2, 12
work page 2024
-
[8]
Taming transformers for high-resolution image synthesis, 2021
Patrick Esser, Robin Rombach, and Bj ¨orn Ommer. Taming transformers for high-resolution image synthesis, 2021. 2, 4, 14
work page 2021
Show all 62 references
-
[9]
Scaling rectified flow trans- formers for high-resolution image synthesis, 2024
Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas M ¨uller, Harry Saini, Yam Levi, Dominik Lorenz, Axel Sauer, Frederic Boesel, Dustin Podell, Tim Dockhorn, Zion English, Kyle Lacey, Alex Goodwin, Yan- nik Marek, and Robin Rombach. Scaling rectified flow tr...
2024
-
[10]
Long video generation with time-agnostic vqgan and time- sensitive transformer, 2022
Songwei Ge, Thomas Hayes, Harry Yang, Xi Yin, Guan Pang, David Jacobs, Jia-Bin Huang, and Devi Parikh. Long video generation with time-agnostic vqgan and time- sensitive transformer, 2022. 2, 8
2022
-
[11]
Maskvit: Masked visual pre-training for video prediction, 2022
Agrim Gupta, Stephen Tian, Yunzhi Zhang, Jiajun Wu, Roberto Mart´ın-Mart´ın, and Li Fei-Fei. Maskvit: Masked visual pre-training for video prediction, 2022. 2
2022
-
[12]
Gaia: Zero- shot talking avatar generation, 2024
Tianyu He, Junliang Guo, Runyi Yu, Yuchi Wang, Jialiang Zhu, Kaikai An, Leyi Li, Xu Tan, Chunyu Wang, Han Hu, HsiangTao Wu, Sheng Zhao, and Jiang Bian. Gaia: Zero- shot talking avatar generation, 2024. 3
2024
-
[13]
Latent video diffusion models for high-fidelity long video generation, 2023
Yingqing He, Tianyu Yang, Yong Zhang, Ying Shan, and Qifeng Chen. Latent video diffusion models for high-fidelity long video generation, 2023. 1
2023
-
[14]
Classifier-free diffusion guidance, 2022
Jonathan Ho and Tim Salimans. Classifier-free diffusion guidance, 2022. 5, 16
2022
-
[15]
Denoising diffu- sion probabilistic models, 2020
Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffu- sion probabilistic models, 2020. 5, 17
2020
-
[16]
Kingma, Ben Poole, Mohammad Norouzi, David J
Jonathan Ho, William Chan, Chitwan Saharia, Jay Whang, Ruiqi Gao, Alexey Gritsenko, Diederik P. Kingma, Ben Poole, Mohammad Norouzi, David J. Fleet, and Tim Sali- mans. Imagen video: High definition video generation with diffusion models, 2022. 2
2022
-
[17]
Jonathan Ho, Tim Salimans, Alexey Gritsenko, William Chan, Mohammad Norouzi, and David J. Fleet. Video diffu- sion models, 2022. 1, 2
2022
-
[18]
Image quality metrics: Psnr vs
Alain Hor ´e and Djemel Ziou. Image quality metrics: Psnr vs. ssim. In 2010 20th International Conference on Pattern Recognition, pages 2366–2369, 2010. 6
2010
-
[19]
Dive: Dit-based video generation with enhanced control, 2024
Junpeng Jiang, Gangyi Hong, Lijun Zhou, Enhui Ma, Heng- tong Hu, Xia Zhou, Jie Xiang, Fan Liu, Kaicheng Yu, Haiyang Sun, Kun Zhan, Peng Jia, and Miao Zhang. Dive: Dit-based video generation with enhanced control, 2024. 2
2024
-
[20]
Video-lavit: Unified video-language pre-training with decoupled visual-motional tokenization, 2024
Yang Jin, Zhicheng Sun, Kun Xu, Kun Xu, Liwei Chen, Hao Jiang, Quzhe Huang, Chengru Song, Yuliang Liu, Di Zhang, Yang Song, Kun Gai, and Yadong Mu. Video-lavit: Unified video-language pre-training with decoupled visual-motional tokenization, 2024. 1, 3, 8
2024
-
[21]
Kingma and Jimmy Ba
Diederik P. Kingma and Jimmy Ba. Adam: A method for stochastic optimization, 2017. 14, 16
2017
-
[22]
Auto-encoding varia- tional bayes, 2022
Diederik P Kingma and Max Welling. Auto-encoding varia- tional bayes, 2022. 2, 17
2022
-
[23]
Mpeg: A video compression standard for multimedia applications
Didier Le Gall. Mpeg: A video compression standard for multimedia applications. Communications of the ACM , 34 (4):46–58, 1991. 3
1991
-
[24]
Disentangled motion modeling for video frame interpolation, 2024
Jaihyun Lew, Jooyoung Choi, Chaehun Shin, Dahuin Jung, and Sungroh Yoon. Disentangled motion modeling for video frame interpolation, 2024. 3
2024
-
[25]
Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models, 2023
Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models, 2023. 2, 4, 15
2023
-
[26]
Wf-vae: Enhancing video vae by wavelet-driven energy flow for latent video diffusion model, 2024
Zongjian Li, Bin Lin, Yang Ye, Liuhan Chen, Xinhua Cheng, Shenghai Yuan, and Li Yuan. Wf-vae: Enhancing video vae by wavelet-driven energy flow for latent video diffusion model, 2024. 2, 12
2024
-
[27]
Open-sora plan: Open-source large video generation model, 2024
Bin Lin, Yunyang Ge, Xinhua Cheng, Zongjian Li, Bin Zhu, Shaodong Wang, Xianyi He, Yang Ye, Shenghai Yuan, Li- uhan Chen, Tanghui Jia, Junwu Zhang, Zhenyu Tang, Ya- tian Pang, Bin She, Cen Yan, Zhiheng Hu, Xiaoyi Dong, Lin Chen, Zhang Pan, Xing Zhou, Shaoling Dong, Yonghong Ti...
2024
-
[28]
Finite scalar quantization: Vq-vae made simple, 2023
Fabian Mentzer, David Minnen, Eirikur Agustsson, and Michael Tschannen. Finite scalar quantization: Vq-vae made simple, 2023. 2
2023
-
[29]
Video generation models as world simulators
OpenAI. Video generation models as world simulators
-
[30]
Scalable diffusion models with transformers, 2023
William Peebles and Saining Xie. Scalable diffusion models with transformers, 2023. 8, 14, 15
2023
-
[31]
Sdxl: Improving latent diffusion models for high-resolution image synthesis, 2023
Dustin Podell, Zion English, Kyle Lacey, Andreas Blattmann, Tim Dockhorn, Jonas M ¨uller, Joe Penna, and Robin Rombach. Sdxl: Improving latent diffusion models for high-resolution image synthesis, 2023. 1, 2
2023
-
[32]
High-resolution image synthesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bj ¨orn Ommer. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10684–10695, 2022. 1, 2
2022
-
[33]
U-net: Convolutional networks for biomedical image segmentation,
Olaf Ronneberger, Philipp Fischer, and Thomas Brox. U-net: Convolutional networks for biomedical image segmentation,
-
[34]
Adversarial diffusion distillation, 2023
Axel Sauer, Dominik Lorenz, Andreas Blattmann, and Robin Rombach. Adversarial diffusion distillation, 2023. 1
2023
-
[35]
Motion-i2v: Consistent and controllable image-to-video generation with explicit motion modeling, 2024
Xiaoyu Shi, Zhaoyang Huang, Fu-Yun Wang, Weikang Bian, Dasong Li, Yi Zhang, Manyuan Zhang, Ka Chun Che- ung, Simon See, Hongwei Qin, Jifeng Dai, and Hongsheng Li. Motion-i2v: Consistent and controllable image-to-video generation with explicit motion modeling, 2024. 3
2024
-
[36]
Make-a-video: Text-to-video generation without text-video data, 2022
Uriel Singer, Adam Polyak, Thomas Hayes, Xi Yin, Jie An, Songyang Zhang, Qiyuan Hu, Harry Yang, Oron Ashual, Oran Gafni, Devi Parikh, Sonal Gupta, and Yaniv Taigman. Make-a-video: Text-to-video generation without text-video data, 2022. 2
2022
-
[37]
Denois- ing diffusion implicit models, 2022
Jiaming Song, Chenlin Meng, and Stefano Ermon. Denois- ing diffusion implicit models, 2022. 5, 16, 17
2022
-
[38]
Ucf101: A dataset of 101 human actions classes from videos in the wild, 2012
Khurram Soomro, Amir Roshan Zamir, and Mubarak Shah. Ucf101: A dataset of 101 human actions classes from videos in the wild, 2012. 2, 5, 8, 15
2012
-
[39]
Vidtok: A versatile and open-source video tokenizer
Anni Tang, Tianyu He, Junliang Guo, Xinle Cheng, Li Song, and Jiang Bian. Vidtok: A versatile and open-source video tokenizer. arXiv preprint arXiv:2412.13061, 2024. 2
2024 arXiv
-
[40]
To- wards accurate generative models of video: A new metric & challenges, 2019
Thomas Unterthiner, Sjoerd van Steenkiste, Karol Kurach, Raphael Marinier, Marcin Michalski, and Sylvain Gelly. To- wards accurate generative models of video: A new metric & challenges, 2019. 6
2019
-
[41]
Neural discrete representation learning,
Aaron van den Oord, Oriol Vinyals, and Koray Kavukcuoglu. Neural discrete representation learning,
-
[42]
Gomez, Lukasz Kaiser, and Illia Polosukhin
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszko- reit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need, 2023. 4, 17
2023
-
[43]
Phenaki: Variable length video generation from open domain textual description, 2022
Ruben Villegas, Mohammad Babaeizadeh, Pieter-Jan Kin- dermans, Hernan Moraldo, Han Zhang, Mohammad Taghi Saffar, Santiago Castro, Julius Kunze, and Dumitru Erhan. Phenaki: Variable length video generation from open domain textual description, 2022. 2
2022
-
[44]
Mcl-jcv: a jnd-based h
Haiqiang Wang, Weihao Gan, Sudeng Hu, Joe Yuchieh Lin, Lina Jin, Longguang Song, Ping Wang, Ioannis Katsavouni- dis, Anne Aaron, and C-C Jay Kuo. Mcl-jcv: a jnd-based h. 264/avc video quality assessment dataset. In 2016 IEEE international conference on image processing (ICIP),...
2016
-
[45]
Bevt: Bert pretraining of video transformers,
Rui Wang, Dongdong Chen, Zuxuan Wu, Yinpeng Chen, Xiyang Dai, Mengchen Liu, Yu-Gang Jiang, Luowei Zhou, and Lu Yuan. Bevt: Bert pretraining of video transformers,
-
[46]
Emu3: Next-token prediction is all you need, 2024
Xinlong Wang, Xiaosong Zhang, Zhengxiong Luo, Quan Sun, Yufeng Cui, Jinsheng Wang, Fan Zhang, Yueze Wang, Zhen Li, Qiying Yu, Yingli Zhao, Yulong Ao, Xuebin Min, Tao Li, Boya Wu, Bo Zhao, Bowen Zhang, Liangdong Wang, Guang Liu, Zheqi He, Xi Yang, Jingjing Liu, Yonghua Lin, Tie...
2024
-
[47]
Instructavatar: Text- guided emotion and motion control for avatar generation,
Yuchi Wang, Junliang Guo, Jianhong Bai, Runyi Yu, Tianyu He, Xu Tan, Xu Sun, and Jiang Bian. Instructavatar: Text- guided emotion and motion control for avatar generation,
-
[48]
Internvid: A large-scale video-text dataset for multimodal understanding and generation, 2024
Yi Wang, Yinan He, Yizhuo Li, Kunchang Li, Jiashuo Yu, Xin Ma, Xinhao Li, Guo Chen, Xinyuan Chen, Yaohui Wang, Conghui He, Ping Luo, Ziwei Liu, Yali Wang, Limin Wang, and Yu Qiao. Internvid: A large-scale video-text dataset for multimodal understanding and generation, 2024. 6
2024
-
[49]
Image quality assessment: from error visibility to structural similarity
Zhou Wang, Alan C Bovik, Hamid R Sheikh, and Eero P Si- moncelli. Image quality assessment: from error visibility to structural similarity. IEEE transactions on image processing, 13(4):600–612, 2004. 6
2004
-
[50]
Janus: Decoupling visual encoding for unified multimodal understanding and genera- tion, 2024
Chengyue Wu, Xiaokang Chen, Zhiyu Wu, Yiyang Ma, Xingchao Liu, Zizheng Pan, Wen Liu, Zhenda Xie, Xingkai Yu, Chong Ruan, and Ping Luo. Janus: Decoupling visual encoding for unified multimodal understanding and genera- tion, 2024. 8
2024
-
[51]
ivideogpt: Interactive videogpts are scalable world models, 2024
Jialong Wu, Shaofeng Yin, Ningya Feng, Xu He, Dong Li, Jianye Hao, and Mingsheng Long. ivideogpt: Interactive videogpts are scalable world models, 2024. 1, 3, 5, 6, 13, 14
2024
-
[52]
Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation
Jay Zhangjie Wu, Yixiao Ge, Xintao Wang, Stan Weixian Lei, Yuchao Gu, Yufei Shi, Wynne Hsu, Ying Shan, Xiaohu Qie, and Mike Zheng Shou. Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation. In Proceedings of the IEEE/CVF International Conference...
2023
-
[53]
Cogvideox: Text-to-video diffusion models with an expert transformer, 2024
Zhuoyi Yang, Jiayan Teng, Wendi Zheng, Ming Ding, Shiyu Huang, Jiazheng Xu, Yuanming Yang, Wenyi Hong, Xiaohan Zhang, Guanyu Feng, Da Yin, Xiaotao Gu, Yuxuan Zhang, Weihan Wang, Yean Cheng, Ting Liu, Bin Xu, Yuxiao Dong, and Jie Tang. Cogvideox: Text-to-video diffusion models ...
2024
-
[54]
Vector-quantized image modeling with improved vqgan, 2022
Jiahui Yu, Xin Li, Jing Yu Koh, Han Zhang, Ruoming Pang, James Qin, Alexander Ku, Yuanzhong Xu, Jason Baldridge, and Yonghui Wu. Vector-quantized image modeling with improved vqgan, 2022. 2
2022
-
[55]
Hauptmann, Ming- Hsuan Yang, Yuan Hao, Irfan Essa, and Lu Jiang
Lijun Yu, Yong Cheng, Kihyuk Sohn, Jos ´e Lezama, Han Zhang, Huiwen Chang, Alexander G. Hauptmann, Ming- Hsuan Yang, Yuan Hao, Irfan Essa, and Lu Jiang. Magvit: Masked generative video transformer, 2023. 1, 2 10
2023
-
[56]
Language model beats diffusion–tokenizer is key to visual generation
Lijun Yu, Jos ´e Lezama, Nitesh B Gundavarapu, Luca Ver- sari, Kihyuk Sohn, David Minnen, Yong Cheng, Vighnesh Birodkar, Agrim Gupta, Xiuye Gu, et al. Language model beats diffusion–tokenizer is key to visual generation. arXiv preprint arXiv:2310.05737, 2023. 1, 2, 5, 6, 8, 13, 14
-
[57]
Make your actor talk: Generalizable and high-fidelity lip sync with motion and appearance disentanglement, 2024
Runyi Yu, Tianyu He, Ailing Zhang, Yuchi Wang, Junliang Guo, Xu Tan, Chang Liu, Jie Chen, and Jiang Bian. Make your actor talk: Generalizable and high-fidelity lip sync with motion and appearance disentanglement, 2024. 3
2024
-
[58]
Efficient video diffusion models via content-frame motion-latent decomposition, 2024
Sihyun Yu, Weili Nie, De-An Huang, Boyi Li, Jinwoo Shin, and Anima Anandkumar. Efficient video diffusion models via content-frame motion-latent decomposition, 2024. 1, 2, 3, 4, 5, 6, 13
2024
-
[59]
Efros, Eli Shecht- man, and Oliver Wang
Richard Zhang, Phillip Isola, Alexei A. Efros, Eli Shecht- man, and Oliver Wang. The unreasonable effectiveness of deep features as a perceptual metric, 2018. 6
2018
-
[60]
Video in-context learning: Autore- gressive transformers are zero-shot video imitators
Wentao Zhang, Junliang Guo, Tianyu He, Li Zhao, Linli Xu, and Jiang Bian. Video in-context learning: Autore- gressive transformers are zero-shot video imitators. In The Thirteenth International Conference on Learning Represen- tations, 2025. 2
2025
-
[61]
Cv-vae: A compatible video vae for latent generative video models,
Sijie Zhao, Yong Zhang, Xiaodong Cun, Shaoshu Yang, Muyao Niu, Xiaoyu Li, Wenbo Hu, and Ying Shan. Cv-vae: A compatible video vae for latent generative video models,
-
[62]
Open-sora: Democratizing efficient video production for all, 2024
Zangwei Zheng, Xiangyu Peng, Tianji Yang, Chenhui Shen, Shenggui Li, Hongxin Liu, Yukun Zhou, Tianyi Li, and Yang You. Open-sora: Democratizing efficient video production for all, 2024. 2, 12 11 A. Additional Experimental Results To enhance the visual experience, we strongly e...
2024
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.