REVIEW 4 major objections 5 minor 2 cited by
Hunyuan-Game: Industrial-grade Intelligent Game Creation Model
T0 review · 4 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read Tencent's Hunyuan-Game applies diffusion transformers to game asset creation across nine image and video generation tasks, with self-reported gains that are partly contradicted by its own evaluation table.
desk verdict A broad industrial systems report with genuinely new task formulations, but its own Table 3 contradicts the abstract's claim that it beats Kling in game scenarios. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
Extended reading notes
Core claim
The abstract states Hunyuan-Game models are 'surpassing competitors like Midjourney, Kling and Wan in game scenarios.' If correct, this would mean a single integrated suite dominates commercial and open-source generative tools for game asset production. However, Table 3 reports a 3.31 overall score for Hunyuan-Game versus 3.47 for Kling 1.6 Pro, directly contradicting the abstract's claim.
Load-bearing premise
The most fragile load-bearing premise is the validity of the self-built evaluation and aesthetic scoring system. The paper trains a proprietary aesthetic scoring model (Section 2.1.2), uses it to filter training data and guide the model's optimization (Section 2.1.3), then evaluates outputs on dimensions that overlap with these same criteria (Section 2.1.6). If these internal standards do not match real player or designer preferences, the claimed state-of-the-art quality and the 'state-of-the-art' conclusion are unsupported. The paper provides no external benchmark or third-party validation.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript describes Hunyuan-Game, an industrial suite from Tencent for game asset generation, comprising four image-generation models (general text-to-image, game visual effects, transparent/seamless images, and character generation) and five video-generation models (image-to-video, 360 A/T pose avatar video, dynamic illustration, generative video super-resolution, and interactive game video). The authors describe large proprietary datasets (billions of images, millions of videos), multi-stage data filtering and captioning pipelines, a proprietary aesthetic scoring system, and training strategies including SFT, quality tuning, and DPO. Quantitative comparisons are reported in Tables 1–4 against commercial and open models such as Midjourney, Flux, Kling, Wan, CogVideoX, and Minimax. The paper claims state-of-the-art performance in game scenarios, particularly in visual fidelity and motion naturalness, and concludes that the suite outperforms existing baselines.
Significance. If the performance claims were fully supported, the paper would be a landmark industrial report: it covers an unusually wide range of game-asset generation tasks, introduces several first-of-kind capabilities (e.g., A/T-pose avatar video, dynamic illustration, interactive game video), and provides detailed descriptions of data engineering, aesthetic scoring, captioning, and training pipelines. The engineering scope is substantial and the qualitative material is rich. However, the central 'surpassing Kling' claim is contradicted by the paper's own Table 3, and the evaluations rest on self-built, self-scored benchmarks whose criteria overlap with the proprietary aesthetic model used to filter training data. The industrial contribution is real, but the state-of-the-art conclusion is not established by the evidence presented.
major comments (4)
- [Abstract and §3.1.5 (Table 3)] The abstract states that the models are 'surpassing competitors like Midjourney, Kling and Wan in game scenarios,' and the conclusion says the suite is 'outperforming existing baselines.' Yet Table 3 reports an overall video score of 3.31 for Hunyuan-Game versus 3.47 for Kling 1.6 Pro, with Kling also higher on image-video alignment (3.92 vs 3.84) and visual quality (3.94 vs 3.86). The text at §3.1.5 concedes that the model 'performs slightly worse than Kling 1.6 Pro.' This is an internal contradiction in the central claim. The image branch leads Table 1, but the video branch does not; the paper must either qualify the headline claims to match Table 3 or supply evidence that the video comparison is not representative.
- [§2.1.2, §2.1.4, §2.1.6, §3.1.2, §3.1.4, §3.1.5] The evaluation is circular in a load-bearing way. The proprietary six-dimensional aesthetic scoring system is trained on the authors' definitions (§2.1.2), used to filter training data and to select QT-stage data (§2.1.4), and the same aesthetic vocabulary reappears in the evaluation criteria (e.g., 'pictorial aesthetics' and 'subject modeling' in Table 1; 'visual quality' and 'motion quality' in Table 3, where the video motion aesthetic operators from §3.1.2 are used for data filtering and model iteration). No external benchmark, third-party annotation, or established correlation with designer or player preferences is provided. To support a state-of-the-art claim, the paper needs an independent evaluation protocol, inter-rater reliability statistics, and ideally a comparison on public benchmarks.
- [§3.4.4 (Table 4)] The generative video super-resolution test set is composed of 40 real videos and 40 videos 'generated by the Hunyuan-Game I2V' model. Using the authors' own generated videos as test data can inflate perceived performance because the model may be biased toward its own output distribution; also, no blinding or separate reporting for the two subsets is described. The paper should report real-video and generated-video scores separately and specify whether annotators were blind to model identity.
- [§2.3.4, §3.1.4, §3.5.1] Several quantitative claims lack supporting evidence: the 60% efficiency improvement for visual effects iteration (§2.3.4) is attributed to 'feedback from designers' with no sample size or measurement method; the prompt-rewriting model's 'consistency rate 98%' (§3.1.4) is not tied to any evaluation protocol; and the '10–20× acceleration... less than 10s per action' (§3.5.1) is reported without benchmark details. If these numbers are meant to support the abstract's efficiency and real-time interactivity claims, they need precise definitions, measurement procedures, and error ranges.
minor comments (5)
- [§3.4.1] The section heading contains a typo: 'Introductrion' should be 'Introduction.'
- [§2.1.3 and §2.1.5] The heading 'Construction and Tiered Filtering of Game Datasets' appears at the start of §2.1.3 and §2.1.5, which is confusing and appears to be a copy-paste artifact; the headings should be made specific to the content that follows.
- [§3.2.2] The lossless codec is written as 'FFV13' and should be 'FFV1,' and 'hdri' should be written consistently as 'HDRI.'
- [Tables 1 and 3] The evaluation tables would be substantially easier to interpret with per-row sample counts and measures of annotator agreement or confidence intervals; as presented, the 5-point-scale means from three annotators cannot be distinguished from noise.
- [Throughout] Several 'first' claims (e.g., §2.2.1, §3.2.1) are asserted without a systematic prior-art search or a clear definition of the comparison scope; please add explicit context or citations to substantiate these novelty statements.
Circularity Check
No construction-level circularity: reported scores come from human juries against external models; the self-referential VSR test set and the abstract/Table 3 contradiction are separate validity and correctness concerns.
full rationale
The claimed derivations do not reduce to their inputs by construction. Hunyuan-Game-Image and Hunyuan-Game-Video are fine-tuned from stated base models (e.g., HunyuanCustom [21], HunyuanVideo [25]), and their reported performances in Tables 1 and 3 are produced by three human annotators comparing outputs of all models side-by-side, not by the proprietary aesthetic scoring operators used internally for data filtering and quality tuning. Although the aesthetic scoring model (Section 2.1.2) selects training data and guides the QT stage (Section 2.1.4), the Table 1 'Aesthetics' score is a human jury rating, so the evaluation is not the fitted operator itself; at most there is a preference-alignment loop, which is a validity concern rather than a circular reduction. The video super-resolution test set includes 40 videos generated by Hunyuan-Game I2V (Section 3.4.4), making that benchmark partially self-referential, but all compared methods are scored by annotators on the same inputs, so the comparison does not follow from the model's fitted parameters by construction. Self-citations to HunyuanCustom and HunyuanVideo are used as base models for fine-tuning, but no uniqueness or ansatz conclusion is imported from them. Separately, the abstract's claim of 'surpassing competitors like Midjourney, Kling and Wan' is internally inconsistent with Table 3, where Kling 1.6 Pro leads overall (3.47 vs 3.31) and in visual quality (3.94 vs 3.86); this is a significant correctness problem in the SOTA claim, but it is not a circularity of the type scored here.
Assumptions & free parameters
free parameters (5)
- Caption length sampling ratio =
1:1:1:7 (short:medium:detailed:comprehensive)
- Image data resolution threshold =
1024x1024 minimum
- Annotator agreement thresholds =
80% agreement; 70%/95% acceptance
- Video aesthetic filtering thresholds =
700K SFT videos, 80K QT videos
- Keyboard-to-camera motion parameters =
pre-defined speed and angle per action
assumptions (4)
- domain assumption Human aesthetic judgments are decomposable into six objective dimensions and learnable by a regression model.
- domain assumption The authors' self-built evaluation sets are representative of game-asset generation quality.
- domain assumption Base models HunyuanVideo and Hunyuan-DiT provide suitable priors for game-specific content.
- domain assumption Designer feedback used to claim a 60% efficiency improvement is reliable and quantified.
invented entities (2)
-
Proprietary six-dimensional aesthetic scoring system
-
Effect-enhanced material data
Cite this review
Pith. "Pith review of Hunyuan-Game: Industrial-grade Intelligent Game Creation Model." pith.science (2026). https://pith.science/paper/2QBOD34L
@misc{pith2026250514135,
author = {Pith},
title = {Pith review of: Hunyuan-Game: Industrial-grade Intelligent Game Creation Model},
year = {2026},
howpublished = {\url{https://pith.science/paper/2QBOD34L}},
note = {Machine review of arXiv:2505.14135}
}
read the original abstract
Intelligent game creation represents a transformative advancement in game development, utilizing generative artificial intelligence to dynamically generate and enhance game content. Despite notable progress in generative models, the comprehensive synthesis of high-quality game assets, including both images and videos, remains a challenging frontier. To create high-fidelity game content that simultaneously aligns with player preferences and significantly boosts designer efficiency, we present Hunyuan-Game, an innovative project designed to revolutionize intelligent game production. Hunyuan-Game encompasses two primary branches: image generation and video generation. The image generation component is built upon a vast dataset comprising billions of game images, leading to the development of a group of customized image generation models tailored for game scenarios: (1) General Text-to-Image Generation. (2) Game Visual Effects Generation, involving text-to-effect and reference image-based game visual effect generation. (3) Transparent Image Generation for characters, scenes, and game visual effects. (4) Game Character Generation based on sketches, black-and-white images, and white models. The video generation component is built upon a comprehensive dataset of millions of game and anime videos, leading to the development of five core algorithmic models, each targeting critical pain points in game development and having robust adaptation to diverse game video scenarios: (1) Image-to-Video Generation. (2) 360 A/T Pose Avatar Video Synthesis. (3) Dynamic Illustration Generation. (4) Generative Video Super-Resolution. (5) Interactive Game Video Generation. These image and video generation models not only exhibit high-level aesthetic expression but also deeply integrate domain-specific knowledge, establishing a systematic understanding of diverse game and anime art styles.
Figures
Figures from the paper (32 more)
Forward citations
Cited by 2 Pith papers
-
Yume: An Interactive World Generation Model
A diffusion-based video model generates extendable, keyboard-controlled walkthroughs from a single input image, using quantized camera actions as text prompts.
-
Hunyuan-GameCraft: High-dynamic Interactive Game Video Generation with Hybrid History Condition
Hunyuan-GameCraft generates long, action-controlled game videos from a single image by unifying keyboard/mouse inputs into a continuous camera space and conditioning on mixed historical context.
Reference graph
Works this paper leans on
-
[1]
Cosmos world foundation model platform for physical ai
Niket Agarwal, Arslan Ali, Maciej Bala, Yogesh Balaji, Erik Barker, Tiffany Cai, Prithvijit Chattopadhyay, Yongxin Chen, Yin Cui, Yifan Ding, et al. Cosmos world foundation model platform for physical ai. arXiv preprint arXiv:2501.03575, 2025
arXiv 2025
-
[2]
Kling ai: Next-generation ai creative studio
Kling AI. Kling ai: Next-generation ai creative studio. https://www.klingai.com/, 2024
2024
-
[3]
Layer ai: Game art without limits
Layer AI. Layer ai: Game art without limits. https://app.layer.ai/, 2025
2025
-
[4]
Blended latent diffusion
Omri Avrahami, Ohad Fried, and Dani Lischinski. Blended latent diffusion. TOG, 42(4):1–11, 2023
2023
-
[5]
Ac3d: Analyzing and improving 3d camera control in video diffusion transformers
Sherwin Bahmani, Ivan Skorokhodov, Guocheng Qian, Aliaksandr Siarohin, Willi Menapace, Andrea Tagliasacchi, David B Lindell, and Sergey Tulyakov. Ac3d: Analyzing and improving 3d camera control in video diffusion transformers. arXiv preprint arXiv:2411.18673, 2024
arXiv 2024
-
[6]
Stable video diffusion: Scaling latent video diffusion models to large datasets, 2023
Andreas Blattmann, Tim Dockhorn, Sumith Kulal, Daniel Mendelevitch, Maciej Kilian, Dominik Lorenz, Yam Levi, Zion English, Vikram V oleti, Adam Letts, Varun Jampani, and Robin Rombach. Stable video diffusion: Scaling latent video diffusion models to large datasets, 2023
2023
-
[7]
Align your latents: High-resolution video synthesis with latent diffusion models
Andreas Blattmann, Robin Rombach, Huan Ling, Tim Dockhorn, Seung Wook Kim, Sanja Fidler, and Karsten Kreis. Align your latents: High-resolution video synthesis with latent diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 22563–22575, 2023
2023
-
[8]
Motionclr: Motion generation and training-free editing via understanding attention mechanisms
Ling-Hao Chen, Wenxun Dai, Xuan Ju, Shunlin Lu, and Lei Zhang. Motionclr: Motion generation and training-free editing via understanding attention mechanisms. arXiv e-prints, pages arXiv–2410, 2024
2024
Show all 95 references
-
[9]
Panic-3d: Stylized single-view 3d reconstruction from portraits of anime characters
Shuhong Chen, Kevin Zhang, Yichun Shi, Heng Wang, Yiheng Zhu, Guoxian Song, Sizhe An, Janus Kristjansson, Xiao Yang, and Matthias Zwicker. Panic-3d: Stylized single-view 3d reconstruction from portraits of anime characters. In Proceedings of the IEEE/CVF Conference on Computer...
2023
-
[10]
Motionlcm: Real-time controllable motion generation via latent consistency model
Wenxun Dai, Ling-Hao Chen, Jingbo Wang, Jinpeng Liu, Bo Dai, and Yansong Tang. Motionlcm: Real-time controllable motion generation via latent consistency model. In ECCV, pages 390–408, 2024
2024
-
[11]
Veo 2: Our state-of-the-art video generation model
Google Deepmind. Veo 2: Our state-of-the-art video generation model. https://deepmind.google/ technologies/veo/veo-2/, 2024
2024
-
[12]
Scaling rectified flow transformers for high-resolution image synthesis
Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas Müller, Harry Saini, Yam Levi, Dominik Lorenz, Axel Sauer, Frederic Boesel, et al. Scaling rectified flow transformers for high-resolution image synthesis. In ICML, 2024
2024
-
[13]
Freemotion: A unified framework for number-free text-to-motion synthesis
Ke Fan, Junshu Tang, Weijian Cao, Ran Yi, Moran Li, Jingyu Gong, Jiangning Zhang, Yabiao Wang, Chengjie Wang, and Lizhuang Ma. Freemotion: A unified framework for number-free text-to-motion synthesis. In ECCV, pages 93–109, 2024
2024
-
[14]
The matrix: Infinite-horizon world generation with real-time moving control
Ruili Feng, Han Zhang, Zhantao Yang, Jie Xiao, Zhilei Shu, Zhiheng Liu, Andy Zheng, Yukun Huang, Yu Liu, and Hongyang Zhang. The matrix: Infinite-horizon world generation with real-time moving control. arXiv preprint arXiv:2412.03568, 2024
2024 arXiv
-
[15]
Seedream 3.0 technical report
Yu Gao, Lixue Gong, Qiushan Guo, Xiaoxia Hou, Zhichao Lai, Fanshi Li, Liang Li, Xiaochen Lian, Chao Liao, Liyang Liu, et al. Seedream 3.0 technical report. arXiv preprint arXiv:2504.11346, 2025
2025 arXiv
-
[16]
Mineworld: a real-time and open-source interactive world model on minecraft
Junliang Guo, Yang Ye, Tianyu He, Haoyu Wu, Yushu Jiang, Tim Pearce, and Jiang Bian. Mineworld: a real-time and open-source interactive world model on minecraft. arXiv preprint arXiv:2504.08388, 2025
2025 arXiv
-
[17]
World models
David Ha and Jürgen Schmidhuber. World models. arXiv preprint arXiv:1803.10122, 2018
2018 arXiv
-
[18]
Cameractrl: Enabling camera control for text-to-video generation
Hao He, Yinghao Xu, Yuwei Guo, Gordon Wetzstein, Bo Dai, Hongsheng Li, and Ceyuan Yang. Cameractrl: Enabling camera control for text-to-video generation. arXiv preprint arXiv:2404.02101, 2024
2024 arXiv
-
[19]
Venhancer: Generative space-time enhancement for video generation
Jingwen He, Tianfan Xue, Dongyang Liu, Xinqi Lin, Peng Gao, Dahua Lin, Yu Qiao, Wanli Ouyang, and Ziwei Liu. Venhancer: Generative space-time enhancement for video generation. arXiv preprint arXiv:2407.07667, 2024
2024 arXiv
-
[20]
Lora: Low-rank adaptation of large language models
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, Weizhu Chen, et al. Lora: Low-rank adaptation of large language models. ICLR, page 3, 2022. 47
2022
-
[21]
Hunyuancustom: A multimodal-driven architecture for customized video generation
Teng Hu, Zhentao Yu, Zhengguang Zhou, Sen Liang, Yuan Zhou, Qin Lin, and Qinglin Lu. Hunyuancustom: A multimodal-driven architecture for customized video generation. arXiv preprint arXiv:2505.04512, 2025
2025 arXiv
-
[22]
Step-video-ti2v technical report: A state-of-the-art text-driven image-to-video generation model
Haoyang Huang, Guoqing Ma, Nan Duan, Xing Chen, Changyi Wan, Ranchen Ming, Tianyu Wang, Bo Wang, Zhiying Lu, Aojie Li, et al. Step-video-ti2v technical report: A state-of-the-art text-driven image-to-video generation model. arXiv preprint arXiv:2503.11251, 2025
2025 arXiv
-
[23]
In-context lora for diffusion transformers
Lianghua Huang, Wei Wang, Zhi-Fan Wu, Yupeng Shi, Huanzhang Dou, Chen Liang, Yutong Feng, Yu Liu, and Jingren Zhou. In-context lora for diffusion transformers. arXiv preprint arXiv:2410.23775, 2024
2024 arXiv
-
[24]
Sport: From zero-shot prompts to real-time motion generation
Bin Ji, Ye Pan, Zhimeng Liu, Shuai Tan, and Xiaokang Yang. Sport: From zero-shot prompts to real-time motion generation. TVCG, 2025
2025
-
[25]
Hunyuanvideo: A systematic framework for large video generative models
Weijie Kong, Qi Tian, Zijian Zhang, Rox Min, Zuozhuo Dai, Jin Zhou, Jiangfeng Xiong, Xin Li, Bo Wu, Jianwei Zhang, et al. Hunyuanvideo: A systematic framework for large video generative models. arXiv preprint arXiv:2412.03603, 2024
2024 arXiv
-
[26]
Matroska and ffv1: One file format for film and video archiving? JFP, (96):41, 2017
Reto Kromer. Matroska and ffv1: One file format for film and video archiving? JFP, (96):41, 2017
2017
-
[27]
Multi-concept customization of text-to-image diffusion
Nupur Kumari, Bingliang Zhang, Richard Zhang, Eli Shechtman, and Jun-Yan Zhu. Multi-concept customization of text-to-image diffusion. In CVPR, pages 1931–1941, 2023
1931
-
[28]
Black Forest Labs. Flux. https://github.com/black-forest-labs/flux , 2024
2024
-
[29]
Hunyuan-dit: A powerful multi-resolution diffusion transformer with fine-grained chinese understanding
Zhimin Li, Jianwei Zhang, Qin Lin, Jiangfeng Xiong, Yanxin Long, Xinchi Deng, Yingfang Zhang, Xingchao Liu, Minbin Huang, Zedong Xiao, et al. Hunyuan-dit: A powerful multi-resolution diffusion transformer with fine-grained chinese understanding. arXiv preprint arXiv:2405.08748, 2024
2024 arXiv
-
[30]
Improved baselines with visual instruction tuning
Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee. Improved baselines with visual instruction tuning. In CVPR, pages 26296–26306, 2024
2024
-
[31]
Visual instruction tuning
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. Visual instruction tuning. NeurIPS, pages 34892–34916, 2023
2023
-
[32]
Plan, posture and go: Towards open-world text-to-motion generation
Jinpeng Liu, Wenxun Dai, Chunyu Wang, Yiji Cheng, Yansong Tang, and Xin Tong. Plan, posture and go: Towards open-world text-to-motion generation. arXiv preprint arXiv:2312.14828, 2023
2023 arXiv
-
[33]
Syncdreamer: Generating multiview-consistent images from a single-view image
Yuan Liu, Cheng Lin, Zijiao Zeng, Xiaoxiao Long, Lingjie Liu, Taku Komura, and Wenping Wang. Syncdreamer: Generating multiview-consistent images from a single-view image. arXiv preprint arXiv:2309.03453, 2023
2023 arXiv
-
[34]
Manganinja: Line art colorization with precise reference following
Zhiheng Liu, Ka Leong Cheng, Xi Chen, Jie Xiao, Hao Ouyang, Kai Zhu, Yu Liu, Yujun Shen, Qifeng Chen, and Ping Luo. Manganinja: Line art colorization with precise reference following. arXiv preprint arXiv:2501.08332, 2025
2025 arXiv
-
[35]
Scamo: Exploring the scaling law in autoregressive motion generation model
Shunlin Lu, Jingbo Wang, Zeyu Lu, Ling-Hao Chen, Wenxun Dai, Junting Dong, Zhiyang Dou, Bo Dai, and Ruimao Zhang. Scamo: Exploring the scaling law in autoregressive motion generation model. arXiv preprint arXiv:2412.14559, 2024
2024 arXiv
-
[36]
Minimax. Hailuo. https://hailuoai.com/video, 2024
2024
-
[37]
OpenAI. Sora. https://openai.com/sora/, 2024
2024
-
[38]
Dinov2: Learning robust visual features without supervision
Maxime Oquab, Timothée Darcet, Théo Moutakanni, Huy V o, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, et al. Dinov2: Learning robust visual features without supervision. arXiv preprint arXiv:2304.07193, 2023
2023 arXiv
-
[39]
Deep blind video super-resolution
Jinshan Pan, Haoran Bai, Jiangxin Dong, Jiawei Zhang, and Jinhui Tang. Deep blind video super-resolution. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 4811–4820, 2021
2021
-
[40]
Tokenhsi: Unified synthesis of physical human-scene interactions through task tokenization
Liang Pan, Zeshi Yang, Zhiyang Dou, Wenjia Wang, Buzhen Huang, Bo Dai, Taku Komura, and Jingbo Wang. Tokenhsi: Unified synthesis of physical human-scene interactions through task tokenization. arXiv preprint arXiv:2503.19901, 2025
2025 arXiv
-
[41]
Genie 2: A large-scale foundation world model
J Parker-Holder, P Ball, J Bruce, V Dasagi, K Holsheimer, C Kaplanis, A Moufarek, G Scully, J Shar, J Shi, et al. Genie 2: A large-scale foundation world model. https://deepmind.google/discover/ blog/genie-2-a-large-scale-foundation-world-model , 2024. 48
2024
-
[42]
Scalable diffusion models with transformers
William Peebles and Saining Xie. Scalable diffusion models with transformers. InICCV, pages 4195–4205, 2023
2023
-
[43]
Charactergen: Efficient 3d character generation from single images with multi-view pose canonicalization
Hao-Yang Peng, Jia-Peng Zhang, Meng-Hao Guo, Yan-Pei Cao, and Shi-Min Hu. Charactergen: Efficient 3d character generation from single images with multi-view pose canonicalization. ACM Transactions on Graphics (TOG), 43(4):1–13, 2024
2024
-
[44]
Pyscenedetect developers
PySceneDetect. Pyscenedetect developers. https://www.scenedetect.com/, 2024
2024
-
[45]
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In ICML, pages 8748–8763. PmLR, 2021
2021
-
[46]
Direct preference optimization: Your language model is secretly a reward model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn. Direct preference optimization: Your language model is secretly a reward model. NeurIPS, pages 53728–53741, 2023
2023
-
[47]
High-resolution image synthesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10684–10695, 2022
2022
-
[48]
Photorealistic text-to- image diffusion models with deep language understanding
Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily L Denton, Kamyar Ghasemipour, Raphael Gontijo Lopes, Burcu Karagol Ayan, Tim Salimans, et al. Photorealistic text-to- image diffusion models with deep language understanding. Advances in neural informatio...
2022
-
[49]
Laion-5b: An open large-scale dataset for training next generation image-text models
Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Wortsman, et al. Laion-5b: An open large-scale dataset for training next generation image-text models. NeurIPS, pages 25278–25294, 2022
2022
-
[50]
Seaweed-7b: Cost-effective training of video generation foundation model
Team Seawead, Ceyuan Yang, Zhijie Lin, Yang Zhao, Shanchuan Lin, Zhibei Ma, Haoyuan Guo, Hao Chen, Lu Qi, Sen Wang, et al. Seaweed-7b: Cost-effective training of video generation foundation model. arXiv preprint arXiv:2504.08685, 2025
2025 arXiv
-
[51]
Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, Y . K. Li, Y . Wu, and Daya Guo. Deepseekmath: Pushing the limits of mathematical reasoning in open language models. arXiv preprint arXiv:2402.03300, 2024
2024 arXiv
-
[52]
Zero123++: a single image to consistent multi-view diffusion base model
Ruoxi Shi, Hansheng Chen, Zhuoyang Zhang, Minghua Liu, Chao Xu, Xinyue Wei, Linghao Chen, Chong Zeng, and Hao Su. Zero123++: a single image to consistent multi-view diffusion base model. arXiv preprint arXiv:2310.15110, 2023
2023 arXiv
-
[53]
Light field networks: Neural scene representations with single-evaluation rendering
Vincent Sitzmann, Semon Rezchikov, Bill Freeman, Josh Tenenbaum, and Fredo Durand. Light field networks: Neural scene representations with single-evaluation rendering. NeurIPS, pages 19313–19325, 2021
2021
-
[54]
Hunyuan-large: An open-source moe model with 52 billion activated parameters by tencent
Xingwu Sun, Yanfeng Chen, Yiqing Huang, Ruobing Xie, Jiaqi Zhu, Kai Zhang, Shuaipeng Li, Zhen Yang, Jonny Han, Xiaobo Shu, et al. Hunyuan-large: An open-source moe model with 52 billion activated parameters by tencent. arXiv preprint arXiv:2411.02265, 2024
2024 arXiv
-
[55]
Make-it-3d: High-fidelity 3d creation from a single image with diffusion prior
Junshu Tang, Tengfei Wang, Bo Zhang, Ting Zhang, Ran Yi, Lizhuang Ma, and Dong Chen. Make-it-3d: High-fidelity 3d creation from a single image with diffusion prior. In Proceedings of the IEEE/CVF international conference on computer vision, pages 22819–22829, 2023
2023
-
[56]
Make-it-vivid: dressing your animatable biped cartoon characters from text
Junshu Tang, Yanhong Zeng, Ke Fan, Xuheng Wang, Bo Dai, Kai Chen, and Lizhuang Ma. Make-it-vivid: dressing your animatable biped cartoon characters from text. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6243–6253, 2024
2024
-
[57]
Learning motion refinement for unsupervised face animation
Jiale Tao, Shuhang Gu, Wen Li, and Lixin Duan. Learning motion refinement for unsupervised face animation. Advances in Neural Information Processing Systems, 36:70483–70496, 2023
2023
-
[58]
Motion transformer for unsupervised image animation
Jiale Tao, Biao Wang, Tiezheng Ge, Yuning Jiang, Wen Li, and Lixin Duan. Motion transformer for unsupervised image animation. In European conference on computer vision, pages 702–719. Springer, 2022
2022
-
[59]
Structure-aware motion transfer with deformable anchor model
Jiale Tao, Biao Wang, Borun Xu, Tiezheng Ge, Yuning Jiang, Wen Li, and Lixin Duan. Structure-aware motion transfer with deformable anchor model. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 3637–3646, June 2022. 49
2022
-
[60]
Instantcharacter: Personalize any characters with a scalable diffusion transformer framework
Jiale Tao, Yanbing Zhang, Qixun Wang, Yiji Cheng, Haofan Wang, Xu Bai, Zhengguang Zhou, Ruihuang Li, Linqing Wang, Chunyu Wang, et al. Instantcharacter: Personalize any characters with a scalable diffusion transformer framework. arXiv preprint arXiv:2504.12395, 2025
2025 arXiv
-
[61]
Midjourney
Midjourney Team. Midjourney. https://www.midjourney.com/, 2023
2023
-
[62]
Generating worlds
World Labs Team. Generating worlds. https://www.worldlabs.ai/blog, 2024
2024
-
[63]
Raft: Recurrent all-pairs field transforms for optical flow
Zachary Teed and Jia Deng. Raft: Recurrent all-pairs field transforms for optical flow. In ECCV, pages 402–419, 2020
2020
-
[64]
Diffusion model alignment using direct preference optimization
Bram Wallace, Meihua Dang, Rafael Rafailov, Linqi Zhou, Aaron Lou, Senthil Purushwalkam, Stefano Ermon, Caiming Xiong, Shafiq Joty, and Nikhil Naik. Diffusion model alignment using direct preference optimization. In CVPR, pages 8228–8238, 2024
2024
-
[65]
Wan: Open and advanced large-scale video generative models
Ang Wang, Baole Ai, Bin Wen, Chaojie Mao, Chen-Wei Xie, Di Chen, Feiwu Yu, Haiming Zhao, Jianxiao Yang, Jianyuan Zeng, et al. Wan: Open and advanced large-scale video generative models. arXiv preprint arXiv:2503.20314, 2025
2025 arXiv
-
[66]
Apisr: anime production inspired real-world anime super-resolution
Boyang Wang, Fengyu Yang, Xihang Yu, Chao Zhang, and Hanbin Zhao. Apisr: anime production inspired real-world anime super-resolution. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 25574–25584, 2024
2024
-
[67]
Phased consistency models
Fu-Yun Wang, Zhaoyang Huang, Alexander Bergman, Dazhong Shen, Peng Gao, Michael Lingelbach, Keqiang Sun, Weikang Bian, Guanglu Song, Yu Liu, et al. Phased consistency models. NeurIPS, pages 83951–84009, 2024
2024
-
[68]
Seedvr: Seeding infinity in diffusion transformer towards generic video restoration
Jianyi Wang, Zhijie Lin, Meng Wei, Yang Zhao, Ceyuan Yang, Fei Xiao, Chen Change Loy, and Lu Jiang. Seedvr: Seeding infinity in diffusion transformer towards generic video restoration. arXiv preprint arXiv:2501.01320, 2025
2025 arXiv
-
[69]
Exploit- ing diffusion prior for real-world image super-resolution
Jianyi Wang, Zongsheng Yue, Shangchen Zhou, Kelvin CK Chan, and Chen Change Loy. Exploit- ing diffusion prior for real-world image super-resolution. International Journal of Computer Vision , 132(12):5929–5949, 2024
2024
-
[70]
Real-esrgan: Training real-world blind super- resolution with pure synthetic data
Xintao Wang, Liangbin Xie, Chao Dong, and Ying Shan. Real-esrgan: Training real-world blind super- resolution with pure synthetic data. In Proceedings of the IEEE/CVF international conference on computer vision, pages 1905–1914, 2021
1905
-
[71]
Image quality assessment: from error visibility to structural similarity
Zhou Wang, Alan C Bovik, Hamid R Sheikh, and Eero P Simoncelli. Image quality assessment: from error visibility to structural similarity. IEEE transactions on image processing, 13(4):600–612, 2004
2004
-
[72]
Motionctrl: A unified and flexible motion controller for video generation
Zhouxia Wang, Ziyang Yuan, Xintao Wang, Yaowei Li, Tianshui Chen, Menghan Xia, Ping Luo, and Ying Shan. Motionctrl: A unified and flexible motion controller for video generation. In SIGGRAPH, pages 1–11, 2024
2024
-
[73]
Vmix: Improving text-to-image diffusion model with cross-attention mixing control
Shaojin Wu, Fei Ding, Mengqi Huang, Wei Liu, and Qian He. Vmix: Improving text-to-image diffusion model with cross-attention mixing control. arXiv preprint arXiv:2412.20800, 2024
2024 arXiv
-
[74]
Motionstreamer: Streaming motion generation via diffusion-based autoregressive model in causal latent space
Lixing Xiao, Shunlin Lu, Huaijin Pi, Ke Fan, Liang Pan, Yueer Zhou, Ziyong Feng, Xiaowei Zhou, Sida Peng, and Jingbo Wang. Motionstreamer: Streaming motion generation via diffusion-based autoregressive model in causal latent space. arXiv preprint arXiv:2503.15451, 2025
2025 arXiv
-
[75]
Move as you like: image animation in e-commerce scenario
Borun Xu, Biao Wang, Jiale Tao, Tiezheng Ge, Yuning Jiang, Wen Li, and Lixin Duan. Move as you like: image animation in e-commerce scenario. In Proceedings of the 29th ACM international conference on multimedia, pages 2759–2761, 2021
2021
-
[76]
Learning semantic latent directions for accurate and controllable human motion prediction
Guowei Xu, Jiale Tao, Wen Li, and Lixin Duan. Learning semantic latent directions for accurate and controllable human motion prediction. InEuropean Conference on Computer Vision, pages 56–73. Springer, 2024
2024
-
[77]
Pandora3d: A comprehensive framework for high-quality 3d shape and texture generation
Jiayu Yang, Taizhang Shang, Weixuan Sun, Xibin Song, Ziang Cheng, Senbo Wang, Shenzhou Chen, Weizhe Liu, Hongdong Li, and Pan Ji. Pandora3d: A comprehensive framework for high-quality 3d shape and texture generation. arXiv preprint arXiv:2502.14247, 2025
2025 arXiv
-
[78]
Learning interactive real-world simulators
Mengjiao Yang, Yilun Du, Kamyar Ghasemipour, Jonathan Tompson, Dale Schuurmans, and Pieter Abbeel. Learning interactive real-world simulators. arXiv preprint arXiv:2310.06114, 1(2):6, 2023. 50
2023 arXiv
-
[79]
Position: video as the new language for real-world decision making
Sherry Yang, Jacob C Walker, Jack Parker-Holder, Yilun Du, Jake Bruce, Andre Barreto, Pieter Abbeel, and Dale Schuurmans. Position: video as the new language for real-world decision making. In Forty-first International Conference on Machine Learning, 2024
2024
-
[80]
Motion-guided latent diffusion for temporally consistent real-world video super-resolution
Xi Yang, Chenhang He, Jianqi Ma, and Lei Zhang. Motion-guided latent diffusion for temporally consistent real-world video super-resolution. In European Conference on Computer Vision, pages 224–242. Springer, 2024
2024
-
[81]
Real-world video super-resolution: A benchmark dataset and a decomposition based learning scheme
Xi Yang, Wangmeng Xiang, Hui Zeng, and Lei Zhang. Real-world video super-resolution: A benchmark dataset and a decomposition based learning scheme. In Proceedings of the IEEE/CVF international conference on computer vision, pages 4781–4790, 2021
2021
-
[82]
Hunyuan3d 1.0: A unified framework for text-to-3d and image-to-3d generation
Xianghui Yang, Huiwen Shi, Bowen Zhang, Fan Yang, Jiacheng Wang, Hongxu Zhao, Xinhai Liu, Xinzhou Wang, Qingxiang Lin, Jiaao Yu, et al. Hunyuan3d 1.0: A unified framework for text-to-3d and image-to-3d generation. arXiv preprint arXiv:2411.02293, 2024
2024 arXiv
-
[83]
Cogvideox: Text-to-video diffusion models with an expert transformer
Zhuoyi Yang, Jiayan Teng, Wendi Zheng, Ming Ding, Shiyu Huang, Jiazheng Xu, Yuanming Yang, Wenyi Hong, Xiaohan Zhang, Guanyu Feng, et al. Cogvideox: Text-to-video diffusion models with an expert transformer. arXiv preprint arXiv:2408.06072, 2024
2024 arXiv
-
[84]
Ip-adapter: Text compatible image prompt adapter for text-to-image diffusion models
Hu Ye, Jun Zhang, Sibo Liu, Xiao Han, and Wei Yang. Ip-adapter: Text compatible image prompt adapter for text-to-image diffusion models. arXiv preprint arXiv:2308.06721, 2023
2023 arXiv
-
[85]
Resshift: Efficient diffusion model for image super- resolution by residual shifting
Zongsheng Yue, Jianyi Wang, and Chen Change Loy. Resshift: Efficient diffusion model for image super- resolution by residual shifting. Advances in Neural Information Processing Systems, 36:13294–13307, 2023
2023
-
[86]
Sigmoid loss for language image pre-training
Xiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, and Lucas Beyer. Sigmoid loss for language image pre-training. In ICCV, pages 11975–11986, 2023
2023
-
[87]
Monst3r: A simple approach for estimating geometry in the presence of motion
Junyi Zhang, Charles Herrmann, Junhwa Hur, Varun Jampani, Trevor Darrell, Forrester Cole, Deqing Sun, and Ming-Hsuan Yang. Monst3r: A simple approach for estimating geometry in the presence of motion. arXiv preprint arXiv:2410.03825, 2024
-
[88]
Clay: A controllable large-scale generative model for creating high-quality 3d assets
Longwen Zhang, Ziyu Wang, Qixuan Zhang, Qiwei Qiu, Anqi Pang, Haoran Jiang, Wei Yang, Lan Xu, and Jingyi Yu. Clay: A controllable large-scale generative model for creating high-quality 3d assets. ACM Transactions on Graphics (TOG), 43(4):1–20, 2024
2024
-
[89]
Transparent image layer diffusion using latent transparency
Lvmin Zhang and Maneesh Agrawala. Transparent image layer diffusion using latent transparency. arXiv preprint arXiv:2402.17113, 2024
2024 arXiv
-
[90]
Packing input frame context in next-frame prediction models for video generation
Lvmin Zhang and Maneesh Agrawala. Packing input frame context in next-frame prediction models for video generation. arXiv preprint arXiv:2504.12626, 2025
2025
-
[91]
Adding conditional control to text-to-image diffusion models
Lvmin Zhang, Anyi Rao, and Maneesh Agrawala. Adding conditional control to text-to-image diffusion models. In ICCV, pages 3836–3847, 2023
2023
-
[92]
Realviformer: Investigating attention for real-world video super-resolution
Yuehan Zhang and Angela Yao. Realviformer: Investigating attention for real-world video super-resolution. In European Conference on Computer Vision, pages 412–428. Springer, 2024
2024
-
[93]
Hunyuan3d 2.0: Scaling diffusion models for high resolution textured 3d assets generation
Zibo Zhao, Zeqiang Lai, Qingxiang Lin, Yunfei Zhao, Haolin Liu, Shuhui Yang, Yifei Feng, Mingxin Yang, Sheng Zhang, Xianghui Yang, et al. Hunyuan3d 2.0: Scaling diffusion models for high resolution textured 3d assets generation. arXiv preprint arXiv:2501.12202, 2025
2025 arXiv
-
[94]
Upscale-a-video: Temporal-consistent diffusion model for real-world video super-resolution
Shangchen Zhou, Peiqing Yang, Jianyi Wang, Yihang Luo, and Chen Change Loy. Upscale-a-video: Temporal-consistent diffusion model for real-world video super-resolution. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2535–2545, 2024
2024
-
[95]
Allegro: Open the black box of commercial-level video generation model
Yuan Zhou, Qiuyue Wang, Yuxuan Cai, and Huan Yang. Allegro: Open the black box of commercial-level video generation model. arXiv preprint arXiv:2410.15458, 2024. 51
2024 arXiv
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.