Pith. sign in

REVIEW 4 major objections 5 minor 2 cited by

Hunyuan-Game: Industrial-grade Intelligent Game Creation Model

T0 review · 4 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read Tencent's Hunyuan-Game applies diffusion transformers to game asset creation across nine image and video generation tasks, with self-reported gains that are partly contradicted by its own evaluation table.

desk verdict A broad industrial systems report with genuinely new task formulations, but its own Table 3 contradicts the abstract's claim that it beats Kling in game scenarios. read the letter →

arxiv 2505.14135 v2 pith:2QBOD34L submitted 2025-05-20 cs.CV

classification cs.CV
keywords gamegenerationvideoimagemodelsdevelopmentgenerativehunyuan-game
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Hunyuan-Game is a set of nine generative AI models developed at Tencent for creating game art and animation. Four models work on images: general text-to-image generation, game visual effects, transparent images, and character generation. Five work on video: image-to-video, 360-degree character rotation, dynamic looping illustrations, video super-resolution, and interactive gameplay generation. The system is built by fine-tuning Tencent's existing Hunyuan models on large internal datasets of game and anime content, and by adding task-specific controls such as depth maps, sketches, and keyboard inputs. The paper's main evidence is qualitative examples plus small human evaluation studies. In the image evaluation, Hunyuan-Game scores highest among four commercial and open models. In the video evaluation, it outperforms Wan 2.1 but scores below Kling 1.6 Pro on overall quality, despite the abstract claiming it surpasses Kling. The evaluation sets were designed by the authors, and the aesthetic scoring system used to filter training data is also used to define what counts as good output. No code, data, or model weights are released. The practical value of such a system, if it works as described, is that game studios could generate concept art, effects, and animated sequences faster. But because the benchmarks are self-built and the results are partly self-contradictory, the state-of-the-art claim is not established by this report.
Extended reading notes

Core claim

The abstract states Hunyuan-Game models are 'surpassing competitors like Midjourney, Kling and Wan in game scenarios.' If correct, this would mean a single integrated suite dominates commercial and open-source generative tools for game asset production. However, Table 3 reports a 3.31 overall score for Hunyuan-Game versus 3.47 for Kling 1.6 Pro, directly contradicting the abstract's claim.

Load-bearing premise

The most fragile load-bearing premise is the validity of the self-built evaluation and aesthetic scoring system. The paper trains a proprietary aesthetic scoring model (Section 2.1.2), uses it to filter training data and guide the model's optimization (Section 2.1.3), then evaluates outputs on dimensions that overlap with these same criteria (Section 2.1.6). If these internal standards do not match real player or designer preferences, the claimed state-of-the-art quality and the 'state-of-the-art' conclusion are unsupported. The paper provides no external benchmark or third-party validation.

Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The manuscript describes Hunyuan-Game, an industrial suite from Tencent for game asset generation, comprising four image-generation models (general text-to-image, game visual effects, transparent/seamless images, and character generation) and five video-generation models (image-to-video, 360 A/T pose avatar video, dynamic illustration, generative video super-resolution, and interactive game video). The authors describe large proprietary datasets (billions of images, millions of videos), multi-stage data filtering and captioning pipelines, a proprietary aesthetic scoring system, and training strategies including SFT, quality tuning, and DPO. Quantitative comparisons are reported in Tables 1–4 against commercial and open models such as Midjourney, Flux, Kling, Wan, CogVideoX, and Minimax. The paper claims state-of-the-art performance in game scenarios, particularly in visual fidelity and motion naturalness, and concludes that the suite outperforms existing baselines.

Significance. If the performance claims were fully supported, the paper would be a landmark industrial report: it covers an unusually wide range of game-asset generation tasks, introduces several first-of-kind capabilities (e.g., A/T-pose avatar video, dynamic illustration, interactive game video), and provides detailed descriptions of data engineering, aesthetic scoring, captioning, and training pipelines. The engineering scope is substantial and the qualitative material is rich. However, the central 'surpassing Kling' claim is contradicted by the paper's own Table 3, and the evaluations rest on self-built, self-scored benchmarks whose criteria overlap with the proprietary aesthetic model used to filter training data. The industrial contribution is real, but the state-of-the-art conclusion is not established by the evidence presented.

major comments (4)
  1. [Abstract and §3.1.5 (Table 3)] The abstract states that the models are 'surpassing competitors like Midjourney, Kling and Wan in game scenarios,' and the conclusion says the suite is 'outperforming existing baselines.' Yet Table 3 reports an overall video score of 3.31 for Hunyuan-Game versus 3.47 for Kling 1.6 Pro, with Kling also higher on image-video alignment (3.92 vs 3.84) and visual quality (3.94 vs 3.86). The text at §3.1.5 concedes that the model 'performs slightly worse than Kling 1.6 Pro.' This is an internal contradiction in the central claim. The image branch leads Table 1, but the video branch does not; the paper must either qualify the headline claims to match Table 3 or supply evidence that the video comparison is not representative.
  2. [§2.1.2, §2.1.4, §2.1.6, §3.1.2, §3.1.4, §3.1.5] The evaluation is circular in a load-bearing way. The proprietary six-dimensional aesthetic scoring system is trained on the authors' definitions (§2.1.2), used to filter training data and to select QT-stage data (§2.1.4), and the same aesthetic vocabulary reappears in the evaluation criteria (e.g., 'pictorial aesthetics' and 'subject modeling' in Table 1; 'visual quality' and 'motion quality' in Table 3, where the video motion aesthetic operators from §3.1.2 are used for data filtering and model iteration). No external benchmark, third-party annotation, or established correlation with designer or player preferences is provided. To support a state-of-the-art claim, the paper needs an independent evaluation protocol, inter-rater reliability statistics, and ideally a comparison on public benchmarks.
  3. [§3.4.4 (Table 4)] The generative video super-resolution test set is composed of 40 real videos and 40 videos 'generated by the Hunyuan-Game I2V' model. Using the authors' own generated videos as test data can inflate perceived performance because the model may be biased toward its own output distribution; also, no blinding or separate reporting for the two subsets is described. The paper should report real-video and generated-video scores separately and specify whether annotators were blind to model identity.
  4. [§2.3.4, §3.1.4, §3.5.1] Several quantitative claims lack supporting evidence: the 60% efficiency improvement for visual effects iteration (§2.3.4) is attributed to 'feedback from designers' with no sample size or measurement method; the prompt-rewriting model's 'consistency rate 98%' (§3.1.4) is not tied to any evaluation protocol; and the '10–20× acceleration... less than 10s per action' (§3.5.1) is reported without benchmark details. If these numbers are meant to support the abstract's efficiency and real-time interactivity claims, they need precise definitions, measurement procedures, and error ranges.
minor comments (5)
  1. [§3.4.1] The section heading contains a typo: 'Introductrion' should be 'Introduction.'
  2. [§2.1.3 and §2.1.5] The heading 'Construction and Tiered Filtering of Game Datasets' appears at the start of §2.1.3 and §2.1.5, which is confusing and appears to be a copy-paste artifact; the headings should be made specific to the content that follows.
  3. [§3.2.2] The lossless codec is written as 'FFV13' and should be 'FFV1,' and 'hdri' should be written consistently as 'HDRI.'
  4. [Tables 1 and 3] The evaluation tables would be substantially easier to interpret with per-row sample counts and measures of annotator agreement or confidence intervals; as presented, the 5-point-scale means from three annotators cannot be distinguished from noise.
  5. [Throughout] Several 'first' claims (e.g., §2.2.1, §3.2.1) are asserted without a systematic prior-art search or a clear definition of the comparison scope; please add explicit context or citations to substantiate these novelty statements.

Circularity Check

0 steps flagged · score 0.0 of 10

No construction-level circularity: reported scores come from human juries against external models; the self-referential VSR test set and the abstract/Table 3 contradiction are separate validity and correctness concerns.

full rationale

The claimed derivations do not reduce to their inputs by construction. Hunyuan-Game-Image and Hunyuan-Game-Video are fine-tuned from stated base models (e.g., HunyuanCustom [21], HunyuanVideo [25]), and their reported performances in Tables 1 and 3 are produced by three human annotators comparing outputs of all models side-by-side, not by the proprietary aesthetic scoring operators used internally for data filtering and quality tuning. Although the aesthetic scoring model (Section 2.1.2) selects training data and guides the QT stage (Section 2.1.4), the Table 1 'Aesthetics' score is a human jury rating, so the evaluation is not the fitted operator itself; at most there is a preference-alignment loop, which is a validity concern rather than a circular reduction. The video super-resolution test set includes 40 videos generated by Hunyuan-Game I2V (Section 3.4.4), making that benchmark partially self-referential, but all compared methods are scored by annotators on the same inputs, so the comparison does not follow from the model's fitted parameters by construction. Self-citations to HunyuanCustom and HunyuanVideo are used as base models for fine-tuning, but no uniqueness or ansatz conclusion is imported from them. Separately, the abstract's claim of 'surpassing competitors like Midjourney, Kling and Wan' is internally inconsistent with Table 3, where Kling 1.6 Pro leads overall (3.47 vs 3.31) and in visual quality (3.94 vs 3.86); this is a significant correctness problem in the SOTA claim, but it is not a circularity of the type scored here.

Assumptions & free parameters 5 free parameters · 4 assumptions · 2 invented entities

The central claim rests on the validity of internal aesthetic scoring and self-constructed benchmarks. The model is trained to match these internal standards, then evaluated on dimensions derived from the same framework, with several hand-chosen thresholds and data volumes that are not justified by external evidence.

free parameters (5)
  • Caption length sampling ratio = 1:1:1:7 (short:medium:detailed:comprehensive)
    Hand-chosen sampling ratio for training captions in Section 2.1.3; affects text-image alignment and is not derived from data.
  • Image data resolution threshold = 1024x1024 minimum
    Filtering criterion in Section 2.1.2; determines which images enter the Gold-tier dataset and thus shapes the model's output resolution and content.
  • Annotator agreement thresholds = 80% agreement; 70%/95% acceptance
    Ad hoc thresholds in Section 2.1.2 for accepting human aesthetic annotations; directly determines the training labels for the aesthetic scoring model.
  • Video aesthetic filtering thresholds = 700K SFT videos, 80K QT videos
    Data volume choices in Section 3.1.4; the quality-tuning subset size affects final model fidelity and motion quality.
  • Keyboard-to-camera motion parameters = pre-defined speed and angle per action
    Section 3.5.3 maps discrete key presses to continuous camera space with unspecified speed/angle values; these parameters control the interactive generation behavior.
assumptions (4)
  • domain assumption Human aesthetic judgments are decomposable into six objective dimensions and learnable by a regression model.
    Section 2.1.2 assumes color harmony, light/shadow, structure, form fluidity, completeness, and composition can be scored on 1-5 scales and predicted by a multimodal model, without external validation.
  • domain assumption The authors' self-built evaluation sets are representative of game-asset generation quality.
    Sections 2.1.6, 3.1.5, and 3.4.4 introduce benchmarks of 268 prompts, 200 images, and 80 videos constructed by the authors; no independent or third-party benchmark is used.
  • domain assumption Base models HunyuanVideo and Hunyuan-DiT provide suitable priors for game-specific content.
    The paper fine-tunes from these internal models (Sections 2.1.4, 3.1.4, 3.4.3); if the base priors encode non-game visual biases, the fine-tuned results inherit them.
  • domain assumption Designer feedback used to claim a 60% efficiency improvement is reliable and quantified.
    Section 2.3.4 states the improvement based on unstated feedback from designers who used the model; no survey methodology or sample size is given.
invented entities (2)
  • Proprietary six-dimensional aesthetic scoring system
    purpose: Filter training data and define quality dimensions that the generation models are optimized toward
    Introduced in Section 2.1.2, this internal scoring instrument is not validated against any external aesthetic benchmark, and it is reused in the evaluation criteria, creating a feedback loop.
  • Effect-enhanced material data
    purpose: Scale up visual effects training data by applying effects to general materials via ControlNet and IP-Adapter
    Section 2.2.2 describes this synthetic data type; its quality is asserted without quantitative validation.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Hunyuan-Game: Industrial-grade Intelligent Game Creation Model." pith.science (2026). https://pith.science/paper/2QBOD34L

@misc{pith2026250514135,
  author       = {Pith},
  title        = {Pith review of: Hunyuan-Game: Industrial-grade Intelligent Game Creation Model},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/2QBOD34L}},
  note         = {Machine review of arXiv:2505.14135}
}
read the original abstract

Intelligent game creation represents a transformative advancement in game development, utilizing generative artificial intelligence to dynamically generate and enhance game content. Despite notable progress in generative models, the comprehensive synthesis of high-quality game assets, including both images and videos, remains a challenging frontier. To create high-fidelity game content that simultaneously aligns with player preferences and significantly boosts designer efficiency, we present Hunyuan-Game, an innovative project designed to revolutionize intelligent game production. Hunyuan-Game encompasses two primary branches: image generation and video generation. The image generation component is built upon a vast dataset comprising billions of game images, leading to the development of a group of customized image generation models tailored for game scenarios: (1) General Text-to-Image Generation. (2) Game Visual Effects Generation, involving text-to-effect and reference image-based game visual effect generation. (3) Transparent Image Generation for characters, scenes, and game visual effects. (4) Game Character Generation based on sketches, black-and-white images, and white models. The video generation component is built upon a comprehensive dataset of millions of game and anime videos, leading to the development of five core algorithmic models, each targeting critical pain points in game development and having robust adaptation to diverse game video scenarios: (1) Image-to-Video Generation. (2) 360 A/T Pose Avatar Video Synthesis. (3) Dynamic Illustration Generation. (4) Generative Video Super-Resolution. (5) Interactive Game Video Generation. These image and video generation models not only exhibit high-level aesthetic expression but also deeply integrate domain-specific knowledge, establishing a systematic understanding of diverse game and anime art styles.

Figures

Figures reproduced from arXiv: 2505.14135 by the authors.

Figure 1
Figure 1. Hunyuan-Game-Image. The image generation capabilities of Hunyuan-Game include text￾to-image generation, text-to-game effects generation, reference-based game visual effects generation, transparent and seamless image generation, and game character/scene generation. These capabilities offer a powerful toolset that significantly reduces the time and resources required for content creation, thereby greatly enhancing the… view at source ↗
Figure 2
Figure 2. Hunyuan-Game-Video. The five key video generation capabilities of Hunyuan-Game are demonstrated as follows: an Image-to-Video generation and 360-degree A/T Pose Avatar Video Synthesis (I2V); Dynamic Illustration Generation based on first and last frame generation (FLF2V); video super-resolution from original video content (V2V); and Interactive Game Video Generation based on text or image input. 3 [PITH_FULL_IMAGE:… view at source ↗
Figure 3
Figure 3. The data filtering pipeline. address this, our study developed a fine-grained, as shown in [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗
Figures from the paper (32 more)
Figure 4
Figure 4. Figure 4: Examples of multi-dimensional aesthetic scores. [PITH_FULL_IMAGE:figures/full_fig_p007_4.png]
Figure 5
Figure 5. Figure 5: Examples of Multi-Length Captions. Construction and Tiered Filtering of Game Datasets In this study, a proprietary captioning model was employed to generate textual annotations for image data. To ensure stable output across prompts of varying lengths, as shown in [PIT…
Figure 6
Figure 6. Figure 6: Prompt rewriting can significantly add content information to the picture, thus enhancing [PITH_FULL_IMAGE:figures/full_fig_p009_6.png]
Figure 7
Figure 7. Figure 7: Visualizations of text-to-image generation results. [PITH_FULL_IMAGE:figures/full_fig_p010_7.png]
Figure 8
Figure 8. Figure 8: The pipeline to create effectualized materials. [PITH_FULL_IMAGE:figures/full_fig_p012_8.png]
Figure 9
Figure 9. Figure 9: Examples of brief, detailed, and comprehensive descriptions. [PITH_FULL_IMAGE:figures/full_fig_p013_9.png]
Figure 10
Figure 10. Figure 10: Rewritten descriptions can significantly enhance the details and texture of generated [PITH_FULL_IMAGE:figures/full_fig_p014_10.png]
Figure 11
Figure 11. Figure 11: Qualitative comparisons with State-of-The-Art methods, [PITH_FULL_IMAGE:figures/full_fig_p015_11.png]
Figure 12
Figure 12. Figure 12: Visualizations of text-to-game visual effects generation results. [PITH_FULL_IMAGE:figures/full_fig_p016_12.png]
Figure 13
Figure 13. Figure 13: Game effects generation results based on (a) black sketch control, (b) color sketch control, [PITH_FULL_IMAGE:figures/full_fig_p018_13.png]
Figure 14
Figure 14. Figure 14: The pipeline of black-and-white draft generation. [PITH_FULL_IMAGE:figures/full_fig_p019_14.png]
Figure 15
Figure 15. Figure 15: Visualizations of image-to-game visual effects generation results. [PITH_FULL_IMAGE:figures/full_fig_p020_15.png]
Figure 16
Figure 16. Figure 16: Hunyuan transparent image generation 22 [PITH_FULL_IMAGE:figures/full_fig_p022_16.png]
Figure 17
Figure 17. Figure 17: Seamless tile image generation. The first row represents the generation of seamless tile [PITH_FULL_IMAGE:figures/full_fig_p023_17.png]
Figure 18
Figure 18. Figure 18: Method of Game Character Generation for lineart to grayscale image and grayscale image [PITH_FULL_IMAGE:figures/full_fig_p025_18.png]
Figure 19
Figure 19. Figure 19: Method of Game Character Generation for character consistency. [PITH_FULL_IMAGE:figures/full_fig_p025_19.png]
Figure 20
Figure 20. Figure 20: Visualizations of lineart→Grayscale→Character image generation. leading game AIGC platform. As illustrated in [PITH_FULL_IMAGE:figures/full_fig_p026_20.png]
Figure 21
Figure 21. Figure 21: Visualizations of consistent game character generation results. [PITH_FULL_IMAGE:figures/full_fig_p027_21.png]
Figure 22
Figure 22. Figure 22: The data filtering pipeline. The filtering methodology was systematically implemented based on the model-generated labels. For 2D animation content, we enforced rigorous motion-based filtering criteria to exclude static or minimally dynamic sequences, retaining exclus…
Figure 23
Figure 23. Figure 23: An example of the structured caption for video clips. The VLM generated the caption sequentially from the most detailed caption to a summarised caption to several labels. It mocks a chain-of-thought process that can reduce the probability of model hallucination. The s…
Figure 24
Figure 24. Figure 24: Qualitative results of Hunyuan-Game. Our method exhibits robust ID preservation and [PITH_FULL_IMAGE:figures/full_fig_p031_24.png]
Figure 25
Figure 25. Figure 25: Examples of the prompt rewriting. Rewriting prompts can reduce the probability of subject distortion and enhance the model’s ability to follow instructions. 3.1.5 Evaluation We developed an evaluation dataset comprising approximately 200 images, which encompasses a di…
Figure 26
Figure 26. Figure 26: Mesh data processing pipeline used to generate dataset used in training. We filter character [PITH_FULL_IMAGE:figures/full_fig_p034_26.png]
Figure 27
Figure 27. Figure 27: Framework fo Hunyuan-Game 360° Character Video Generation [PITH_FULL_IMAGE:figures/full_fig_p034_27.png]
Figure 28
Figure 28. Figure 28: Qualitative comparisons with State-of-The-Art methods, to align the results, for the Kling [PITH_FULL_IMAGE:figures/full_fig_p035_28.png]
Figure 29
Figure 29. Figure 29: Qualitative results of 360° Character Video Generation. Our method demonstrates excellent character consistency and rotation robustness, and it is capable of generating reasonable clothing and texture details from different viewpoints. 36 [PITH_FULL_IMAGE:figures/ful…
Figure 30
Figure 30. Figure 30: The dynamic illustration training data we collected is divided into three levels: [PITH_FULL_IMAGE:figures/full_fig_p038_30.png]
Figure 31
Figure 31. Figure 31: Qualitative comparisons with State-of-The-Art methods, [PITH_FULL_IMAGE:figures/full_fig_p039_31.png]
Figure 32
Figure 32. Figure 32: Framework of Hunyuan-Game-Generative Video Super-Resolution. [PITH_FULL_IMAGE:figures/full_fig_p041_32.png]
Figure 33
Figure 33. Figure 33: Qualitative results of Hunyuan-Game-Video Super-Resolution. We compare our method [PITH_FULL_IMAGE:figures/full_fig_p042_33.png]
Figure 34
Figure 34. Figure 34: Overall framework of Hunyuan-GameCraft. as candidate split points. After partition, we reconstruct six-degree-of-freedom (6-DoF) camera trajectories using MonST3R [87], enabling precise modeling of viewpoint dynamics. For game video annotation, we follow the annotatio…
Figure 35
Figure 35. Figure 35: Qualitative results of interactive game video generation. Given discrete keyboard action [PITH_FULL_IMAGE:figures/full_fig_p045_35.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Yume: An Interactive World Generation Model

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A diffusion-based video model generates extendable, keyboard-controlled walkthroughs from a single input image, using quantized camera actions as text prompts.

  2. Hunyuan-GameCraft: High-dynamic Interactive Game Video Generation with Hybrid History Condition

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Hunyuan-GameCraft generates long, action-controlled game videos from a single image by unifying keyboard/mouse inputs into a continuous camera space and conditioning on mixed historical context.

Reference graph

Works this paper leans on

95 extracted references · 29 canonical work pages · cited by 2 Pith papers

  1. [1]

    Cosmos world foundation model platform for physical ai

    Niket Agarwal, Arslan Ali, Maciej Bala, Yogesh Balaji, Erik Barker, Tiffany Cai, Prithvijit Chattopadhyay, Yongxin Chen, Yin Cui, Yifan Ding, et al. Cosmos world foundation model platform for physical ai. arXiv preprint arXiv:2501.03575, 2025

  2. [2]

    Kling ai: Next-generation ai creative studio

    Kling AI. Kling ai: Next-generation ai creative studio. https://www.klingai.com/, 2024

  3. [3]

    Layer ai: Game art without limits

    Layer AI. Layer ai: Game art without limits. https://app.layer.ai/, 2025

  4. [4]

    Blended latent diffusion

    Omri Avrahami, Ohad Fried, and Dani Lischinski. Blended latent diffusion. TOG, 42(4):1–11, 2023

  5. [5]

    Ac3d: Analyzing and improving 3d camera control in video diffusion transformers

    Sherwin Bahmani, Ivan Skorokhodov, Guocheng Qian, Aliaksandr Siarohin, Willi Menapace, Andrea Tagliasacchi, David B Lindell, and Sergey Tulyakov. Ac3d: Analyzing and improving 3d camera control in video diffusion transformers. arXiv preprint arXiv:2411.18673, 2024

  6. [6]

    Stable video diffusion: Scaling latent video diffusion models to large datasets, 2023

    Andreas Blattmann, Tim Dockhorn, Sumith Kulal, Daniel Mendelevitch, Maciej Kilian, Dominik Lorenz, Yam Levi, Zion English, Vikram V oleti, Adam Letts, Varun Jampani, and Robin Rombach. Stable video diffusion: Scaling latent video diffusion models to large datasets, 2023

  7. [7]

    Align your latents: High-resolution video synthesis with latent diffusion models

    Andreas Blattmann, Robin Rombach, Huan Ling, Tim Dockhorn, Seung Wook Kim, Sanja Fidler, and Karsten Kreis. Align your latents: High-resolution video synthesis with latent diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 22563–22575, 2023

  8. [8]

    Motionclr: Motion generation and training-free editing via understanding attention mechanisms

    Ling-Hao Chen, Wenxun Dai, Xuan Ju, Shunlin Lu, and Lei Zhang. Motionclr: Motion generation and training-free editing via understanding attention mechanisms. arXiv e-prints, pages arXiv–2410, 2024

Show all 95 references
  1. [9]

    Panic-3d: Stylized single-view 3d reconstruction from portraits of anime characters

    Shuhong Chen, Kevin Zhang, Yichun Shi, Heng Wang, Yiheng Zhu, Guoxian Song, Sizhe An, Janus Kristjansson, Xiao Yang, and Matthias Zwicker. Panic-3d: Stylized single-view 3d reconstruction from portraits of anime characters. In Proceedings of the IEEE/CVF Conference on Computer...

  2. [10]

    Motionlcm: Real-time controllable motion generation via latent consistency model

    Wenxun Dai, Ling-Hao Chen, Jingbo Wang, Jinpeng Liu, Bo Dai, and Yansong Tang. Motionlcm: Real-time controllable motion generation via latent consistency model. In ECCV, pages 390–408, 2024

  3. [11]

    Veo 2: Our state-of-the-art video generation model

    Google Deepmind. Veo 2: Our state-of-the-art video generation model. https://deepmind.google/ technologies/veo/veo-2/, 2024

  4. [12]

    Scaling rectified flow transformers for high-resolution image synthesis

    Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas Müller, Harry Saini, Yam Levi, Dominik Lorenz, Axel Sauer, Frederic Boesel, et al. Scaling rectified flow transformers for high-resolution image synthesis. In ICML, 2024

  5. [13]

    Freemotion: A unified framework for number-free text-to-motion synthesis

    Ke Fan, Junshu Tang, Weijian Cao, Ran Yi, Moran Li, Jingyu Gong, Jiangning Zhang, Yabiao Wang, Chengjie Wang, and Lizhuang Ma. Freemotion: A unified framework for number-free text-to-motion synthesis. In ECCV, pages 93–109, 2024

  6. [14]

    The matrix: Infinite-horizon world generation with real-time moving control

    Ruili Feng, Han Zhang, Zhantao Yang, Jie Xiao, Zhilei Shu, Zhiheng Liu, Andy Zheng, Yukun Huang, Yu Liu, and Hongyang Zhang. The matrix: Infinite-horizon world generation with real-time moving control. arXiv preprint arXiv:2412.03568, 2024

  7. [15]

    Seedream 3.0 technical report

    Yu Gao, Lixue Gong, Qiushan Guo, Xiaoxia Hou, Zhichao Lai, Fanshi Li, Liang Li, Xiaochen Lian, Chao Liao, Liyang Liu, et al. Seedream 3.0 technical report. arXiv preprint arXiv:2504.11346, 2025

  8. [16]

    Mineworld: a real-time and open-source interactive world model on minecraft

    Junliang Guo, Yang Ye, Tianyu He, Haoyu Wu, Yushu Jiang, Tim Pearce, and Jiang Bian. Mineworld: a real-time and open-source interactive world model on minecraft. arXiv preprint arXiv:2504.08388, 2025

  9. [17]

    World models

    David Ha and Jürgen Schmidhuber. World models. arXiv preprint arXiv:1803.10122, 2018

  10. [18]

    Cameractrl: Enabling camera control for text-to-video generation

    Hao He, Yinghao Xu, Yuwei Guo, Gordon Wetzstein, Bo Dai, Hongsheng Li, and Ceyuan Yang. Cameractrl: Enabling camera control for text-to-video generation. arXiv preprint arXiv:2404.02101, 2024

  11. [19]

    Venhancer: Generative space-time enhancement for video generation

    Jingwen He, Tianfan Xue, Dongyang Liu, Xinqi Lin, Peng Gao, Dahua Lin, Yu Qiao, Wanli Ouyang, and Ziwei Liu. Venhancer: Generative space-time enhancement for video generation. arXiv preprint arXiv:2407.07667, 2024

  12. [20]

    Lora: Low-rank adaptation of large language models

    Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, Weizhu Chen, et al. Lora: Low-rank adaptation of large language models. ICLR, page 3, 2022. 47

  13. [21]

    Hunyuancustom: A multimodal-driven architecture for customized video generation

    Teng Hu, Zhentao Yu, Zhengguang Zhou, Sen Liang, Yuan Zhou, Qin Lin, and Qinglin Lu. Hunyuancustom: A multimodal-driven architecture for customized video generation. arXiv preprint arXiv:2505.04512, 2025

  14. [22]

    Step-video-ti2v technical report: A state-of-the-art text-driven image-to-video generation model

    Haoyang Huang, Guoqing Ma, Nan Duan, Xing Chen, Changyi Wan, Ranchen Ming, Tianyu Wang, Bo Wang, Zhiying Lu, Aojie Li, et al. Step-video-ti2v technical report: A state-of-the-art text-driven image-to-video generation model. arXiv preprint arXiv:2503.11251, 2025

  15. [23]

    In-context lora for diffusion transformers

    Lianghua Huang, Wei Wang, Zhi-Fan Wu, Yupeng Shi, Huanzhang Dou, Chen Liang, Yutong Feng, Yu Liu, and Jingren Zhou. In-context lora for diffusion transformers. arXiv preprint arXiv:2410.23775, 2024

  16. [24]

    Sport: From zero-shot prompts to real-time motion generation

    Bin Ji, Ye Pan, Zhimeng Liu, Shuai Tan, and Xiaokang Yang. Sport: From zero-shot prompts to real-time motion generation. TVCG, 2025

  17. [25]

    Hunyuanvideo: A systematic framework for large video generative models

    Weijie Kong, Qi Tian, Zijian Zhang, Rox Min, Zuozhuo Dai, Jin Zhou, Jiangfeng Xiong, Xin Li, Bo Wu, Jianwei Zhang, et al. Hunyuanvideo: A systematic framework for large video generative models. arXiv preprint arXiv:2412.03603, 2024

  18. [26]

    Matroska and ffv1: One file format for film and video archiving? JFP, (96):41, 2017

    Reto Kromer. Matroska and ffv1: One file format for film and video archiving? JFP, (96):41, 2017

  19. [27]

    Multi-concept customization of text-to-image diffusion

    Nupur Kumari, Bingliang Zhang, Richard Zhang, Eli Shechtman, and Jun-Yan Zhu. Multi-concept customization of text-to-image diffusion. In CVPR, pages 1931–1941, 2023

  20. [28]

    Black Forest Labs. Flux. https://github.com/black-forest-labs/flux , 2024

  21. [29]

    Hunyuan-dit: A powerful multi-resolution diffusion transformer with fine-grained chinese understanding

    Zhimin Li, Jianwei Zhang, Qin Lin, Jiangfeng Xiong, Yanxin Long, Xinchi Deng, Yingfang Zhang, Xingchao Liu, Minbin Huang, Zedong Xiao, et al. Hunyuan-dit: A powerful multi-resolution diffusion transformer with fine-grained chinese understanding. arXiv preprint arXiv:2405.08748, 2024

  22. [30]

    Improved baselines with visual instruction tuning

    Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee. Improved baselines with visual instruction tuning. In CVPR, pages 26296–26306, 2024

  23. [31]

    Visual instruction tuning

    Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. Visual instruction tuning. NeurIPS, pages 34892–34916, 2023

  24. [32]

    Plan, posture and go: Towards open-world text-to-motion generation

    Jinpeng Liu, Wenxun Dai, Chunyu Wang, Yiji Cheng, Yansong Tang, and Xin Tong. Plan, posture and go: Towards open-world text-to-motion generation. arXiv preprint arXiv:2312.14828, 2023

  25. [33]

    Syncdreamer: Generating multiview-consistent images from a single-view image

    Yuan Liu, Cheng Lin, Zijiao Zeng, Xiaoxiao Long, Lingjie Liu, Taku Komura, and Wenping Wang. Syncdreamer: Generating multiview-consistent images from a single-view image. arXiv preprint arXiv:2309.03453, 2023

  26. [34]

    Manganinja: Line art colorization with precise reference following

    Zhiheng Liu, Ka Leong Cheng, Xi Chen, Jie Xiao, Hao Ouyang, Kai Zhu, Yu Liu, Yujun Shen, Qifeng Chen, and Ping Luo. Manganinja: Line art colorization with precise reference following. arXiv preprint arXiv:2501.08332, 2025

  27. [35]

    Scamo: Exploring the scaling law in autoregressive motion generation model

    Shunlin Lu, Jingbo Wang, Zeyu Lu, Ling-Hao Chen, Wenxun Dai, Junting Dong, Zhiyang Dou, Bo Dai, and Ruimao Zhang. Scamo: Exploring the scaling law in autoregressive motion generation model. arXiv preprint arXiv:2412.14559, 2024

  28. [36]

    Minimax. Hailuo. https://hailuoai.com/video, 2024

  29. [37]

    OpenAI. Sora. https://openai.com/sora/, 2024

  30. [38]

    Dinov2: Learning robust visual features without supervision

    Maxime Oquab, Timothée Darcet, Théo Moutakanni, Huy V o, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, et al. Dinov2: Learning robust visual features without supervision. arXiv preprint arXiv:2304.07193, 2023

  31. [39]

    Deep blind video super-resolution

    Jinshan Pan, Haoran Bai, Jiangxin Dong, Jiawei Zhang, and Jinhui Tang. Deep blind video super-resolution. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 4811–4820, 2021

  32. [40]

    Tokenhsi: Unified synthesis of physical human-scene interactions through task tokenization

    Liang Pan, Zeshi Yang, Zhiyang Dou, Wenjia Wang, Buzhen Huang, Bo Dai, Taku Komura, and Jingbo Wang. Tokenhsi: Unified synthesis of physical human-scene interactions through task tokenization. arXiv preprint arXiv:2503.19901, 2025

  33. [41]

    Genie 2: A large-scale foundation world model

    J Parker-Holder, P Ball, J Bruce, V Dasagi, K Holsheimer, C Kaplanis, A Moufarek, G Scully, J Shar, J Shi, et al. Genie 2: A large-scale foundation world model. https://deepmind.google/discover/ blog/genie-2-a-large-scale-foundation-world-model , 2024. 48

  34. [42]

    Scalable diffusion models with transformers

    William Peebles and Saining Xie. Scalable diffusion models with transformers. InICCV, pages 4195–4205, 2023

  35. [43]

    Charactergen: Efficient 3d character generation from single images with multi-view pose canonicalization

    Hao-Yang Peng, Jia-Peng Zhang, Meng-Hao Guo, Yan-Pei Cao, and Shi-Min Hu. Charactergen: Efficient 3d character generation from single images with multi-view pose canonicalization. ACM Transactions on Graphics (TOG), 43(4):1–13, 2024

  36. [44]

    Pyscenedetect developers

    PySceneDetect. Pyscenedetect developers. https://www.scenedetect.com/, 2024

  37. [45]

    Learning transferable visual models from natural language supervision

    Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In ICML, pages 8748–8763. PmLR, 2021

  38. [46]

    Direct preference optimization: Your language model is secretly a reward model

    Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn. Direct preference optimization: Your language model is secretly a reward model. NeurIPS, pages 53728–53741, 2023

  39. [47]

    High-resolution image synthesis with latent diffusion models

    Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10684–10695, 2022

  40. [48]

    Photorealistic text-to- image diffusion models with deep language understanding

    Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily L Denton, Kamyar Ghasemipour, Raphael Gontijo Lopes, Burcu Karagol Ayan, Tim Salimans, et al. Photorealistic text-to- image diffusion models with deep language understanding. Advances in neural informatio...

  41. [49]

    Laion-5b: An open large-scale dataset for training next generation image-text models

    Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Wortsman, et al. Laion-5b: An open large-scale dataset for training next generation image-text models. NeurIPS, pages 25278–25294, 2022

  42. [50]

    Seaweed-7b: Cost-effective training of video generation foundation model

    Team Seawead, Ceyuan Yang, Zhijie Lin, Yang Zhao, Shanchuan Lin, Zhibei Ma, Haoyuan Guo, Hao Chen, Lu Qi, Sen Wang, et al. Seaweed-7b: Cost-effective training of video generation foundation model. arXiv preprint arXiv:2504.08685, 2025

  43. [51]

    Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, Y . K. Li, Y . Wu, and Daya Guo. Deepseekmath: Pushing the limits of mathematical reasoning in open language models. arXiv preprint arXiv:2402.03300, 2024

  44. [52]

    Zero123++: a single image to consistent multi-view diffusion base model

    Ruoxi Shi, Hansheng Chen, Zhuoyang Zhang, Minghua Liu, Chao Xu, Xinyue Wei, Linghao Chen, Chong Zeng, and Hao Su. Zero123++: a single image to consistent multi-view diffusion base model. arXiv preprint arXiv:2310.15110, 2023

  45. [53]

    Light field networks: Neural scene representations with single-evaluation rendering

    Vincent Sitzmann, Semon Rezchikov, Bill Freeman, Josh Tenenbaum, and Fredo Durand. Light field networks: Neural scene representations with single-evaluation rendering. NeurIPS, pages 19313–19325, 2021

  46. [54]

    Hunyuan-large: An open-source moe model with 52 billion activated parameters by tencent

    Xingwu Sun, Yanfeng Chen, Yiqing Huang, Ruobing Xie, Jiaqi Zhu, Kai Zhang, Shuaipeng Li, Zhen Yang, Jonny Han, Xiaobo Shu, et al. Hunyuan-large: An open-source moe model with 52 billion activated parameters by tencent. arXiv preprint arXiv:2411.02265, 2024

  47. [55]

    Make-it-3d: High-fidelity 3d creation from a single image with diffusion prior

    Junshu Tang, Tengfei Wang, Bo Zhang, Ting Zhang, Ran Yi, Lizhuang Ma, and Dong Chen. Make-it-3d: High-fidelity 3d creation from a single image with diffusion prior. In Proceedings of the IEEE/CVF international conference on computer vision, pages 22819–22829, 2023

  48. [56]

    Make-it-vivid: dressing your animatable biped cartoon characters from text

    Junshu Tang, Yanhong Zeng, Ke Fan, Xuheng Wang, Bo Dai, Kai Chen, and Lizhuang Ma. Make-it-vivid: dressing your animatable biped cartoon characters from text. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6243–6253, 2024

  49. [57]

    Learning motion refinement for unsupervised face animation

    Jiale Tao, Shuhang Gu, Wen Li, and Lixin Duan. Learning motion refinement for unsupervised face animation. Advances in Neural Information Processing Systems, 36:70483–70496, 2023

  50. [58]

    Motion transformer for unsupervised image animation

    Jiale Tao, Biao Wang, Tiezheng Ge, Yuning Jiang, Wen Li, and Lixin Duan. Motion transformer for unsupervised image animation. In European conference on computer vision, pages 702–719. Springer, 2022

  51. [59]

    Structure-aware motion transfer with deformable anchor model

    Jiale Tao, Biao Wang, Borun Xu, Tiezheng Ge, Yuning Jiang, Wen Li, and Lixin Duan. Structure-aware motion transfer with deformable anchor model. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 3637–3646, June 2022. 49

  52. [60]

    Instantcharacter: Personalize any characters with a scalable diffusion transformer framework

    Jiale Tao, Yanbing Zhang, Qixun Wang, Yiji Cheng, Haofan Wang, Xu Bai, Zhengguang Zhou, Ruihuang Li, Linqing Wang, Chunyu Wang, et al. Instantcharacter: Personalize any characters with a scalable diffusion transformer framework. arXiv preprint arXiv:2504.12395, 2025

  53. [61]

    Midjourney

    Midjourney Team. Midjourney. https://www.midjourney.com/, 2023

  54. [62]

    Generating worlds

    World Labs Team. Generating worlds. https://www.worldlabs.ai/blog, 2024

  55. [63]

    Raft: Recurrent all-pairs field transforms for optical flow

    Zachary Teed and Jia Deng. Raft: Recurrent all-pairs field transforms for optical flow. In ECCV, pages 402–419, 2020

  56. [64]

    Diffusion model alignment using direct preference optimization

    Bram Wallace, Meihua Dang, Rafael Rafailov, Linqi Zhou, Aaron Lou, Senthil Purushwalkam, Stefano Ermon, Caiming Xiong, Shafiq Joty, and Nikhil Naik. Diffusion model alignment using direct preference optimization. In CVPR, pages 8228–8238, 2024

  57. [65]

    Wan: Open and advanced large-scale video generative models

    Ang Wang, Baole Ai, Bin Wen, Chaojie Mao, Chen-Wei Xie, Di Chen, Feiwu Yu, Haiming Zhao, Jianxiao Yang, Jianyuan Zeng, et al. Wan: Open and advanced large-scale video generative models. arXiv preprint arXiv:2503.20314, 2025

  58. [66]

    Apisr: anime production inspired real-world anime super-resolution

    Boyang Wang, Fengyu Yang, Xihang Yu, Chao Zhang, and Hanbin Zhao. Apisr: anime production inspired real-world anime super-resolution. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 25574–25584, 2024

  59. [67]

    Phased consistency models

    Fu-Yun Wang, Zhaoyang Huang, Alexander Bergman, Dazhong Shen, Peng Gao, Michael Lingelbach, Keqiang Sun, Weikang Bian, Guanglu Song, Yu Liu, et al. Phased consistency models. NeurIPS, pages 83951–84009, 2024

  60. [68]

    Seedvr: Seeding infinity in diffusion transformer towards generic video restoration

    Jianyi Wang, Zhijie Lin, Meng Wei, Yang Zhao, Ceyuan Yang, Fei Xiao, Chen Change Loy, and Lu Jiang. Seedvr: Seeding infinity in diffusion transformer towards generic video restoration. arXiv preprint arXiv:2501.01320, 2025

  61. [69]

    Exploit- ing diffusion prior for real-world image super-resolution

    Jianyi Wang, Zongsheng Yue, Shangchen Zhou, Kelvin CK Chan, and Chen Change Loy. Exploit- ing diffusion prior for real-world image super-resolution. International Journal of Computer Vision , 132(12):5929–5949, 2024

  62. [70]

    Real-esrgan: Training real-world blind super- resolution with pure synthetic data

    Xintao Wang, Liangbin Xie, Chao Dong, and Ying Shan. Real-esrgan: Training real-world blind super- resolution with pure synthetic data. In Proceedings of the IEEE/CVF international conference on computer vision, pages 1905–1914, 2021

  63. [71]

    Image quality assessment: from error visibility to structural similarity

    Zhou Wang, Alan C Bovik, Hamid R Sheikh, and Eero P Simoncelli. Image quality assessment: from error visibility to structural similarity. IEEE transactions on image processing, 13(4):600–612, 2004

  64. [72]

    Motionctrl: A unified and flexible motion controller for video generation

    Zhouxia Wang, Ziyang Yuan, Xintao Wang, Yaowei Li, Tianshui Chen, Menghan Xia, Ping Luo, and Ying Shan. Motionctrl: A unified and flexible motion controller for video generation. In SIGGRAPH, pages 1–11, 2024

  65. [73]

    Vmix: Improving text-to-image diffusion model with cross-attention mixing control

    Shaojin Wu, Fei Ding, Mengqi Huang, Wei Liu, and Qian He. Vmix: Improving text-to-image diffusion model with cross-attention mixing control. arXiv preprint arXiv:2412.20800, 2024

  66. [74]

    Motionstreamer: Streaming motion generation via diffusion-based autoregressive model in causal latent space

    Lixing Xiao, Shunlin Lu, Huaijin Pi, Ke Fan, Liang Pan, Yueer Zhou, Ziyong Feng, Xiaowei Zhou, Sida Peng, and Jingbo Wang. Motionstreamer: Streaming motion generation via diffusion-based autoregressive model in causal latent space. arXiv preprint arXiv:2503.15451, 2025

  67. [75]

    Move as you like: image animation in e-commerce scenario

    Borun Xu, Biao Wang, Jiale Tao, Tiezheng Ge, Yuning Jiang, Wen Li, and Lixin Duan. Move as you like: image animation in e-commerce scenario. In Proceedings of the 29th ACM international conference on multimedia, pages 2759–2761, 2021

  68. [76]

    Learning semantic latent directions for accurate and controllable human motion prediction

    Guowei Xu, Jiale Tao, Wen Li, and Lixin Duan. Learning semantic latent directions for accurate and controllable human motion prediction. InEuropean Conference on Computer Vision, pages 56–73. Springer, 2024

  69. [77]

    Pandora3d: A comprehensive framework for high-quality 3d shape and texture generation

    Jiayu Yang, Taizhang Shang, Weixuan Sun, Xibin Song, Ziang Cheng, Senbo Wang, Shenzhou Chen, Weizhe Liu, Hongdong Li, and Pan Ji. Pandora3d: A comprehensive framework for high-quality 3d shape and texture generation. arXiv preprint arXiv:2502.14247, 2025

  70. [78]

    Learning interactive real-world simulators

    Mengjiao Yang, Yilun Du, Kamyar Ghasemipour, Jonathan Tompson, Dale Schuurmans, and Pieter Abbeel. Learning interactive real-world simulators. arXiv preprint arXiv:2310.06114, 1(2):6, 2023. 50

  71. [79]

    Position: video as the new language for real-world decision making

    Sherry Yang, Jacob C Walker, Jack Parker-Holder, Yilun Du, Jake Bruce, Andre Barreto, Pieter Abbeel, and Dale Schuurmans. Position: video as the new language for real-world decision making. In Forty-first International Conference on Machine Learning, 2024

  72. [80]

    Motion-guided latent diffusion for temporally consistent real-world video super-resolution

    Xi Yang, Chenhang He, Jianqi Ma, and Lei Zhang. Motion-guided latent diffusion for temporally consistent real-world video super-resolution. In European Conference on Computer Vision, pages 224–242. Springer, 2024

  73. [81]

    Real-world video super-resolution: A benchmark dataset and a decomposition based learning scheme

    Xi Yang, Wangmeng Xiang, Hui Zeng, and Lei Zhang. Real-world video super-resolution: A benchmark dataset and a decomposition based learning scheme. In Proceedings of the IEEE/CVF international conference on computer vision, pages 4781–4790, 2021

  74. [82]

    Hunyuan3d 1.0: A unified framework for text-to-3d and image-to-3d generation

    Xianghui Yang, Huiwen Shi, Bowen Zhang, Fan Yang, Jiacheng Wang, Hongxu Zhao, Xinhai Liu, Xinzhou Wang, Qingxiang Lin, Jiaao Yu, et al. Hunyuan3d 1.0: A unified framework for text-to-3d and image-to-3d generation. arXiv preprint arXiv:2411.02293, 2024

  75. [83]

    Cogvideox: Text-to-video diffusion models with an expert transformer

    Zhuoyi Yang, Jiayan Teng, Wendi Zheng, Ming Ding, Shiyu Huang, Jiazheng Xu, Yuanming Yang, Wenyi Hong, Xiaohan Zhang, Guanyu Feng, et al. Cogvideox: Text-to-video diffusion models with an expert transformer. arXiv preprint arXiv:2408.06072, 2024

  76. [84]

    Ip-adapter: Text compatible image prompt adapter for text-to-image diffusion models

    Hu Ye, Jun Zhang, Sibo Liu, Xiao Han, and Wei Yang. Ip-adapter: Text compatible image prompt adapter for text-to-image diffusion models. arXiv preprint arXiv:2308.06721, 2023

  77. [85]

    Resshift: Efficient diffusion model for image super- resolution by residual shifting

    Zongsheng Yue, Jianyi Wang, and Chen Change Loy. Resshift: Efficient diffusion model for image super- resolution by residual shifting. Advances in Neural Information Processing Systems, 36:13294–13307, 2023

  78. [86]

    Sigmoid loss for language image pre-training

    Xiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, and Lucas Beyer. Sigmoid loss for language image pre-training. In ICCV, pages 11975–11986, 2023

  79. [87]

    Monst3r: A simple approach for estimating geometry in the presence of motion

    Junyi Zhang, Charles Herrmann, Junhwa Hur, Varun Jampani, Trevor Darrell, Forrester Cole, Deqing Sun, and Ming-Hsuan Yang. Monst3r: A simple approach for estimating geometry in the presence of motion. arXiv preprint arXiv:2410.03825, 2024

  80. [88]

    Clay: A controllable large-scale generative model for creating high-quality 3d assets

    Longwen Zhang, Ziyu Wang, Qixuan Zhang, Qiwei Qiu, Anqi Pang, Haoran Jiang, Wei Yang, Lan Xu, and Jingyi Yu. Clay: A controllable large-scale generative model for creating high-quality 3d assets. ACM Transactions on Graphics (TOG), 43(4):1–20, 2024

  81. [89]

    Transparent image layer diffusion using latent transparency

    Lvmin Zhang and Maneesh Agrawala. Transparent image layer diffusion using latent transparency. arXiv preprint arXiv:2402.17113, 2024

  82. [90]

    Packing input frame context in next-frame prediction models for video generation

    Lvmin Zhang and Maneesh Agrawala. Packing input frame context in next-frame prediction models for video generation. arXiv preprint arXiv:2504.12626, 2025

  83. [91]

    Adding conditional control to text-to-image diffusion models

    Lvmin Zhang, Anyi Rao, and Maneesh Agrawala. Adding conditional control to text-to-image diffusion models. In ICCV, pages 3836–3847, 2023

  84. [92]

    Realviformer: Investigating attention for real-world video super-resolution

    Yuehan Zhang and Angela Yao. Realviformer: Investigating attention for real-world video super-resolution. In European Conference on Computer Vision, pages 412–428. Springer, 2024

  85. [93]

    Hunyuan3d 2.0: Scaling diffusion models for high resolution textured 3d assets generation

    Zibo Zhao, Zeqiang Lai, Qingxiang Lin, Yunfei Zhao, Haolin Liu, Shuhui Yang, Yifei Feng, Mingxin Yang, Sheng Zhang, Xianghui Yang, et al. Hunyuan3d 2.0: Scaling diffusion models for high resolution textured 3d assets generation. arXiv preprint arXiv:2501.12202, 2025

  86. [94]

    Upscale-a-video: Temporal-consistent diffusion model for real-world video super-resolution

    Shangchen Zhou, Peiqing Yang, Jianyi Wang, Yihang Luo, and Chen Change Loy. Upscale-a-video: Temporal-consistent diffusion model for real-world video super-resolution. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2535–2545, 2024

  87. [95]

    Allegro: Open the black box of commercial-level video generation model

    Yuan Zhou, Qiuyue Wang, Yuxuan Cai, and Huan Yang. Allegro: Open the black box of commercial-level video generation model. arXiv preprint arXiv:2410.15458, 2024. 51

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.