Pith. sign in

REVIEW 2 major objections 5 minor 295 references

Generative AI for Character Animation: A Comprehensive Survey of Techniques, Applications, and Future Directions

T0 review · 2 major / 5 minor · reviewed 2026-08-16 · deepseek-v4-flash

Pith's one-line read This survey claims that generative character animation—faces, expressions, images, avatars, gestures, motion, objects, and textures—forms one coherent technical landscape, and it supplies a single taxonomy, datasets, metrics, and…

desk verdict A genuinely broad and useful survey of generative character animation, but the FLAME citation mix-up means the reference list needs a careful pass before this can be trusted as a map. read the letter →

arxiv 2504.19056 v1 pith:QPLVPA67 submitted 2025-04-27 cs.CV cs.AIcs.CLcs.LGcs.MM

classification cs.CVcs.AIcs.CLcs.LGcs.MM
keywords surveygenerativeAIcharacteranimationdiffusionmodelsfacialgesturegenerationmotionsynthesistext-to-3D
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This survey claims that the many strands of generative character animation—face, expression, image, avatar, gesture, motion, object, and texture—form a single technical landscape, and it provides one map of that landscape in the form of a systematic taxonomy. The contribution is not a new model but a structured review: for each component it catalogues representative models, datasets, evaluation metrics, and applications, and it adds a background primer on the foundational architectures and metrics. A sympathetic reading sees the paper as answering a practical problem: the field's pace has made it hard to keep a coherent view, so a newcomer would otherwise have to stitch together several specialized surveys. If the map is faithful, a reader can enter any subfield and see how techniques such as diffusion-based synthesis, CLIP conditioning, and parametric body modeling recur across the others.

What carries the argument

The load-bearing device is the taxonomy in Figure 2, which partitions character animation into eight components and further subdivides each into named model families (for example, diffusion-guided score distillation for objects, VQ-VAE-based generation for motion, inpainting-based pipelines for texture). The taxonomy does the argumentative work: it is what unites previously separate literatures into a single map and what supports the survey's cross-component observations, such as the reappearance of diffusion and CLIP-based conditioning in both expression synthesis and gesture generation. Around the taxonomy, the survey assembles supporting machinery—a background section on foundational models and metrics, and per-component tables of datasets and evaluation measures—that gives the map its practical footing.

What would settle it

Audit the taxonomy against its references: reference [18], cited in Sections 3 and 6.1 as the FLAME parametric face model, is actually "FLAME: Free-form language-based motion synthesis & editing," a motion-generation model; if a broader audit of the taxonomy's citations finds numerous mismatches of this kind, the survey's unified map and recommendations are not reliable.

Watch

Extended reading notes

Core claim

On the paper's own terms, the discovery is that the generative AI applications for character animation are best understood as one interconnected field rather than as separate literatures on faces, gestures, avatars, and so on. The survey operationalizes that claim by organizing current work into a component taxonomy—face, expression, image, avatar, gesture, motion, object, and texture—and, within each component, classifying models by methodological family such as GAN-based, diffusion-based, transformer-based, or hybrid. It then attaches to each component its datasets, evaluation metrics, and applications, so the taxonomy functions as a single navigational tool for the whole area. The central claim is that this unified perspective is both accurate and useful: it exposes interconnections across subfields, supports newcomers with background material, and gives the field a shared agenda of open problems and future directions.

Load-bearing premise

The survey's entire value depends on its annotations of the cited literature being correct, because a unified map built on misidentified references would mislead readers instead of guiding them.

Editorial extensions

If this is right

  • A newcomer can enter any of the eight subfields from one document, using the background primer and the per-component model, dataset, and metric tables as a starting point.
  • Cross-component transfer becomes visible: the same conditioning and architecture families recur across face, gesture, and motion, which suggests that methods from one component can be reused in another.
  • Evaluation is standardized per component, so future work can report against comparable metrics (for example, FID and CLIP Score for visual quality, FGD for gestures, R-Precision for text-to-motion).
  • The open-problems section gives the field a concrete agenda, including real-time efficiency, controllability, multimodal integration, generalization across styles, identity preservation, and ethical safeguards.
  • The shared resource repository provides the datasets, benchmarks, models, and tools that the survey identifies, turning the map into a practical entry point.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the taxonomy is kept current, it could evolve into a living roadmap for the field; a static snapshot, by contrast, will age quickly given the pace the survey itself describes.
  • The unified framing implies that a method from one component might transfer to a neighbor more directly than isolated surveys suggest; a testable extension would be to take a motion diffusion model's conditioning scheme and apply it to facial-expression generation.
  • The survey's cross-component structure implicitly argues for harmonized evaluation—for instance, using distributional metrics like FGD across subfields—though the paper does not itself propose such unification.
  • Because the survey's claims are about the literature, the map is mechanically auditable: checking each taxonomy entry against its cited reference would settle whether the unified perspective is faithful.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. This manuscript is a comprehensive survey of generative AI methods for character animation, covering facial animation, expression synthesis, image generation, avatar creation, gesture modeling, motion synthesis, object generation, and texture synthesis. It provides per-topic discussions of datasets, evaluation metrics, and representative models, organizes the field into a taxonomy (Figure 2), and concludes with open problems and research directions. The stated contribution is a unified, cross-domain perspective intended to serve as an entry point for newcomers, supported by a public GitHub repository of resources.

Significance. If the survey accurately represents the cited literature, it fills a genuine gap: most prior surveys focus on a single subfield (faces, gestures, avatars, or motion), whereas this work attempts to connect them. The paper compiles a large amount of organized information—dataset tables, model taxonomies, evaluation metrics, and application discussions—that could be useful to researchers entering the field. The authors also provide a public resource repository, which adds practical value. However, the survey's value is entirely contingent on the faithfulness of its annotations of prior work, and the citation problems identified below directly affect that reliability.

major comments (2)
  1. [§3, §3.3.2, §6, §6.1, §6.3.4, §8.3.3] Reference [18] is used inconsistently and incorrectly. In §3 and §6 the text identifies [18] as the parametric face model FLAME ('Faces Learned with an Articulated Model'), and §3.3.2 and §6.1 rely on that identification when describing AlbedoGAN, the Hybrid Generator, and RenderMe-360 annotations. However, the reference list entry [18] is Kim et al. 2022, 'FLAME: Free-form language-based motion synthesis & editing,' which is a text-to-motion diffusion model, not the FLAME face model of Li et al. 2017. The same reference number is used correctly in §8.3.3 for the motion model. Thus one bibliographic number refers to two different papers, and readers of the face/avatar sections are directed to the wrong work. Because a survey's central claim is to provide a trustworthy map of the literature, this is a load-bearing error, not a cosmetic typo. I recommend reassigning the face-model FLAME citation (the authors themselves use the correct reference as [253] in §6.3.4) and performing a systematic pass to eliminate such dual-use reference numbers.
  2. [§7.3.3] The description of DiM-Gesture states that it 'utilizes an adaptive layer normalization mechanism called Mamba-2 [287].' Mamba-2 is a selective state-space architecture, not a layer normalization mechanism. As written, this misrepresents the cited method and, together with the FLAME issue, suggests that the annotation of individual papers may not be reliable in other places. The authors should verify each model description against its primary source and correct this passage, or clarify what specific mechanism from the Mamba-2 paper is being used.
minor comments (5)
  1. [Figure 1 caption] The final sentence of the caption is duplicated verbatim ('Generative AI techniques, such as transformer-based and diffusion-based models, contribute to these components, significantly enhancing quality and streamlining content creation.' appears twice).
  2. [Table 3] The entry for LAION-5B reads '5,85 billion'; the decimal comma should be a decimal point ('5.85 billion').
  3. [§3.3.3, Table 1] M3Face [34] and M3CelebA [34] are presented as key methods and datasets without disclosure that they are the authors' own work. For transparency, the survey should explicitly flag self-citations of this kind.
  4. [§6.3.3] The text states that 'Neural Body introduced structured latent codes anchored to SMPL model vertices' but provides no citation for Neural Body; please add the reference.
  5. [§11.6] The FID metric is cited as [278] along with a reference to a GAN paper; please verify that the citation points to the original FID source (Heusel et al.) or to a work that introduced the metric.

Circularity Check

0 steps flagged · score 2.0 of 10

No significant circularity: the survey derives nothing; a minor self-citation and a citation error do not form a logical loop.

full rationale

This is a survey, not a derivation: it makes no predictions, fits no parameters, and proves no theorems. Its central claim—offering a single, comprehensive perspective on generative AI for character animation—is an organizational claim supported by the taxonomy in Figure 2 and by the section-by-section aggregation of external literature, none of which depends on the authors' own prior results. The only self-citation is M3Face/M3CelebA [34], featured in Section 3.3.3 and Table 1 as one of several text-to-face methods and a dataset; removing it would not alter the survey's structure or conclusions, so it is not load-bearing. I also checked for imported uniqueness theorems, ansatz-by-citation, and renaming of known results; none occur. One non-circular defect should be noted: reference [18] (Kim et al., 'FLAME: Free-form language-based motion synthesis & editing') is used in Sections 3 and 6 as if it were the parametric face model FLAME ('Faces Learned with an Articulated Model'), while Section 8.3.3 correctly uses [18] for the motion-diffusion method Flame. The face-model FLAME paper is apparently missing from the bibliography. This mis-citation undermines the survey's reliability claim but is an annotation error, not a circular reduction. Accordingly, the circularity score is minimal.

Assumptions & free parameters 0 free parameters · 2 assumptions · 0 invented entities

The survey's central value rests on accurate curation of external literature and on the meaningfulness of its taxonomy. No free parameters or invented entities are involved.

assumptions (2)
  • domain assumption The cited papers are accurately summarized and correctly referenced.
    The survey's utility depends on faithful representation of hundreds of external papers; the FLAME [18] conflation in Section 3 demonstrates this assumption can fail.
  • domain assumption The eight component categories (face, expression, image, avatar, gesture, motion, object, texture) form a complete decomposition of character animation.
    The taxonomy is the organizing framework; if it omits important subfields or mis-assigns methods, the survey's coverage claim weakens. No formal justification is provided.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Generative AI for Character Animation: A Comprehensive Survey of Techniques, Applications, and Future Directions." pith.science (2026). https://pith.science/paper/QPLVPA67

@misc{pith2026250419056,
  author       = {Pith},
  title        = {Pith review of: Generative AI for Character Animation: A Comprehensive Survey of Techniques, Applications, and Future Directions},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/QPLVPA67}},
  note         = {Machine review of arXiv:2504.19056}
}
read the original abstract

Generative AI is reshaping art, gaming, and most notably animation. Recent breakthroughs in foundation and diffusion models have reduced the time and cost of producing animated content. Characters are central animation components, involving motion, emotions, gestures, and facial expressions. The pace and breadth of advances in recent months make it difficult to maintain a coherent view of the field, motivating the need for an integrative review. Unlike earlier overviews that treat avatars, gestures, or facial animation in isolation, this survey offers a single, comprehensive perspective on all the main generative AI applications for character animation. We begin by examining the state-of-the-art in facial animation, expression rendering, image synthesis, avatar creation, gesture modeling, motion synthesis, object generation, and texture synthesis. We highlight leading research, practical deployments, commonly used datasets, and emerging trends for each area. To support newcomers, we also provide a comprehensive background section that introduces foundational models and evaluation metrics, equipping readers with the knowledge needed to enter the field. We discuss open challenges and map future research directions, providing a roadmap to advance AI-driven character-animation technologies. This survey is intended as a resource for researchers and developers entering the field of generative AI animation or adjacent fields. Resources are available at: https://github.com/llm-lab-org/Generative-AI-for-Character-Animation-Survey.

Figures

Figures reproduced from arXiv: 2504.19056 by the authors.

Figure 1
Figure 1. Overview of different components in animated character generation. Each aspect, including face, [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Taxonomy of recent advances in generative AI for character animation, organized by key components [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. Overview of the Dual-Generator (DG) [23] network, which consists of two generators: the ID-preserving Shape Generator (IDSG) and the Reenacted Face Generator (RFG). Given a source face Is and a reference face Ir, the IDSG transforms the reference’s actions into landmarks ˆlt. Using these landmarks and Is, the RFG produces a reenacted face ˆIt that matches the pose and expression of Ir while preserving the identity o… view at source ↗
Figures from the paper (18 more)
Figure 4
Figure 4. Figure 4: Overview of InViTe [187] in three stages: (a) capturing a user’s face to produce a personalized 3D model, (b) generating intermediate outputs for rendering, and (c) performing 3D face manipulation (such as makeup style changes) on a mobile device. Reprinted from [187].…
Figure 5
Figure 5. Figure 5: Overview of AdaMesh [44] model: (a) The expression adapter integrates MoLoRA [206] parameters (striped patches) into pre-trained encoders and the decoder to enable efficient adaptation for facial expressions. (b) Architecture of the Conformer block [207], showcasing it…
Figure 6
Figure 6. Figure 6: Overview of Neural Face Rigging (NFR) [51]. The model extracts an expression code ze from an unrigged mesh with an unknown expression (yellow) and an identity code zi from a target neutral mesh (cyan). The extracted codes are combined to generate a retargeted mesh (blu…
Figure 7
Figure 7. Figure 7: The Cut-Mix-Unmix data augmentation technique for multi-subject generation, implemented in [PITH_FULL_IMAGE:figures/full_fig_p017_7.png]
Figure 8
Figure 8. Figure 8: Examples of image editing using the DiffusionDisentanglement model [ [PITH_FULL_IMAGE:figures/full_fig_p019_8.png]
Figure 9
Figure 9. Figure 9: Visual ChatGPT architecture [67] showing the workflow from user query to output. The system processes complex instructions on a flower image through a prompt manager that coordinates visual foundation models (BLIP, Stable Diffusion, ControlNet, etc.). The example demon…
Figure 10
Figure 10. Figure 10: DreamAvatar’s dual-observation-space architecture shows how text prompts and SMPL parameters [PITH_FULL_IMAGE:figures/full_fig_p023_10.png]
Figure 11
Figure 11. Figure 11: Pipeline of DreamHuman [84]: the model takes a text prompt p (e.g., a woman wearing a dress) and generates a 3D animatable avatar using a deformable, pose-conditioned NeRF constrained by the imGHUM [252] body model. The pipeline incorporates semantic zooming for criti…
Figure 12
Figure 12. Figure 12: GestureDiffuCLIP [106] integrates a CLIP encoder for semantic alignment and a diffusion process for refining gesture sequences. The model uses multi-head causal and semantic-aware attention mechanisms and adaptive instance normalization (AdaIN) to generate expressive …
Figure 13
Figure 13. Figure 13: The architecture of DiffuseStyleGesture+ [111], showcasing its multimodal integration for co-speech gesture generation. The model incorporates speaker IDs, seed gestures, audio, and text through feature extraction and representation modules. A diffusion process is emp…
Figure 14
Figure 14. Figure 14: The architecture of the TMR [120] framework shows the use of dual encoders for text and motion, and a joint embedding space for similarity-based retrieval. Reprinted from [120]. but autoregressively across the sequence of actions specified by the input texts. This all…
Figure 15
Figure 15. Figure 15: An overview of MotionGPT [21], which uses a frozen VQ-VAE and LLM, with LoRA [320] applied to fine-tune the LLM for generating motion tokens. Reprinted from [21]. Longer and more complex motion sequences are addressed in T2LM [124], which maps multi-sentence text into…
Figure 16
Figure 16. Figure 16: DreamFusion [130] generates 3D objects from text prompts like “a DSLR photo of a peacock on a surfboard." It trains a Neural Radiance Field (NeRF) from scratch for each caption, using shading from normals and a frozen Imagen model [57] to guide updates for improved ge…
Figure 17
Figure 17. Figure 17: Pipeline of Text2Tex [145], featuring a two-stage process for generating high-quality, multi-view consistent textures. Stage I generates textures via depth-to-image diffusion from predefined viewpoints. Stage II refines the results by automatically selecting additiona…
Figure 18
Figure 18. Figure 18: Meta 3D TextureGen [153] employs a two-stage architecture: a geometry-aware diffusion process generates multi-view images, followed by UV-space inpainting using incidence-aware blending. This design enables seamless, high-resolution (up to 4K) textures with minimal ar…
Figure 19
Figure 19. Figure 19: Comparison of SMPL [1], SMPL+H [424], and SMPL-X [305] models: SMPL captures basic body shapes, SMPL+H adds detailed hand poses, and SMPL-X includes body, hands, and facial expressions for a more complete and realistic representation. Reprinted from [305]. A.1.6 Multi…
Figure 20
Figure 20. Figure 20: An overview of volumetric rendering in NeRF. Reprinted from [9]. [PITH_FULL_IMAGE:figures/full_fig_p092_20.png]
Figure 21
Figure 21. Figure 21: Before ControlNet (left): A neural network block processes input [PITH_FULL_IMAGE:figures/full_fig_p095_21.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

295 extracted references · 27 canonical work pages

  1. [18]

    Flame: Free-form language-based motion synthesis & editing,

    J. Kim, J. Kim, and S. Choi, “Flame: Free-form language-based motion synthesis & editing,”arXiv preprint arXiv:2209.00349, 2022

  2. [34]

    M 3 face: A unified multi-modal multilingual framework for human face generation and editing,

    M. Mofayezi, R. Alipour, M. A. Kakavand, and E. Asgari, “M 3 face: A unified multi-modal multilingual framework for human face generation and editing,” 2024. [Online]. Available: https://arxiv.org/abs/2402.02369

  3. [253]

    Detailed, accurate, human shape estimation from clothed 3d scan sequences,

    C. Zhang, S. Pujades, M. J. Black, and G. Pons-Moll, “Detailed, accurate, human shape estimation from clothed 3d scan sequences,” inThe IEEE Conference on Computer Vision and Pattern Recognition (CVPR), July 2017

  4. [287]

    Transformers are ssms: Generalized models and efficient algorithms through structured state space duality,

    T. Dao and A. Gu, “Transformers are ssms: Generalized models and efficient algorithms through structured state space duality,” 2024. [Online]. Available: https://arxiv.org/abs/2405.21060

  5. [1]

    SMPL: A skinned multi-person linear model,

    M. Loper, N. Mahmood, J. Romero, G. Pons-Moll, and M. J. Black, “SMPL: A skinned multi-person linear model,”ACM Trans. Graphics (Proc. SIGGRAPH Asia), vol. 34, no. 6, pp. 248:1–248:16, Oct. 2015

  6. [2]

    Temporal convolutional networks: A unified approach to action segmentation,

    C. Lea, R. Vidal, A. Reiter, and G. D. Hager, “Temporal convolutional networks: A unified approach to action segmentation,” inComputer Vision – ECCV 2016 Workshops, G. Hua and H. Jégou, Eds. Cham: Springer International Publishing, 2016, pp. 47–54

  7. [3]

    Generative adversarial nets,

    I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, “Generative adversarial nets,”Advances in neural information processing systems, vol. 27, 2014

  8. [4]

    Auto-encoding variational bayes,

    D. P. Kingma and M. Welling, “Auto-encoding variational bayes,” 2022. [Online]. Available: https://arxiv.org/abs/1312.6114

Show all 295 references
  1. [5]

    Attention is all you need,

    A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,”Advances in neural information processing systems, vol. 30, 2017

  2. [6]

    Denoising diffusion probabilistic models,

    J. Ho, A. Jain, and P. Abbeel, “Denoising diffusion probabilistic models,”Advances in neural information processing systems, vol. 33, pp. 6840–6851, 2020

  3. [7]

    Learning transferable visual models from natural language supervision,

    A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever, “Learning transferable visual models from natural language supervision,” in Proceedings of the 38th International Conference on Machine L...

  4. [8]

    Adding conditional control to text-to-image diffusion models,

    L. Zhang, A. Rao, and M. Agrawala, “Adding conditional control to text-to-image diffusion models,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 3836–3847

  5. [9]

    Nerf: representing scenes as neural radiance fields for view synthesis,

    B. Mildenhall, P. P. Srinivasan, M. Tancik, J. T. Barron, R. Ramamoorthi, and R. Ng, “Nerf: representing scenes as neural radiance fields for view synthesis,”Commun. ACM, vol. 65, no. 1, p. 99–106, Dec. 2021. [Online]. Available: https://doi.org/10.1145/3503250 50

  6. [10]

    3d gaussian splatting for real-time radiance field rendering,

    B. Kerbl, G. Kopanas, T. Leimkühler, and G. Drettakis, “3d gaussian splatting for real-time radiance field rendering,” ACM Transactions on Graphics, vol. 42, no. 4, July 2023. [Online]. Available: https://repo-sam.inria.fr/fungraph/3d-gaussian-splatting/

  7. [11]

    A style-based generator architecture for generative adversarial networks,

    T. Karras, S. Laine, and T. Aila, “A style-based generator architecture for generative adversarial networks,” IEEE Transactions on Pattern Analysis & Machine Intelligence, vol. 43, no. 12, pp. 4217–4228, dec 2021

  8. [12]

    Diffsheg: A diffusion-based approach for real-time speech-driven holistic 3d expression and gesture generation,

    J. Chen, Y. Liu, J. Wang, A. Zeng, Y. Li, and Q. Chen, “Diffsheg: A diffusion-based approach for real-time speech-driven holistic 3d expression and gesture generation,” 2024. [Online]. Available: https://arxiv.org/abs/2401.04747

  9. [13]

    Laion-5b: An open large-scale dataset for training next generation image-text models,

    C. Schuhmann, R. Beaumont, R. Vencu, C. Gordon, R. Wightman, M. Cherti, T. Coombes, A. Katta, C. Mullis, M. Wortsmanet al., “Laion-5b: An open large-scale dataset for training next generation image-text models,”Advances in Neural Information Processing Systems, vol. 35, pp. 25...

  10. [14]

    Microsoft coco: Common objects in context,

    T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick, “Microsoft coco: Common objects in context,” in Computer Vision–ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part V 13. Springer, 201...

  11. [15]

    Semantic understanding of scenes through the ade20k dataset,

    B. Zhou, H. Zhao, X. Puig, T. Xiao, S. Fidler, A. Barriuso, and A. Torralba, “Semantic understanding of scenes through the ade20k dataset,”International Journal of Computer Vision, vol. 127, pp. 302–321, 2019

  12. [16]

    Wildavatar: Web-scale in-the- wild video dataset for 3d avatar creation,

    Z. Huang, S. Hu, G. Wang, T. Liu, Y. Zang, Z. Cao, W. Li, and Z. Liu, “Wildavatar: Web-scale in-the- wild video dataset for 3d avatar creation,” 2024. [Online]. Available: https://arxiv.org/abs/2407.02165

  13. [17]

    Renderme-360: A large digital asset library and benchmarks towards high-fidelity head avatars,

    D. Pan, L. Zhuo, J. Piao, H. Luo, W. Cheng, Y. Wang, S. Fan, S. Liu, L. Yang, B. Dai, Z. Liu, C. C. Loy, C. Qian, W. Wu, D. Lin, and K.-Y. Lin, “Renderme-360: A large digital asset library and benchmarks towards high-fidelity head avatars,”Advances in Neural Information Proces...

  14. [19]

    Beat: A large- scale semantic and emotional multi-modal dataset for conversational gestures synthesis,

    H. Liu, Z. Zhu, N. Iwamoto, Y. Peng, Z. Li, Y. Zhou, E. Bozkurt, and B. Zheng, “Beat: A large- scale semantic and emotional multi-modal dataset for conversational gestures synthesis,” inEuropean conference on computer vision. Springer, 2022, pp. 612–630

  15. [20]

    Expressgesture: Expressive gesture generation from speech through database matching,

    Y. Ferstl, M. Neff, and R. McDonnell, “Expressgesture: Expressive gesture generation from speech through database matching,”Computer Animation and Virtual Worlds, vol. 32, 05 2021

  16. [21]

    Motiongpt: Finetuned llms are general-purpose motion generators,

    Y. Zhang, D. Huang, B. Liu, S. Tang, Y. Lu, L. Chen, L. Bai, Q. Chu, N. Yu, and W. Ouyang, “Motiongpt: Finetuned llms are general-purpose motion generators,” 2024. [Online]. Available: https://arxiv.org/abs/2306.10900

  17. [22]

    Motiondiffuse: Text-driven human motion generation with diffusion model,

    M. Zhang, Z. Cai, L. Pan, F. Hong, X. Guo, L. Yang, and Z. Liu, “Motiondiffuse: Text-driven human motion generation with diffusion model,”arXiv preprint arXiv:2208.15001, 2022

  18. [23]

    Dual-generator face reenactment,

    G.-S. Hsu, C.-H. Tsai, and H.-Y. Wu, “Dual-generator face reenactment,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2022, pp. 642–650

  19. [24]

    Designing one unified framework for high-fidelity face reenactment and swapping,

    C. Xu, J. Zhang, Y. Han, G. Tian, X. Zeng, Y. Tai, Y. Wang, C. Wang, and Y. Liu, “Designing one unified framework for high-fidelity face reenactment and swapping,” inEuropean Conference on Computer Vision, 2022. [Online]. Available: https://api.semanticscholar.org/CorpusID:253270179

  20. [25]

    Finding directions in gan’s latent space for neural face reenactment,

    S. Bounareli, V. Argyriou, and G. Tzimiropoulos, “Finding directions in gan’s latent space for neural face reenactment,” CoRR, vol. abs/2202.00046, 2022. [Online]. Available: https://arxiv.org/abs/2202.00046

  21. [26]

    Face editing based on facial recognition features,

    X. Ning, S. Xu, F. Nan, Q. Zeng, C. Wang, W. Cai, W. Li, and Y. Jiang, “Face editing based on facial recognition features,”IEEE Transactions on Cognitive and Developmental Systems, vol. 15, no. 2, pp. 774–783, 2023. 51

  22. [27]

    Hiface: High-fidelity 3d face reconstruction by learning static and dynamic details,

    Z. Chai, T. Zhang, T. He, X. Tan, T. Baltrusaitis, H. Wu, R. Li, S. Zhao, C. Yuan, and J. Bian, “Hiface: High-fidelity 3d face reconstruction by learning static and dynamic details,” inProceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), October 2023...

  23. [28]

    Controllable 3d generative adversarial face model via disentangling shape and appearance,

    F. Taherkhani, A. Rai, Q. Gao, S. Srivastava, X. Chen, F. de la Torre, S. Song, A. Prakash, and D. Kim, “Controllable 3d generative adversarial face model via disentangling shape and appearance,” in Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Visi...

  24. [29]

    Towards realistic generative 3d face models,

    A. Rai, H. Gupta, A. Pandey, F. V. Carrasco, S. J. Takagi, A. Aubel, D. Kim, A. Prakash, and F. De la Torre, “Towards realistic generative 3d face models,”arXiv preprint arXiv:2304.12483, 2023

  25. [30]

    Gsmoothface: Generalized smooth talking face generation via fine grained 3d face guidance,

    H. Zhang, Z. Yuan, C. Zheng, X. Yan, B. Wang, G. Li, S. Wu, S. Cui, and Z. Li, “Gsmoothface: Generalized smooth talking face generation via fine grained 3d face guidance,” ArXiv, vol. abs/2312.07385, 2023. [Online]. Available: https://api.semanticscholar.org/CorpusID:266174481

  26. [32]

    Adaptive nonlinear latent transformation for conditional face editing,

    Z. Huang, S. Ma, J. Zhang, and H. Shan, “Adaptive nonlinear latent transformation for conditional face editing,” inICCV, 2023

  27. [33]

    Stylet2i: Toward compositional and high-fidelity text-to-image synthesis,

    Z. Li, M. R. Min, K. Li, and C. Xu, “Stylet2i: Toward compositional and high-fidelity text-to-image synthesis,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022

  28. [35]

    Guidedstyle: Attribute knowledge guided style manipulation for semantic face editing,

    X. Hou, X. Zhang, H. Liang, L. Shen, Z. Lai, and J. Wan, “Guidedstyle: Attribute knowledge guided style manipulation for semantic face editing,”Neural Networks, vol. 145, pp. 209–220, 2022. [Online]. Available: https://www.sciencedirect.com/science/article/pii/S0893608021004081

  29. [36]

    Anyface: Free-style text-to-face synthesis and manipulation,

    J. Sun, Q. Deng, Q. Li, M. Sun, M. Ren, and Z. Sun, “Anyface: Free-style text-to-face synthesis and manipulation,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2022, pp. 18687–18696

  30. [37]

    Joint audio-text model for expressive speech-driven 3d facial animation,

    Y. Fan, Z. Lin, J. Saito, W. Wang, and T. Komura, “Joint audio-text model for expressive speech-driven 3d facial animation,”Proceedings of the ACM on Computer Graphics and Interactive Techniques, vol. 5, pp. 1 – 15, 2021. [Online]. Available: https://api.semanticscholar.org/Co...

  31. [38]

    Capture, learning, and synthesis of 3d speaking styles,

    D. Cudeiro, T. Bolkart, C. Laidlaw, A. Ranjan, and M. J. Black, “Capture, learning, and synthesis of 3d speaking styles,”2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 10093–10103, 2019. [Online]. Available: https://api.semanticscholar.org/Corp...

  32. [39]

    Meshtalk: 3d face animation from speech using cross-modality disentanglement,

    A. Richard, M. Zollhöfer, Y. Wen, F. de la Torre, and Y. Sheikh, “Meshtalk: 3d face animation from speech using cross-modality disentanglement,” inProceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), October 2021, pp. 1173–1182

  33. [40]

    Cstalk: Correlation supervised speech-driven 3d emotional facial animation generation,

    X. Liang, W. Zhuang, T. Wang, G. Geng, G. Geng, H. Xia, and S. Xia, “Cstalk: Correlation supervised speech-driven 3d emotional facial animation generation,” 2024. [Online]. Available: https://arxiv.org/abs/2404.18604

  34. [41]

    Expclip: Bridging text and facial expressions via semantic alignment,

    Y. Zhong, H. Wei, P. Yang, and Z. Wang, “Expclip: Bridging text and facial expressions via semantic alignment,” 2023. [Online]. Available: https://arxiv.org/abs/2308.14448

  35. [42]

    Personalized speech-driven expressive 3d facial animation synthesis with style control,

    E. Bozkurt, “Personalized speech-driven expressive 3d facial animation synthesis with style control,”

  36. [43]

    Faceformer: Speech-driven 3d facial animation with transformers,

    Y. Fan, Z. Lin, J. Saito, W. Wang, and T. Komura, “Faceformer: Speech-driven 3d facial animation with transformers,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022. 52

  37. [44]

    Adamesh: Personalized facial expressions and head poses for adaptive speech-driven 3d facial animation,

    L. Chen, W. Bao, S. Lei, B. Tang, Z. Wu, S. Kang, H. Huang, and H. Meng, “Adamesh: Personalized facial expressions and head poses for adaptive speech-driven 3d facial animation,” 2024. [Online]. Available: https://arxiv.org/abs/2310.07236

  38. [45]

    Geneface: Generalized and high-fidelity audio-driven 3d talking face synthesis,

    Z. Ye, Z. Jiang, Y. Ren, J. Liu, J. He, and Z. Zhao, “Geneface: Generalized and high-fidelity audio-driven 3d talking face synthesis,”arXiv preprint arXiv:2301.13430, 2023

  39. [46]

    Imitator: Personalized speech-driven 3d facial animation,

    B. Thambiraja, I. Habibie, S. Aliakbarian, D. Cosker, C. Theobalt, and J. Thies, “Imitator: Personalized speech-driven 3d facial animation,” inProceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), October 2023, pp. 20621–20631

  40. [47]

    Data-driven expressive 3d facial animation synthesis for digital humans,

    K. I. Haque, “Data-driven expressive 3d facial animation synthesis for digital humans,” inSIGGRAPH Asia 2023 Doctoral Consortium, ser. SA ’23. New York, NY, USA: Association for Computing Machinery, 2023. [Online]. Available: https://doi.org/10.1145/3623053.3623369

  41. [48]

    Facexhubert: Text-less speech-driven e(x)pressive 3d facial animation synthesis using self-supervised speech representation learning,

    K. I. Haque and Z. Yumak, “Facexhubert: Text-less speech-driven e(x)pressive 3d facial animation synthesis using self-supervised speech representation learning,” inINTERNATIONAL CONFERENCE ON MULTIMODAL INTERACTION (ICMI ’23). New York, NY, USA: ACM, 2023. [Online]. Available:...

  42. [49]

    Facediffuser: Speech-driven 3d facial animation synthesis using diffusion,

    S. Stan, K. I. Haque, and Z. Yumak, “Facediffuser: Speech-driven 3d facial animation synthesis using diffusion,” inProceedings of the 16th ACM SIGGRAPH Conference on Motion, Interaction and Games, ser. MIG ’23. New York, NY, USA: Association for Computing Machinery, 2023. [Onl...

  43. [50]

    Diffusestylegesture: Stylized audio-driven co-speech gesture generation with diffusion models,

    S. Yang, Z. Wu, M. Li, Z. Zhang, L. Hao, W. Bao, M. Cheng, and L. Xiao, “Diffusestylegesture: Stylized audio-driven co-speech gesture generation with diffusion models,” in Proceedings of the Thirty-Second International Joint Conference on Artificial Intelligence, IJCAI-23. Int...

  44. [51]

    Neural face rigging for animating and retargeting facial meshes in the wild,

    D. Qin, J. Saito, N. Aigerman, T. Groueix, and T. Komura, “Neural face rigging for animating and retargeting facial meshes in the wild,” inACM SIGGRAPH 2023 Conference Proceedings, ser. SIGGRAPH ’23. New York, NY, USA: Association for Computing Machinery, 2023. [Online]. Avail...

  45. [52]

    Magicpose: Realistic human poses and facial expressions retargeting with identity-aware diffusion,

    D. Chang, Y. Shi, Q. Gao, J. Fu, H. Xu, G. Song, Q. Yan, Y. Zhu, X. Yang, and M. Soleymani, “Magicpose: Realistic human poses and facial expressions retargeting with identity-aware diffusion,”

  46. [53]

    Dreampose: Fashion image-to- video synthesis via stable diffusion,

    J. Karras, A. Holynski, T.-C. Wang, and I. Kemelmacher-Shlizerman, “Dreampose: Fashion image-to- video synthesis via stable diffusion,” 2023. [Online]. Available: https://arxiv.org/abs/2304.06025

  47. [54]

    Disco: Disentangled control for realistic human dance generation,

    T. Wang, L. Li, K. Lin, Y. Zhai, C.-C. Lin, Z. Yang, H. Zhang, Z. Liu, and L. Wang, “Disco: Disentangled control for realistic human dance generation,”arXiv preprint arXiv:2307.00040, 2023

  48. [55]

    Generating holistic 3d human motion from speech,

    H. Yi, H. Liang, Y. Liu, Q. Cao, Y. Wen, T. Bolkart, D. Tao, and M. J. Black, “Generating holistic 3d human motion from speech,”2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 469–480, 2022. [Online]. Available: https://api.semanticscholar.org/C...

  49. [57]

    Photorealistic text-to-image diffusion models with deep language understanding,

    C. Saharia, W. Chan, S. Saxena, L. Li, J. Whang, E. L. Denton, K. Ghasemipour, R. Gontijo Lopes, B. Karagol Ayan, T. Salimanset al., “Photorealistic text-to-image diffusion models with deep language understanding,” Advances in neural information processing systems, vol. 35, pp...

  50. [58]

    Sdxl: Improving latent diffusion models for high-resolution image synthesis,

    D. Podell, Z. English, K. Lacey, A. Blattmann, T. Dockhorn, J. Müller, J. Penna, and R. Rombach, “Sdxl: Improving latent diffusion models for high-resolution image synthesis,”arXiv preprint arXiv:2307.01952, 2023. 53

  51. [59]

    High-resolution image synthesis with latent diffusion models,

    R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer, “High-resolution image synthesis with latent diffusion models,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 10684–10695

  52. [60]

    Svdiff: Compact parameter space for diffusion fine-tuning,

    L. Han, Y. Li, H. Zhang, P. Milanfar, D. Metaxas, and F. Yang, “Svdiff: Compact parameter space for diffusion fine-tuning,” inProceedings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 7323–7334

  53. [61]

    Uncovering the disentanglement capability in text-to-image diffusion models,

    Q. Wu, Y. Liu, H. Zhao, A. Kale, T. Bui, T. Yu, Z. Lin, Y. Zhang, and S. Chang, “Uncovering the disentanglement capability in text-to-image diffusion models,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2023, pp. 1900–1910

  54. [62]

    Instructpix2pix: Learning to follow image editing instructions,

    T. Brooks, A. Holynski, and A. A. Efros, “Instructpix2pix: Learning to follow image editing instructions,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 18392–18402

  55. [63]

    Sine: Single image editing with text-to-image diffusion models,

    Z. Zhang, L. Han, A. Ghosh, D. N. Metaxas, and J. Ren, “Sine: Single image editing with text-to-image diffusion models,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 6027–6037

  56. [64]

    Null-text inversion for editing real images using guided diffusion models,

    R. Mokady, A. Hertz, K. Aberman, Y. Pritch, and D. Cohen-Or, “Null-text inversion for editing real images using guided diffusion models,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 6038–6047

  57. [65]

    Imagic: Text-based real image editing with diffusion models,

    B. Kawar, S. Zada, O. Lang, O. Tov, H. Chang, T. Dekel, I. Mosseri, and M. Irani, “Imagic: Text-based real image editing with diffusion models,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 6007–6017

  58. [66]

    Unified concept editing in diffusion models,

    R. Gandikota, H. Orgad, Y. Belinkov, J. Materzyńska, and D. Bau, “Unified concept editing in diffusion models,” inProceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, 2024, pp. 5111–5120

  59. [67]

    Visual chatgpt: Talking, drawing and editing with visual foundation models,

    C. Wu, S. Yin, W. Qi, X. Wang, Z. Tang, and N. Duan, “Visual chatgpt: Talking, drawing and editing with visual foundation models,”arXiv preprint arXiv:2303.04671, 2023

  60. [68]

    Kosmos-g: Generating images in context with multimodal large language models,

    X. Pan, L. Dong, S. Huang, Z. Peng, W. Chen, and F. Wei, “Kosmos-g: Generating images in context with multimodal large language models,”arXiv preprint arXiv:2310.02992, 2023

  61. [69]

    Mm-react: Prompting chatgpt for multimodal reasoning and action,

    Z. Yang, L. Li, J. Wang, K. Lin, E. Azarnasab, F. Ahmed, Z. Liu, C. Liu, M. Zeng, and L. Wang, “Mm-react: Prompting chatgpt for multimodal reasoning and action,”arXiv preprint arXiv:2303.11381, 2023

  62. [70]

    Avatarclip: zero-shot text-driven generation and animation of 3d avatars,

    F. Hong, M. Zhang, L. Pan, Z. Cai, L. Yang, and Z. Liu, “Avatarclip: zero-shot text-driven generation and animation of 3d avatars,”ACM Trans. Graph., vol. 41, no. 4, Jul. 2022. [Online]. Available: https://doi.org/10.1145/3528223.3530094

  63. [71]

    Neus: learning neural implicit surfaces by volume rendering for multi-view reconstruction,

    P. Wang, L. Liu, Y. Liu, C. Theobalt, T. Komura, and W. Wang, “Neus: learning neural implicit surfaces by volume rendering for multi-view reconstruction,” inProceedings of the 35th International Conference on Neural Information Processing Systems, ser. NIPS ’21. Red Hook, NY, ...

  64. [72]

    Zero-shot text-guided object generation with dream fields,

    A. Jain, B. Mildenhall, J. T. Barron, P. Abbeel, and B. Poole, “Zero-shot text-guided object generation with dream fields,” 2022

  65. [73]

    Text2mesh: Text-driven neural stylization for meshes,

    O. Michel, R. Bar-On, R. Liu, S. Benaim, and R. Hanocka, “Text2mesh: Text-driven neural stylization for meshes,”arXiv preprint arXiv:2112.03221, 2021

  66. [74]

    Pifu: Pixel-aligned implicit function for high-resolution clothed human digitization,

    S. Saito, , Z. Huang, R. Natsume, S. Morishima, A. Kanazawa, and H. Li, “Pifu: Pixel-aligned implicit function for high-resolution clothed human digitization,”arXiv preprint arXiv:1905.05172, 2019

  67. [75]

    Pifuhd: Multi-level pixel-aligned implicit function for high-resolution 3d human digitization,

    S. Saito, T. Simon, J. Saragih, and H. Joo, “Pifuhd: Multi-level pixel-aligned implicit function for high-resolution 3d human digitization,” inProceedings of the IEEE Conference on Computer Vision and Pattern Recognition, June 2020

  68. [76]

    Arch: Animatable reconstruction of clothed humans,

    Z. Huang, Y. Xu, C. Lassner, H. Li, and T. Tung, “Arch: Animatable reconstruction of clothed humans,” 2020. 54

  69. [77]

    Arch++: Animation-ready clothed human reconstruction revisited,

    T. He, Y. Xu, S. Saito, S. Soatto, and T. Tung, “Arch++: Animation-ready clothed human reconstruction revisited,” 2022. [Online]. Available: https://arxiv.org/abs/2108.07845

  70. [78]

    Pamir: Parametric model-conditioned implicit representation for image-based human reconstruction,

    Z. Zheng, T. Yu, Y. Liu, and Q. Dai, “Pamir: Parametric model-conditioned implicit representation for image-based human reconstruction,” 2020. [Online]. Available: https://arxiv.org/abs/2007.03858

  71. [79]

    Tada! text to animatable digital avatars,

    T. Liao, H. Yi, Y. Xiu, J. Tang, Y. Huang, J. Thies, and M. J. Black, “Tada! text to animatable digital avatars,” 2023. [Online]. Available: https://arxiv.org/abs/2308.10899

  72. [80]

    Getavatar: Generative textured meshes for animatable human avatars,

    X. Zhang, J. Zhang, C. Rohan, H. Xu, G. Song, Y. Yang, and J. Feng, “Getavatar: Generative textured meshes for animatable human avatars,” inICCV, 2023

  73. [81]

    Rodinhd: High-fidelity 3d avatar generation with diffusion models,

    B. Zhang, Y. Cheng, C. Wang, T. Zhang, J. Yang, Y. Tang, F. Zhao, D. Chen, and B. Guo, “Rodinhd: High-fidelity 3d avatar generation with diffusion models,”arXiv preprint arXiv:2407.06938, 2024

  74. [82]

    Humannerf: Free-viewpoint rendering of moving people from monocular video,

    C.-Y. Weng, B. Curless, P. P. Srinivasan, J. T. Barron, and I. Kemelmacher-Shlizerman, “Humannerf: Free-viewpoint rendering of moving people from monocular video,” inProceedings of the IEEE/CVF conference on computer vision and pattern Recognition, 2022, pp. 16210–16220

  75. [83]

    Vid2avatar: 3d avatar reconstruction from videos in the wild via self-supervised scene decomposition,

    C. Guo, T. Jiang, X. Chen, J. Song, and O. Hilliges, “Vid2avatar: 3d avatar reconstruction from videos in the wild via self-supervised scene decomposition,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 12858–12868

  76. [84]

    Dreamhuman: Animatable 3d avatars from text,

    N. Kolotouros, T. Alldieck, A. Zanfir, E. G. Bazavan, M. Fieraru, and C. Sminchisescu, “Dreamhuman: Animatable 3d avatars from text,” 2023

  77. [85]

    Text-conditional contextualized avatars for zero-shot personalization,

    S. Azadi, T. Hayes, A. Shah, G. Pang, D. Parikh, and S. Gupta, “Text-conditional contextualized avatars for zero-shot personalization,” 2023. [Online]. Available: https://arxiv.org/abs/2304.07410

  78. [86]

    Articulated 3d head avatar generation using text-to-image diffusion models,

    A. W. Bergman, W. Yifan, and G. Wetzstein, “Articulated 3d head avatar generation using text-to-image diffusion models,” ArXiv, vol. abs/2307.04859, 2023. [Online]. Available: https://api.semanticscholar.org/CorpusID:259766646

  79. [87]

    Make-your-anchor: Adiffusion-based 2d avatar generation framework,

    Z.Huang, F.Tang, Y.Zhang, X.Cun, J.Cao, J.Li, andT.-Y.Lee, “Make-your-anchor: Adiffusion-based 2d avatar generation framework,”arXiv preprint arXiv:2403.16510, 2024

  80. [88]

    Dreamavatar: Text-and-shape guided 3d human avatar generation via diffusion models,

    Y. Cao, Y.-P. Cao, K. Han, Y. Shan, and K.-Y. K. Wong, “Dreamavatar: Text-and-shape guided 3d human avatar generation via diffusion models,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 958–968

  81. [89]

    Dreamwaltz: Make a scene with complex 3d animatable avatars,

    Y. Huang, J. Wang, A. Zeng, H. Cao, X. Qi, Y. Shi, Z.-J. Zha, and L. Zhang, “Dreamwaltz: Make a scene with complex 3d animatable avatars,” 2023

  82. [90]

    Evaluating emotive character animations created with procedural animation,

    Y.-H. Lin, C.-Y. Liu, H.-W. Lee, S.-L. Huang, and T.-Y. Li, “Evaluating emotive character animations created with procedural animation,” 09 2009, pp. 308–315

  83. [91]

    Practice and Theory of Blendshape Facial Models,

    J. P. Lewis, K. Anjyo, T. Rhee, M. Zhang, F. Pighin, and Z. Deng, “Practice and Theory of Blendshape Facial Models,” inEurographics 2014 - State of the Art Reports, S. Lefebvre and M. Spagnuolo, Eds. The Eurographics Association, 2014

  84. [92]

    Beat: the behavior expression animation toolkit,

    J. Cassell, H. H. Vilhjálmsson, and T. Bickmore, “Beat: the behavior expression animation toolkit,” in Proceedings of the 28th annual conference on Computer graphics and interactive techniques, 2001, pp. 477–486

  85. [93]

    Gesturegan for hand gesture-to-gesture translation in the wild,

    H. Tang, W. Wang, D. Xu, Y. Yan, and N. Sebe, “Gesturegan for hand gesture-to-gesture translation in the wild,” 2019. [Online]. Available: https://arxiv.org/abs/1808.04859

  86. [94]

    Learning individual styles of conversational gesture,

    S. Ginosar, A. Bar, G. Kohavi, C. Chan, A. Owens, and J. Malik, “Learning individual styles of conversational gesture,” inComputer Vision and Pattern Recognition (CVPR). IEEE, Jun. 2019

  87. [95]

    Taming diffusion models for audio-driven co-speech gesture generation,

    L. Zhu, X. Liu, X. Liu, R. Qian, Z. Liu, and L. Yu, “Taming diffusion models for audio-driven co-speech gesture generation,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 10544–10553

  88. [96]

    Gesturemaster: Graph-based speech-driven gesture generation,

    C. Zhou, T. Bian, and K. Chen, “Gesturemaster: Graph-based speech-driven gesture generation,” in Proceedings of the 2022 International Conference on Multimodal Interaction, ser. ICMI ’22. New York, NY, USA: Association for Computing Machinery, 2022, p. 764–770. [Online]. Avail...

  89. [97]

    Dim-gesture: Co-speech gesture generation with adaptive layer normalization mamba-2 framework,

    F. Zhang, N. Ji, F. Gao, B. Zhao, J. Wu, Y. Jiang, H. Du, Z. Ye, J. Zhu, W. Zhong, L. Yan, and X. Ma, “Dim-gesture: Co-speech gesture generation with adaptive layer normalization mamba-2 framework,”

  90. [98]

    AMUSE: Emotional speech-driven 3D body animation via disentangled latent diffusion,

    K. Chhatre, R. Daněček, N. Athanasiou, G. Becherini, C. Peters, M. J. Black, and T. Bolkart, “AMUSE: Emotional speech-driven 3D body animation via disentangled latent diffusion,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2...

  91. [99]

    Freetalker: Controllable speech and text-driven gesture generation based on diffusion models for enhanced speaker naturalness,

    S. Yang, Z. Xu, H. Xue, Y. Cheng, S. Huang, M. Gong, and Z. Wu, “Freetalker: Controllable speech and text-driven gesture generation based on diffusion models for enhanced speaker naturalness,” inICASSP 2024 - 2024 IEEE International Conference on Acoustics, Speech and Signal P...

  92. [100]

    Available: https://arxiv.org/abs/2408.00370

    [Online]. Available: https://arxiv.org/abs/2408.00370

  93. [102]

    Gesticulator: A framework for semantically-aware speech-driven gesture generation,

    T. Kucherenko, P. Jonell, S. van Waveren, G. E. Henter, S. Alexandersson, I. Leite, and H. Kjellström, “Gesticulator: A framework for semantically-aware speech-driven gesture generation,” inProceedings of the 2020 International Conference on Multimodal Interaction, ser. ICMI ’...

  94. [103]

    Diffugesture: Generating human gesture from two-person dialogue with diffusion models,

    W. Zhao, L. Hu, and S. Zhang, “Diffugesture: Generating human gesture from two-person dialogue with diffusion models,” inCompanion Publication of the 25th International Conference on Multimodal Interaction, ser. ICMI ’23 Companion. New York, NY, USA: Association for Computing ...

  95. [104]

    Style-controllable speech-driven gesture synthesis using normalising flows,

    S. Alexanderson, G. E. Henter, T. Kucherenko, and J. Beskow, “Style-controllable speech-driven gesture synthesis using normalising flows,”Computer Graphics Forum, vol. 39, no. 2, pp. 487–496, 2020. [Online]. Available: https://diglib.eg.org/handle/10.1111/cgf13946

  96. [105]

    Zeroeggs: Zero-shot example- based gesture generation from speech,

    S. Ghorbani, Y. Ferstl, D. Holden, N. F. Troje, and M.-A. Carbonneau, “Zeroeggs: Zero-shot example- based gesture generation from speech,” 2022. [Online]. Available: https://arxiv.org/abs/2209.07556

  97. [106]

    Vitpose: Simple vision transformer baselines for human pose estimation,

    Y. Xu, J. Zhang, Q. Zhang, and D. Tao, “Vitpose: Simple vision transformer baselines for human pose estimation,” 2022. [Online]. Available: https://arxiv.org/abs/2204.12484

  98. [107]

    Zs-mstm: Zero-shot style transfer for gesture animation driven by text and speech using adversarial disentanglement of multimodal style encoding,

    M. Fares, C. Pelachaud, and N. Obin, “Zs-mstm: Zero-shot style transfer for gesture animation driven by text and speech using adversarial disentanglement of multimodal style encoding,” 2023. [Online]. Available: https://arxiv.org/abs/2305.12887

  99. [108]

    Augmented co-speech gesture generation: Including form and meaning features to guide learning-based gesture synthesis,

    H. Voß and S. Kopp, “Augmented co-speech gesture generation: Including form and meaning features to guide learning-based gesture synthesis,” inProceedings of the 23rd ACM International Conference on Intelligent Virtual Agents, ser. IVA ’23. New York, NY, USA: Association for C...

  100. [109]

    Gesturediffuclip: Gesture diffusion model with clip latents,

    T. Ao, Z. Zhang, and L. Liu, “Gesturediffuclip: Gesture diffusion model with clip latents,” 2023. [Online]. Available: https://arxiv.org/abs/2303.14613

  101. [110]

    Cocogesture: Toward coherent co-speech 3d gesture generation in the wild,

    X. Qi, H. Zhang, Y. Wang, J. Pan, C. Liu, P. Li, X. Chi, M. Li, Q. Zhang, W. Xue, S. Zhang, W. Luo, Q. Liu, and Y. Guo, “Cocogesture: Toward coherent co-speech 3d gesture generation in the wild,” 2024. [Online]. Available: https://arxiv.org/abs/2405.16874

  102. [112]

    Available: https://doi.org/10.1145/3570945.3607337

    [Online]. Available: https://doi.org/10.1145/3570945.3607337

  103. [113]

    C2g2: Controllable co-speech gesture generation with latent diffusion model,

    L. Ji, P. Wei, Y. Ren, J. Liu, C. Zhang, and X. Yin, “C2g2: Controllable co-speech gesture generation with latent diffusion model,” 2023. [Online]. Available: https://arxiv.org/abs/2308.15016

  104. [114]

    Emage: Towards unified holistic co-speech gesture generation via expressive masked audio gesture modeling,

    H. Liu, Z. Zhu, G. Becherini, Y. Peng, M. Su, Y. Zhou, X. Zhe, N. Iwamoto, B. Zheng, and M. J. Black, “Emage: Towards unified holistic co-speech gesture generation via expressive masked audio gesture modeling,” 2024. [Online]. Available: https://arxiv.org/abs/2401.00374

  105. [115]

    Mambatalk: Efficient holistic gesture synthesis with selective state space models,

    Z. Xu, Y. Lin, H. Han, S. Yang, R. Li, Y. Zhang, and X. Li, “Mambatalk: Efficient holistic gesture synthesis with selective state space models,” 2025. [Online]. Available: https://arxiv.org/abs/2403.09471

  106. [117]

    Style transfer for co-speech gesture animation: A multi-speaker conditional-mixture approach,

    C. Ahuja, D. W. Lee, Y. I. Nakano, and L.-P. Morency, “Style transfer for co-speech gesture animation: A multi-speaker conditional-mixture approach,” 2020. [Online]. Available: https://arxiv.org/abs/2007.12553

  107. [118]

    Teach: Temporal action composition for 3d humans,

    N. Athanasiou, M. Petrovich, M. J. Black, and G. Varol, “Teach: Temporal action composition for 3d humans,” 2022. [Online]. Available: https://arxiv.org/abs/2209.04066

  108. [119]

    Generating diverse and natural 3d human motions from text,

    Guo, Chuan, Zou, Shihao, Zuo, Xinxin, Wang, Sen, Ji, Wei, Li, Xingyu, Cheng, and Li, “Generating diverse and natural 3d human motions from text,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2022, pp. 5152–5161

  109. [120]

    Action-conditioned 3d human motion synthesis with transformer VAE,

    M. Petrovich, M. J. Black, and G. Varol, “Action-conditioned 3d human motion synthesis with transformer VAE,” CoRR, vol. abs/2104.05670, 2021. [Online]. Available: https: //arxiv.org/abs/2104.05670

  110. [121]

    TEMOS: Generating diverse human motions from textual descriptions,

    ——, “TEMOS: Generating diverse human motions from textual descriptions,” inEuropean Conference on Computer Vision (ECCV), 2022

  111. [122]

    Diversemotion: Towards diverse human motion generation via discrete diffusion,

    Y. Lou, L. Zhu, Y. Wang, X. Wang, and Y. Yang, “Diversemotion: Towards diverse human motion generation via discrete diffusion,” 2023. [Online]. Available: https://arxiv.org/abs/2309.01372

  112. [123]

    Momask: Generative masked modeling of 3d human motions,

    C. Guo, Y. Mu, M. G. Javed, S. Wang, and L. Cheng, “Momask: Generative masked modeling of 3d human motions,” 2023. [Online]. Available: https://arxiv.org/abs/2312.00063

  113. [124]

    Tmr: Text-to-motion retrieval using contrastive 3d human motion synthesis,

    M. Petrovich, M. J. Black, and G. Varol, “Tmr: Text-to-motion retrieval using contrastive 3d human motion synthesis,” 2023. [Online]. Available: https://arxiv.org/abs/2305.00976

  114. [125]

    T2m-gpt: Generating human motion from textual descriptions with discrete representations,

    J. Zhang, Y. Zhang, X. Cun, S. Huang, Y. Zhang, H. Zhao, H. Lu, and X. Shen, “T2m-gpt: Generating human motion from textual descriptions with discrete representations,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023

  115. [126]

    Human motion diffusion model,

    G. Tevet, S. Raab, B. Gordon, Y. Shafir, D. Cohen-or, and A. H. Bermano, “Human motion diffusion model,” inThe Eleventh International Conference on Learning Representations, 2023. [Online]. Available: https://openreview.net/forum?id=SJ1kSyO2jwu

  116. [127]

    Make-an-animation: Large-scale text-conditional 3d human motion generation,

    S. Azadi, A. Shah, T. Hayes, D. Parikh, and S. Gupta, “Make-an-animation: Large-scale text-conditional 3d human motion generation,” 2023. [Online]. Available: https://arxiv.org/abs/2305.09662

  117. [128]

    T2lm: Long-term 3d human motion generation from multiple sentences,

    T. Lee, F. Baradel, T. Lucas, K. M. Lee, and G. Rogez, “T2lm: Long-term 3d human motion generation from multiple sentences,” 2024. [Online]. Available: https://arxiv.org/abs/2406.00636

  118. [129]

    Motion anything: Any to motion generation,

    Z. Zhang, Y. Wang, W. Mao, D. Li, R. Zhao, B. Wu, Z. Song, B. Zhuang, I. Reid, and R. Hartley, “Motion anything: Any to motion generation,”arXiv preprint arXiv:2503.06955, 2025

  119. [130]

    Dreamfusion: Text-to-3d using 2d diffusion,

    B. Poole, A. Jain, J. T. Barron, and B. Mildenhall, “Dreamfusion: Text-to-3d using 2d diffusion,”arXiv preprint arXiv:2209.14988, 2022. 57

  120. [131]

    Magic3d: High-resolution text-to-3d content creation,

    C.-H. Lin, J. Gao, L. Tang, T. Takikawa, X. Zeng, X. Huang, K. Kreis, S. Fidler, M.-Y. Liu, and T.-Y. Lin, “Magic3d: High-resolution text-to-3d content creation,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 300–309

  121. [132]

    Guided motion diffusion for controllable human motion synthesis,

    K. Karunratanakul, K. Preechakul, S. Suwajanakorn, and S. Tang, “Guided motion diffusion for controllable human motion synthesis,” inProceedings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 2151–2162

  122. [133]

    Omnicontrol: Control any joint at any time for human motion generation,

    Y. Xie, V. Jampani, L. Zhong, D. Sun, and H. Jiang, “Omnicontrol: Control any joint at any time for human motion generation,” inThe Twelfth International Conference on Learning Representations, 2024

  123. [134]

    Dream3d: Zero-shot text-to-3d synthesis using 3d shape prior and text-to-image diffusion models,

    J. Xu, X. Wang, W. Cheng, Y.-P. Cao, Y. Shan, X. Qie, and S. Gao, “Dream3d: Zero-shot text-to-3d synthesis using 3d shape prior and text-to-image diffusion models,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 20908–20918

  124. [135]

    Sdfusion: Multimodal 3d shape completion, reconstruction, and generation,

    Cheng, Yen-Chi, Lee, Hsin-Ying, Tulyakov, Sergey, Schwing, A. G, Gui, and Liang-Yan, “Sdfusion: Multimodal 3d shape completion, reconstruction, and generation,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 4456–4465

  125. [136]

    Dreamtime: An improved optimization strategy for diffusion-guided 3d generation,

    Y. Huang, J. Wang, Y. Shi, B. Tang, X. Qi, and L. Zhang, “Dreamtime: An improved optimization strategy for diffusion-guided 3d generation,” 2024. [Online]. Available: https://arxiv.org/abs/2306.12422

  126. [137]

    Hd-fusion: Detailed text-to-3d generation leveraging multiple noise estimation,

    J. Wu, X. Gao, X. Liu, Z. Shen, C. Zhao, H. Feng, J. Liu, and E. Ding, “Hd-fusion: Detailed text-to-3d generation leveraging multiple noise estimation,” 2023. [Online]. Available: https://arxiv.org/abs/2307.16183

  127. [138]

    It3d: Improved text-to-3d generation with explicit view synthesis,

    Y. Chen, C. Zhang, X. Yang, Z. Cai, G. Yu, L. Yang, and G. Lin, “It3d: Improved text-to-3d generation with explicit view synthesis,” inProceedings of the AAAI Conference on Artificial Intelligence, vol. 38, no. 2, 2024, pp. 1237–1244

  128. [139]

    Fantasia3d: Disentangling geometry and appearance for high-quality text-to-3d content creation,

    R. Chen, Y. Chen, N. Jiao, and K. Jia, “Fantasia3d: Disentangling geometry and appearance for high-quality text-to-3d content creation,” inProceedings of the IEEE/CVF international conference on computer vision, 2023, pp. 22246–22256

  129. [140]

    Att3d: Amortized text-to-3d object synthesis,

    J. Lorraine, K. Xie, X. Zeng, C.-H. Lin, T. Takikawa, N. Sharp, T.-Y. Lin, M.-Y. Liu, S. Fidler, and J. Lucas, “Att3d: Amortized text-to-3d object synthesis,”The International Conference on Computer Vision (ICCV), 2023

  130. [141]

    Latte3d: Large-scale amortized text-to-enhanced3d synthesis,

    K. Xie, J. Lorraine, T. Cao, J. Gao, J. Lucas, A. Torralba, S. Fidler, and X. Zeng, “Latte3d: Large-scale amortized text-to-enhanced3d synthesis,”The 18th European Conference on Computer Vision (ECCV), 2024

  131. [142]

    Meta 3d assetgen: Text-to-mesh generation with high-quality geometry, texture, and pbr materials,

    Y. Siddiqui, T. Monnier, F. Kokkinos, M. Kariya, Y. Kleiman, E. Garreau, O. Gafni, N. Neverova, A. Vedaldi, R. Shapovalov, and D. Novotny, “Meta 3d assetgen: Text-to-mesh generation with high-quality geometry, texture, and pbr materials,” 2024. [Online]. Available: https://arx...

  132. [143]

    Ipdreamer: Appearance-controllable 3d object generation with image prompts,

    B. Zeng, S. Li, Y. Feng, L. Yang, H. Li, S. Gao, J. Liu, C. He, W. Zhang, J. Liu, B. Zhang, and S. Yan, “Ipdreamer: Appearance-controllable 3d object generation with image prompts,”arXiv preprint arXiv:2310.05375, 2023

  133. [144]

    Vox-e: Text-guided voxel editing of 3d objects,

    E. Sella, G. Fiebelman, P. Hedman, and H. Averbuch-Elor, “Vox-e: Text-guided voxel editing of 3d objects,” inProceedings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 430–440

  134. [145]

    Lgm: Large multi-view gaussian model for high-resolution 3d content creation,

    J. Tang, Z. Chen, X. Chen, T. Wang, G. Zeng, and Z. Liu, “Lgm: Large multi-view gaussian model for high-resolution 3d content creation,” inComputer Vision – ECCV 2024: 18th European Conference, Milan, Italy, September 29–October 4, 2024, Proceedings, Part IV. Berlin, Heidelber...

  135. [146]

    Texture: Text-guided texturing of 3d shapes,

    E. Richardson, G. Metzer, Y. Alaluf, R. Giryes, and D. Cohen-Or, “Texture: Text-guided texturing of 3d shapes,” 2023. [Online]. Available: https://arxiv.org/abs/2302.01721 58

  136. [147]

    Paint-it: Text-to-texture synthesis via deep convolutional texture map optimization and physically-based rendering,

    K. Youwang, T.-H. Oh, and G. Pons-Moll, “Paint-it: Text-to-texture synthesis via deep convolutional texture map optimization and physically-based rendering,” 2024. [Online]. Available: https://arxiv.org/abs/2312.11360

  137. [148]

    TAPS3D: Text-Guided 3D Textured Shape Generation from Pseudo Supervision ,

    J. Wei, H. Wang, J. Feng, G. Lin, and K.-H. Yap, “ TAPS3D: Text-Guided 3D Textured Shape Generation from Pseudo Supervision ,” in2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Los Alamitos, CA, USA: IEEE Computer Society, Jun. 2023, pp. 16805–16815...

  138. [149]

    Text2tex: Text-driven texture synthesis via diffusion models,

    D. Z. Chen, Y. Siddiqui, H.-Y. Lee, S. Tulyakov, and M. Nießner, “Text2tex: Text-driven texture synthesis via diffusion models,” 2023. [Online]. Available: https://arxiv.org/abs/2303.11396

  139. [150]

    Genesistex: Adapting image denoising diffusion to texture space,

    C. Gao, B. Jiang, X. Li, Y. Zhang, and Q. Yu, “Genesistex: Adapting image denoising diffusion to texture space,” 2024. [Online]. Available: https://arxiv.org/abs/2403.17782

  140. [151]

    Texture generation on 3d meshes with point-uv diffusion,

    X. Yu, P. Dai, W. Li, L. Ma, Z. Liu, and X. Qi, “Texture generation on 3d meshes with point-uv diffusion,” 2023. [Online]. Available: https://arxiv.org/abs/2308.10490

  141. [152]

    Texpainter: Generative mesh texturing with multi-view consistency,

    H. Zhang, Z. Pan, C. Zhang, L. Zhu, and X. Gao, “Texpainter: Generative mesh texturing with multi-view consistency,” 2024. [Online]. Available: https://arxiv.org/abs/2406.18539

  142. [153]

    Texfusion: Synthesizing 3d textures with text-guided image diffusion models,

    T. Cao, K. Kreis, S. Fidler, N. Sharp, and K. Yin, “Texfusion: Synthesizing 3d textures with text-guided image diffusion models,” 2023. [Online]. Available: https://arxiv.org/abs/2310.13772

  143. [154]

    Vcd-texture: Variance alignment based 3d-2d co-denoising for text-guided texturing,

    S. Liu, C. Yu, C. Cao, W. Qian, and F. Wang, “Vcd-texture: Variance alignment based 3d-2d co-denoising for text-guided texturing,” 2024. [Online]. Available: https://arxiv.org/abs/2407.04461

  144. [155]

    Generative Adversarial Networks for Face Generation: A Survey,

    A. Kammoun, R. Slama, H. Tabia, T. Ouni, and M. Abid, “Generative Adversarial Networks for Face Generation: A Survey,”ACM Computing Surveys, vol. 55, no. 5, pp. 1–37, 2022

  145. [156]

    Consistency2: Consistent and fast 3d painting with latent consistency models,

    T. Wang, A. Obukhov, and K. Schindler, “Consistency2: Consistent and fast 3d painting with latent consistency models,” 2024. [Online]. Available: https://arxiv.org/abs/2406.11202

  146. [157]

    Meta 3d texturegen: Fast and consistent texture generation for 3d objects,

    R. Bensadoun, Y. Kleiman, I. Azuri, O. Harosh, A. Vedaldi, N. Neverova, and O. Gafni, “Meta 3d texturegen: Fast and consistent texture generation for 3d objects,” 2024. [Online]. Available: https://arxiv.org/abs/2407.02430

  147. [158]

    A Comprehensive Review of Data-Driven Co-Speech Gesture Generation,

    S. Nyatsanga, T. Kucherenko, C. Ahuja, G. E. Henter, and M. Neff, “A Comprehensive Review of Data-Driven Co-Speech Gesture Generation,”Computer Graphics Forum, vol. 42, no. 2, pp. 569–596, 2023

  148. [159]

    Human Motion Generation: A Survey,

    W. Zhu, X. Ma, D. Ro, H. Ci, J. Zhang, J. Shi, F. Gao, Q. Tian, and Y. Wang, “Human Motion Generation: A Survey,”IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 46, no. 11, pp. 2430–2449, 2024

  149. [160]

    Talking human face generation: A survey,

    M. A. Toshpulatov, W. Lee, and S. Lee, “Talking human face generation: A survey,”Expert Systems with Applications, vol. 219, p. 119678, 2023

  150. [161]

    A survey on deep learning based reenactment methods for deepfake applications,

    R. Dhanyalakshmi, C.-I. Popirlan, and D. J. Hemanth, “A survey on deep learning based reenactment methods for deepfake applications,”IET Image Processing, 2024

  151. [162]

    A Comprehensive Survey of Image Generation Models Based on Deep Learning,

    J. Li, C. Zhang, W. Zhu, and Y. Ren, “A Comprehensive Survey of Image Generation Models Based on Deep Learning,”Annals of Data Science, vol. 12, no. 1, pp. 141–170, 2024

  152. [163]

    Presentation and validation of the radboud faces database,

    O. Langner, R. Dotsch, G. Bijlstra, D. H. J. Wigboldus, S. T. Hawk, and A. van Knippenberg, “Presentation and validation of the radboud faces database,”Cognition and Emotion, vol. 24, no. 8, pp. 1377–1388, 2010. [Online]. Available: https://doi.org/10.1080/02699930903485076

  153. [164]

    A Survey on 3D Human Avatar Modeling – From Reconstruction to Generation,

    R. Wang, Y. Cao, K. Han, and K.-Y. K. Wong, “A Survey on 3D Human Avatar Modeling – From Reconstruction to Generation,”arXiv preprint arXiv:2406.04253, 2024

  154. [165]

    A survey of deep learning-based 3D shape generation,

    Q.-C. Xu, T.-J. Mu, and Y.-L. Yang, “A survey of deep learning-based 3D shape generation,”Compu- tational Visual Media, vol. 9, no. 3, pp. 407–442, 2023

  155. [166]

    Progressive growing of gans for improved quality, stability, and variation,

    T. Karras, T. Aila, S. Laine, and J. Lehtinen, “Progressive growing of gans for improved quality, stability, and variation,”International Conference on Learning Representations, Feb. 2018. [Online]. Available: https://dblp.uni-trier.de/db/journals/corr/corr1710.html#abs-1710-10196

  156. [167]

    Faceforensics: A large-scale video dataset for forgery detection in human faces,

    A. Rössler, D. Cozzolino, L. Verdoliva, C. Riess, J. Thies, and M. Nießner, “Faceforensics: A large-scale video dataset for forgery detection in human faces,”arXiv, 2018. 59

  157. [168]

    Multi-pie,

    R. Gross, I. Matthews, J. Cohn, T. Kanade, and S. Baker, “Multi-pie,”Image and Vision Computing, vol. 28, no. 5, p. 807–813, May 2010. [Online]. Available: https://doi.org/10.1016/j.imavis.2009.08.002

  158. [169]

    Voxceleb: a large-scale speaker identification dataset,

    A. Nagrani, J. S. Chung, and A. Zisserman, “Voxceleb: a large-scale speaker identification dataset,” in INTERSPEECH, 2017

  159. [170]

    Offline deformable face tracking in arbitrary videos,

    G. G. Chrysos, E. Antonakos, S. Zafeiriou, and P. Snape, “Offline deformable face tracking in arbitrary videos,” in2015 IEEE International Conference on Computer Vision Workshop (ICCVW), 2015, pp. 954–962

  160. [171]

    Affectnet: a database for facial expression, valence, and arousal computing in the wild,

    A. Mollahosseini, B. Hasani, and M. H. Mahoor, “Affectnet: a database for facial expression, valence, and arousal computing in the wild,”IEEE Transactions on Affective Computing, vol. 10, no. 1, p. 18–31, Jan. 2019. [Online]. Available: https://doi.org/10.1109/taffc.2017.2740923

  161. [172]

    The first facial land- mark tracking in-the-wild challenge: Benchmark and results,

    J. Shen, S. Zafeiriou, G. G. Chrysos, J. Kossaifi, G. Tzimiropoulos, and M. Pantic, “The first facial land- mark tracking in-the-wild challenge: Benchmark and results,” in2015 IEEE International Conference on Computer Vision Workshop (ICCVW), 2015, pp. 1003–1011

  162. [173]

    Project-out cascaded regression with an application to face alignment,

    G. Tzimiropoulos, “Project-out cascaded regression with an application to face alignment,” in2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2015, pp. 3659–3667

  163. [174]

    How far are we from solving the 2d & 3d face alignment problem? (and a dataset of 230,000 3d facial landmarks),

    A. Bulat and G. Tzimiropoulos, “How far are we from solving the 2d & 3d face alignment problem? (and a dataset of 230,000 3d facial landmarks),” inInternational Conference on Computer Vision, 2017

  164. [175]

    Luvli face alignment: Estimating landmarks’ location, uncertainty, and visibility likelihood,

    A. Kumar, T. K. Marks, W. Mou, Y. Wang, M. Jones, A. Cherian, T. Koike-Akino, X. Liu, and C. Feng, “Luvli face alignment: Estimating landmarks’ location, uncertainty, and visibility likelihood,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020

  165. [176]

    Caltech-ucsd birds-200-2011,

    C. Wah, S. Branson, P. Welinder, P. Perona, and S. Belongie, “Caltech-ucsd birds-200-2011,” California Institute of Technology, Tech. Rep. CNS-TR-2011-001, 2011

  166. [177]

    Talk-to-edit: Fine-grained facial editing via dialog,

    Y. Jiang, Z. Huang, X. Pan, C. C. Loy, and Z. Liu, “Talk-to-edit: Fine-grained facial editing via dialog,” in Proceedings of International Conference on Computer Vision (ICCV), 2021

  167. [178]

    Analyzing and improving the image quality of StyleGAN,

    T. Karras, S. Laine, M. Aittala, J. Hellsten, J. Lehtinen, and T. Aila, “Analyzing and improving the image quality of StyleGAN,” inProc. CVPR, 2020

  168. [179]

    Alias-free generative adversarial networks,

    T. Karras, M. Aittala, S. Laine, E. Härkönen, J. Hellsten, J. Lehtinen, and T. Aila, “Alias-free generative adversarial networks,” inProc. NeurIPS, 2021

  169. [180]

    Face alignment across large poses: A 3d solution,

    X. Zhu, Z. Lei, X. Liu, H. Shi, and S. Z. Li, “Face alignment across large poses: A 3d solution,”CoRR, vol. abs/1511.07212, 2015. [Online]. Available: http://arxiv.org/abs/1511.07212

  170. [181]

    Facescape: 3d facial dataset and benchmark for single-view 3d face reconstruction,

    H. Zhu, H. Yang, L. Guo, Y. Zhang, Y. Wang, M. Huang, M. Wu, Q. Shen, R. Yang, and X. Cao, “Facescape: 3d facial dataset and benchmark for single-view 3d face reconstruction,”IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), 2023

  171. [182]

    Stargan v2: Diverse image synthesis for multiple domains,

    Y. Choi, Y. Uh, J. Yoo, and J.-W. Ha, “Stargan v2: Diverse image synthesis for multiple domains,”

  172. [183]

    Arbitrary style transfer in real-time with adaptive instance normalization,

    X. Huang and S. Belongie, “Arbitrary style transfer in real-time with adaptive instance normalization,”

  173. [184]

    Efficient geometry-aware 3D generative adversarial networks,

    E. R. Chan, C. Z. Lin, M. A. Chan, K. Nagano, B. Pan, S. D. Mello, O. Gallo, L. Guibas, J. Tremblay, S. Khamis, T. Karras, and G. Wetzstein, “Efficient geometry-aware 3D generative adversarial networks,” in arXiv, 2021

  174. [185]

    Deep residual learning for image recognition,

    K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016, pp. 770–778

  175. [186]

    Taming transformers for high-resolution image synthesis,

    P. Esser, R. Rombach, and B. Ommer, “Taming transformers for high-resolution image synthesis,”

  176. [187]

    Invite: individual virtual transfer for personalized 3d face generation system,

    M. Jang, K. Lee, S. Lee, H. Tong, J. Chung, Y. Ro, and S. Lee, “Invite: individual virtual transfer for personalized 3d face generation system,” in Proceedings of the Thirty-Third International Joint Conference on Artificial Intelligence, ser. IJCAI ’24, 2024. [Online]. Availa...

  177. [188]

    Generative adversarial networks,

    I. J. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, “Generative adversarial networks,” 2014. [Online]. Available: https://arxiv.org/abs/1406.2661

  178. [189]

    Avatargen: a 3d generative model for animatable human avatars,

    J. Zhang, Z. Jiang, D. Yang, H. Xu, Y. Shi, G. Song, Z. Xu, X. Wang, and J. Feng, “Avatargen: a 3d generative model for animatable human avatars,” 2022. [Online]. Available: https://arxiv.org/abs/2208.00561

  179. [190]

    Conditional generative adversarial nets,

    M. Mirza and S. Osindero, “Conditional generative adversarial nets,” 2014. [Online]. Available: https://arxiv.org/abs/1411.1784

  180. [191]

    Muse: Text-to-image generation via masked generative transformers,

    H. Chang, H. Zhang, J. Barber, A. Maschinot, J. Lezama, L. Jiang, M.-H. Yang, K. Murphy, W. T. Freeman, M. Rubinstein, Y. Li, and D. Krishnan, “Muse: Text-to-image generation via masked generative transformers,” 2023. [Online]. Available: https://arxiv.org/abs/2301.00704

  181. [192]

    Facial action coding system,

    P. Ekman and W. V. Friesen, “Facial action coding system,”Environmental Psychology & Nonverbal Behavior, 1978

  182. [193]

    A comprehensive survey on deep facial expression recognition: challenges, applications, and future guidelines,

    M. Sajjad, F. U. M. Ullah, M. Ullah, G. Christodoulou, F. Alaya Cheikh, M. Hijji, K. Muhammad, and J. J. Rodrigues, “A comprehensive survey on deep facial expression recognition: challenges, applications, and future guidelines,” Alexandria Engineering Journal, vol. 68, pp. 817...

  183. [194]

    Mead: A large-scale audio-visual dataset for emotional talking-face generation,

    K. Wang, Q. Wu, L. Song, Z. Yang, W. Wu, C. Qian, R. He, Y. Qiao, and C. C. Loy, “Mead: A large-scale audio-visual dataset for emotional talking-face generation,” inComputer Vision – ECCV 2020, A. Vedaldi, H. Bischof, T. Brox, and J.-M. Frahm, Eds. Cham: Springer International...

  184. [195]

    The japanese female facial expression (jaffe) dataset,

    M. J. Lyons, M. G. Kamachi, and J. Gyoba, “The japanese female facial expression (jaffe) dataset,”

  185. [196]

    Web-based database for facial expression analysis,

    M. Pantic, M. Valstar, R. Rademaker, and L. Maat, “Web-based database for facial expression analysis,” in 2005 IEEE International Conference on Multimedia and Expo, 2005, pp. 5 pp.–

  186. [197]

    Single image, any face: Generalisable 3d face generation,

    W. Wang, H. Yang, J. Kittler, and X. Zhu, “Single image, any face: Generalisable 3d face generation,”

  187. [198]

    Learning formation of physically-based face attributes,

    R. Li, K. Bladin, Y. Zhao, C. Chinara, O. Ingraham, P. Xiang, X. Ren, P. B. Prasad, B. Kishore, J. Xing, and H. Li, “Learning formation of physically-based face attributes,” 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 3407–3416, 2020. [Onlin...

  188. [199]

    Chapter 5 - measuring emotions: A survey of cutting edge methodologies used in computer-based learning environment research,

    J. M. Harley, “Chapter 5 - measuring emotions: A survey of cutting edge methodologies used in computer-based learning environment research,” inEmotions, Technology, Design, and Learning, ser. Emotions and Technology, S. Y. Tettegah and M. Gartmeier, Eds. San Diego: Academic Pr...

  189. [200]

    Everybody dance now,

    C. Chan, S. Ginosar, T. Zhou, and A. A. Efros, “Everybody dance now,” inIEEE International Conference on Computer Vision (ICCV), 2019

  190. [201]

    Synthesizing obama: learning lip sync from audio,

    S. Suwajanakorn, S. M. Seitz, and I. Kemelmacher-Shlizerman, “Synthesizing obama: learning lip sync from audio,” ACM Trans. Graph., vol. 36, no. 4, jul 2017. [Online]. Available: https://doi.org/10.1145/3072959.3073640

  191. [202]

    VoxCeleb2: Deep Speaker Recognition,

    J. S. Chung, A. Nagrani, and A. Zisserman, “VoxCeleb2: Deep Speaker Recognition,” inProc. Interspeech 2018, 2018, pp. 1086–1090

  192. [203]

    Real time head pose estimation from consumer depth cameras,

    G. Fanelli, T. Weise, J. Gall, and L. Van Gool, “Real time head pose estimation from consumer depth cameras,” inPattern Recognition, R. Mester and M. Felsberg, Eds. Berlin, Heidelberg: Springer Berlin Heidelberg, 2011, pp. 101–110

  193. [204]

    Expressive body capture: 3d hands, face, and body from a single image,

    G. Pavlakos, V. Choutas, N. Ghorbani, T. Bolkart, A. A. A. Osman, D. Tzionas, and M. J. Black, “Expressive body capture: 3d hands, face, and body from a single image,”2019 IEEE/CVF Conference 61 on Computer Vision and Pattern Recognition (CVPR), pp. 10967–10977, 2019. [Online]...

  194. [205]

    Image qualityassessment: From errorvisibilitytostructural similarity,

    B. WangZhou, H. Sheikhet al., “Image qualityassessment: From errorvisibilitytostructural similarity,” IEEE Transon ImageProcessing, vol. 13, no. 4, p. 600, 2004

  195. [206]

    Multiface: A dataset for neural face rendering,

    C. hsin Wuu, N. Zheng, S. Ardisson, R. Bali, D. Belko, E. Brockmeyer, L. Evans, T. Godisart, H. Ha, X. Huang, A. Hypes, T. Koska, S. Krenn, S. Lombardi, X. Luo, K. McPhail, L. Millerschoen, M. Perdoch, M. Pitts, A. Richard, J. Saragih, J. Saragih, T. Shiratori, T. Simon, M. St...

  196. [207]

    Conformer-based speech recognition with linear nyström attention and rotary position embedding,

    L. Samarakoon and T.-Y. Leung, “Conformer-based speech recognition with linear nyström attention and rotary position embedding,” inICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2022, pp. 8012–8016

  197. [208]

    Learning high fidelity depths of dressed humans by watching social media dance videos,

    Y. Jafarian and H. S. Park, “Learning high fidelity depths of dressed humans by watching social media dance videos,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2021, pp. 12753–12762

  198. [209]

    Approximating facial expression effects on diagnostic accuracy via generative ai in medical genetics,

    T. Patel, A. A. Othman, Ö. Sümer, F. Hellman, P. Krawitz, E. André, M. E. Ripper, C. Fortney, S. Persky, P. Hu, C. Tekendo-Ngongang, S. L. Hanchard, K. A. Flaharty, R. L. Waikel, D. Duong, and B. D. Solomon, “Approximating facial expression effects on diagnostic accuracy via g...

  199. [210]

    Generative adversarial nets,

    I. J. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, “Generative adversarial nets,” inAdvances in Neural Information Processing Systems, Z. Ghahramani, M. Welling, C. Cortes, N. Lawrence, and K. Weinberger, Eds., vol. 27....

  200. [211]

    Denoising diffusion probabilistic models,

    J. Ho, A. Jain, and P. Abbeel, “Denoising diffusion probabilistic models,” in Advances in Neural Information Processing Systems, H. Larochelle, M. Ranzato, R. Hadsell, M. Balcan, and H. Lin, Eds., vol. 33. Curran Associates, Inc., 2020, pp. 6840–6851. [Online]. Available: http...

  201. [212]

    Realtimegen: An intervenable ai image generation system for commercial digital art asset creators,

    Z. Li, Y. Zhang, S. Zhou, Q. Liu, J. Zhang, H. Xu, S. Chen, X. Chen, and L. Sun, “Realtimegen: An intervenable ai image generation system for commercial digital art asset creators,”International Journal of Human–Computer Interaction, pp. 1–24, 2024

  202. [213]

    Using text-to-image generation for architectural design ideation,

    V. Paananen, J. Oppenlaender, and A. Visuri, “Using text-to-image generation for architectural design ideation,” International Journal of Architectural Computing, vol. 22, no. 3, pp. 458–474, 2024

  203. [214]

    A compensation method of two-stage image generation for human-ai collaborated in-situ fashion design in augmented reality environment,

    Z. Zhao and X. Ma, “A compensation method of two-stage image generation for human-ai collaborated in-situ fashion design in augmented reality environment,” in2018 IEEE international conference on artificial intelligence and virtual reality (AIVR). IEEE, 2018, pp. 76–83

  204. [215]

    Pushing mixture of experts to the limit: Extremely parameter efficient moe for instruction tuning,

    T. Zadouri, A. Üstün, A. Ahmadian, B. Ermiş, A. Locatelli, and S. Hooker, “Pushing mixture of experts to the limit: Extremely parameter efficient moe for instruction tuning,”arXiv preprint arXiv:2309.05444, 2023

  205. [216]

    Dcface: Synthetic face generation with dual condition diffusion model,

    M. Kim, F. Liu, A. Jain, and X. Liu, “Dcface: Synthetic face generation with dual condition diffusion model,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2023, pp. 12715–12725

  206. [217]

    Impact of deep learning approaches on facial expression recognition in healthcare industries,

    C. Bisogni, A. Castiglione, S. Hossain, F. Narducci, and S. Umer, “Impact of deep learning approaches on facial expression recognition in healthcare industries,”IEEE Transactions on Industrial Informatics, vol. 18, no. 8, pp. 5619–5627, 2022

  207. [218]

    Laion-400m: Open dataset of clip-filtered 400 million image-text pairs,

    C. Schuhmann, R. Vencu, R. Beaumont, R. Kaczmarczyk, C. Mullis, A. Katta, T. Coombes, J. Jitsev, and A. Komatsuzaki, “Laion-400m: Open dataset of clip-filtered 400 million image-text pairs,”arXiv preprint arXiv:2111.02114, 2021

  208. [219]

    The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale,

    A. Kuznetsova, H. Rom, N. Alldrin, J. Uijlings, I. Krasin, J. Pont-Tuset, S. Kamali, S. Popov, M. Malloci, A. Kolesnikov, T. Duerig, and V. Ferrari, “The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale,”IJCV, 2020. 62

  209. [220]

    Coyo-700m: Image-text pair dataset,

    M. Byeon, B. Park, H. Kim, S. Lee, W. Baek, and S. Kim, “Coyo-700m: Image-text pair dataset,” https://github.com/kakaobrain/coyo-dataset, 2022

  210. [221]

    Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning,

    P. Sharma, N. Ding, S. Goodman, and R. Soricut, “Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning,” inProceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), I. Gurevych a...

  211. [222]

    Sharegpt4v: Improving large multi-modal models with better captions,

    L. Chen, J. Li, X. Dong, P. Zhang, C. He, J. Wang, F. Zhao, and D. Lin, “Sharegpt4v: Improving large multi-modal models with better captions,” inEuropean Conference on Computer Vision. Springer, 2025, pp. 370–387

  212. [223]

    (2025) Free high-resolution photos

    Unsplash. (2025) Free high-resolution photos. Accessed: 27 March 2025. [Online]. Available: https://unsplash.com

  213. [224]

    Image generation step by step: animation generation-image translation,

    B. Jing, H. Ding, Z. Yang, B. Li, and Q. Liu, “Image generation step by step: animation generation-image translation,” Applied Intelligence, pp. 1–14, 2022

  214. [225]

    Image quality assessment: From error visibility to structural similarity,

    Z. Wang, A. C. Bovik, H. R. Sheikh, and E. P. Simoncelli, “Image quality assessment: From error visibility to structural similarity,”IEEE Transactions on Image Processing, vol. 13, no. 4, pp. 600–612, 2004

  215. [226]

    Charactergan: Few-shot keypoint character animation and reposing,

    T. Hinz, M. Fisher, O. Wang, E. Shechtman, and S. Wermter, “Charactergan: Few-shot keypoint character animation and reposing,” inProceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, 2022, pp. 1988–1997

  216. [227]

    Transforming sketches into realistic images: leveraging machine learning and image processing for enhanced architectural visualization,

    İ. Karadağ, “Transforming sketches into realistic images: leveraging machine learning and image processing for enhanced architectural visualization,”Sakarya University Journal of Science, vol. 27, no. 6, pp. 1209–1216, 2023

  217. [228]

    Generating architectural floor plans through conditional large diffusion model,

    Z. He, X. Li, P. Wu, L. Fan, H. J. Wang, N. Wang, M. Li, and Y. Chen, “Generating architectural floor plans through conditional large diffusion model,” inInternational Conference on Human-Computer Interaction. Springer, 2024, pp. 53–63

  218. [229]

    Llama- adapter v2: Parameter-efficient visual instruction model,

    P. Gao, J. Han, R. Zhang, Z. Lin, S. Geng, A. Zhou, W. Zhang, P. Lu, C. He, X. Yueet al., “Llama- adapter v2: Parameter-efficient visual instruction model,”arXiv preprint arXiv:2304.15010, 2023

  219. [230]

    Dreambooth: Fine tuning text- to-image diffusion models for subject-driven generation,

    N. Ruiz, Y. Li, V. Jampani, Y. Pritch, M. Rubinstein, and K. Aberman, “Dreambooth: Fine tuning text- to-image diffusion models for subject-driven generation,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2023, pp. 22500–22510

  220. [231]

    U-net: Convolutional networks for biomedical image segmentation,

    O. Ronneberger, P. Fischer, and T. Brox, “U-net: Convolutional networks for biomedical image segmentation,” CoRR, vol. abs/1505.04597, 2015. [Online]. Available: http://arxiv.org/abs/1505.04597

  221. [232]

    Applications of computer vision in entertainment and media industry,

    M. Hasan, K. S. Athrey, A. Khalid, D. Xie, E. Younessian, and T. Braskich, “Applications of computer vision in entertainment and media industry,”Computer Vision: Challenges, Trends, and Opportunities, p. 205, 2024

  222. [233]

    (2025) Free images and royalty free stock photos

    Pixabay. (2025) Free images and royalty free stock photos. Accessed: 27 March 2025. [Online]. Available: https://pixabay.com

  223. [234]

    Exploring the role of text-to-image ai in concept generation,

    R. Brisco, L. Hay, and S. Dhami, “Exploring the role of text-to-image ai in concept generation,” Proceedings of the Design Society, vol. 3, pp. 1835–1844, 2023

  224. [235]

    Sketch-to-architecture: Generative ai-aided architectural design,

    P. Li, B. Li, and Z. Li, “Sketch-to-architecture: Generative ai-aided architectural design,”arXiv preprint arXiv:2403.20186, 2024

  225. [236]

    Sketchar: Supporting character design and illustration prototyping using generative ai,

    L. Long, C. Xinyi, W. Ruoyu, L. Toby Jia-Jun, and L. Ray, “Sketchar: Supporting character design and illustration prototyping using generative ai,”Proceedings of the ACM on Human-Computer Interaction, vol. 8, no. CHI PLAY, p. 337, 2024

  226. [237]

    Innovative 3d character model texture mapping solution based on artificial intelligence image generation model,

    J. Yin and B. Song, “Innovative 3d character model texture mapping solution based on artificial intelligence image generation model,”International Journal of Contents, vol. 20, no. 4, pp. 14–21, 2024. 63

  227. [238]

    Virtualmodel: Generating object-id-retentive human-object interaction image by diffusion model for e-commerce marketing,

    B. Chen, C. Zhong, W. Xiang, Y. Geng, and X. Xie, “Virtualmodel: Generating object-id-retentive human-object interaction image by diffusion model for e-commerce marketing,” arXiv preprint arXiv:2405.09985, 2024

  228. [239]

    A new chapter for medical image generation: the stable diffusion method,

    L. X. Nguyen, P. S. Aung, H. Q. Le, S.-B. Park, and C. S. Hong, “A new chapter for medical image generation: the stable diffusion method,” in2023 International Conference on Information Networking (ICOIN). IEEE, 2023, pp. 483–486

  229. [240]

    Enhanced visual instruction tuning with synthesized image-dialogue data,

    Y. Li, C. Zhang, G. Yu, W. Yang, Z. Wang, B. Fu, G. Lin, C. Shen, L. Chen, and Y. Wei, “Enhanced visual instruction tuning with synthesized image-dialogue data,” inFindings of the Association for Computational Linguistics ACL 2024, 2024, pp. 14512–14531

  230. [241]

    Humman: Multi-modal 4d human dataset for versatile sensing and modeling,

    Z. Cai, D. Ren, A. Zeng, Z. Lin, T. Yu, W. Wang, X. Fan, Y. Gao, Y. Yu, L. Pan, F. Hong, M. Zhang, C. C. Loy, L. Yang, and Z. Liu, “Humman: Multi-modal 4d human dataset for versatile sensing and modeling,” inComputer Vision – ECCV 2022, S. Avidan, G. Brostow, M. Cissé, G. M. F...

  231. [242]

    The chosen one: Consistent characters in text-to-image diffusion models,

    O. Avrahami, A. Hertz, Y. Vinker, M. Arar, S. Fruchter, O. Fried, D. Cohen-Or, and D. Lischinski, “The chosen one: Consistent characters in text-to-image diffusion models,” inACM SIGGRAPH 2024 conference papers, 2024, pp. 1–12

  232. [243]

    Neural body: Implicit neural representations with structured latent codes for novel view synthesis of dynamic humans,

    S. Peng, Y. Zhang, Y. Xu, Q. Wang, Q. Shuai, H. Bao, and X. Zhou, “Neural body: Implicit neural representations with structured latent codes for novel view synthesis of dynamic humans,” inCVPR, 2021

  233. [244]

    Progressive background images generation,

    Y.-C. Chung, J.-M. Wang, and S.-W. Chen, “Progressive background images generation,” inProc. of 15th IPPR Conf. on Computer Vision, Graphics and Image Processing, 2002, pp. 858–865

  234. [245]

    Recovering accurate 3d human pose in the wild using imus and a moving camera,

    T. von Marcard, R. Henschel, M. Black, B. Rosenhahn, and G. Pons-Moll, “Recovering accurate 3d human pose in the wild using imus and a moving camera,” inEuropean Conference on Computer Vision (ECCV), sep 2018

  235. [246]

    Learn to dance with aist++: Music conditioned 3d dance generation,

    R. Li, S. Yang, D. A. Ross, and A. Kanazawa, “Learn to dance with aist++: Music conditioned 3d dance generation,” 2021

  236. [247]

    Puzzleavatar: Assembling 3d avatars from personal albums,

    Y. Xiu, Y. Ye, Z. Liu, D. Tzionas, and M. J. Black, “Puzzleavatar: Assembling 3d avatars from personal albums,” ACM Transactions on Graphics (TOG), 2024

  237. [248]

    Gaussianavatar-editor: Photorealistic animatable gaussian head avatar editor,

    X. Liu, K. Luo, H. Li, Q. Zhang, Y. Liu, L. Yi, and P. Tan, “Gaussianavatar-editor: Photorealistic animatable gaussian head avatar editor,”arXiv preprint arXiv:2501.09978, 2025

  238. [249]

    Avatarcraft: Transforming text into neural human avatars with parameterized shape and pose control,

    R. Jiang, C. Wang, J. Zhang, M. Chai, M. He, D. Chen, and J. Liao, “Avatarcraft: Transforming text into neural human avatars with parameterized shape and pose control,”arXiv preprint arXiv:2303.17606, 2023

  239. [250]

    Avatarverse: High-quality & stable 3d avatar creation from text and pose,

    H. Zhang, B. Chen, H. Yang, L. Qu, X. Wang, L. Chen, C. Long, F. Zhu, K. Du, and M. Zheng, “Avatarverse: High-quality & stable 3d avatar creation from text and pose,” 2023

  240. [251]

    AMASS: Archive of motion capture as surface shapes,

    N. Mahmood, N. Ghorbani, N. F. Troje, G. Pons-Moll, and M. J. Black, “AMASS: Archive of motion capture as surface shapes,” inInternational Conference on Computer Vision, Oct. 2019, pp. 5442–5451

  241. [252]

    imghum: Implicit generative models of 3d human shape and articulated pose,

    T. Alldieck, H. Xu, and C. Sminchisescu, “imghum: Implicit generative models of 3d human shape and articulated pose,”2021 IEEE/CVF International Conference on Computer Vision (ICCV), pp. 5441–5450, 2021. [Online]. Available: https://api.semanticscholar.org/CorpusID:237278058

  242. [254]

    Report on methods and applications for crafting 3d humans,

    L. Liu and K. Zhao, “Report on methods and applications for crafting 3d humans,” 2024. [Online]. Available: https://arxiv.org/abs/2406.01223

  243. [255]

    Avatarbooth: High-quality and customizable 3d human avatar generation,

    Y. Zeng, Y. Lu, X. Ji, Y. Yao, H. Zhu, and X. Cao, “Avatarbooth: High-quality and customizable 3d human avatar generation,” 2023. 64

  244. [256]

    A comprehensive review of data-driven co-speech gesture generation,

    S. Nyatsanga, T. Kucherenko, C. Ahuja, G. E. Henter, and M. Neff, “A comprehensive review of data-driven co-speech gesture generation,”Computer Graphics Forum, vol. 42, no. 2, p. 569–596, May

  245. [257]

    Hop: Heterogeneous topology- based multimodal entanglement for co-speech gesture generation,

    H. Cheng, T. Wang, G. Shi, Z. Zhao, and Y. Fu, “Hop: Heterogeneous topology- based multimodal entanglement for co-speech gesture generation,” 2025. [Online]. Available: https://arxiv.org/abs/2503.01175

  246. [258]

    Learning speech-driven 3d conversational gestures from video,

    I. Habibie, W. Xu, D. Mehta, L. Liu, H.-P. Seidel, G. Pons-Moll, M. Elgharib, and C. Theobalt, “Learning speech-driven 3d conversational gestures from video,” inProceedings of the 21st ACM International Conference on Intelligent Virtual Agents, 2021, pp. 101–108

  247. [259]

    Analyzing input and output representations for speech-driven gesture generation,

    T. Kucherenko, D. Hasegawa, G. E. Henter, N. Kaneko, and H. Kjellström, “Analyzing input and output representations for speech-driven gesture generation,” inProceedings of the 19th ACM International Conference on Intelligent Virtual Agents, ser. IVA ’19. ACM, Jul. 2019. [Onlin...

  248. [260]

    Latent-nerf for shape-guided generation of 3d shapes and textures,

    G. Metzer, E. Richardson, O. Patashnik, R. Giryes, and D. Cohen-Or, “Latent-nerf for shape-guided generation of 3d shapes and textures,”arXiv preprint arXiv:2211.07600, 2022

  249. [261]

    Style-controllable speech-driven gesture synthesis using normalising flows,

    S. Alexanderson, G. Henter, T. Kucherenko, and J. Beskow, “Style-controllable speech-driven gesture synthesis using normalising flows,”Computer Graphics Forum, vol. 39, pp. 487–496, 07 2020

  250. [262]

    Learning a model of facial shape and expression from 4d scans,

    T. Li, T. Bolkart, M. J. Black, H. Li, and J. Romero, “Learning a model of facial shape and expression from 4d scans,” ACM Trans. Graph., vol. 36, no. 6, Nov. 2017. [Online]. Available: https://doi.org/10.1145/3130800.3130813

  251. [263]

    The bielefeld speech and gesture alignment corpus (saga),

    A. Lücking, K. Bergmann, F. Hahn, S. Kopp, and H. Rieser, “The bielefeld speech and gesture alignment corpus (saga),” 01 2010

  252. [264]

    The usc creativeit database: A multimodal database of theatrical improvisation,

    A. Metallinou, C.-C. Lee, C. Busso, S. Carnicke, and S. Narayanan, “The usc creativeit database: A multimodal database of theatrical improvisation,” 05 2010

  253. [265]

    Panoptic studio: A massively multiview system for social interaction capture,

    H. Joo, T. Simon, X. Li, H. Liu, L. Tan, L. Gui, S. Banerjee, T. S. Godisart, B. Nabbe, I. Matthews, T. Kanade, S. Nobuhara, and Y. Sheikh, “Panoptic studio: A massively multiview system for social interaction capture,”IEEE Transactions on Pattern Analysis and Machine Intellig...

  254. [266]

    Available: http://dx.doi.org/10.1111/cgf.14776

    [Online]. Available: http://dx.doi.org/10.1111/cgf.14776

  255. [267]

    Talking with hands 16.2m: A large-scale dataset of synchronized body-finger motion and audio for conversational motion analysis and synthesis,

    G. Lee, Z. Deng, S. Ma, T. Shiratori, S. Srinivasa, and Y. Sheikh, “Talking with hands 16.2m: A large-scale dataset of synchronized body-finger motion and audio for conversational motion analysis and synthesis,” 10 2019, pp. 763–772

  256. [268]

    No gestures left behind: Learning relationships between spoken language and freeform gestures,

    C. Ahuja, D. W. Lee, R. Ishii, and L.-P. Morency, “No gestures left behind: Learning relationships between spoken language and freeform gestures,” inProceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: Findings, 2020, pp. 1884–1895

  257. [269]

    Speech2Properties2Gestures: Gesture-property prediction as a tool for generating represen- tational gestures from speech,

    T. Kucherenko, R. Nagy, P. Jonell, M. Neff, H. Kjellström, and G. E. Henter, “Speech2Properties2Gestures: Gesture-property prediction as a tool for generating represen- tational gestures from speech,” inProceedings of the 21th ACM International Conference on Intelligent Virtua...

  258. [270]

    Multi-objective adversarial gesture generation (mig 2019),

    Y. Ferstl, M. Neff, and R. McDonnell, “Multi-objective adversarial gesture generation (mig 2019),” 10 2019

  259. [271]

    Semantic gesticulator: Semantics- aware co-speech gesture synthesis,

    Z. Zhang, T. Ao, Y. Zhang, Q. Gao, C. Lin, B. Chen, and L. Liu, “Semantic gesticulator: Semantics- aware co-speech gesture synthesis,” 2024. [Online]. Available: https://arxiv.org/abs/2405.09814

  260. [272]

    Iemocap: Interactive emotional dyadic motion capture database,

    C. Busso, M. Bulut, C.-C. Lee, A. Kazemzadeh, E. Mower Provost, S. Kim, J. Chang, S. Lee, and S. Narayanan, “Iemocap: Interactive emotional dyadic motion capture database,”Language Resources and Evaluation, vol. 42, pp. 335–359, 12 2008

  261. [273]

    Motionlab: Unified human motion generation and editing via the motion-condition-motion paradigm,

    Z. Guo, Z. Hu, N. Zhao, and D. W. Soh, “Motionlab: Unified human motion generation and editing via the motion-condition-motion paradigm,” 2025. [Online]. Available: https://arxiv.org/abs/2502.02358

  262. [274]

    Ai choreographer: Music conditioned 3d dance generation with aist++,

    R. Li, S. Yang, D. A. Ross, and A. Kanazawa, “Ai choreographer: Music conditioned 3d dance generation with aist++,” 2021. [Online]. Available: https://arxiv.org/abs/2101.08779

  263. [275]

    Articulated human detection with flexible mixtures of parts,

    Y. Yang and D. Ramanan, “Articulated human detection with flexible mixtures of parts,”IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 35, no. 12, pp. 2878–2890, 2013

  264. [276]

    Learning Individual Styles of Conversational Gesture,

    S. Ginosar, A. Bar, G. Kohavi, C. Chan, A. Owens, and J. Malik, “Learning Individual Styles of Conversational Gesture,” in2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, Jun. 2019

  265. [277]

    The unreasonable effectiveness of deep features as a perceptual metric,

    R. Zhang, P. Isola, A. A. Efros, E. Shechtman, and O. Wang, “The unreasonable effectiveness of deep features as a perceptual metric,” in2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2018, pp. 586–595

  266. [278]

    Gans trained by a two time-scale update rule converge to a local nash equilibrium,

    M. Heusel, H. Ramsauer, T. Unterthiner, B. Nessler, and S. Hochreiter, “Gans trained by a two time-scale update rule converge to a local nash equilibrium,” 2018. [Online]. Available: https://arxiv.org/abs/1706.08500

  267. [279]

    Towards accurate generative models of video: A new metric & challenges,

    T. Unterthiner, S. van Steenkiste, K. Kurach, R. Marinier, M. Michalski, and S. Gelly, “Towards accurate generative models of video: A new metric & challenges,” 2019. [Online]. Available: https://arxiv.org/abs/1812.01717

  268. [280]

    Robutti, C

    O. Robutti, C. Sabena, C. Krause, C. Soldano, and F. Arzarello,Gestures in Mathematics Thinking and Learning, 11 2022, pp. 685–726

  269. [281]

    The genea challenge 2023: A large-scale evaluation of gesture generation models in monadic and dyadic settings,

    T. Kucherenko, R. Nagy, Y. Yoon, J. Woo, T. Nikolov, M. Tsakov, and G. E. Henter, “The genea challenge 2023: A large-scale evaluation of gesture generation models in monadic and dyadic settings,” inProceedings of the 25th International Conference on Multimodal Interaction, ser...

  270. [282]

    Convofusion: Multi-modal conversational diffusion for co-speech gesture synthesis,

    M. H. Mughal, R. Dabral, I. Habibie, L. Donatelli, M. Habermann, and C. Theobalt, “Convofusion: Multi-modal conversational diffusion for co-speech gesture synthesis,” 2024. [Online]. Available: https://arxiv.org/abs/2403.17936 65

  271. [283]

    Learning internal representations by error propagation,

    D. E. Rumelhart, G. E. Hinton, R. J. Williamset al., “Learning internal representations by error propagation,” 1985

  272. [284]

    Long short-term memory,

    S. Hochreiter and J. Schmidhuber, “Long short-term memory,”Neural Computation, vol. 9, pp. 1735– 1780, 11 1997

  273. [285]

    The graph neural network model,

    F. Scarselli, M. Gori, A. C. Tsoi, M. Hagenbuchner, and G. Monfardini, “The graph neural network model,” IEEE transactions on neural networks, vol. 20, no. 1, pp. 61–80, 2008

  274. [286]

    Speech gesture generation from the trimodal context of text, audio, and speaker identity,

    Y. Yoon, B. Cha, J.-H. Lee, M. Jang, J. Lee, J. Kim, and G. Lee, “Speech gesture generation from the trimodal context of text, audio, and speaker identity,”ACM Transactions on Graphics, vol. 39, no. 6, p. 1–16, Nov. 2020. [Online]. Available: http://dx.doi.org/10.1145/3414685.3417838

  275. [288]

    Switch transformers: scaling to trillion parameter models with simple and efficient sparsity,

    W. Fedus, B. Zoph, and N. Shazeer, “Switch transformers: scaling to trillion parameter models with simple and efficient sparsity,”J. Mach. Learn. Res., vol. 23, no. 1, Jan. 2022

  276. [289]

    Robust estimation of a location parameter,

    P. J. Huber, “Robust estimation of a location parameter,” inBreakthroughs in statistics: Methodology and distribution. Springer, 1992, pp. 492–518

  277. [290]

    Generating holistic 3d human motion from speech,

    H. Yi, H. Liang, Y. Liu, Q. Cao, Y. Wen, T. Bolkart, D. Tao, and M. J. Black, “Generating holistic 3d human motion from speech,” 2023. [Online]. Available: https://arxiv.org/abs/2212.04420

  278. [292]

    Towards a genea leaderboard – an extended, living benchmark for evaluating and advancing conversational motion synthesis,

    R. Nagy, H. Voss, Y. Yoon, T. Kucherenko, T. Nikolov, T. Hoang-Minh, R. McDonnell, S. Kopp, M. Neff, and G. E. Henter, “Towards a genea leaderboard – an extended, living benchmark for evaluating and advancing conversational motion synthesis,” 2024. [Online]. Available: https:/...

  279. [296]

    Denoising diffusion probabilistic models,

    J. Ho, A. Jain, and P. Abbeel, “Denoising diffusion probabilistic models,” 2020. [Online]. Available: https://arxiv.org/abs/2006.11239

  280. [300]

    Neural discrete representation learning,

    A. van den Oord, O. Vinyals, and K. Kavukcuoglu, “Neural discrete representation learning,” 2018. [Online]. Available: https://arxiv.org/abs/1711.00937

  281. [1998]

    Available: https://api.semanticscholar.org/CorpusID:231996102

    [Online]. Available: https://api.semanticscholar.org/CorpusID:231996102

  282. [2017]

    Available: https://arxiv.org/abs/1703.06868

    [Online]. Available: https://arxiv.org/abs/1703.06868

  283. [2020]

    Available: https://arxiv.org/abs/1912.01865

    [Online]. Available: https://arxiv.org/abs/1912.01865

  284. [2021]

    Available: https://arxiv.org/abs/2012.09841

    [Online]. Available: https://arxiv.org/abs/2012.09841

  285. [2023]

    Available: https://arxiv.org/abs/2310.17011

    [Online]. Available: https://arxiv.org/abs/2310.17011

  286. [2024]

    Available: https://arxiv.org/abs/2311.12052

    [Online]. Available: https://arxiv.org/abs/2311.12052

  287. [2025]

    Available: https://arxiv.org/abs/2409.16990

    [Online]. Available: https://arxiv.org/abs/2409.16990

Pith tools

Reviewed August 16, 2026 · model on record in the stance chip above.