Pith. sign in

REVIEW 4 major objections 5 minor 1 cited by

Multimodal Generative AI with Autoregressive LLMs for Human Motion Understanding and Generation: A Way Forward

T0 review · 4 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read This survey claims to be the first to map all of human motion understanding and generation, organizing text-conditioned synthesis into autoregressive LLMs, diffusion models, GANs, VAEs, and unified AR-diffusion frameworks.

desk verdict A useful survey of text-to-motion with a broad scope, but its central comparison table contains impossible numbers and the paper is not trustworthy as printed. read the letter →

arxiv 2506.03191 v1 pith:FZVU5FFM submitted 2025-05-31 cs.CV cs.AI

classification cs.CVcs.AI
keywords humanmotiongenerationtext-to-motionmultimodallargelanguagemodelsdiffusionautoregressiveunderstandinggenerativeAIsurveyunifiedframework
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that human motion understanding and generation (HMUG) is best viewed as a multimodal problem in which text conditions motion, with two main engines: autoregressive large language models and diffusion models. It claims to be the first survey to cover all of HMUG in one place—data representation, generation architectures, unified AR+diffusion frameworks, datasets, evaluation metrics, and applications—where earlier surveys covered only subtopics such as prediction, image or video generation, or generative methods without multimodal LLMs. The practical value the authors see is that a researcher entering text-to-motion can use this survey to choose an architecture, dataset, and metric without assembling the literature from scattered papers. The paper also argues that the next step is a unified model that plans, understands, and generates motion, together with a unified evaluation metric to compare such models fairly.

What carries the argument

The organizing device is an architecture taxonomy: autoregressive LLMs, which predict the next token and therefore need motion converted into discrete tokens by codebook-based vector quantizers such as VQ-VAE and VQ-GAN; diffusion models, which add and remove Gaussian noise in a forward-reverse process that ideally runs in a compressed latent space; and unified frameworks that align text and motion in a shared embedding space, using a multimodal transformer or a mixture-of-experts connector to let the autoregressive and diffusion paradigms cooperate. Each mechanism does a specific job in the survey: it explains why methods behave differently, since autoregressive models capture long-range text-motion dependencies but lose fine detail at tokenization, diffusion models produce high-fidelity motion at the cost of many denoising steps, and unified models aim to get both understanding and generation from a single training objective.

What would settle it

A reader can settle the uniqueness claim by checking the surveys listed in the paper's own Table 1: if any of them already covers all nine scope columns the paper counts for itself, the 'first and unique' assertion is false. The benchmark tables are independently falsifiable, since recomputing entries against the cited papers will expose an error whenever a reported Precision exceeds the [0,1] bounds the paper itself lists (for example, the 5.400 value in Table 6).

Watch

Extended reading notes

Core claim

In the authors' telling, the central discovery is that text-conditioned human motion generation has converged on two complementary paradigms—autoregressive LLMs that treat motion as a foreign language of discrete tokens (via VQ-VAE or VQ-GAN tokenizers) and diffusion models that refine noisy motion in continuous or latent space—and that recent work is combining them into unified frameworks. The survey's own contribution is the claim of uniqueness: no prior survey, it says, covers both motion understanding and generation together with multimodal generative AI and autoregressive LLMs, because previous reviews stop at prediction, at image or video generation, or at general generative models without LLMs. It supports this claim by categorizing methods by architecture and backbone, comparing them on the HumanML3D and KIT-ML benchmarks with fidelity, diversity, and consistency metrics, and mapping them onto applications from healthcare to autonomous driving.

Load-bearing premise

The load-bearing premise is that the comparison-table numbers copied from the cited papers are accurate and were measured under comparable protocols; if they are wrong or incomparable, the survey's value as a cross-method benchmark is compromised.

Editorial extensions

If this is right

  • A newcomer to text-to-motion generation can select an architecture—autoregressive LLM, diffusion, GAN, VAE, or unified—directly from the survey's taxonomy, with matching datasets and metrics.
  • Autoregressive LLM approaches that treat motion as discrete tokens are mature enough to generate, caption, retrieve, and reason about motion, not merely synthesize sequences.
  • Unified AR+diffusion models trained in a shared embedding space with a multimodal transformer or mixture-of-experts connector are presented as the direction that combines understanding and generation.
  • The absence of a standardized evaluation framework is identified as a real bottleneck, and the survey argues for a unified metric that combines fidelity, diversity, and consistency, citing work toward that goal.
  • Real-world deployment in rehabilitation, VR/AR, gaming, robotics, surveillance, and autonomous vehicles depends on making motion generation efficient, controllable, and long-form, which the survey lists as the main open challenges.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If discrete motion tokens become standard, general-purpose LLM tooling—instruction tuning, prompting, retrieval, and even adversarial red-teaming—would apply to motion almost directly, which is why the survey's unified AR+diffusion direction is a plausible successor to diffusion-only text-to-motion models.
  • A concrete testable extension of the survey's call for a unified metric would be to build a single score that normalizes fidelity, diversity, and consistency across HumanML3D and KIT-ML and validates it against human perceptual judgments.
  • The taxonomy implies that motion understanding and generation are converging: once the same tokenizer feeds both an LLM and a diffusion decoder, captioning and synthesis become two directions of one mapping, an idea the survey hints at but does not develop.
  • The paper's comparison tables are best treated as a starting point to be checked against the original sources before being reused as a benchmark, since cross-method comparability is assumed rather than demonstrated.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. This manuscript is a survey of text-conditioned human motion understanding and generation using multimodal generative AI and autoregressive large language models. It reviews data representations, text and motion tokenizers, LLM-based motion models, GAN/VAE/diffusion text-to-motion methods, unified frameworks, datasets, evaluation metrics, applications, and open challenges. The authors claim that this is the first survey to cover all areas of Human Motion Understanding and Generation (HMUG), and they support the survey with comparative tables, including a quantitative comparison of methods on HumanML3D and KIT-ML.

Significance. If the comparative tables and citation mappings were accurate, this survey would be a useful reference because it gathers a large body of recent work and organizes it by architectural family and task. The breadth is real: the paper reviews more than 250 references and provides useful taxonomy tables and figures that map models, backbones, datasets, and tasks. However, the survey's central value as a cross-method benchmark depends on Table 6 and on the traceability of Tables 2-4, and those components are currently not reliable as printed. The paper is a survey, so it contains no machine-checked proofs or code, but its factual claims about other papers must be verifiable against the cited sources; at present they are not.

major comments (4)
  1. [Table 5 and Table 6] Table 5 defines Precision with bounds 0≤Precision≤1 and MM Score with bounds 0≤MMS≤1, but Table 6 reports Precision values of 5.400 (AlertMotion), 2.550 (ActFormer), and 8.210 (UDE) on HumanML3D, and MM values above 1 in nearly every row. Since Table 6 is the only quantitative cross-method comparison in the survey, these internally inconsistent numbers cannot support the paper's comparative conclusions. The authors must re-extract every value from the original papers, state the metric definition used, and ensure consistency with Table 5.
  2. [Table 6 and reference list] Several rows in Table 6 do not point to the cited work. The row 'VQ-VAE Mot [103]' cites [103], Ghosh et al., 'Synthesis of Compositional Animations from Textual Descriptions,' which is not a VQ-VAE method; 'Unify MoGPT [146]' cites [146], Ma et al., which is the MoFusion diffusion paper; and 'WalkLLM [221]' cites [221], PackDiT, not the WalkLLM pedestrian-motion paper. In addition, there are two distinct papers named MotionLLM ([62] and [74]), and Table 6 lists MotionLLM [62] twice with different FID and Precision values. The reference-to-row mapping must be corrected throughout the tables.
  3. [Sections 6 and 7] Sections 6 and 7 have identical titles ('Datasets and Evaluation Metrics') and identical opening paragraphs, with Section 7 containing the actual subsections while Section 6 is left empty in substance. This duplication is a structural error that must be fixed by merging the content into a single section and renumbering the subsequent sections before the survey can be read as a coherent document.
  4. [Introduction and Table 1] The claim that this is 'the first, and unique attempt that covers all the areas of Human Motion Understanding and Generation (HMUG)' is not substantiated by Table 1 as presented. The checkmark criteria in Table 1 are undefined, and several prior surveys (e.g., [22] and [31]) overlap substantially with the stated scope. The authors should either provide an operational definition of 'all areas of HMUG' and demonstrate specific coverage gaps relative to each prior survey, or soften the uniqueness claim.
minor comments (5)
  1. [Equation (14)] Equation (14) is written as FID(P1,P2)^2 = ..., while the surrounding text refers to 'the FID score'; please adopt the standard convention FID = ||μ1-μ2||^2 + Tr(Σ1+Σ2-2(Σ1Σ2)^{1/2}) to avoid ambiguity.
  2. [Table 3] Table 3 lists 'ActFormer [95]' and 'FineMoGen [93]', but the cited references are the DCGAN paper and 'To Create What You Tell', respectively; the correct citations appear to be [92] and [89]. Please verify and correct these mappings.
  3. [Scope, Abstract and Section 7.3] The abstract and Section 1 state that the survey focuses exclusively on text and motion modalities, but Section 7.3 and Table 4 include audio, speech, music, and emotion datasets and methods; please reconcile the stated scope with the included material.
  4. [Table 2] Table 2 uses the overlapping names 'MotionGPT-3 [78]', 'MotionGPT-2 [137]', and 'T2M-GPT [77]' while the text and reference list use similar names for different papers; please adopt a consistent naming and citation scheme so that readers can map each row to its source.
  5. [Table 6 and evaluation protocols] The FID and other metric values in Table 6 are reported without the evaluation protocol used (e.g., number of samples, text prompt set, and whether metrics come from the original papers or from re-evaluation), so even after correcting the transcription errors the values may not be directly comparable across methods; please state the protocol explicitly.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity found: this paper is a review that derives no new predictions or fitted results, so the circularity patterns do not apply.

full rationale

The manuscript is a survey of text-conditioned human motion generation and understanding. It contains no derivation chain from fitted parameters to predictions, no claimed first-principles result, and no new empirical benchmark produced by the authors. The paper's stated contribution is organizational: it summarizes more than 250 prior works and provides comparative tables. A survey can be inaccurate, incomplete, or internally inconsistent, but those are correctness and reliability concerns, not circularity. The only quantitative synthesis is Table 6, which transcribes reported metrics from external papers. Even if some Precision values exceed 1 and conflict with Table 5's stated bounds, this would be a transcription or provenance error, not a case where the output is equivalent to the input by construction. The 'first and unique attempt' claim is a scope assertion, not a derivation. The paper does cite prior work by others, and there is no load-bearing self-citation chain or imported uniqueness theorem. Accordingly, the circularity score is 0.

Assumptions & free parameters 0 free parameters · 0 assumptions · 0 invented entities

The paper is a literature review and introduces no derivations, fitted parameters, or new entities. Its contribution is organizational, not derivational.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Multimodal Generative AI with Autoregressive LLMs for Human Motion Understanding and Generation: A Way Forward." pith.science (2026). https://pith.science/paper/FZVU5FFM

@misc{pith2026250603191,
  author       = {Pith},
  title        = {Pith review of: Multimodal Generative AI with Autoregressive LLMs for Human Motion Understanding and Generation: A Way Forward},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/FZVU5FFM}},
  note         = {Machine review of arXiv:2506.03191}
}
read the original abstract

This paper presents an in-depth survey on the use of multimodal Generative Artificial Intelligence (GenAI) and autoregressive Large Language Models (LLMs) for human motion understanding and generation, offering insights into emerging methods, architectures, and their potential to advance realistic and versatile motion synthesis. Focusing exclusively on text and motion modalities, this research investigates how textual descriptions can guide the generation of complex, human-like motion sequences. The paper explores various generative approaches, including autoregressive models, diffusion models, Generative Adversarial Networks (GANs), Variational Autoencoders (VAEs), and transformer-based models, by analyzing their strengths and limitations in terms of motion quality, computational efficiency, and adaptability. It highlights recent advances in text-conditioned motion generation, where textual inputs are used to control and refine motion outputs with greater precision. The integration of LLMs further enhances these models by enabling semantic alignment between instructions and motion, improving coherence and contextual relevance. This systematic survey underscores the transformative potential of text-to-motion GenAI and LLM architectures in applications such as healthcare, humanoids, gaming, animation, and assistive technologies, while addressing ongoing challenges in generating efficient and realistic human motion.

Figures

Figures reproduced from arXiv: 2506.03191 by the authors.

Figure 1
Figure 1. The overall flow diagram illustrating variants of multimodal generative models, multimodal LLM [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. The architecture of a visual tokenizer typically used in generative AI. The model converts visual [PITH_FULL_IMAGE:figures/full_fig_p008_2.png] view at source ↗
Figure 3
Figure 3. An illustration of various architectures used in motion understanding and generation, including [PITH_FULL_IMAGE:figures/full_fig_p010_3.png] view at source ↗
Figures from the paper (7 more)
Figure 4
Figure 4. Figure 4: A representation of early fusion and alignment architectures employing contrastive learning. The [PITH_FULL_IMAGE:figures/full_fig_p011_4.png]
Figure 5
Figure 5. Figure 5: An illustration of generative model architectures, including GAN, VAE, and Di [PITH_FULL_IMAGE:figures/full_fig_p014_5.png]
Figure 6
Figure 6. Figure 6: An illustration of multimodal diffusion GAN and VAE variants in the architecture. The figure highlights how each variant is integrated into the architecture, with the diffusion GAN enabling generative learning across multiple modalities and the VAE variants facilitatin…
Figure 7
Figure 7. Figure 7: An illustration of a) the cooperative capabilities of an AR and Diffusion Model in a single unified framework, and b) the Dense versus Mixture of Experts alignment methodologies. generation. However, these methods may falter when text conditions fail to represent expli…
Figure 8
Figure 8. Figure 8: A detailed illustration of a) multimodal LLM, b) multimodal Diffusion, c) multimodal Trans￾former, and unified approaches with an encoder-decoder backbone architecture. enables adaptive token selection, thereby optimizing fine-grained understanding and generation tasks…
Figure 9
Figure 9. Figure 9: A comparative analysis of methods based on [PITH_FULL_IMAGE:figures/full_fig_p029_9.png]
Figure 10
Figure 10. Figure 10: A comparative analysis of different methodologies based on the FID metric on the HU￾MANML3D dataset. 8.2 Virtual Reality (VR) and Augmented Reality (AR) VR and AR applications are significantly enhanced through the generation of realistic human movements that adapt se…

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. IRG-MotionLLM: Interleaving Motion Generation, Assessment and Refinement for Text-to-Motion Generation

    cs.CV 2025-12 conditional novelty 6.0 of 10

    Interleaving motion generation with text-motion assessment and refinement improves alignment between generated human motion and goal text.

Reference graph

Works this paper leans on

215 extracted references · 42 canonical work pages · cited by 1 Pith paper

  1. [103]

    Synthesis of Compositional Animations from Textual Descriptions

    A. Ghosh, N. Cheema, C. Oguz, C. Theobalt, and P . Slusallek, "Synthesis of Compositional Anima- tions from Textual Descriptions," 2021, arXiv. doi: 10.48550/ARXIV.2103.14675

  2. [146]

    Pretrained Diffusion Models for Unified Human Motion Synthesis

    J. Ma, S. Bai, and C. Zhou, “Pretrained Diffusion Models for Unified Human Motion Synthesis,” Dec. 06, 2022, arXiv: arXiv:2212.02837. doi: 10.48550/arXiv.2212.02837

  3. [221]

    PackDiT: Joint Human Motion and Text Generation via Mutual Prompting

    Kiang, Z., Chai, W., Zhou, Z., Yang, C. Y., Huang, H. W., & Hwang, J. N. (2025). PackDiT: Joint Human Motion and Text Generation via Mutual Prompting.arXiv preprint arXiv:2501.16551. 46

  4. [74]

    MotionLLM: Understanding Human Behaviors from Human Motions and Videos,

    L.-H. Chen et al., "MotionLLM: Understanding Human Behaviors from Human Motions and Videos," arXiv, 2024, doi: 10.48550/ARXIV.2405.20340

  5. [22]

    Multi-Modal Generative AI: Multi-modal LLM, Diffusion and Beyond,

    H. Chen et al., “Multi-Modal Generative AI: Multi-modal LLM, Diffusion and Beyond,” 2024, arXiv. doi: 10.48550/ARXIV.2409.14993

  6. [31]

    Xiong, J., Liu, G., Huang, L., Wu, C., Wu, T., Mu, Y., & Wong, N. (2024). Autoregressive Models in Vision: A Survey. arXiv preprint arXiv:2411.05902

  7. [1]

    Towards artificial general intelligence via a multimodal foundation model

    N. Fei et al., “Towards artificial general intelligence via a multimodal foundation model,” 2021, doi: 10.48550/ARXIV.2110.14378

  8. [2]

    Video generation models as world simulators OpenAI SORA

    T. Brooks, B. Peebles, C. Holmes, W. DePue, Y. Guo, L. Jing, D. Schnurr, J. Taylor, T. Luhman, E. Luh- man, C. Ng, R. Wang, and A. Ramesh, “Video generation models as world simulators OpenAI SORA.” [Online]. Available: https://openai.com/research/video-generation-models-as-world-simulators

Show all 215 references
  1. [3]

    GPT-4 Technical Report,

    OpenAI et al., “GPT-4 Technical Report,” 2023, arXiv. doi: 10.48550/ARXIV.2303.08774

  2. [4]

    DALL·E 3 understands significantly more nuance and detail than our previous systems, allowing you to easily translate your ideas into exceptionally accurate images

    Jong Wook Kim, Alex Nichol, Yang Song, Lijuan Wang, Tao Xu, “DALL·E 3 understands significantly more nuance and detail than our previous systems, allowing you to easily translate your ideas into exceptionally accurate images.”

  3. [5]

    MotionDiffuse: Text-Driven Human Motion Generation with Diffusion Model,

    M. Zhang et al., “MotionDiffuse: Text-Driven Human Motion Generation with Diffusion Model,” 2022, arXiv. doi: 10.48550/ARXIV.2208.15001

  4. [6]

    MoCoGAN: Decomposing Motion and Content for Video Generation,

    S. Tulyakov, M.-Y. Liu, X. Yang, and J. Kautz, “MoCoGAN: Decomposing Motion and Content for Video Generation,” 2017, arXiv. doi: 10.48550/ARXIV.1707.04993

  5. [7]

    Deep Generative Modelling: A Compar- ative Review of V AEs, GANs, Normalizing Flows, Energy-Based and Autoregressive Models,

    S. Bond-Taylor, A. Leach, Y. Long, and C. G. Willcocks, “Deep Generative Modelling: A Compar- ative Review of V AEs, GANs, Normalizing Flows, Energy-Based and Autoregressive Models,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 44, no. 11, pp. 7327–7347, Nov. 2022, doi: 10.11...

  6. [8]

    Graph-based Normalizing Flow for Human Motion Generation and Reconstruction,

    W. Yin, H. Yin, D. Kragic, and M. Bjorkman, “Graph-based Normalizing Flow for Human Motion Generation and Reconstruction,” in 2021 30th IEEE International Conference on Robot & Human Interactive Communication (RO-MAN), Vancouver, BC, Canada: IEEE, Aug. 2021, pp. 641–648. doi: ...

  7. [9]

    BAMM: Bidirectional Autoregressive Motion Model,

    E. Pinyoanuntapong, M. U. Saleem, P . Wang, M. Lee, S. Das, and C. Chen, “BAMM: Bidirectional Autoregressive Motion Model,” Apr. 01, 2024, arXiv: arXiv:2403.19435. Accessed: Sep. 22, 2024. [Online]. Available: http://arxiv.org/abs/2403.19435

  8. [10]

    Executing your Commands via Motion Diffusion in Latent Space,

    X. Chen et al., “Executing your Commands via Motion Diffusion in Latent Space,” 2022, arXiv. doi: 10.48550/ARXIV.2212.04048

  9. [11]

    Denoising Diffusion Probabilistic Models,

    J. Ho, A. Jain, and P . Abbeel, “Denoising Diffusion Probabilistic Models,” 2020, arXiv. doi: 10.48550/ARXIV.2006.11239

  10. [12]

    A survey of GPT-3 family large language models including ChatGPT and GPT-4,

    K. S. Kalyan, “A survey of GPT-3 family large language models including ChatGPT and GPT-4,” Nat. Lang. Process. J., vol. 6, p. 100048, Mar. 2024, doi: 10.1016/j.nlp.2023.100048

  11. [13]

    ChatGPT vs Gemini vs LLaMA on Multilingual Sentiment Anal- ysis,

    A. Buscemi and D. Proverbio, “ChatGPT vs Gemini vs LLaMA on Multilingual Sentiment Anal- ysis,” Jan. 25, 2024, arXiv: arXiv:2402.01715. Accessed: May 06, 2024. [Online]. Available: http://arxiv.org/abs/2402.01715

  12. [14]

    Text-to-Motion Retrieval: Towards Joint Un- derstanding of Human Motion Data and Natural Language,

    N. Messina, J. Sedmidubsky, F. Falchi, and T. Rebok, “Text-to-Motion Retrieval: Towards Joint Un- derstanding of Human Motion Data and Natural Language,” in Proceedings of the 46th Interna- tional ACM SIGIR Conference on Research and Development in Information Retrieval, Jul. ...

  13. [15]

    Sentence-T5: Scalable Sentence Encoders from Pre-trained Text-to-Text Models,

    J. Ni et al., “Sentence-T5: Scalable Sentence Encoders from Pre-trained Text-to-Text Models,” 2021, arXiv. doi: 10.48550/ARXIV.2108.08877

  14. [16]

    HP-GAN: Probabilistic 3D human motion prediction via GAN,

    E. Barsoum, J. Kender, and Z. Liu, “HP-GAN: Probabilistic 3D human motion prediction via GAN,” 2017, arXiv. doi: 10.48550/ARXIV.1711.09561. 34

  15. [17]

    RIVQ-V AE: Discrete Rotation-Invariant 3D Rep- resentation Learning,

    M. Mezghanni, M. Boulkenafed, and M. Ovsjanikov, “RIVQ-V AE: Discrete Rotation-Invariant 3D Rep- resentation Learning,” in 2024 International Conference on 3D Vision (3DV), Davos, Switzerland: IEEE, Mar. 2024, pp. 1382–1391. doi: 10.1109/3DV62453.2024.00129

  16. [18]

    A Survey of Cross-Modal Visual Content Generation,

    F. Nazarieh, Z. Feng, M. Awais, W. Wang, and J. Kittler, “A Survey of Cross-Modal Visual Content Generation,” IEEE Trans. Circuits Syst. Video Technol., vol. 34, no. 8, pp. 6814–6832, Aug. 2024, doi: 10.1109/TCSVT.2024.3351601

  17. [19]

    Human Image Generation: A Comprehensive Survey,

    Z. Jia, Z. Zhang, L. Wang, and T. Tan, “Human Image Generation: A Comprehensive Survey,” ACM Comput. Surv., vol. 56, no. 11, pp. 1–39, Nov. 2024, doi: 10.1145/3665869

  18. [20]

    3D human motion prediction: A survey,

    K. Lyu, H. Chen, Z. Liu, B. Zhang, and R. Wang, “3D human motion prediction: A survey,” Neuro- computing, vol. 489, pp. 345–365, Jun. 2022, doi: 10.1016/j.neucom.2022.02.045

  19. [21]

    Human Motion Generation: A Survey,

    W. Zhu et al., “Human Motion Generation: A Survey,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 46, no. 4, pp. 2430–2449, Apr. 2024, doi: 10.1109/TPAMI.2023.3330935

  20. [23]

    Deep generative models on 3d representations: A survey

    Shi, Zifan, et al. "Deep generative models on 3d representations: A survey." arXiv preprint arXiv:2210.15663 (2022)

  21. [24]

    Gavrila, D. M. (1999). The visual analysis of human movement: A survey. Computer vision and image understanding, 73(1), 82-98

  22. [25]

    Wang, L., Hu, W., & Tan, T. (2003). Recent developments in human motion analysis. Pattern recogni- tion, 36(3), 585-601

  23. [26]

    B., Hilton, A., & Krüger, V

    Moeslund, T. B., Hilton, A., & Krüger, V. (2006). A survey of advances in vision-based human motion capture and analysis. Computer vision and image understanding, 104(2-3), 90-126

  24. [27]

    Sedmidubsky, J., Elias, P ., Budikova, P ., & Zezula, P . (2021). Content-based management of human motion data: survey and challenges. IEEE Access, 9, 64241-64255

  25. [28]

    Xue, H., Luo, X., Hu, Z., Zhang, X., Xiang, X., Dai, Y., & Yu, F. R. (2024). Human motion video genera- tion: A survey. Authorea Preprints

  26. [29]

    Joshi, I., Grimmer, M., Rathgeb, C., Busch, C., Bremond, F., & Dantcheva, A. (2024). Synthetic data in human analysis: A survey. IEEE Transactions on Pattern Analysis and Machine Intelligence, 46(7), 4957-4976

  27. [30]

    Liao, F., Zou, X., & Wong, W. (2024). Appearance and Pose-guided Human Generation: A Survey. ACM Computing Surveys, 56(5), 1-35

  28. [32]

    Ismail-Fawaz, A., Devanne, M., Berretti, S., Weber, J., & Forestier, G. (2024). Establishing a Unified Evaluation Framework for Human Motion Generation: A Comparative Analysis of Metrics. arXiv preprint arXiv:2405.07680

  29. [33]

    Human Motion Prediction via Learning Local Structure Representations and Temporal Dependencies,

    X. Guo and J. Choi, “Human Motion Prediction via Learning Local Structure Representations and Temporal Dependencies,” Proc. AAAI Conf. Artif. Intell., vol. 33, no. 01, pp. 2580–2587, Jul. 2019, doi: 10.1609/aaai.v33i01.33012580

  30. [34]

    RNN -based Human Motion Prediction via Differ- ential Sequence Representation,

    Y. Wang, X. Wang, P . Jiang, and F. Wang, “RNN -based Human Motion Prediction via Differ- ential Sequence Representation,” in 2019 IEEE 6th International Conference on Cloud Comput- ing and Intelligence Systems (CCIS), Singapore: IEEE, Dec. 2019, pp. 138–143. doi: 10.1109/C- C...

  31. [35]

    On Human Motion Prediction Using Recurrent Neural Net- works,

    J. Martinez, M. J. Black, and J. Romero, “On Human Motion Prediction Using Recurrent Neural Net- works,” in 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI: IEEE, Jul. 2017, pp. 4674–4683. doi: 10.1109/CVPR.2017.497

  32. [36]

    Recurrent Network Models for Human Dynamics,

    K. Fragkiadaki, S. Levine, P . Felsen, and J. Malik, “Recurrent Network Models for Human Dynamics,” 2015, arXiv. doi: 10.48550/ARXIV.1508.00271

  33. [37]

    Efficient convolutional hierarchical autoencoder for human motion prediction,

    Y. Li et al., “Efficient convolutional hierarchical autoencoder for human motion prediction,” Vis. Com- put., vol. 35, no. 6–8, pp. 1143–1156, Jun. 2019, doi: 10.1007/s00371-019-01692-9

  34. [38]

    Learning Multiscale Correlations for Human Motion Pre- diction,

    H. Zhou, C. Guo, H. Zhang, and Y. Wang, “Learning Multiscale Correlations for Human Motion Pre- diction,” 2021, arXiv. doi: 10.48550/ARXIV.2103.10674

  35. [39]

    Aggregated Multi-GANs for Controlled 3D Human Motion Prediction,

    Z. Liu, K. Lyu, S. Wu, H. Chen, Y. Hao, and S. Ji, “Aggregated Multi-GANs for Controlled 3D Human Motion Prediction,” Proc. AAAI Conf. Artif. Intell., vol. 35, no. 3, pp. 2225–2232, May 2021, doi: 10.1609/aaai.v35i3.16321

  36. [40]

    THUNDR: Transformer-based 3D HUmaN Reconstruction with Markers,

    M. Zanfir, A. Zanfir, E. G. Bazavan, W. T. Freeman, R. Sukthankar, and C. Sminchisescu, “THUNDR: Transformer-based 3D HUmaN Reconstruction with Markers,” 2021, arXiv. doi: 10.48550/ARXIV.2106.09336

  37. [42]

    3D Human Mesh Estimation from Virtual Markers,

    X. Ma, J. Su, C. Wang, W. Zhu, and Y. Wang, "3D Human Mesh Estimation from Virtual Markers," 2023, arXiv. doi: 10.48550/ARXIV.2303.11726

  38. [43]

    SMPL: a skinned multi-person linear model,

    M. Loper, N. Mahmood, J. Romero, G. Pons-Moll, and M. J. Black, "SMPL: a skinned multi-person linear model," ACM Trans. Graph., vol. 34, no. 6, pp. 1–16, Nov. 2015, doi: 10.1145/2816795.2818013

  39. [44]

    Expressive Body Capture: 3D Hands, Face, and Body from a Single Image,

    G. Pavlakos et al., "Expressive Body Capture: 3D Hands, Face, and Body from a Single Image," 2019, arXiv. doi: 10.48550/ARXIV.1904.05866

  40. [45]

    STAR: Sparse Trained Articulated Human Body Regres- sor,

    A. A. A. Osman, T. Bolkart, and M. J. Black, "STAR: Sparse Trained Articulated Human Body Regres- sor," 2020, doi: 10.48550/ARXIV.2008.08535

  41. [47]

    QuaterNet: A Quaternion-based Recurrent Model for Human Motion,

    D. Pavllo, D. Grangier, and M. Auli, "QuaterNet: A Quaternion-based Recurrent Model for Human Motion," 2018, arXiv. doi: 10.48550/ARXIV.1805.06485

  42. [48]

    Towards Natural and Accurate Future Motion Prediction of Humans and Animals,

    Z. Liu et al., "Towards Natural and Accurate Future Motion Prediction of Humans and Animals," in 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA: IEEE, Jun. 2019, pp. 9996–10004. doi: 10.1109/CVPR.2019.01024

  43. [49]

    Motion Prediction via Joint Dependency Modeling in Phase Space,

    P . Su, Z. Liu, S. Wu, L. Zhu, Y. Yin, and X. Shen, "Motion Prediction via Joint Dependency Modeling in Phase Space," in Proceedings of the 29th ACM International Conference on Multimedia, Virtual Event China: ACM, Oct. 2021, pp. 713–721. doi: 10.1145/3474085.3475237

  44. [50]

    Multiscale Spatio-Temporal Graph Neu- ral Networks for 3D Skeleton-Based Motion Prediction,

    M. Li, S. Chen, Y. Zhao, Y. Zhang, Y. Wang, and Q. Tian, "Multiscale Spatio-Temporal Graph Neu- ral Networks for 3D Skeleton-Based Motion Prediction," IEEE Trans. Image Process., vol. 30, pp. 7760–7775, 2021, doi: 10.1109/TIP .2021.3108708

  45. [51]

    Vedaldi, Computer Vision - ECCV 2020: 16th European Conference, Glasgow, UK, August 23-28, 2020, Proceedings, Part XIV

    A. Vedaldi, Computer Vision - ECCV 2020: 16th European Conference, Glasgow, UK, August 23-28, 2020, Proceedings, Part XIV. in Lecture Notes in Computer Science Ser, no. v. 12359. Cham: Springer International Publishing AG, 2020. 36

  46. [53]

    Dynamic Multiscale Graph Neural Net- works for 3D Skeleton Based Human Motion Prediction,

    M. Li, S. Chen, Y. Zhao, Y. Zhang, Y. Wang, and Q. Tian, "Dynamic Multiscale Graph Neural Net- works for 3D Skeleton Based Human Motion Prediction," in 2020 IEEE/CVF Conference on Com- puter Vision and Pattern Recognition (CVPR), Seattle, WA, USA: IEEE, Jun. 2020, pp. 211–220....

  47. [54]

    Learning Trajectory Dependencies for Human Motion Prediction,

    W. Mao, M. Liu, M. Salzmann, and H. Li, "Learning Trajectory Dependencies for Human Motion Prediction," in 2019 IEEE/CVF International Conference on Computer Vision (ICCV), Seoul, Korea (South): IEEE, Oct. 2019, pp. 9488–9496. doi: 10.1109/ICCV.2019.00958

  48. [55]

    TrajectoryCNN: A New Spatio-Temporal Feature Learning Network for Human Motion Prediction,

    X. Liu, J. Yin, J. Liu, P . Ding, J. Liu, and H. Liu, "TrajectoryCNN: A New Spatio-Temporal Feature Learning Network for Human Motion Prediction," IEEE Trans. Circuits Syst. Video Technol., vol. 31, no. 6, pp. 2133–2146, Jun. 2021, doi: 10.1109/TCSVT.2020.3021409

  49. [56]

    Structural-RNN: Deep Learning on Spatio-Temporal Graphs,

    A. Jain, A. R. Zamir, S. Savarese, and A. Saxena, "Structural-RNN: Deep Learning on Spatio-Temporal Graphs," in 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA: IEEE, Jun. 2016, pp. 5308–5317. doi: 10.1109/CVPR.2016.573

  50. [57]

    Towards Accurate 3D Human Motion Prediction from Incomplete Observations,

    Q. Cui and H. Sun, "Towards Accurate 3D Human Motion Prediction from Incomplete Observations," in 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA: IEEE, Jun. 2021, pp. 4799–4808. doi: 10.1109/CVPR46437.2021.00477

  51. [58]

    Ferrari, Computer Vision - ECCV 2018: 15th European Conference, Munich, Germany, September 8-14, 2018, Proceedings, Part IV

    V. Ferrari, Computer Vision - ECCV 2018: 15th European Conference, Munich, Germany, September 8-14, 2018, Proceedings, Part IV. in Lecture Notes in Computer Science Ser, no. v. 11208. Cham: Springer International Publishing AG, 2018

  52. [59]

    word2vec, node2vec, graph2vec, X2vec: Towards a Theory of Vector Embeddings of Struc- tured Data,

    M. Grohe, "word2vec, node2vec, graph2vec, X2vec: Towards a Theory of Vector Embeddings of Struc- tured Data," in Proceedings of the 39th ACM SIGMOD-SIGACT-SIGAI Symposium on Principles of Database Systems, Portland OR USA: ACM, Jun. 2020, pp. 1–16. doi: 10.1145/3375395.3387641

  53. [60]

    Glove: Global Vectors for Word Representation,

    J. Pennington, R. Socher, and C. Manning, "Glove: Global Vectors for Word Representation," in Pro- ceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), Doha, Qatar: Association for Computational Linguistics, 2014, pp. 1532–1543. doi: 10....

  54. [61]

    BERT: A Review of Applications in Natural Language Processing and Understanding,

    M. V. Koroteev, "BERT: A Review of Applications in Natural Language Processing and Understanding," 2021, arXiv. doi: 10.48550/ARXIV.2103.11943

  55. [63]

    MotionGPT: Finetuned LLMs Are General-Purpose Motion Generators,

    Y. Zhang et al., "MotionGPT: Finetuned LLMs Are General-Purpose Motion Generators," Proc. AAAI Conf. Artif. Intell., vol. 38, no. 7, pp. 7368–7376, Mar. 2024, doi: 10.1609/aaai.v38i7.28567

  56. [64]

    Human3.6M: Large Scale Datasets and Pre- dictive Methods for 3D Human Sensing in Natural Environments,

    C. Ionescu, D. Papava, V. Olaru, and C. Sminchisescu, "Human3.6M: Large Scale Datasets and Pre- dictive Methods for 3D Human Sensing in Natural Environments," IEEE Trans. Pattern Anal. Mach. Intell., vol. 36, no. 7, pp. 1325–1339, Jul. 2014, doi: 10.1109/TPAMI.2013.248

  57. [65]

    MoSh: motion and shape capture from sparse markers,

    M. Loper, N. Mahmood, and M. J. Black, "MoSh: motion and shape capture from sparse markers," ACM Trans. Graph., vol. 33, no. 6, pp. 1–13, Nov. 2014, doi: 10.1145/2661229.2661273

  58. [66]

    Learnable Triangulation of Human Pose,

    K. Iskakov, E. Burkov, V. Lempitsky, and Y. Malkov, "Learnable Triangulation of Human Pose," arXiv, 2019, doi: 10.48550/ARXIV.1905.05754

  59. [67]

    VoxelPose: Towards Multi-Camera 3D Human Pose Estimation in Wild Environment,

    H. Tu, C. Wang, and W. Zeng, "VoxelPose: Towards Multi-Camera 3D Human Pose Estimation in Wild Environment," arXiv, 2020, doi: 10.48550/ARXIV.2004.06239. 37

  60. [68]

    Faster VoxelPose: Real-time 3D Human Pose Estimation by Orthographic Projection,

    H. Ye, W. Zhu, C. Wang, R. Wu, and Y. Wang, "Faster VoxelPose: Real-time 3D Human Pose Estimation by Orthographic Projection," arXiv, 2022, doi: 10.48550/ARXIV.2207.10955

  61. [69]

    Realtime Multi-Person 2D Pose Estimation using Part Affinity Fields,

    Z. Cao, T. Simon, S.-E. Wei, and Y. Sheikh, "Realtime Multi-Person 2D Pose Estimation using Part Affinity Fields," arXiv, 2016, doi: 10.48550/ARXIV.1611.08050

  62. [70]

    3D human pose estimation in video with tem- poral convolutions and semi-supervised training,

    D. Pavllo, C. Feichtenhofer, D. Grangier, and M. Auli, "3D human pose estimation in video with tem- poral convolutions and semi-supervised training," arXiv, 2018, doi: 10.48550/ARXIV.1811.11742

  63. [71]

    Liu, J., Dai, W., Wang, C., Cheng, Y., Tang, Y., & Tong, X. (2023). Plan, posture and go: Towards open-world text-to-motion generation. arXiv preprint arXiv:2312.14828

  64. [72]

    Recovering 3D Human Mesh From Monocular Images: A Survey,

    Y. Tian, H. Zhang, Y. Liu, and L. Wang, "Recovering 3D Human Mesh From Monocular Images: A Survey," IEEE Trans. Pattern Anal. Mach. Intell., vol. 45, no. 12, pp. 15406–15425, Dec. 2023, doi: 10.1109/TPAMI.2023.3298850

  65. [73]

    H., Tan, H., Bansal, M., Rohrbach, A., Chang, K

    Shen, S., Li, L. H., Tan, H., Bansal, M., Rohrbach, A., Chang, K. W., ... & Keutzer, K. (2021). How much can clip benefit vision-and-language tasks?.arXiv preprint arXiv:2107.06383

  66. [75]

    Shen, X., Li, X., & Elhoseiny, M. (2023). Mostgan-v: Video generation with temporal motion styles. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 5652- 5661)

  67. [76]

    AvatarGPT: All-in-One Framework for Motion Understanding, Plan- ning, Generation and Beyond,

    Z. Zhou, Y. Wan, and B. Wang, "AvatarGPT: All-in-One Framework for Motion Understanding, Plan- ning, Generation and Beyond," arXiv, 2023, doi: 10.48550/ARXIV.2311.16468

  68. [77]

    MotionGPT: Human Motion as a Foreign Lan- guage,

    B. Jiang, X. Chen, W. Liu, J. Yu, G. Yu, and T. Chen, "MotionGPT: Human Motion as a Foreign Lan- guage," arXiv, Jul. 19, 2023. Accessed: Sep. 22, 2024. [Online].http://arxiv.org/abs/2306.14795

  69. [78]

    MotionGPT: Human Motion Synthesis with Improved Diversity and Realism via GPT-3 Prompting,

    J. Ribeiro-Gomes et al., "MotionGPT: Human Motion Synthesis with Improved Diversity and Realism via GPT-3 Prompting," in 2024 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), Waikoloa, HI, USA: IEEE, Jan. 2024, pp. 5058–5068, doi: 10.1109/WACV57701.2024.00499

  70. [79]

    MotionScript: Natural Language Descriptions for Ex- pressive 3D Human Motions,

    P . J. Yazdian, E. Liu, L. Cheng, and A. Lim, "MotionScript: Natural Language Descriptions for Ex- pressive 3D Human Motions," arXiv, Dec. 19, 2023. Accessed: Sep. 29, 2024. [Online]. Available: http://arxiv.org/abs/2312.12634

  71. [80]

    T2M-GPT: Generating Human Motion from Textual Descriptions with Dis- crete Representations,

    J. Zhang et al., "T2M-GPT: Generating Human Motion from Textual Descriptions with Dis- crete Representations," arXiv, Sep. 24, 2023. Accessed: Sep. 29, 2024. [Online]. Avail- able:http://arxiv.org/abs/2301.06052

  72. [81]

    MotionChain: Conversational Motion Controllers via Multimodal Prompts,

    B. Jiang et al., "MotionChain: Conversational Motion Controllers via Multimodal Prompts," arXiv, 2024, doi: 10.48550/ARXIV.2404.01700

  73. [82]

    Autonomous LLM-Enhanced Adversarial Attack for Text-to-Motion,

    H. Miao, F. Ma, R. Quan, K. Zhan, and Y. Yang, "Autonomous LLM-Enhanced Adversarial Attack for Text-to-Motion," arXiv, Aug. 01, 2024. Accessed: Sep. 22, 2024. [Online].http://arxiv.org/abs/ 2408.00352

  74. [83]

    Walk-the-Talk: LLM driven pedestrian motion generation,

    M. Ramesh and F. B. Flohr, "Walk-the-Talk: LLM driven pedestrian motion generation," in 2024 IEEE Intelligent Vehicles Symposium (IV), Jeju Island, Korea, Republic of: IEEE, Jun. 2024, pp. 3057–3062, doi: 10.1109/IV55156.2024.10588860

  75. [84]

    Motion Generation from Fine-grained Textual Descriptions,

    K. Li and Y. Feng, "Motion Generation from Fine-grained Textual Descriptions," arXiv, Mar. 26, 2024. Accessed: Sep. 22, 2024. [Online].http://arxiv.org/abs/2403.13518

  76. [85]

    TAAT: Think and Act from Arbitrary Texts in Text2Motion,

    R. Wang, C. Ma, G. Li, and Z. Wang, "TAAT: Think and Act from Arbitrary Texts in Text2Motion," arXiv, 2024, doi: 10.48550/ARXIV.2404.14745. 38

  77. [86]

    (2024, September)

    Shrestha, A., Liu, P ., Ros, G., Yuan, K., & Fern, A. (2024, September). Generating Physically Realis- tic and Directable Human Motions from Multi-modal Inputs. In European Conference on Computer Vision (pp. 1-17). Cham: Springer Nature Switzerland

  78. [87]

    Generation of Walking Motions Based on Whole-Body Poses and QP Control,

    R. Grimm, A. Kheddar and T. Asfour, "Generation of Walking Motions Based on Whole-Body Poses and QP Control," 2018 IEEE-RAS 18th International Conference on Humanoid Robots (Humanoids), Beijing, China, 2018, pp. 510-515, doi: 10.1109/HUMANOIDS.2018.8624913

  79. [88]

    Luo, Z., Cao, J., Merel, J., Winkler, A., Huang, J., Kitani, K., & Xu, W. (2023). Universal humanoid motion representations for physics-based control. arXiv preprint arXiv:2310.04582

  80. [89]

    FineMoGen: Fine-Grained Spatio-Temporal Motion Generation and Editing,

    M. Zhang, H. Li, Z. Cai, J. Ren, L. Yang, and Z. Liu, "FineMoGen: Fine-Grained Spatio-Temporal Motion Generation and Editing," arXiv

  81. [90]

    A., Dashtipour, K., Zahid, A., Abbasi, Q

    Taylor, W., Shah, S. A., Dashtipour, K., Zahid, A., Abbasi, Q. H., & Imran, M. A. (2020). An Intel- ligent Non-Invasive Real-Time Human Activity Recognition System for Next-Generation Healthcare. Sensors, 20(9), 2653. https://doi.org/10.3390/s20092653

  82. [91]

    Generative adversarial networks,

    I. Goodfellow et al., "Generative adversarial networks," Commun. ACM, vol. 63, no. 11, pp. 139–144, Oct. 2020, doi: 10.1145/3422622

  83. [92]

    ActFormer: A GAN-based Transformer towards General Action-Conditioned 3D Human Motion Generation,

    L. Xu et al., "ActFormer: A GAN-based Transformer towards General Action-Conditioned 3D Human Motion Generation," arXiv, 2022, doi: 10.48550/ARXIV.2203.07706

  84. [93]

    To Create What You Tell: Generating Videos from Captions,

    Y. Pan, Z. Qiu, T. Yao, H. Li, and T. Mei, "To Create What You Tell: Generating Videos from Captions," arXiv, 2018, doi: 10.48550/ARXIV.1804.08264

  85. [94]

    IRC-GAN: Introspective Recurrent Convolutional GAN for Text-to-video Generation,

    K. Deng, T. Fei, X. Huang, and Y. Peng, "IRC-GAN: Introspective Recurrent Convolutional GAN for Text-to-video Generation," in Proceedings of the Twenty-Eighth International Joint Conference on Ar- tificial Intelligence, Macao, China: International Joint Conferences on Artifici...

  86. [95]

    Unsupervised Representation Learning with Deep Convolu- tional Generative Adversarial Networks,

    A. Radford, L. Metz, and S. Chintala, "Unsupervised Representation Learning with Deep Convolu- tional Generative Adversarial Networks," arXiv, 2015, doi: 10.48550/ARXIV.1511.06434

  87. [96]

    Progressive Growing of GANs for Improved Quality, Stability, and Variation,

    T. Karras, T. Aila, S. Laine, and J. Lehtinen, "Progressive Growing of GANs for Improved Quality, Stability, and Variation," arXiv, 2017, doi: 10.48550/ARXIV.1710.10196

  88. [97]

    C., Azghadi, M

    Huang, T., Liu, J., Zhou, X., Nguyen, D. C., Azghadi, M. R., Xia, Y., & Sun, S. (2023). V2X cooperative perception for autonomous driving: Recent advances and challenges. arXiv preprint arXiv:2310.03525

  89. [98]

    A Style-Based Generator Architecture for Generative Adversarial Networks,

    T. Karras, S. Laine, and T. Aila, "A Style-Based Generator Architecture for Generative Adversarial Networks," in 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA: IEEE, Jun. 2019, pp. 4396–4405, doi: 10.1109/CVPR.2019.00453

  90. [99]

    Analyzing and Improving the Image Quality of StyleGAN,

    T. Karras, S. Laine, M. Aittala, J. Hellsten, J. Lehtinen, and T. Aila, "Analyzing and Improving the Image Quality of StyleGAN," 2019, arXiv. doi: 10.48550/ARXIV.1912.04958

  91. [100]

    Anatomically-Informed Vector Quantization Variational Auto-Encoder for Text to Motion Generation,

    L. Chen, Z. Niu, Q. Liu, J. Wang, J. Xue, and K. Lu, "Anatomically-Informed Vector Quantization Variational Auto-Encoder for Text to Motion Generation," in 2024 IEEE International Conference on Multimedia and Expo Workshops (ICMEW), Niagara Falls, ON, Canada: IEEE, Jul. 2024, ...

  92. [101]

    Human motion generation with StyleGAN,

    K. Yamamoto and M. Murakami, "Human motion generation with StyleGAN," in Seventh Interna- tional Conference on Computer Graphics and Virtuality (ICCGV 2024), J. Li, Ed., Hangzhou, China: SPIE, May 2024, p. 6. doi: 10.1117/12.3029447

  93. [102]

    Text2Action: Generative Adversarial Synthesis from Language to Action,

    H. Ahn, T. Ha, Y. Choi, H. Yoo, and S. Oh, "Text2Action: Generative Adversarial Synthesis from Language to Action," 2017, arXiv. doi: 10.48550/ARXIV.1710.05298. 39

  94. [104]

    TransCGan-based human motion generator,

    W. Yu, "TransCGan-based human motion generator," in Fifth International Conference on Computer Information Science and Artificial Intelligence (CISAI 2022), Y. Zhong, Ed., Chongqing, China: SPIE, Mar. 2023, p. 188. doi: 10.1117/12.2668277

  95. [105]

    MSFF-GAN: A Refined Method for Human Video Motion Transfer,

    Y. Xue and L. Chai, "MSFF-GAN: A Refined Method for Human Video Motion Transfer," in 2024 36th Chinese Control and Decision Conference (CCDC), Xi’an, China: IEEE, May 2024, pp. 4416–4421. doi: 10.1109/CCDC62350.2024.10587557

  96. [106]

    Text-guided 3D Human Motion Generation with Keyframe-based Parallel Skip Transformer,

    Z. Geng, C. Han, Z. Hayder, J. Liu, M. Shah, and A. Mian, "Text-guided 3D Human Motion Generation with Keyframe-based Parallel Skip Transformer," May 24, 2024, arXiv: arXiv:2405.15439. Accessed: Sep. 22, 2024. [Online]. Available: http://arxiv.org/abs/2405.15439

  97. [107]

    AvatarCLIP: Zero-Shot Text-Driven Genera- tion and Animation of 3D Avatars,

    F. Hong, M. Zhang, L. Pan, Z. Cai, L. Yang, and Z. Liu, "AvatarCLIP: Zero-Shot Text-Driven Genera- tion and Animation of 3D Avatars," 2022, arXiv. doi: 10.48550/ARXIV.2205.08535

  98. [108]

    Amballa, A., Akkinapalli, G., & Muralikrishnan, V. (2024). LS-GAN: Human Motion Synthesis with Latent-space GANs. arXiv preprint arXiv:2501.01449

  99. [109]

    Auto-Encoding Variational Bayes,

    D. P . Kingma and M. Welling, "Auto-Encoding Variational Bayes," 2013, arXiv. doi: 10.48550/ARXIV.1312.6114

  100. [110]

    (2022, October)

    Lu, Q., Zhang, Y., Lu, M., & Roychowdhury, V. (2022, October). Action-conditioned on-demand mo- tion generation. In Proceedings of the 30th ACM International Conference on Multimedia (pp. 2249- 2257)

  101. [111]

    (2024, September)

    Jin, P ., Li, H., Cheng, Z., Li, K., Yu, R., Liu, C., & Chen, J. (2024, September). Local action-guided motion diffusion model for text-to-motion generation. In European Conference on Computer Vision (pp. 392-409). Cham: Springer Nature Switzerland

  102. [112]

    Action2Motion: Conditioned Generation of 3D Human Motions,

    C. Guo et al., "Action2Motion: Conditioned Generation of 3D Human Motions," 2020, doi: 10.48550/ARXIV.2007.15240

  103. [114]

    Contact-aware Human Motion Generation from Textual De- scriptions,

    S. Ma, Q. Cao, J. Zhang, and D. Tao, "Contact-aware Human Motion Generation from Textual De- scriptions," 2024, arXiv. doi: 10.48550/ARXIV.2403.15709

  104. [115]

    UDE: A Unified Driving Engine for Human Motion Generation,

    Z. Zhou and B. Wang, "UDE: A Unified Driving Engine for Human Motion Generation," in 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vancouver, BC, Canada: IEEE, Jun. 2023, pp. 5632–5641. doi: 10.1109/CVPR52729.2023.00545

  105. [116]

    Priority-Centric Human Motion Generation in Discrete Latent Space,

    H. Kong, K. Gong, D. Lian, M. B. Mi, and X. Wang, "Priority-Centric Human Motion Generation in Discrete Latent Space," 2023, arXiv. doi: 10.48550/ARXIV.2308.14480

  106. [117]

    BAD: Bidirectional Auto-regressive Diffusion for Text-to-Motion Generation,

    S. R. Hosseyni, A. A. Rahmani, S. J. Seyedmohammadi, S. Seyedin, and A. Mohammadi, "BAD: Bidirectional Auto-regressive Diffusion for Text-to-Motion Generation," 2024, arXiv. doi: 10.48550/ARXIV.2409.10847

  107. [118]

    Motion-Agent: A Conversational Frame- work for Human Motion Generation with LLMs,

    Q. Wu, Y. Zhao, Y. Wang, X. Liu, Y.-W. Tai, and C.-K. Tang, "Motion-Agent: A Conversational Frame- work for Human Motion Generation with LLMs," 2024, arXiv. doi: 10.48550/ARXIV.2405.17013

  108. [119]

    UniMuMo: Unified Text, Music and Motion Generation,

    H. Yang et al., "UniMuMo: Unified Text, Music and Motion Generation," 2024, arXiv. doi: 10.48550/ARXIV.2410.04534. 40

  109. [120]

    MoMask: Generative Masked Modeling of 3D Human Motions,

    C. Guo, Y. Mu, M. G. Javed, S. Wang, and L. Cheng, "MoMask: Generative Masked Modeling of 3D Human Motions," 2023, arXiv. doi: 10.48550/ARXIV.2312.00063

  110. [121]

    InfiniMotion: Mamba Boosts Memory in Transformer for Arbitrary Long Motion Generation,

    Z. Zhang et al., "InfiniMotion: Mamba Boosts Memory in Transformer for Arbitrary Long Motion Generation," Jul. 13, 2024, arXiv: arXiv:2407.10061. Accessed: Sep. 22, 2024. [Online]. Available: http://arxiv.org/abs/2407.10061

  111. [122]

    T2M-HiFiGPT: Generating High Quality Human Motion from Textual Descriptions with Residual Discrete Representations,

    C. Wang, "T2M-HiFiGPT: Generating High Quality Human Motion from Textual Descriptions with Residual Discrete Representations," 2023, arXiv. doi: 10.48550/ARXIV.2312.10628

  112. [123]

    ControlMM: Controllable Masked Motion Generation,

    E. Pinyoanuntapong et al., "ControlMM: Controllable Masked Motion Generation," 2024, arXiv. doi: 10.48550/ARXIV.2410.10780

  113. [124]

    Enabling Synergistic Full-Body Control in Prompt- Based Co-Speech Motion Generation,

    B. Chen, Y. Li, Y.-X. Ding, T. Shao, and K. Zhou, "Enabling Synergistic Full-Body Control in Prompt- Based Co-Speech Motion Generation," 2024, arXiv. doi: 10.48550/ARXIV.2410.00464

  114. [125]

    AttT2M: Text-Driven Human Motion Generation with Multi- Perspective Attention Mechanism,

    C. Zhong, L. Hu, Z. Zhang, and S. Xia, "AttT2M: Text-Driven Human Motion Generation with Multi- Perspective Attention Mechanism," Sep. 01, 2023, arXiv: arXiv:2309.00796. Accessed: Sep. 22, 2024. [Online]. Available: http://arxiv.org/abs/2309.00796

  115. [126]

    HumanTOMATO: Text-aligned Whole-body Motion Generation,

    S. Lu et al., "HumanTOMATO: Text-aligned Whole-body Motion Generation," Oct. 19, 2023, arXiv: arXiv:2310.12978. Accessed: Sep. 22, 2024. [Online]. Available: http://arxiv.org/abs/2310.12978

  116. [127]

    Language2Pose: Natural Language Grounded Pose Forecasting,

    C. Ahuja and L.-P . Morency, "Language2Pose: Natural Language Grounded Pose Forecasting," 2019, arXiv. doi: 10.48550/ARXIV.1907.01108

  117. [128]

    Hierarchical Text-Conditional Image Generation with CLIP Latents,

    A. Ramesh, P . Dhariwal, A. Nichol, C. Chu, and M. Chen, "Hierarchical Text-Conditional Image Generation with CLIP Latents," 2022, arXiv. doi: 10.48550/ARXIV.2204.06125

  118. [129]

    MotionCLIP: Exposing Human Motion Generation to CLIP Space,

    G. Tevet, B. Gordon, A. Hertz, A. H. Bermano, and D. Cohen-Or, "MotionCLIP: Exposing Human Motion Generation to CLIP Space," in Computer Vision – ECCV 2022, vol. 13682, S. Avidan, G. Bros- tow, M. Cissé, G. M. Farinella, and T. Hassner, Eds., in Lecture Notes in Computer Scien...

  119. [130]

    FLAME: Free-form Language-based Motion Synthesis & Editing,

    J. Kim, J. Kim, and S. Choi, "FLAME: Free-form Language-based Motion Synthesis & Editing," 2022, arXiv. doi: 10.48550/ARXIV.2209.00349

  120. [131]

    Diffusion Motion: Generate Text-Guided 3D Human Motion by Diffusion Model,

    Z. Ren, Z. Pan, X. Zhou, and L. Kang, "Diffusion Motion: Generate Text-Guided 3D Human Motion by Diffusion Model," 2022, arXiv. doi: 10.48550/ARXIV.2210.12315

  121. [132]

    Guided Motion Diffusion for Controllable Human Motion Synthesis,

    K. Karunratanakul, K. Preechakul, S. Suwajanakorn, and S. Tang, “Guided Motion Diffusion for Controllable Human Motion Synthesis,” in 2023 IEEE/CVF International Conference on Computer Vision (ICCV), Paris, France: IEEE, Oct. 2023, pp. 2151–2162. doi: 10.1109/ICCV51070.2023.00205

  122. [133]

    ReMoDiffuse: Retrieval-Augmented Motion Diffusion Model,

    M. Zhang et al., “ReMoDiffuse: Retrieval-Augmented Motion Diffusion Model,” 2023, arXiv. doi: 10.48550/ARXIV.2304.01116

  123. [134]

    MotionFix: Text-Driven 3D Human Motion Editing,

    N. Athanasiou, A. Ceske, M. Diomataris, M. J. Black, and G. Varol, “MotionFix: Text-Driven 3D Human Motion Editing,” Sep. 19, 2024, arXiv: arXiv:2408.00712. Accessed: Sep. 22, 2024. [Online]. Available: http://arxiv.org/abs/2408.00712

  124. [135]

    GUESS: GradUally Enriching SyntheSis for Text-Driven Human Motion Generation,

    X. Gao, Y. Yang, Z. Xie, S. Du, Z. Sun, and Y. Wu, “GUESS: GradUally Enriching SyntheSis for Text-Driven Human Motion Generation,” IEEE Trans. Vis. Comput. Graph., pp. 1–13, 2024, doi: 10.1109/TVCG.2024.3352002

  125. [136]

    MotionLLaMA: A Unified Frame- work for Motion Synthesis and Comprehension,

    Z. Ling, B. Han, S. Li, H. Shen, J. Cheng, and C. Zou, “MotionLLaMA: A Unified Frame- work for Motion Synthesis and Comprehension,” Nov. 26, 2024, arXiv: arXiv:2411.17335. doi: 10.48550/arXiv.2411.17335. 41

  126. [137]

    MotionGPT-2: A General-Purpose Motion-Language Model for Motion Generation and Understanding,

    Y. Wang et al., “MotionGPT-2: A General-Purpose Motion-Language Model for Motion Generation and Understanding,” Oct. 29, 2024, arXiv: arXiv:2410.21747. doi: 10.48550/arXiv.2410.21747

  127. [138]

    Learning Generalizable Human Motion Generator with Reinforcement Learning,

    Y. Mao, X. Liu, W. Zhou, Z. Lu, and H. Li, “Learning Generalizable Human Motion Generator with Reinforcement Learning,” May 24, 2024, arXiv: arXiv:2405.15541. Accessed: Sep. 29, 2024. [Online]. Available: http://arxiv.org/abs/2405.15541

  128. [139]

    AMD: Autoregressive Motion Diffusion,

    B. Han, H. Peng, M. Dong, Y. Ren, Y. Shen, and C. Xu, “AMD: Autoregressive Motion Diffusion,” Proc. AAAI Conf. Artif. Intell., vol. 38, no. 3, pp. 2022–2030, Mar. 2024, doi: 10.1609/aaai.v38i3.27973

  129. [140]

    AAMDM: Accelerated Auto-Regressive Motion Diffusion Model,

    T. Li, C. Qiao, G. Ren, K. Yin, and S. Ha, “AAMDM: Accelerated Auto-Regressive Motion Diffusion Model,” in 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA: IEEE, Jun. 2024, pp. 1813–1823. doi: 10.1109/CVPR52733.2024.00178

  130. [141]

    Large Motion Model for Unified Multi-modal Motion Generation,

    M. Zhang et al., “Large Motion Model for Unified Multi-modal Motion Generation,” in Computer Vision – ECCV 2024, vol. 15071, A. Leonardis, E. Ricci, S. Roth, O. Russakovsky, T. Sattler, and G. Varol, Eds., in Lecture Notes in Computer Science, vol. 15071. , Cham: Springer Natu...

  131. [142]

    M2D2M: Multi-Motion Generation from Text with Discrete Diffusion Models,

    S. Chi et al., “M2D2M: Multi-Motion Generation from Text with Discrete Diffusion Models,” in Com- puter Vision – ECCV 2024, vol. 15072, A. Leonardis, E. Ricci, S. Roth, O. Russakovsky, T. Sattler, and G. Varol, Eds., in Lecture Notes in Computer Science, vol. 15072. , Cham: Sp...

  132. [143]

    L., Wu, W., Loy, C

    Jiang, Y., Yang, S., Koh, T. L., Wu, W., Loy, C. C., & Liu, Z. (2023). Text2performer: Text-driven human video generation. In Proceedings of the IEEE/CVF International Conference on Computer Vision (pp. 22747-22757)

  133. [144]

    & Yang, W

    Huang, Z., Yu, Y., Yang, L., Qin, C., Zheng, B., Zheng, X., ... & Yang, W. (2024, October). Motion-aware latent diffusion models for video frame interpolation. In Proceedings of the 32nd ACM International Conference on Multimedia (pp. 1043-1052)

  134. [145]

    DART: A Diffusion-Based Autoregressive Motion Model for Real-Time Text-Driven Motion Control,

    K. Zhao, G. Li, and S. Tang, “DART: A Diffusion-Based Autoregressive Motion Model for Real-Time Text-Driven Motion Control,” 2024, arXiv. doi: 10.48550/ARXIV.2410.05260

  135. [147]

    Plappert, M., Mandery, C., & Asfour, T. (2016). The kit motion-language dataset. Big data, 4(4), 236- 252

  136. [148]

    T., & Zheng, W

    Ji, Y., Xu, F., Yang, Y., Shen, F., Shen, H. T., & Zheng, W. S. (2018, October). A large-scale RGB-D database for arbitrary-view human action recognition. In Proceedings of the 26th ACM international Conference on Multimedia (pp. 1510-1518)

  137. [149]

    Ntu rgb+d 120: A large-scale benchmark for 3d human activity understanding,

    J. Liu, A. Shahroudy, M. Perez, G. Wang, L.-Y. Duan, and A. C. Kot, “Ntu rgb+d 120: A large-scale benchmark for 3d human activity understanding,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 42, no. 10, pp. 2684–2701, 2020

  138. [150]

    Ntu rgb+d: A large scale dataset for 3d human activity analysis,

    A. Shahroudy, J. Liu, T.-T. Ng, and G. Wang, “Ntu rgb+d: A large scale dataset for 3d human activity analysis,” in Proc. IEEE Conf. Comput. Vis. Pattern Recognit., 2016, pp. 1010–1019

  139. [151]

    Action2motion: Condi- tioned generation of 3d human motions,

    C. Guo, X. Zuo, S. Wang, S. Zou, Q. Sun, A. Deng, M. Gong, and L. Cheng, “Action2motion: Condi- tioned generation of 3d human motions,” in Proc. ACM Int. Conf. Multimedia, 2020, pp. 2021–2029

  140. [152]

    3d human shape reconstruction from a polarization image,

    S. Zou, X. Zuo, Y. Qian, S. Wang, C. Xu, M. Gong, and L. Cheng, “3d human shape reconstruction from a polarization image,” in Proc. Eur. Conf. Comput. Vis., 2020, pp. 351–368. 42

  141. [153]

    Babel: bodies, action and behavior with english labels,

    A. R. Punnakkal, A. Chandrasekaran, N. Athanasiou, A. QuirosRamirez, and M. J. Black, “Babel: bodies, action and behavior with english labels,” in Proc. IEEE Conf. Comput. Vis. Pattern Recognit., 2021, pp. 722–731

  142. [154]

    A Unified 3D Human Motion Synthesis Model via Conditional Variational Auto- Encoder,

    Y. Cai et al., "A Unified 3D Human Motion Synthesis Model via Conditional Variational Auto- Encoder," 2021 IEEE/CVF International Conference on Computer Vision (ICCV), Montreal, QC, Canada, 2021, pp. 11625-11635

  143. [155]

    Zhou, Z., Wan, Y., & Wang, B. (2023). A unified framework for multimodal, multi-part human motion synthesis. arXiv preprint arXiv:2311.16471

  144. [156]

    AMASS: Archive of motion capture as surface shapes,

    N. Mahmood, N. Ghorbani, N. F. Troje, G. Pons-Moll, and M. J. Black, “AMASS: Archive of motion capture as surface shapes,” in Proc. Int. Conf. Comput. Vis., Oct. 2019, pp. 5441–5450

  145. [157]

    Generating diverse and natural 3d human motions from text,

    C. Guo, S. Zou, X. Zuo, S. Wang, W. Ji, X. Li, and L. Cheng, “Generating diverse and natural 3d human motions from text,” in Proc. IEEE Conf. Comput. Vis. Pattern Recognit., 2022, pp. 5152–5161

  146. [158]

    Human3.6m: Large scale datasets and predic- tive methods for 3d human sensing in natural environments,

    C. Ionescu, D. Papava, V. Olaru, and C. Sminchisescu, “Human3.6m: Large scale datasets and predic- tive methods for 3d human sensing in natural environments,” IEEE Trans. Pattern Anal. Mach. Intell., 2014

  147. [159]

    Cmu graphics lab motion capture database,

    J. Hodgins, “Cmu graphics lab motion capture database,” 2015

  148. [160]

    Humman: Multi-modal 4d human dataset for versatile sensing and modeling,

    Z. Cai, D. Ren, A. Zeng, Z. Lin, T. Yu, W. Wang, X. Fan, Y. Gao, Y. Yu, L. Pan, F. Hong, M. Zhang, C. C. Loy, L. Yang, and Z. Liu, “Humman: Multi-modal 4d human dataset for versatile sensing and modeling,” in Proc. Eur. Conf. Comput. Vis., October 2022

  149. [161]

    HUMANISE: Language-conditioned human motion generation in 3d scenes,

    Z. Wang, Y. Chen, T. Liu, Y. Zhu, W. Liang, and S. Huang, “HUMANISE: Language-conditioned human motion generation in 3d scenes,” in Proc. Adv. Neural Inform. Process. Syst., 2022

  150. [162]

    Robots learn social skills: End-to-end learning of co-speech gesture generation for humanoid robots,

    Y. Yoon, W.-R. Ko, M. Jang, J. Lee, J. Kim, and G. Lee, “Robots learn social skills: End-to-end learning of co-speech gesture generation for humanoid robots,” in Int. Conf. on Robot. and Automa., 2019, pp. 4303–4309

  151. [163]

    Speech gesture generation from the trimodal context of text, audio, and speaker identity,

    Y. Yoon, B. Cha, J.-H. Lee, M. Jang, J. Lee, J. Kim, and G. Lee, “Speech gesture generation from the trimodal context of text, audio, and speaker identity,” ACM Trans. Graph., vol. 39, no. 6, 2020

  152. [164]

    Style transfer for co-speech gesture animation: A multi-speaker conditionalmixture approach,

    C. Ahuja, D. W. Lee, Y. I. Nakano, and L.-P . Morency, “Style transfer for co-speech gesture animation: A multi-speaker conditionalmixture approach,” in Proc. Eur. Conf. Comput. Vis., 2020

  153. [165]

    Beat: A large-scale semantic and emotional multimodal dataset for conversational gestures synthesis,

    H. Liu, Z. Zhu, N. Iwamoto, Y. Peng, Z. Li, Y. Zhou, E. Bozkurt, and B. Zheng, “Beat: A large-scale semantic and emotional multimodal dataset for conversational gestures synthesis,” Proc. Eur. Conf. Comput. Vis., 2022

  154. [166]

    Implicit neural representations for variable length human motion generation,

    P . Cervantes, Y. Sekikawa, I. Sato, and K. Shinoda, “Implicit neural representations for variable length human motion generation,” in Proc. Eur. Conf. Comput. Vis., 2022, pp. 356–372

  155. [167]

    Multiact: Long-term 3d human motion generation from multiple action labels,

    T. Lee, G. Moon, and K. M. Lee, “Multiact: Long-term 3d human motion generation from multiple action labels,” in Proc. Assoc. Advance. Artif. Intell., 2023, pp. 1231–1239

  156. [168]

    Executing your commands via motion diffusion in latent space,

    C. Xin, B. Jiang, W. Liu, Z. Huang, B. Fu, T. Chen, J. Yu, and G. Yu, “Executing your commands via motion diffusion in latent space,” in Proc. IEEE Conf. Comput. Vis. Pattern Recognit., June 2023, pp. 18 000–18 010

  157. [169]

    Motionclip: Exposing human motion generation to clip space,

    G. Tevet, B. Gordon, A. Hertz, A. H. Bermano, and D. Cohen-Or, “Motionclip: Exposing human motion generation to clip space,” in Proc. Eur. Conf. Comput. Vis., 2022, pp. 358–374

  158. [170]

    Tm2t: Stochastic and tokenized modeling for the reciprocal generation of 3d human motions and texts,

    C. Guo, X. Zuo, S. Wang, and L. Cheng, “Tm2t: Stochastic and tokenized modeling for the reciprocal generation of 3d human motions and texts,” in Proc. Eur. Conf. Comput. Vis., 2022, pp. 580–597. 43

  159. [171]

    T2m-gpt: Generating human motion from textual descriptions with discrete representations,

    J. Zhang, Y. Zhang, X. Cun, Y. Zhang, H. Zhao, H. Lu, X. Shen, and Y. Shan, “T2m-gpt: Generating human motion from textual descriptions with discrete representations,” in Proc. IEEE Conf. Comput. Vis. Pattern Recognit., June 2023, pp. 14 730–14 740

  160. [172]

    Being comes from not-being: Open- vocabulary text-to-motion generation with wordless training,

    J. Lin, J. Chang, L. Liu, G. Li, L. Lin, Q. Tian, and C.-W. Chen, “Being comes from not-being: Open- vocabulary text-to-motion generation with wordless training,” in Proc. IEEE Conf. Comput. Vis. Pat- tern Recognit., June 2023, pp. 23 222–23 231

  161. [173]

    Ude: A unified driving engine for human motion generation,

    Z. Zhou and B. Wang, “Ude: A unified driving engine for human motion generation,” in Proc. IEEE Conf. Comput. Vis. Pattern Recognit., June 2023, pp. 5632–5641

  162. [174]

    Mofusion: A framework for denoising- diffusion-based motion synthesis,

    R. Dabral, M. H. Mughal, V. Golyanik, and C. Theobalt, “Mofusion: A framework for denoising- diffusion-based motion synthesis,” in Proc. IEEE Conf. Comput. Vis. Pattern Recognit., June 2023, pp. 97609770

  163. [175]

    Human motion diffusion model,

    G. Tevet, S. Raab, B. Gordon, Y. Shafir, D. Cohen-or, and A. H. Bermano, “Human motion diffusion model,” in Proc. Int. Conf. Learn. Represent., 2023

  164. [176]

    Action-conditioned 3D human motion synthesis with trans- former V AE,

    M. Petrovich, M. J. Black, and G. Varol, “Action-conditioned 3D human motion synthesis with trans- former V AE,” in Proc. Int. Conf. Comput. Vis., 2021

  165. [177]

    Actionconditioned on-demand motion generation,

    Q. Lu, Y. Zhang, M. Lu, and V. Roychowdhury, “Actionconditioned on-demand motion generation,” in Proc. ACM Int. Conf. Multimedia, 2022, pp. 2249–2257

  166. [178]

    Gans trained by a two time-scale update rule converge to a local nash equilibrium

    Heusel, M., Ramsauer, H., Unterthiner, T., Nessler, B., Hochreiter, S., 2017. Gans trained by a two time-scale update rule converge to a local nash equilibrium. Advances in neural information process- ing systems 30

  167. [179]

    Wang, H., Zhu, W., Miao, L., Xu, Y., Gao, F., Tian, Q., & Wang, Y. (2024). Aligning Human Motion Generation with Human Perceptions. arXiv preprint arXiv:2407.02272

  168. [180]

    (2023, December)

    Voas, J., Wang, Y., Huang, Q., & Mooney, R. (2023, December). What is the best automated metric for text to motion generation?. In SIGGRAPH Asia 2023 Conference Papers (pp. 1-11)

  169. [181]

    Qazi, A., & Iqbal, A. (2024). ExerAIde: AI-assisted Multimodal Diagnosis for Enhanced Sports Per- formance and Personalised Rehabilitation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 3430-3438)

  170. [182]

    Ekambaram, D., & Ponnusamy, V. (2023). AI-assisted Physical Therapy for Post-injury Rehabilitation: Current State of the Art. IEIE Transactions on Smart Processing & Computing, 12(3), 234-242

  171. [183]

    Emerging Role of Artificial Intelligence and Robotics in Physiotherapy: Past, Present, and Future Perspective

  172. [184]

    Kaur, J. (2024). FutureCare: AI Robots Revolutionizing Health and Healing. In Revolutionizing the Healthcare Sector with AI (pp. 311-340). IGI Global

  173. [185]

    J., & Wilken, J

    Darter, B. J., & Wilken, J. M. (2011). Gait training with virtual reality–based real-time feedback: improving gait performance following transfemoral amputation. Physical Therapy, 91(9), 1385-1394

  174. [186]

    T., Begg, R

    Lai, D. T., Begg, R. K., & Palaniswami, M. (2009). Computational intelligence in gait research: a per- spective on current applications and future challenges. IEEE Transactions on Information Technology in Biomedicine, 13(5), 687-702

  175. [187]

    A., Rodriguez, C., Frizera-Neto, A., Bastos-Filho, T

    Cifuentes, C. A., Rodriguez, C., Frizera-Neto, A., Bastos-Filho, T. F., & Carelli, R. (2014). Multimodal human–robot interaction for walker-assisted gait. IEEE Systems Journal, 10(3), 933-943

  176. [188]

    B., Hou, Z

    Cui, C., Bian, G. B., Hou, Z. G., Zhao, J., & Zhou, H. (2017). A multimodal framework based on inte- gration of cortical and muscular activities for decoding human intentions about lower limb motions. IEEE transactions on biomedical circuits and systems, 11(4), 889-899. 44

  177. [189]

    P ., Paul, D., & Baker, R

    Vakanski, A., Jun, H. P ., Paul, D., & Baker, R. (2018). A data set of human body movements for physical rehabilitation exercises. Data, 3(1), 2

  178. [190]

    Sun, L., Wang, Y., & Qin, W. (2024). A language-directed virtual human motion generation approach based on musculoskeletal models. Computer Animation and Virtual Worlds, 35(3), e2257

  179. [191]

    D., Johannes, M

    Katyal, K. D., Johannes, M. S., McGee, T. G., Harris, A. J., Armiger, R. S., Firpi, A. H., ... & Wester, B. A. (2013, November). HARMONIE: A multimodal control framework for human assistive robotics. In 2013 6th International IEEE/EMBS Conference on Neural Engineering (NER) (p...

  180. [192]

    & Chen, J

    Liu, Y., Cao, X., Chen, T., Jiang, Y., You, J., Wu, M., ... & Chen, J. (2025). A Survey of Embodied AI in Healthcare: Techniques, Applications, and Opportunities. arXiv preprint arXiv:2501.07468

  181. [193]

    Park, M., Cho, Y., Na, G., & Kim, J. (2024). Application of virtual avatar using motion capture in immersive virtual environment. International Journal of Human–Computer Interaction, 40(20), 6344- 6358

  182. [194]

    Ma, S. (2024). 3D Human Motion Recovery and Generation: From Noisy Pose to Language Descrip- tion (Doctoral dissertation)

  183. [195]

    Teaching Robots to Pre- dict Human Motion,

    L. -Y. Gui, K. Zhang, Y. -X. Wang, X. Liang, J. M. F. Moura and M. Veloso, "Teaching Robots to Pre- dict Human Motion,"2018 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Madrid, Spain, 2018, pp. 562-567, doi: 10.1109/IROS.2018.8594452

  184. [196]

    & Zhang, Y

    Fan, Z., Dai, P ., Su, Z., Gao, X., Lv, Z., Zhang, J., ... & Zhang, Y. (2024). Emhi: A multimodal egocentric human motion dataset with hmd and body-worn imus. arXiv preprint arXiv:2408.17168

  185. [197]

    Armanto, H., & Rosyid, H. A. (2024). Improved Non-Player Character (NPC) behavior using evolu- tionary algorithm—A systematic review. Entertainment Computing, 100875

  186. [198]

    J., Iyer, H., Jeong, H., & Guo, S

    Macwan, N., Hude, A. J., Iyer, H., Jeong, H., & Guo, S. (2024, September). High-fidelity worker motion simulation with generative AI. In Proceedings of the Human Factors and Ergonomics Society Annual Meeting (Vol. 68, No. 1, pp. 1540-1541). Sage CA: Los Angeles, CA: SAGE Publications

  187. [199]

    Zhao, J., Weng, D., Du, Q., & Tian, Z. (2024). Motion Generation Review: Exploring Deep Learning for Lifelike Animation with Manifold. arXiv preprint arXiv:2412.10458

  188. [201]

    Menapace, W., Siarohin, A., Lathuilière, S., Achlioptas, P ., Golyanik, V., Tulyakov, S., & Ricci, E. (2024). Promptable game models: Text-guided game simulation via masked diffusion models. ACM Transactions on Graphics, 43(2), 1-16

  189. [202]

    Yang, H., Li, C., Wu, Z., Li, G., Wang, J., Yu, J., ... & Xu, L. (2024). SMGDiff: Soccer Motion Generation using diffusion probabilistic models. arXiv preprint arXiv:2411.16216

  190. [203]

    & Wang, W

    Wang, J., Liu, Y., Dou, Z., Yu, Z., Liang, Y., Lin, C., ... & Wang, W. (2025). Disentangled clothed avatar generation from text descriptions. In European Conference on Computer Vision (pp. 381-401). Springer, Cham

  191. [204]

    (2024, March)

    Huang, Y., Yi, H., Xiu, Y., Liao, T., Tang, J., Cai, D., & Thies, J. (2024, March). Tech: Text-guided reconstruction of lifelike clothed humans. In 2024 International Conference on 3D Vision (3DV) (pp. 1531-1542). IEEE

  192. [205]

    ClothFit: Cloth-Human-Attribute Guided Virtual Try-on Network Using 3D Simulated Dataset,

    Y. Cho, L. S. S. Ray, K. S. P . Thota, S. Suh and P . Lukowicz, "ClothFit: Cloth-Human-Attribute Guided Virtual Try-on Network Using 3D Simulated Dataset," 2023 IEEE International Conference on Image Processing (ICIP), Kuala Lumpur, Malaysia, 2023, pp. 3484-3488. 45

  193. [206]

    Xintong Han, Zuxuan Wu, Zhe Wu, Ruichi Yu, and Larry S. Davis. 2018. Viton: An image-based vir- tual try-on network. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recog- nition. 7543–7552

  194. [207]

    Jia, Z., Zhang, Z., Wang, L., & Tan, T. (2024). Human image generation: A comprehensive survey. ACM Computing Surveys, 56(11), 1-39

  195. [208]

    (2009, October)

    Kim, S., Kim, C., You, B., & Oh, S. (2009, October). Stable whole-body motion generation for hu- manoid robots to imitate human motions. In 2009 IEEE/RSJ International Conference on Intelligent Robots and Systems (pp. 2518-2524). IEEE

  196. [209]

    Y., Zhang, K., Wang, Y

    Gui, L. Y., Zhang, K., Wang, Y. X., Liang, X., Moura, J. M., & Veloso, M. (2018, October). Teaching robots to predict human motion. In 2018 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) (pp. 562-567). IEEE

  197. [210]

    (2018, May)

    Bütepage, J., Kjellström, H., & Kragic, D. (2018, May). Anticipating many futures: Online human motion prediction and generation for human-robot interaction. In 2018 IEEE international conference on robotics and automation (ICRA) (pp. 4563-4570). IEEE

  198. [211]

    S., Islam, M

    Yasar, M. S., Islam, M. M., & Iqbal, T. (2024, March). PoseTron: Enabling Close-Proximity Human- Robot Collaboration Through Multi-human Motion Prediction. In Proceedings of the 2024 ACM/IEEE International Conference on Human-Robot Interaction (pp. 830-839)

  199. [212]

    S., & Baek, S

    Cha, J., Kim, J., Yoon, J. S., & Baek, S. (2024). Text2HOI: Text-guided 3D Motion Generation for Hand-Object Interaction. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pat- tern Recognition (pp. 1577-1585)

  200. [213]

    & Zhang, L

    Ren, T., Liu, S., Zeng, A., Lin, J., Li, K., Cao, H., ... & Zhang, L. (2024). Grounded sam: Assembling open-world models for diverse visual tasks. arXiv preprint arXiv:2401.14159

  201. [214]

    Kaur, N., Rani, S., & Kaur, S. (2024). Real-time video surveillance based human fall detection system using hybrid haar cascade classifier. Multimedia Tools and Applications, 1-19

  202. [215]

    (2024, July)

    Rangelov, D., Knotter, J., & Miltchev, R. (2024, July). 3D Reconstruction in Crime Scenes Inves- tigation: Impacts, Benefits, and Limitations. In Intelligent Systems Conference (pp. 46-64). Cham: Springer Nature Switzerland

  203. [216]

    Maksymowicz, K., Kuzan, A., & Tunikowski, W. (2024). 3D reconstruction of events: Search for a spatial correlation between injuries and the geometry of the body discovery site. Forensic Science International, 357, 111970

  204. [217]

    de Vette, V., Hutchinson, K., Mugge, W., Loeve, A., & van Zandwijk, J. P . (2024). Applicability of the Madymo Pedestrian Model for forensic fall analysis. Forensic science international, 112068

  205. [219]

    Ramesh, M., & Flohr, F. B. (2024, June). Walk-the-Talk: LLM driven pedestrian motion generation. In 2024 IEEE Intelligent Vehicles Symposium (IV) (pp. 3057-3062). IEEE

  206. [220]

    J., Peng, X

    Yi, H., Thies, J., Black, M. J., Peng, X. B., & Rempe, D. (2025). Generating human interaction motions in scenes with text control. In European Conference on Computer Vision (pp. 246-263). Springer, Cham

  207. [2024]

    Available: http://arxiv.org/abs/2405.17013

    [Online]. Available: http://arxiv.org/abs/2405.17013

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.