REVIEW 4 major objections 5 minor 1 cited by
Multimodal Generative AI with Autoregressive LLMs for Human Motion Understanding and Generation: A Way Forward
T0 review · 4 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read This survey claims to be the first to map all of human motion understanding and generation, organizing text-conditioned synthesis into autoregressive LLMs, diffusion models, GANs, VAEs, and unified AR-diffusion frameworks.
desk verdict A useful survey of text-to-motion with a broad scope, but its central comparison table contains impossible numbers and the paper is not trustworthy as printed. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The organizing device is an architecture taxonomy: autoregressive LLMs, which predict the next token and therefore need motion converted into discrete tokens by codebook-based vector quantizers such as VQ-VAE and VQ-GAN; diffusion models, which add and remove Gaussian noise in a forward-reverse process that ideally runs in a compressed latent space; and unified frameworks that align text and motion in a shared embedding space, using a multimodal transformer or a mixture-of-experts connector to let the autoregressive and diffusion paradigms cooperate. Each mechanism does a specific job in the survey: it explains why methods behave differently, since autoregressive models capture long-range text-motion dependencies but lose fine detail at tokenization, diffusion models produce high-fidelity motion at the cost of many denoising steps, and unified models aim to get both understanding and generation from a single training objective.
What would settle it
A reader can settle the uniqueness claim by checking the surveys listed in the paper's own Table 1: if any of them already covers all nine scope columns the paper counts for itself, the 'first and unique' assertion is false. The benchmark tables are independently falsifiable, since recomputing entries against the cited papers will expose an error whenever a reported Precision exceeds the [0,1] bounds the paper itself lists (for example, the 5.400 value in Table 6).
Extended reading notes
Core claim
In the authors' telling, the central discovery is that text-conditioned human motion generation has converged on two complementary paradigms—autoregressive LLMs that treat motion as a foreign language of discrete tokens (via VQ-VAE or VQ-GAN tokenizers) and diffusion models that refine noisy motion in continuous or latent space—and that recent work is combining them into unified frameworks. The survey's own contribution is the claim of uniqueness: no prior survey, it says, covers both motion understanding and generation together with multimodal generative AI and autoregressive LLMs, because previous reviews stop at prediction, at image or video generation, or at general generative models without LLMs. It supports this claim by categorizing methods by architecture and backbone, comparing them on the HumanML3D and KIT-ML benchmarks with fidelity, diversity, and consistency metrics, and mapping them onto applications from healthcare to autonomous driving.
Load-bearing premise
The load-bearing premise is that the comparison-table numbers copied from the cited papers are accurate and were measured under comparable protocols; if they are wrong or incomparable, the survey's value as a cross-method benchmark is compromised.
Editorial extensions
If this is right
- A newcomer to text-to-motion generation can select an architecture—autoregressive LLM, diffusion, GAN, VAE, or unified—directly from the survey's taxonomy, with matching datasets and metrics.
- Autoregressive LLM approaches that treat motion as discrete tokens are mature enough to generate, caption, retrieve, and reason about motion, not merely synthesize sequences.
- Unified AR+diffusion models trained in a shared embedding space with a multimodal transformer or mixture-of-experts connector are presented as the direction that combines understanding and generation.
- The absence of a standardized evaluation framework is identified as a real bottleneck, and the survey argues for a unified metric that combines fidelity, diversity, and consistency, citing work toward that goal.
- Real-world deployment in rehabilitation, VR/AR, gaming, robotics, surveillance, and autonomous vehicles depends on making motion generation efficient, controllable, and long-form, which the survey lists as the main open challenges.
Reading between the lines
- If discrete motion tokens become standard, general-purpose LLM tooling—instruction tuning, prompting, retrieval, and even adversarial red-teaming—would apply to motion almost directly, which is why the survey's unified AR+diffusion direction is a plausible successor to diffusion-only text-to-motion models.
- A concrete testable extension of the survey's call for a unified metric would be to build a single score that normalizes fidelity, diversity, and consistency across HumanML3D and KIT-ML and validates it against human perceptual judgments.
- The taxonomy implies that motion understanding and generation are converging: once the same tokenizer feeds both an LLM and a diffusion decoder, captioning and synthesis become two directions of one mapping, an idea the survey hints at but does not develop.
- The paper's comparison tables are best treated as a starting point to be checked against the original sources before being reused as a benchmark, since cross-method comparability is assumed rather than demonstrated.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This manuscript is a survey of text-conditioned human motion understanding and generation using multimodal generative AI and autoregressive large language models. It reviews data representations, text and motion tokenizers, LLM-based motion models, GAN/VAE/diffusion text-to-motion methods, unified frameworks, datasets, evaluation metrics, applications, and open challenges. The authors claim that this is the first survey to cover all areas of Human Motion Understanding and Generation (HMUG), and they support the survey with comparative tables, including a quantitative comparison of methods on HumanML3D and KIT-ML.
Significance. If the comparative tables and citation mappings were accurate, this survey would be a useful reference because it gathers a large body of recent work and organizes it by architectural family and task. The breadth is real: the paper reviews more than 250 references and provides useful taxonomy tables and figures that map models, backbones, datasets, and tasks. However, the survey's central value as a cross-method benchmark depends on Table 6 and on the traceability of Tables 2-4, and those components are currently not reliable as printed. The paper is a survey, so it contains no machine-checked proofs or code, but its factual claims about other papers must be verifiable against the cited sources; at present they are not.
major comments (4)
- [Table 5 and Table 6] Table 5 defines Precision with bounds 0≤Precision≤1 and MM Score with bounds 0≤MMS≤1, but Table 6 reports Precision values of 5.400 (AlertMotion), 2.550 (ActFormer), and 8.210 (UDE) on HumanML3D, and MM values above 1 in nearly every row. Since Table 6 is the only quantitative cross-method comparison in the survey, these internally inconsistent numbers cannot support the paper's comparative conclusions. The authors must re-extract every value from the original papers, state the metric definition used, and ensure consistency with Table 5.
- [Table 6 and reference list] Several rows in Table 6 do not point to the cited work. The row 'VQ-VAE Mot [103]' cites [103], Ghosh et al., 'Synthesis of Compositional Animations from Textual Descriptions,' which is not a VQ-VAE method; 'Unify MoGPT [146]' cites [146], Ma et al., which is the MoFusion diffusion paper; and 'WalkLLM [221]' cites [221], PackDiT, not the WalkLLM pedestrian-motion paper. In addition, there are two distinct papers named MotionLLM ([62] and [74]), and Table 6 lists MotionLLM [62] twice with different FID and Precision values. The reference-to-row mapping must be corrected throughout the tables.
- [Sections 6 and 7] Sections 6 and 7 have identical titles ('Datasets and Evaluation Metrics') and identical opening paragraphs, with Section 7 containing the actual subsections while Section 6 is left empty in substance. This duplication is a structural error that must be fixed by merging the content into a single section and renumbering the subsequent sections before the survey can be read as a coherent document.
- [Introduction and Table 1] The claim that this is 'the first, and unique attempt that covers all the areas of Human Motion Understanding and Generation (HMUG)' is not substantiated by Table 1 as presented. The checkmark criteria in Table 1 are undefined, and several prior surveys (e.g., [22] and [31]) overlap substantially with the stated scope. The authors should either provide an operational definition of 'all areas of HMUG' and demonstrate specific coverage gaps relative to each prior survey, or soften the uniqueness claim.
minor comments (5)
- [Equation (14)] Equation (14) is written as FID(P1,P2)^2 = ..., while the surrounding text refers to 'the FID score'; please adopt the standard convention FID = ||μ1-μ2||^2 + Tr(Σ1+Σ2-2(Σ1Σ2)^{1/2}) to avoid ambiguity.
- [Table 3] Table 3 lists 'ActFormer [95]' and 'FineMoGen [93]', but the cited references are the DCGAN paper and 'To Create What You Tell', respectively; the correct citations appear to be [92] and [89]. Please verify and correct these mappings.
- [Scope, Abstract and Section 7.3] The abstract and Section 1 state that the survey focuses exclusively on text and motion modalities, but Section 7.3 and Table 4 include audio, speech, music, and emotion datasets and methods; please reconcile the stated scope with the included material.
- [Table 2] Table 2 uses the overlapping names 'MotionGPT-3 [78]', 'MotionGPT-2 [137]', and 'T2M-GPT [77]' while the text and reference list use similar names for different papers; please adopt a consistent naming and citation scheme so that readers can map each row to its source.
- [Table 6 and evaluation protocols] The FID and other metric values in Table 6 are reported without the evaluation protocol used (e.g., number of samples, text prompt set, and whether metrics come from the original papers or from re-evaluation), so even after correcting the transcription errors the values may not be directly comparable across methods; please state the protocol explicitly.
Circularity Check
No significant circularity found: this paper is a review that derives no new predictions or fitted results, so the circularity patterns do not apply.
full rationale
The manuscript is a survey of text-conditioned human motion generation and understanding. It contains no derivation chain from fitted parameters to predictions, no claimed first-principles result, and no new empirical benchmark produced by the authors. The paper's stated contribution is organizational: it summarizes more than 250 prior works and provides comparative tables. A survey can be inaccurate, incomplete, or internally inconsistent, but those are correctness and reliability concerns, not circularity. The only quantitative synthesis is Table 6, which transcribes reported metrics from external papers. Even if some Precision values exceed 1 and conflict with Table 5's stated bounds, this would be a transcription or provenance error, not a case where the output is equivalent to the input by construction. The 'first and unique attempt' claim is a scope assertion, not a derivation. The paper does cite prior work by others, and there is no load-bearing self-citation chain or imported uniqueness theorem. Accordingly, the circularity score is 0.
Assumptions & free parameters
Cite this review
Pith. "Pith review of Multimodal Generative AI with Autoregressive LLMs for Human Motion Understanding and Generation: A Way Forward." pith.science (2026). https://pith.science/paper/FZVU5FFM
@misc{pith2026250603191,
author = {Pith},
title = {Pith review of: Multimodal Generative AI with Autoregressive LLMs for Human Motion Understanding and Generation: A Way Forward},
year = {2026},
howpublished = {\url{https://pith.science/paper/FZVU5FFM}},
note = {Machine review of arXiv:2506.03191}
}
read the original abstract
This paper presents an in-depth survey on the use of multimodal Generative Artificial Intelligence (GenAI) and autoregressive Large Language Models (LLMs) for human motion understanding and generation, offering insights into emerging methods, architectures, and their potential to advance realistic and versatile motion synthesis. Focusing exclusively on text and motion modalities, this research investigates how textual descriptions can guide the generation of complex, human-like motion sequences. The paper explores various generative approaches, including autoregressive models, diffusion models, Generative Adversarial Networks (GANs), Variational Autoencoders (VAEs), and transformer-based models, by analyzing their strengths and limitations in terms of motion quality, computational efficiency, and adaptability. It highlights recent advances in text-conditioned motion generation, where textual inputs are used to control and refine motion outputs with greater precision. The integration of LLMs further enhances these models by enabling semantic alignment between instructions and motion, improving coherence and contextual relevance. This systematic survey underscores the transformative potential of text-to-motion GenAI and LLM architectures in applications such as healthcare, humanoids, gaming, animation, and assistive technologies, while addressing ongoing challenges in generating efficient and realistic human motion.
Figures
Figures from the paper (7 more)
Forward citations
Cited by 1 Pith paper
-
IRG-MotionLLM: Interleaving Motion Generation, Assessment and Refinement for Text-to-Motion Generation
Interleaving motion generation with text-motion assessment and refinement improves alignment between generated human motion and goal text.
Reference graph
Works this paper leans on
-
[103]
Synthesis of Compositional Animations from Textual Descriptions
A. Ghosh, N. Cheema, C. Oguz, C. Theobalt, and P . Slusallek, "Synthesis of Compositional Anima- tions from Textual Descriptions," 2021, arXiv. doi: 10.48550/ARXIV.2103.14675
work page Pith review arXiv doi:10.48550/arxiv.2103.14675 2021
-
[146]
Pretrained Diffusion Models for Unified Human Motion Synthesis
J. Ma, S. Bai, and C. Zhou, “Pretrained Diffusion Models for Unified Human Motion Synthesis,” Dec. 06, 2022, arXiv: arXiv:2212.02837. doi: 10.48550/arXiv.2212.02837
work page Pith review arXiv doi:10.48550/arxiv.2212.02837 2022
-
[221]
PackDiT: Joint Human Motion and Text Generation via Mutual Prompting
Kiang, Z., Chai, W., Zhou, Z., Yang, C. Y., Huang, H. W., & Hwang, J. N. (2025). PackDiT: Joint Human Motion and Text Generation via Mutual Prompting.arXiv preprint arXiv:2501.16551. 46
work page Pith review arXiv 2025
-
[74]
MotionLLM: Understanding Human Behaviors from Human Motions and Videos,
L.-H. Chen et al., "MotionLLM: Understanding Human Behaviors from Human Motions and Videos," arXiv, 2024, doi: 10.48550/ARXIV.2405.20340
-
[22]
Multi-Modal Generative AI: Multi-modal LLM, Diffusion and Beyond,
H. Chen et al., “Multi-Modal Generative AI: Multi-modal LLM, Diffusion and Beyond,” 2024, arXiv. doi: 10.48550/ARXIV.2409.14993
-
[31]
Xiong, J., Liu, G., Huang, L., Wu, C., Wu, T., Mu, Y., & Wong, N. (2024). Autoregressive Models in Vision: A Survey. arXiv preprint arXiv:2411.05902
arXiv 2024
-
[1]
Towards artificial general intelligence via a multimodal foundation model
N. Fei et al., “Towards artificial general intelligence via a multimodal foundation model,” 2021, doi: 10.48550/ARXIV.2110.14378
work page Pith review arXiv doi:10.48550/arxiv.2110.14378 2021
-
[2]
Video generation models as world simulators OpenAI SORA
T. Brooks, B. Peebles, C. Holmes, W. DePue, Y. Guo, L. Jing, D. Schnurr, J. Taylor, T. Luhman, E. Luh- man, C. Ng, R. Wang, and A. Ramesh, “Video generation models as world simulators OpenAI SORA.” [Online]. Available: https://openai.com/research/video-generation-models-as-world-simulators
Show all 215 references
- [3]
-
[4]
DALL·E 3 understands significantly more nuance and detail than our previous systems, allowing you to easily translate your ideas into exceptionally accurate images
Jong Wook Kim, Alex Nichol, Yang Song, Lijuan Wang, Tao Xu, “DALL·E 3 understands significantly more nuance and detail than our previous systems, allowing you to easily translate your ideas into exceptionally accurate images.”
- [5]
- [6]
-
[7]
Deep Generative Modelling: A Compar- ative Review of V AEs, GANs, Normalizing Flows, Energy-Based and Autoregressive Models,
S. Bond-Taylor, A. Leach, Y. Long, and C. G. Willcocks, “Deep Generative Modelling: A Compar- ative Review of V AEs, GANs, Normalizing Flows, Energy-Based and Autoregressive Models,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 44, no. 11, pp. 7327–7347, Nov. 2022, doi: 10.11...
2022
-
[8]
Graph-based Normalizing Flow for Human Motion Generation and Reconstruction,
W. Yin, H. Yin, D. Kragic, and M. Bjorkman, “Graph-based Normalizing Flow for Human Motion Generation and Reconstruction,” in 2021 30th IEEE International Conference on Robot & Human Interactive Communication (RO-MAN), Vancouver, BC, Canada: IEEE, Aug. 2021, pp. 641–648. doi: ...
2021
-
[9]
BAMM: Bidirectional Autoregressive Motion Model,
E. Pinyoanuntapong, M. U. Saleem, P . Wang, M. Lee, S. Das, and C. Chen, “BAMM: Bidirectional Autoregressive Motion Model,” Apr. 01, 2024, arXiv: arXiv:2403.19435. Accessed: Sep. 22, 2024. [Online]. Available: http://arxiv.org/abs/2403.19435
2024 arXiv
- [10]
- [11]
-
[12]
A survey of GPT-3 family large language models including ChatGPT and GPT-4,
K. S. Kalyan, “A survey of GPT-3 family large language models including ChatGPT and GPT-4,” Nat. Lang. Process. J., vol. 6, p. 100048, Mar. 2024, doi: 10.1016/j.nlp.2023.100048
2024
-
[13]
ChatGPT vs Gemini vs LLaMA on Multilingual Sentiment Anal- ysis,
A. Buscemi and D. Proverbio, “ChatGPT vs Gemini vs LLaMA on Multilingual Sentiment Anal- ysis,” Jan. 25, 2024, arXiv: arXiv:2402.01715. Accessed: May 06, 2024. [Online]. Available: http://arxiv.org/abs/2402.01715
2024 arXiv
-
[14]
Text-to-Motion Retrieval: Towards Joint Un- derstanding of Human Motion Data and Natural Language,
N. Messina, J. Sedmidubsky, F. Falchi, and T. Rebok, “Text-to-Motion Retrieval: Towards Joint Un- derstanding of Human Motion Data and Natural Language,” in Proceedings of the 46th Interna- tional ACM SIGIR Conference on Research and Development in Information Retrieval, Jul. ...
2023
- [15]
- [16]
-
[17]
RIVQ-V AE: Discrete Rotation-Invariant 3D Rep- resentation Learning,
M. Mezghanni, M. Boulkenafed, and M. Ovsjanikov, “RIVQ-V AE: Discrete Rotation-Invariant 3D Rep- resentation Learning,” in 2024 International Conference on 3D Vision (3DV), Davos, Switzerland: IEEE, Mar. 2024, pp. 1382–1391. doi: 10.1109/3DV62453.2024.00129
2024
-
[18]
A Survey of Cross-Modal Visual Content Generation,
F. Nazarieh, Z. Feng, M. Awais, W. Wang, and J. Kittler, “A Survey of Cross-Modal Visual Content Generation,” IEEE Trans. Circuits Syst. Video Technol., vol. 34, no. 8, pp. 6814–6832, Aug. 2024, doi: 10.1109/TCSVT.2024.3351601
2024
-
[19]
Human Image Generation: A Comprehensive Survey,
Z. Jia, Z. Zhang, L. Wang, and T. Tan, “Human Image Generation: A Comprehensive Survey,” ACM Comput. Surv., vol. 56, no. 11, pp. 1–39, Nov. 2024, doi: 10.1145/3665869
2024 doi
-
[20]
3D human motion prediction: A survey,
K. Lyu, H. Chen, Z. Liu, B. Zhang, and R. Wang, “3D human motion prediction: A survey,” Neuro- computing, vol. 489, pp. 345–365, Jun. 2022, doi: 10.1016/j.neucom.2022.02.045
2022 doi
-
[21]
Human Motion Generation: A Survey,
W. Zhu et al., “Human Motion Generation: A Survey,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 46, no. 4, pp. 2430–2449, Apr. 2024, doi: 10.1109/TPAMI.2023.3330935
2024
-
[23]
Deep generative models on 3d representations: A survey
Shi, Zifan, et al. "Deep generative models on 3d representations: A survey." arXiv preprint arXiv:2210.15663 (2022)
2022 arXiv
-
[24]
Gavrila, D. M. (1999). The visual analysis of human movement: A survey. Computer vision and image understanding, 73(1), 82-98
1999
-
[25]
Wang, L., Hu, W., & Tan, T. (2003). Recent developments in human motion analysis. Pattern recogni- tion, 36(3), 585-601
2003
-
[26]
B., Hilton, A., & Krüger, V
Moeslund, T. B., Hilton, A., & Krüger, V. (2006). A survey of advances in vision-based human motion capture and analysis. Computer vision and image understanding, 104(2-3), 90-126
2006
-
[27]
Sedmidubsky, J., Elias, P ., Budikova, P ., & Zezula, P . (2021). Content-based management of human motion data: survey and challenges. IEEE Access, 9, 64241-64255
2021
-
[28]
Xue, H., Luo, X., Hu, Z., Zhang, X., Xiang, X., Dai, Y., & Yu, F. R. (2024). Human motion video genera- tion: A survey. Authorea Preprints
2024
-
[29]
Joshi, I., Grimmer, M., Rathgeb, C., Busch, C., Bremond, F., & Dantcheva, A. (2024). Synthetic data in human analysis: A survey. IEEE Transactions on Pattern Analysis and Machine Intelligence, 46(7), 4957-4976
2024
-
[30]
Liao, F., Zou, X., & Wong, W. (2024). Appearance and Pose-guided Human Generation: A Survey. ACM Computing Surveys, 56(5), 1-35
2024
-
[32]
Ismail-Fawaz, A., Devanne, M., Berretti, S., Weber, J., & Forestier, G. (2024). Establishing a Unified Evaluation Framework for Human Motion Generation: A Comparative Analysis of Metrics. arXiv preprint arXiv:2405.07680
2024 arXiv
-
[33]
Human Motion Prediction via Learning Local Structure Representations and Temporal Dependencies,
X. Guo and J. Choi, “Human Motion Prediction via Learning Local Structure Representations and Temporal Dependencies,” Proc. AAAI Conf. Artif. Intell., vol. 33, no. 01, pp. 2580–2587, Jul. 2019, doi: 10.1609/aaai.v33i01.33012580
2019 doi
-
[34]
RNN -based Human Motion Prediction via Differ- ential Sequence Representation,
Y. Wang, X. Wang, P . Jiang, and F. Wang, “RNN -based Human Motion Prediction via Differ- ential Sequence Representation,” in 2019 IEEE 6th International Conference on Cloud Comput- ing and Intelligence Systems (CCIS), Singapore: IEEE, Dec. 2019, pp. 138–143. doi: 10.1109/C- C...
2019
-
[35]
On Human Motion Prediction Using Recurrent Neural Net- works,
J. Martinez, M. J. Black, and J. Romero, “On Human Motion Prediction Using Recurrent Neural Net- works,” in 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI: IEEE, Jul. 2017, pp. 4674–4683. doi: 10.1109/CVPR.2017.497
2017 doi
- [36]
-
[37]
Efficient convolutional hierarchical autoencoder for human motion prediction,
Y. Li et al., “Efficient convolutional hierarchical autoencoder for human motion prediction,” Vis. Com- put., vol. 35, no. 6–8, pp. 1143–1156, Jun. 2019, doi: 10.1007/s00371-019-01692-9
2019 doi
- [38]
-
[39]
Aggregated Multi-GANs for Controlled 3D Human Motion Prediction,
Z. Liu, K. Lyu, S. Wu, H. Chen, Y. Hao, and S. Ji, “Aggregated Multi-GANs for Controlled 3D Human Motion Prediction,” Proc. AAAI Conf. Artif. Intell., vol. 35, no. 3, pp. 2225–2232, May 2021, doi: 10.1609/aaai.v35i3.16321
2021 doi
- [40]
- [42]
-
[43]
SMPL: a skinned multi-person linear model,
M. Loper, N. Mahmood, J. Romero, G. Pons-Moll, and M. J. Black, "SMPL: a skinned multi-person linear model," ACM Trans. Graph., vol. 34, no. 6, pp. 1–16, Nov. 2015, doi: 10.1145/2816795.2818013
2015
- [44]
- [45]
- [47]
-
[48]
Towards Natural and Accurate Future Motion Prediction of Humans and Animals,
Z. Liu et al., "Towards Natural and Accurate Future Motion Prediction of Humans and Animals," in 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA: IEEE, Jun. 2019, pp. 9996–10004. doi: 10.1109/CVPR.2019.01024
2019
-
[49]
Motion Prediction via Joint Dependency Modeling in Phase Space,
P . Su, Z. Liu, S. Wu, L. Zhu, Y. Yin, and X. Shen, "Motion Prediction via Joint Dependency Modeling in Phase Space," in Proceedings of the 29th ACM International Conference on Multimedia, Virtual Event China: ACM, Oct. 2021, pp. 713–721. doi: 10.1145/3474085.3475237
2021
-
[50]
Multiscale Spatio-Temporal Graph Neu- ral Networks for 3D Skeleton-Based Motion Prediction,
M. Li, S. Chen, Y. Zhao, Y. Zhang, Y. Wang, and Q. Tian, "Multiscale Spatio-Temporal Graph Neu- ral Networks for 3D Skeleton-Based Motion Prediction," IEEE Trans. Image Process., vol. 30, pp. 7760–7775, 2021, doi: 10.1109/TIP .2021.3108708
2021
-
[51]
Vedaldi, Computer Vision - ECCV 2020: 16th European Conference, Glasgow, UK, August 23-28, 2020, Proceedings, Part XIV
A. Vedaldi, Computer Vision - ECCV 2020: 16th European Conference, Glasgow, UK, August 23-28, 2020, Proceedings, Part XIV. in Lecture Notes in Computer Science Ser, no. v. 12359. Cham: Springer International Publishing AG, 2020. 36
2020
-
[53]
Dynamic Multiscale Graph Neural Net- works for 3D Skeleton Based Human Motion Prediction,
M. Li, S. Chen, Y. Zhao, Y. Zhang, Y. Wang, and Q. Tian, "Dynamic Multiscale Graph Neural Net- works for 3D Skeleton Based Human Motion Prediction," in 2020 IEEE/CVF Conference on Com- puter Vision and Pattern Recognition (CVPR), Seattle, WA, USA: IEEE, Jun. 2020, pp. 211–220....
2020
-
[54]
Learning Trajectory Dependencies for Human Motion Prediction,
W. Mao, M. Liu, M. Salzmann, and H. Li, "Learning Trajectory Dependencies for Human Motion Prediction," in 2019 IEEE/CVF International Conference on Computer Vision (ICCV), Seoul, Korea (South): IEEE, Oct. 2019, pp. 9488–9496. doi: 10.1109/ICCV.2019.00958
2019
-
[55]
TrajectoryCNN: A New Spatio-Temporal Feature Learning Network for Human Motion Prediction,
X. Liu, J. Yin, J. Liu, P . Ding, J. Liu, and H. Liu, "TrajectoryCNN: A New Spatio-Temporal Feature Learning Network for Human Motion Prediction," IEEE Trans. Circuits Syst. Video Technol., vol. 31, no. 6, pp. 2133–2146, Jun. 2021, doi: 10.1109/TCSVT.2020.3021409
2021
-
[56]
Structural-RNN: Deep Learning on Spatio-Temporal Graphs,
A. Jain, A. R. Zamir, S. Savarese, and A. Saxena, "Structural-RNN: Deep Learning on Spatio-Temporal Graphs," in 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA: IEEE, Jun. 2016, pp. 5308–5317. doi: 10.1109/CVPR.2016.573
2016 doi
-
[57]
Towards Accurate 3D Human Motion Prediction from Incomplete Observations,
Q. Cui and H. Sun, "Towards Accurate 3D Human Motion Prediction from Incomplete Observations," in 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA: IEEE, Jun. 2021, pp. 4799–4808. doi: 10.1109/CVPR46437.2021.00477
2021
-
[58]
Ferrari, Computer Vision - ECCV 2018: 15th European Conference, Munich, Germany, September 8-14, 2018, Proceedings, Part IV
V. Ferrari, Computer Vision - ECCV 2018: 15th European Conference, Munich, Germany, September 8-14, 2018, Proceedings, Part IV. in Lecture Notes in Computer Science Ser, no. v. 11208. Cham: Springer International Publishing AG, 2018
2018
-
[59]
word2vec, node2vec, graph2vec, X2vec: Towards a Theory of Vector Embeddings of Struc- tured Data,
M. Grohe, "word2vec, node2vec, graph2vec, X2vec: Towards a Theory of Vector Embeddings of Struc- tured Data," in Proceedings of the 39th ACM SIGMOD-SIGACT-SIGAI Symposium on Principles of Database Systems, Portland OR USA: ACM, Jun. 2020, pp. 1–16. doi: 10.1145/3375395.3387641
2020
-
[60]
Glove: Global Vectors for Word Representation,
J. Pennington, R. Socher, and C. Manning, "Glove: Global Vectors for Word Representation," in Pro- ceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), Doha, Qatar: Association for Computational Linguistics, 2014, pp. 1532–1543. doi: 10....
2014 doi
- [61]
-
[63]
MotionGPT: Finetuned LLMs Are General-Purpose Motion Generators,
Y. Zhang et al., "MotionGPT: Finetuned LLMs Are General-Purpose Motion Generators," Proc. AAAI Conf. Artif. Intell., vol. 38, no. 7, pp. 7368–7376, Mar. 2024, doi: 10.1609/aaai.v38i7.28567
2024 doi
-
[64]
Human3.6M: Large Scale Datasets and Pre- dictive Methods for 3D Human Sensing in Natural Environments,
C. Ionescu, D. Papava, V. Olaru, and C. Sminchisescu, "Human3.6M: Large Scale Datasets and Pre- dictive Methods for 3D Human Sensing in Natural Environments," IEEE Trans. Pattern Anal. Mach. Intell., vol. 36, no. 7, pp. 1325–1339, Jul. 2014, doi: 10.1109/TPAMI.2013.248
2014 doi
-
[65]
MoSh: motion and shape capture from sparse markers,
M. Loper, N. Mahmood, and M. J. Black, "MoSh: motion and shape capture from sparse markers," ACM Trans. Graph., vol. 33, no. 6, pp. 1–13, Nov. 2014, doi: 10.1145/2661229.2661273
2014
- [66]
- [67]
- [68]
- [69]
-
[70]
3D human pose estimation in video with tem- poral convolutions and semi-supervised training,
D. Pavllo, C. Feichtenhofer, D. Grangier, and M. Auli, "3D human pose estimation in video with tem- poral convolutions and semi-supervised training," arXiv, 2018, doi: 10.48550/ARXIV.1811.11742
-
[71]
Liu, J., Dai, W., Wang, C., Cheng, Y., Tang, Y., & Tong, X. (2023). Plan, posture and go: Towards open-world text-to-motion generation. arXiv preprint arXiv:2312.14828
2023 arXiv
-
[72]
Recovering 3D Human Mesh From Monocular Images: A Survey,
Y. Tian, H. Zhang, Y. Liu, and L. Wang, "Recovering 3D Human Mesh From Monocular Images: A Survey," IEEE Trans. Pattern Anal. Mach. Intell., vol. 45, no. 12, pp. 15406–15425, Dec. 2023, doi: 10.1109/TPAMI.2023.3298850
2023
-
[73]
H., Tan, H., Bansal, M., Rohrbach, A., Chang, K
Shen, S., Li, L. H., Tan, H., Bansal, M., Rohrbach, A., Chang, K. W., ... & Keutzer, K. (2021). How much can clip benefit vision-and-language tasks?.arXiv preprint arXiv:2107.06383
2021 arXiv
-
[75]
Shen, X., Li, X., & Elhoseiny, M. (2023). Mostgan-v: Video generation with temporal motion styles. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 5652- 5661)
2023
- [76]
-
[77]
MotionGPT: Human Motion as a Foreign Lan- guage,
B. Jiang, X. Chen, W. Liu, J. Yu, G. Yu, and T. Chen, "MotionGPT: Human Motion as a Foreign Lan- guage," arXiv, Jul. 19, 2023. Accessed: Sep. 22, 2024. [Online].http://arxiv.org/abs/2306.14795
2023 arXiv
-
[78]
MotionGPT: Human Motion Synthesis with Improved Diversity and Realism via GPT-3 Prompting,
J. Ribeiro-Gomes et al., "MotionGPT: Human Motion Synthesis with Improved Diversity and Realism via GPT-3 Prompting," in 2024 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), Waikoloa, HI, USA: IEEE, Jan. 2024, pp. 5058–5068, doi: 10.1109/WACV57701.2024.00499
2024
-
[79]
MotionScript: Natural Language Descriptions for Ex- pressive 3D Human Motions,
P . J. Yazdian, E. Liu, L. Cheng, and A. Lim, "MotionScript: Natural Language Descriptions for Ex- pressive 3D Human Motions," arXiv, Dec. 19, 2023. Accessed: Sep. 29, 2024. [Online]. Available: http://arxiv.org/abs/2312.12634
2023
-
[80]
T2M-GPT: Generating Human Motion from Textual Descriptions with Dis- crete Representations,
J. Zhang et al., "T2M-GPT: Generating Human Motion from Textual Descriptions with Dis- crete Representations," arXiv, Sep. 24, 2023. Accessed: Sep. 29, 2024. [Online]. Avail- able:http://arxiv.org/abs/2301.06052
2023 arXiv
- [81]
-
[82]
Autonomous LLM-Enhanced Adversarial Attack for Text-to-Motion,
H. Miao, F. Ma, R. Quan, K. Zhan, and Y. Yang, "Autonomous LLM-Enhanced Adversarial Attack for Text-to-Motion," arXiv, Aug. 01, 2024. Accessed: Sep. 22, 2024. [Online].http://arxiv.org/abs/ 2408.00352
2024 arXiv
-
[83]
Walk-the-Talk: LLM driven pedestrian motion generation,
M. Ramesh and F. B. Flohr, "Walk-the-Talk: LLM driven pedestrian motion generation," in 2024 IEEE Intelligent Vehicles Symposium (IV), Jeju Island, Korea, Republic of: IEEE, Jun. 2024, pp. 3057–3062, doi: 10.1109/IV55156.2024.10588860
2024
-
[84]
Motion Generation from Fine-grained Textual Descriptions,
K. Li and Y. Feng, "Motion Generation from Fine-grained Textual Descriptions," arXiv, Mar. 26, 2024. Accessed: Sep. 22, 2024. [Online].http://arxiv.org/abs/2403.13518
2024 arXiv
- [85]
-
[86]
(2024, September)
Shrestha, A., Liu, P ., Ros, G., Yuan, K., & Fern, A. (2024, September). Generating Physically Realis- tic and Directable Human Motions from Multi-modal Inputs. In European Conference on Computer Vision (pp. 1-17). Cham: Springer Nature Switzerland
2024
-
[87]
Generation of Walking Motions Based on Whole-Body Poses and QP Control,
R. Grimm, A. Kheddar and T. Asfour, "Generation of Walking Motions Based on Whole-Body Poses and QP Control," 2018 IEEE-RAS 18th International Conference on Humanoid Robots (Humanoids), Beijing, China, 2018, pp. 510-515, doi: 10.1109/HUMANOIDS.2018.8624913
2018
-
[88]
Luo, Z., Cao, J., Merel, J., Winkler, A., Huang, J., Kitani, K., & Xu, W. (2023). Universal humanoid motion representations for physics-based control. arXiv preprint arXiv:2310.04582
2023 arXiv
-
[89]
FineMoGen: Fine-Grained Spatio-Temporal Motion Generation and Editing,
M. Zhang, H. Li, Z. Cai, J. Ren, L. Yang, and Z. Liu, "FineMoGen: Fine-Grained Spatio-Temporal Motion Generation and Editing," arXiv
-
[90]
A., Dashtipour, K., Zahid, A., Abbasi, Q
Taylor, W., Shah, S. A., Dashtipour, K., Zahid, A., Abbasi, Q. H., & Imran, M. A. (2020). An Intel- ligent Non-Invasive Real-Time Human Activity Recognition System for Next-Generation Healthcare. Sensors, 20(9), 2653. https://doi.org/10.3390/s20092653
2020 doi
-
[91]
Generative adversarial networks,
I. Goodfellow et al., "Generative adversarial networks," Commun. ACM, vol. 63, no. 11, pp. 139–144, Oct. 2020, doi: 10.1145/3422622
2020 doi
- [92]
- [93]
-
[94]
IRC-GAN: Introspective Recurrent Convolutional GAN for Text-to-video Generation,
K. Deng, T. Fei, X. Huang, and Y. Peng, "IRC-GAN: Introspective Recurrent Convolutional GAN for Text-to-video Generation," in Proceedings of the Twenty-Eighth International Joint Conference on Ar- tificial Intelligence, Macao, China: International Joint Conferences on Artifici...
2019 doi
- [95]
- [96]
-
[97]
C., Azghadi, M
Huang, T., Liu, J., Zhou, X., Nguyen, D. C., Azghadi, M. R., Xia, Y., & Sun, S. (2023). V2X cooperative perception for autonomous driving: Recent advances and challenges. arXiv preprint arXiv:2310.03525
2023 arXiv
-
[98]
A Style-Based Generator Architecture for Generative Adversarial Networks,
T. Karras, S. Laine, and T. Aila, "A Style-Based Generator Architecture for Generative Adversarial Networks," in 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA: IEEE, Jun. 2019, pp. 4396–4405, doi: 10.1109/CVPR.2019.00453
2019
- [99]
-
[100]
Anatomically-Informed Vector Quantization Variational Auto-Encoder for Text to Motion Generation,
L. Chen, Z. Niu, Q. Liu, J. Wang, J. Xue, and K. Lu, "Anatomically-Informed Vector Quantization Variational Auto-Encoder for Text to Motion Generation," in 2024 IEEE International Conference on Multimedia and Expo Workshops (ICMEW), Niagara Falls, ON, Canada: IEEE, Jul. 2024, ...
2024
-
[101]
Human motion generation with StyleGAN,
K. Yamamoto and M. Murakami, "Human motion generation with StyleGAN," in Seventh Interna- tional Conference on Computer Graphics and Virtuality (ICCGV 2024), J. Li, Ed., Hangzhou, China: SPIE, May 2024, p. 6. doi: 10.1117/12.3029447
2024 doi
- [102]
-
[104]
TransCGan-based human motion generator,
W. Yu, "TransCGan-based human motion generator," in Fifth International Conference on Computer Information Science and Artificial Intelligence (CISAI 2022), Y. Zhong, Ed., Chongqing, China: SPIE, Mar. 2023, p. 188. doi: 10.1117/12.2668277
2022 doi
-
[105]
MSFF-GAN: A Refined Method for Human Video Motion Transfer,
Y. Xue and L. Chai, "MSFF-GAN: A Refined Method for Human Video Motion Transfer," in 2024 36th Chinese Control and Decision Conference (CCDC), Xi’an, China: IEEE, May 2024, pp. 4416–4421. doi: 10.1109/CCDC62350.2024.10587557
2024
-
[106]
Text-guided 3D Human Motion Generation with Keyframe-based Parallel Skip Transformer,
Z. Geng, C. Han, Z. Hayder, J. Liu, M. Shah, and A. Mian, "Text-guided 3D Human Motion Generation with Keyframe-based Parallel Skip Transformer," May 24, 2024, arXiv: arXiv:2405.15439. Accessed: Sep. 22, 2024. [Online]. Available: http://arxiv.org/abs/2405.15439
2024 arXiv
- [107]
-
[108]
Amballa, A., Akkinapalli, G., & Muralikrishnan, V. (2024). LS-GAN: Human Motion Synthesis with Latent-space GANs. arXiv preprint arXiv:2501.01449
2024 arXiv
- [109]
-
[110]
(2022, October)
Lu, Q., Zhang, Y., Lu, M., & Roychowdhury, V. (2022, October). Action-conditioned on-demand mo- tion generation. In Proceedings of the 30th ACM International Conference on Multimedia (pp. 2249- 2257)
2022
-
[111]
(2024, September)
Jin, P ., Li, H., Cheng, Z., Li, K., Yu, R., Liu, C., & Chen, J. (2024, September). Local action-guided motion diffusion model for text-to-motion generation. In European Conference on Computer Vision (pp. 392-409). Cham: Springer Nature Switzerland
2024
- [112]
- [114]
-
[115]
UDE: A Unified Driving Engine for Human Motion Generation,
Z. Zhou and B. Wang, "UDE: A Unified Driving Engine for Human Motion Generation," in 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vancouver, BC, Canada: IEEE, Jun. 2023, pp. 5632–5641. doi: 10.1109/CVPR52729.2023.00545
2023
- [116]
- [117]
- [118]
- [119]
- [120]
-
[121]
InfiniMotion: Mamba Boosts Memory in Transformer for Arbitrary Long Motion Generation,
Z. Zhang et al., "InfiniMotion: Mamba Boosts Memory in Transformer for Arbitrary Long Motion Generation," Jul. 13, 2024, arXiv: arXiv:2407.10061. Accessed: Sep. 22, 2024. [Online]. Available: http://arxiv.org/abs/2407.10061
2024 arXiv
-
[122]
T2M-HiFiGPT: Generating High Quality Human Motion from Textual Descriptions with Residual Discrete Representations,
C. Wang, "T2M-HiFiGPT: Generating High Quality Human Motion from Textual Descriptions with Residual Discrete Representations," 2023, arXiv. doi: 10.48550/ARXIV.2312.10628
-
[123]
ControlMM: Controllable Masked Motion Generation,
E. Pinyoanuntapong et al., "ControlMM: Controllable Masked Motion Generation," 2024, arXiv. doi: 10.48550/ARXIV.2410.10780
2024 doi
- [124]
-
[125]
AttT2M: Text-Driven Human Motion Generation with Multi- Perspective Attention Mechanism,
C. Zhong, L. Hu, Z. Zhang, and S. Xia, "AttT2M: Text-Driven Human Motion Generation with Multi- Perspective Attention Mechanism," Sep. 01, 2023, arXiv: arXiv:2309.00796. Accessed: Sep. 22, 2024. [Online]. Available: http://arxiv.org/abs/2309.00796
2023 arXiv
-
[126]
HumanTOMATO: Text-aligned Whole-body Motion Generation,
S. Lu et al., "HumanTOMATO: Text-aligned Whole-body Motion Generation," Oct. 19, 2023, arXiv: arXiv:2310.12978. Accessed: Sep. 22, 2024. [Online]. Available: http://arxiv.org/abs/2310.12978
2023 arXiv
- [127]
- [128]
-
[129]
MotionCLIP: Exposing Human Motion Generation to CLIP Space,
G. Tevet, B. Gordon, A. Hertz, A. H. Bermano, and D. Cohen-Or, "MotionCLIP: Exposing Human Motion Generation to CLIP Space," in Computer Vision – ECCV 2022, vol. 13682, S. Avidan, G. Bros- tow, M. Cissé, G. M. Farinella, and T. Hassner, Eds., in Lecture Notes in Computer Scien...
2022 doi
- [130]
- [131]
-
[132]
Guided Motion Diffusion for Controllable Human Motion Synthesis,
K. Karunratanakul, K. Preechakul, S. Suwajanakorn, and S. Tang, “Guided Motion Diffusion for Controllable Human Motion Synthesis,” in 2023 IEEE/CVF International Conference on Computer Vision (ICCV), Paris, France: IEEE, Oct. 2023, pp. 2151–2162. doi: 10.1109/ICCV51070.2023.00205
2023
- [133]
-
[134]
MotionFix: Text-Driven 3D Human Motion Editing,
N. Athanasiou, A. Ceske, M. Diomataris, M. J. Black, and G. Varol, “MotionFix: Text-Driven 3D Human Motion Editing,” Sep. 19, 2024, arXiv: arXiv:2408.00712. Accessed: Sep. 22, 2024. [Online]. Available: http://arxiv.org/abs/2408.00712
2024 arXiv
-
[135]
GUESS: GradUally Enriching SyntheSis for Text-Driven Human Motion Generation,
X. Gao, Y. Yang, Z. Xie, S. Du, Z. Sun, and Y. Wu, “GUESS: GradUally Enriching SyntheSis for Text-Driven Human Motion Generation,” IEEE Trans. Vis. Comput. Graph., pp. 1–13, 2024, doi: 10.1109/TVCG.2024.3352002
2024
- [136]
- [137]
-
[138]
Learning Generalizable Human Motion Generator with Reinforcement Learning,
Y. Mao, X. Liu, W. Zhou, Z. Lu, and H. Li, “Learning Generalizable Human Motion Generator with Reinforcement Learning,” May 24, 2024, arXiv: arXiv:2405.15541. Accessed: Sep. 29, 2024. [Online]. Available: http://arxiv.org/abs/2405.15541
2024 arXiv
-
[139]
AMD: Autoregressive Motion Diffusion,
B. Han, H. Peng, M. Dong, Y. Ren, Y. Shen, and C. Xu, “AMD: Autoregressive Motion Diffusion,” Proc. AAAI Conf. Artif. Intell., vol. 38, no. 3, pp. 2022–2030, Mar. 2024, doi: 10.1609/aaai.v38i3.27973
2022 doi
-
[140]
AAMDM: Accelerated Auto-Regressive Motion Diffusion Model,
T. Li, C. Qiao, G. Ren, K. Yin, and S. Ha, “AAMDM: Accelerated Auto-Regressive Motion Diffusion Model,” in 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA: IEEE, Jun. 2024, pp. 1813–1823. doi: 10.1109/CVPR52733.2024.00178
2024
-
[141]
Large Motion Model for Unified Multi-modal Motion Generation,
M. Zhang et al., “Large Motion Model for Unified Multi-modal Motion Generation,” in Computer Vision – ECCV 2024, vol. 15071, A. Leonardis, E. Ricci, S. Roth, O. Russakovsky, T. Sattler, and G. Varol, Eds., in Lecture Notes in Computer Science, vol. 15071. , Cham: Springer Natu...
2024 doi
-
[142]
M2D2M: Multi-Motion Generation from Text with Discrete Diffusion Models,
S. Chi et al., “M2D2M: Multi-Motion Generation from Text with Discrete Diffusion Models,” in Com- puter Vision – ECCV 2024, vol. 15072, A. Leonardis, E. Ricci, S. Roth, O. Russakovsky, T. Sattler, and G. Varol, Eds., in Lecture Notes in Computer Science, vol. 15072. , Cham: Sp...
2024 doi
-
[143]
L., Wu, W., Loy, C
Jiang, Y., Yang, S., Koh, T. L., Wu, W., Loy, C. C., & Liu, Z. (2023). Text2performer: Text-driven human video generation. In Proceedings of the IEEE/CVF International Conference on Computer Vision (pp. 22747-22757)
2023
-
[144]
& Yang, W
Huang, Z., Yu, Y., Yang, L., Qin, C., Zheng, B., Zheng, X., ... & Yang, W. (2024, October). Motion-aware latent diffusion models for video frame interpolation. In Proceedings of the 32nd ACM International Conference on Multimedia (pp. 1043-1052)
2024
- [145]
-
[147]
Plappert, M., Mandery, C., & Asfour, T. (2016). The kit motion-language dataset. Big data, 4(4), 236- 252
2016
-
[148]
T., & Zheng, W
Ji, Y., Xu, F., Yang, Y., Shen, F., Shen, H. T., & Zheng, W. S. (2018, October). A large-scale RGB-D database for arbitrary-view human action recognition. In Proceedings of the 26th ACM international Conference on Multimedia (pp. 1510-1518)
2018
-
[149]
Ntu rgb+d 120: A large-scale benchmark for 3d human activity understanding,
J. Liu, A. Shahroudy, M. Perez, G. Wang, L.-Y. Duan, and A. C. Kot, “Ntu rgb+d 120: A large-scale benchmark for 3d human activity understanding,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 42, no. 10, pp. 2684–2701, 2020
2020
-
[150]
Ntu rgb+d: A large scale dataset for 3d human activity analysis,
A. Shahroudy, J. Liu, T.-T. Ng, and G. Wang, “Ntu rgb+d: A large scale dataset for 3d human activity analysis,” in Proc. IEEE Conf. Comput. Vis. Pattern Recognit., 2016, pp. 1010–1019
2016
-
[151]
Action2motion: Condi- tioned generation of 3d human motions,
C. Guo, X. Zuo, S. Wang, S. Zou, Q. Sun, A. Deng, M. Gong, and L. Cheng, “Action2motion: Condi- tioned generation of 3d human motions,” in Proc. ACM Int. Conf. Multimedia, 2020, pp. 2021–2029
2020
-
[152]
3d human shape reconstruction from a polarization image,
S. Zou, X. Zuo, Y. Qian, S. Wang, C. Xu, M. Gong, and L. Cheng, “3d human shape reconstruction from a polarization image,” in Proc. Eur. Conf. Comput. Vis., 2020, pp. 351–368. 42
2020
-
[153]
Babel: bodies, action and behavior with english labels,
A. R. Punnakkal, A. Chandrasekaran, N. Athanasiou, A. QuirosRamirez, and M. J. Black, “Babel: bodies, action and behavior with english labels,” in Proc. IEEE Conf. Comput. Vis. Pattern Recognit., 2021, pp. 722–731
2021
-
[154]
A Unified 3D Human Motion Synthesis Model via Conditional Variational Auto- Encoder,
Y. Cai et al., "A Unified 3D Human Motion Synthesis Model via Conditional Variational Auto- Encoder," 2021 IEEE/CVF International Conference on Computer Vision (ICCV), Montreal, QC, Canada, 2021, pp. 11625-11635
2021
-
[155]
Zhou, Z., Wan, Y., & Wang, B. (2023). A unified framework for multimodal, multi-part human motion synthesis. arXiv preprint arXiv:2311.16471
2023 arXiv
-
[156]
AMASS: Archive of motion capture as surface shapes,
N. Mahmood, N. Ghorbani, N. F. Troje, G. Pons-Moll, and M. J. Black, “AMASS: Archive of motion capture as surface shapes,” in Proc. Int. Conf. Comput. Vis., Oct. 2019, pp. 5441–5450
2019
-
[157]
Generating diverse and natural 3d human motions from text,
C. Guo, S. Zou, X. Zuo, S. Wang, W. Ji, X. Li, and L. Cheng, “Generating diverse and natural 3d human motions from text,” in Proc. IEEE Conf. Comput. Vis. Pattern Recognit., 2022, pp. 5152–5161
2022
-
[158]
Human3.6m: Large scale datasets and predic- tive methods for 3d human sensing in natural environments,
C. Ionescu, D. Papava, V. Olaru, and C. Sminchisescu, “Human3.6m: Large scale datasets and predic- tive methods for 3d human sensing in natural environments,” IEEE Trans. Pattern Anal. Mach. Intell., 2014
2014
-
[159]
Cmu graphics lab motion capture database,
J. Hodgins, “Cmu graphics lab motion capture database,” 2015
2015
-
[160]
Humman: Multi-modal 4d human dataset for versatile sensing and modeling,
Z. Cai, D. Ren, A. Zeng, Z. Lin, T. Yu, W. Wang, X. Fan, Y. Gao, Y. Yu, L. Pan, F. Hong, M. Zhang, C. C. Loy, L. Yang, and Z. Liu, “Humman: Multi-modal 4d human dataset for versatile sensing and modeling,” in Proc. Eur. Conf. Comput. Vis., October 2022
2022
-
[161]
HUMANISE: Language-conditioned human motion generation in 3d scenes,
Z. Wang, Y. Chen, T. Liu, Y. Zhu, W. Liang, and S. Huang, “HUMANISE: Language-conditioned human motion generation in 3d scenes,” in Proc. Adv. Neural Inform. Process. Syst., 2022
2022
-
[162]
Robots learn social skills: End-to-end learning of co-speech gesture generation for humanoid robots,
Y. Yoon, W.-R. Ko, M. Jang, J. Lee, J. Kim, and G. Lee, “Robots learn social skills: End-to-end learning of co-speech gesture generation for humanoid robots,” in Int. Conf. on Robot. and Automa., 2019, pp. 4303–4309
2019
-
[163]
Speech gesture generation from the trimodal context of text, audio, and speaker identity,
Y. Yoon, B. Cha, J.-H. Lee, M. Jang, J. Lee, J. Kim, and G. Lee, “Speech gesture generation from the trimodal context of text, audio, and speaker identity,” ACM Trans. Graph., vol. 39, no. 6, 2020
2020
-
[164]
Style transfer for co-speech gesture animation: A multi-speaker conditionalmixture approach,
C. Ahuja, D. W. Lee, Y. I. Nakano, and L.-P . Morency, “Style transfer for co-speech gesture animation: A multi-speaker conditionalmixture approach,” in Proc. Eur. Conf. Comput. Vis., 2020
2020
-
[165]
Beat: A large-scale semantic and emotional multimodal dataset for conversational gestures synthesis,
H. Liu, Z. Zhu, N. Iwamoto, Y. Peng, Z. Li, Y. Zhou, E. Bozkurt, and B. Zheng, “Beat: A large-scale semantic and emotional multimodal dataset for conversational gestures synthesis,” Proc. Eur. Conf. Comput. Vis., 2022
2022
-
[166]
Implicit neural representations for variable length human motion generation,
P . Cervantes, Y. Sekikawa, I. Sato, and K. Shinoda, “Implicit neural representations for variable length human motion generation,” in Proc. Eur. Conf. Comput. Vis., 2022, pp. 356–372
2022
-
[167]
Multiact: Long-term 3d human motion generation from multiple action labels,
T. Lee, G. Moon, and K. M. Lee, “Multiact: Long-term 3d human motion generation from multiple action labels,” in Proc. Assoc. Advance. Artif. Intell., 2023, pp. 1231–1239
2023
-
[168]
Executing your commands via motion diffusion in latent space,
C. Xin, B. Jiang, W. Liu, Z. Huang, B. Fu, T. Chen, J. Yu, and G. Yu, “Executing your commands via motion diffusion in latent space,” in Proc. IEEE Conf. Comput. Vis. Pattern Recognit., June 2023, pp. 18 000–18 010
2023
-
[169]
Motionclip: Exposing human motion generation to clip space,
G. Tevet, B. Gordon, A. Hertz, A. H. Bermano, and D. Cohen-Or, “Motionclip: Exposing human motion generation to clip space,” in Proc. Eur. Conf. Comput. Vis., 2022, pp. 358–374
2022
-
[170]
Tm2t: Stochastic and tokenized modeling for the reciprocal generation of 3d human motions and texts,
C. Guo, X. Zuo, S. Wang, and L. Cheng, “Tm2t: Stochastic and tokenized modeling for the reciprocal generation of 3d human motions and texts,” in Proc. Eur. Conf. Comput. Vis., 2022, pp. 580–597. 43
2022
-
[171]
T2m-gpt: Generating human motion from textual descriptions with discrete representations,
J. Zhang, Y. Zhang, X. Cun, Y. Zhang, H. Zhao, H. Lu, X. Shen, and Y. Shan, “T2m-gpt: Generating human motion from textual descriptions with discrete representations,” in Proc. IEEE Conf. Comput. Vis. Pattern Recognit., June 2023, pp. 14 730–14 740
2023
-
[172]
Being comes from not-being: Open- vocabulary text-to-motion generation with wordless training,
J. Lin, J. Chang, L. Liu, G. Li, L. Lin, Q. Tian, and C.-W. Chen, “Being comes from not-being: Open- vocabulary text-to-motion generation with wordless training,” in Proc. IEEE Conf. Comput. Vis. Pat- tern Recognit., June 2023, pp. 23 222–23 231
2023
-
[173]
Ude: A unified driving engine for human motion generation,
Z. Zhou and B. Wang, “Ude: A unified driving engine for human motion generation,” in Proc. IEEE Conf. Comput. Vis. Pattern Recognit., June 2023, pp. 5632–5641
2023
-
[174]
Mofusion: A framework for denoising- diffusion-based motion synthesis,
R. Dabral, M. H. Mughal, V. Golyanik, and C. Theobalt, “Mofusion: A framework for denoising- diffusion-based motion synthesis,” in Proc. IEEE Conf. Comput. Vis. Pattern Recognit., June 2023, pp. 97609770
2023
-
[175]
Human motion diffusion model,
G. Tevet, S. Raab, B. Gordon, Y. Shafir, D. Cohen-or, and A. H. Bermano, “Human motion diffusion model,” in Proc. Int. Conf. Learn. Represent., 2023
2023
-
[176]
Action-conditioned 3D human motion synthesis with trans- former V AE,
M. Petrovich, M. J. Black, and G. Varol, “Action-conditioned 3D human motion synthesis with trans- former V AE,” in Proc. Int. Conf. Comput. Vis., 2021
2021
-
[177]
Actionconditioned on-demand motion generation,
Q. Lu, Y. Zhang, M. Lu, and V. Roychowdhury, “Actionconditioned on-demand motion generation,” in Proc. ACM Int. Conf. Multimedia, 2022, pp. 2249–2257
2022
-
[178]
Gans trained by a two time-scale update rule converge to a local nash equilibrium
Heusel, M., Ramsauer, H., Unterthiner, T., Nessler, B., Hochreiter, S., 2017. Gans trained by a two time-scale update rule converge to a local nash equilibrium. Advances in neural information process- ing systems 30
2017
-
[179]
Wang, H., Zhu, W., Miao, L., Xu, Y., Gao, F., Tian, Q., & Wang, Y. (2024). Aligning Human Motion Generation with Human Perceptions. arXiv preprint arXiv:2407.02272
2024 arXiv
-
[180]
(2023, December)
Voas, J., Wang, Y., Huang, Q., & Mooney, R. (2023, December). What is the best automated metric for text to motion generation?. In SIGGRAPH Asia 2023 Conference Papers (pp. 1-11)
2023
-
[181]
Qazi, A., & Iqbal, A. (2024). ExerAIde: AI-assisted Multimodal Diagnosis for Enhanced Sports Per- formance and Personalised Rehabilitation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 3430-3438)
2024
-
[182]
Ekambaram, D., & Ponnusamy, V. (2023). AI-assisted Physical Therapy for Post-injury Rehabilitation: Current State of the Art. IEIE Transactions on Smart Processing & Computing, 12(3), 234-242
2023
-
[183]
Emerging Role of Artificial Intelligence and Robotics in Physiotherapy: Past, Present, and Future Perspective
-
[184]
Kaur, J. (2024). FutureCare: AI Robots Revolutionizing Health and Healing. In Revolutionizing the Healthcare Sector with AI (pp. 311-340). IGI Global
2024
-
[185]
J., & Wilken, J
Darter, B. J., & Wilken, J. M. (2011). Gait training with virtual reality–based real-time feedback: improving gait performance following transfemoral amputation. Physical Therapy, 91(9), 1385-1394
2011
-
[186]
T., Begg, R
Lai, D. T., Begg, R. K., & Palaniswami, M. (2009). Computational intelligence in gait research: a per- spective on current applications and future challenges. IEEE Transactions on Information Technology in Biomedicine, 13(5), 687-702
2009
-
[187]
A., Rodriguez, C., Frizera-Neto, A., Bastos-Filho, T
Cifuentes, C. A., Rodriguez, C., Frizera-Neto, A., Bastos-Filho, T. F., & Carelli, R. (2014). Multimodal human–robot interaction for walker-assisted gait. IEEE Systems Journal, 10(3), 933-943
2014
-
[188]
B., Hou, Z
Cui, C., Bian, G. B., Hou, Z. G., Zhao, J., & Zhou, H. (2017). A multimodal framework based on inte- gration of cortical and muscular activities for decoding human intentions about lower limb motions. IEEE transactions on biomedical circuits and systems, 11(4), 889-899. 44
2017
-
[189]
P ., Paul, D., & Baker, R
Vakanski, A., Jun, H. P ., Paul, D., & Baker, R. (2018). A data set of human body movements for physical rehabilitation exercises. Data, 3(1), 2
2018
-
[190]
Sun, L., Wang, Y., & Qin, W. (2024). A language-directed virtual human motion generation approach based on musculoskeletal models. Computer Animation and Virtual Worlds, 35(3), e2257
2024
-
[191]
D., Johannes, M
Katyal, K. D., Johannes, M. S., McGee, T. G., Harris, A. J., Armiger, R. S., Firpi, A. H., ... & Wester, B. A. (2013, November). HARMONIE: A multimodal control framework for human assistive robotics. In 2013 6th International IEEE/EMBS Conference on Neural Engineering (NER) (p...
2013
-
[192]
& Chen, J
Liu, Y., Cao, X., Chen, T., Jiang, Y., You, J., Wu, M., ... & Chen, J. (2025). A Survey of Embodied AI in Healthcare: Techniques, Applications, and Opportunities. arXiv preprint arXiv:2501.07468
2025 arXiv
-
[193]
Park, M., Cho, Y., Na, G., & Kim, J. (2024). Application of virtual avatar using motion capture in immersive virtual environment. International Journal of Human–Computer Interaction, 40(20), 6344- 6358
2024
-
[194]
Ma, S. (2024). 3D Human Motion Recovery and Generation: From Noisy Pose to Language Descrip- tion (Doctoral dissertation)
2024
-
[195]
Teaching Robots to Pre- dict Human Motion,
L. -Y. Gui, K. Zhang, Y. -X. Wang, X. Liang, J. M. F. Moura and M. Veloso, "Teaching Robots to Pre- dict Human Motion,"2018 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Madrid, Spain, 2018, pp. 562-567, doi: 10.1109/IROS.2018.8594452
2018
-
[196]
& Zhang, Y
Fan, Z., Dai, P ., Su, Z., Gao, X., Lv, Z., Zhang, J., ... & Zhang, Y. (2024). Emhi: A multimodal egocentric human motion dataset with hmd and body-worn imus. arXiv preprint arXiv:2408.17168
2024
-
[197]
Armanto, H., & Rosyid, H. A. (2024). Improved Non-Player Character (NPC) behavior using evolu- tionary algorithm—A systematic review. Entertainment Computing, 100875
2024
-
[198]
J., Iyer, H., Jeong, H., & Guo, S
Macwan, N., Hude, A. J., Iyer, H., Jeong, H., & Guo, S. (2024, September). High-fidelity worker motion simulation with generative AI. In Proceedings of the Human Factors and Ergonomics Society Annual Meeting (Vol. 68, No. 1, pp. 1540-1541). Sage CA: Los Angeles, CA: SAGE Publications
2024
-
[199]
Zhao, J., Weng, D., Du, Q., & Tian, Z. (2024). Motion Generation Review: Exploring Deep Learning for Lifelike Animation with Manifold. arXiv preprint arXiv:2412.10458
2024 arXiv
-
[201]
Menapace, W., Siarohin, A., Lathuilière, S., Achlioptas, P ., Golyanik, V., Tulyakov, S., & Ricci, E. (2024). Promptable game models: Text-guided game simulation via masked diffusion models. ACM Transactions on Graphics, 43(2), 1-16
2024
-
[202]
Yang, H., Li, C., Wu, Z., Li, G., Wang, J., Yu, J., ... & Xu, L. (2024). SMGDiff: Soccer Motion Generation using diffusion probabilistic models. arXiv preprint arXiv:2411.16216
2024 arXiv
-
[203]
& Wang, W
Wang, J., Liu, Y., Dou, Z., Yu, Z., Liang, Y., Lin, C., ... & Wang, W. (2025). Disentangled clothed avatar generation from text descriptions. In European Conference on Computer Vision (pp. 381-401). Springer, Cham
2025
-
[204]
(2024, March)
Huang, Y., Yi, H., Xiu, Y., Liao, T., Tang, J., Cai, D., & Thies, J. (2024, March). Tech: Text-guided reconstruction of lifelike clothed humans. In 2024 International Conference on 3D Vision (3DV) (pp. 1531-1542). IEEE
2024
-
[205]
ClothFit: Cloth-Human-Attribute Guided Virtual Try-on Network Using 3D Simulated Dataset,
Y. Cho, L. S. S. Ray, K. S. P . Thota, S. Suh and P . Lukowicz, "ClothFit: Cloth-Human-Attribute Guided Virtual Try-on Network Using 3D Simulated Dataset," 2023 IEEE International Conference on Image Processing (ICIP), Kuala Lumpur, Malaysia, 2023, pp. 3484-3488. 45
2023
-
[206]
Xintong Han, Zuxuan Wu, Zhe Wu, Ruichi Yu, and Larry S. Davis. 2018. Viton: An image-based vir- tual try-on network. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recog- nition. 7543–7552
2018
-
[207]
Jia, Z., Zhang, Z., Wang, L., & Tan, T. (2024). Human image generation: A comprehensive survey. ACM Computing Surveys, 56(11), 1-39
2024
-
[208]
(2009, October)
Kim, S., Kim, C., You, B., & Oh, S. (2009, October). Stable whole-body motion generation for hu- manoid robots to imitate human motions. In 2009 IEEE/RSJ International Conference on Intelligent Robots and Systems (pp. 2518-2524). IEEE
2009
-
[209]
Y., Zhang, K., Wang, Y
Gui, L. Y., Zhang, K., Wang, Y. X., Liang, X., Moura, J. M., & Veloso, M. (2018, October). Teaching robots to predict human motion. In 2018 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) (pp. 562-567). IEEE
2018
-
[210]
(2018, May)
Bütepage, J., Kjellström, H., & Kragic, D. (2018, May). Anticipating many futures: Online human motion prediction and generation for human-robot interaction. In 2018 IEEE international conference on robotics and automation (ICRA) (pp. 4563-4570). IEEE
2018
-
[211]
S., Islam, M
Yasar, M. S., Islam, M. M., & Iqbal, T. (2024, March). PoseTron: Enabling Close-Proximity Human- Robot Collaboration Through Multi-human Motion Prediction. In Proceedings of the 2024 ACM/IEEE International Conference on Human-Robot Interaction (pp. 830-839)
2024
-
[212]
S., & Baek, S
Cha, J., Kim, J., Yoon, J. S., & Baek, S. (2024). Text2HOI: Text-guided 3D Motion Generation for Hand-Object Interaction. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pat- tern Recognition (pp. 1577-1585)
2024
-
[213]
& Zhang, L
Ren, T., Liu, S., Zeng, A., Lin, J., Li, K., Cao, H., ... & Zhang, L. (2024). Grounded sam: Assembling open-world models for diverse visual tasks. arXiv preprint arXiv:2401.14159
2024 arXiv
-
[214]
Kaur, N., Rani, S., & Kaur, S. (2024). Real-time video surveillance based human fall detection system using hybrid haar cascade classifier. Multimedia Tools and Applications, 1-19
2024
-
[215]
(2024, July)
Rangelov, D., Knotter, J., & Miltchev, R. (2024, July). 3D Reconstruction in Crime Scenes Inves- tigation: Impacts, Benefits, and Limitations. In Intelligent Systems Conference (pp. 46-64). Cham: Springer Nature Switzerland
2024
-
[216]
Maksymowicz, K., Kuzan, A., & Tunikowski, W. (2024). 3D reconstruction of events: Search for a spatial correlation between injuries and the geometry of the body discovery site. Forensic Science International, 357, 111970
2024
-
[217]
de Vette, V., Hutchinson, K., Mugge, W., Loeve, A., & van Zandwijk, J. P . (2024). Applicability of the Madymo Pedestrian Model for forensic fall analysis. Forensic science international, 112068
2024
-
[219]
Ramesh, M., & Flohr, F. B. (2024, June). Walk-the-Talk: LLM driven pedestrian motion generation. In 2024 IEEE Intelligent Vehicles Symposium (IV) (pp. 3057-3062). IEEE
2024
-
[220]
J., Peng, X
Yi, H., Thies, J., Black, M. J., Peng, X. B., & Rempe, D. (2025). Generating human interaction motions in scenes with text control. In European Conference on Computer Vision (pp. 246-263). Springer, Cham
2025
-
[2024]
Available: http://arxiv.org/abs/2405.17013
[Online]. Available: http://arxiv.org/abs/2405.17013
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.