Pith. sign in

REVIEW 4 major objections 5 minor 49 references

Human-Centric Foundation Models: Perception, Generation and Agentic Modeling

T0 review · 4 major / 5 minor · reviewed 2026-08-08 · deepseek-v4-flash

Pith's one-line read This survey proposes that human-centric foundation models divide into four families: perception, generation, unified perception-and-generation, and agentic models.

desk verdict Useful organizational survey of human-centric foundation models, but the first-survey claim and missing methodology keep it from being a fully trustworthy roadmap. read the letter →

arxiv 2502.08556 v1 pith:DRQF42MA submitted 2025-02-12 cs.CV cs.AIcs.LGcs.MM

classification cs.CVcs.AIcs.LGcs.MM
keywords human-centricfoundationmodelstaxonomyperceptionAIGCgenerationunifiedperception-generationagenticmultimodallargelanguagehumanoidembodiedAI
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Human-centric foundation models are generalist models trained to handle many tasks involving human bodies, faces, motion, and behavior within one framework. This paper proposes that those models fall into four families: perception, AI-generated content, unified perception-and-generation, and agentic models, grouped by the downstream tasks they support. The authors claim this is the first survey to organize the area with this taxonomy, and they support it by mapping representative methods, learning paradigms, and open challenges onto each family. The payoff, if the taxonomy holds, is a shared map that lets researchers compare models, spot missing capabilities, and choose starting points for new work.

What carries the argument

The load-bearing structure is the taxonomy itself, a two-level classification: four task-based families, each with paradigm-based subcategories. The organizing signal is downstream task support, so perception, generation, unified perception-generation, and agentic behavior are the four buckets. Subcategories include contrastive learning and masked image modeling under perception; GANs with style modulation and diffusion models under generation; fixed-vocabulary and extended-vocabulary LLM integration under unified models; and vision-language versus vision-language-action architectures under agents. The taxonomy does the work: it turns a scattered literature into a coordinate system for comparison.

What would settle it

Take the most recent 50 human-centric foundation models from top computer-vision and robotics venues and ask two independent human annotators to place each into exactly one of the four categories. If a substantial share cannot be placed, or if annotators disagree on which category a model belongs to, the taxonomy's exhaustiveness and clarity fail. A sharper check: find one model that performs perception, generation, and physical action control in one shared parameter set without treating any of these as a foreign language in an LLM; that model would sit outside all four categories as stated.

Watch

Extended reading notes

Core claim

The central claim is that the sprawling set of human-centric foundation models can be organized by what they do with people: perceiving them, generating them, doing both through a language-model hub, or acting in the world with human-like embodiment. Perception models learn fine-grained human representations for tasks like re-identification, parsing, pose estimation, mesh recovery, and action recognition; AIGC models synthesize human-focused images, videos, and avatars with high fidelity; unified models treat human-centric cues such as skeletons, body-model parameters, motion tokens, or audio as 'foreign languages' attached to large language models, so one model can both understand and generate; agentic models take vision, language, and other sensor signals and map them to humanoid motor behavior. The paper further splits each family by training paradigm and presents the four families as an interconnected roadmap for future digital-human and humanoid-embodiment research.

Load-bearing premise

The four-category taxonomy carries the whole survey, so it must be both exhaustive and fairly representative; the paper does not report a systematic search or selection protocol, so if a significant class of human-centric foundation models falls outside the four families or the representative list skews, the roadmap misleads readers.

Editorial extensions

If this is right

  • New or existing human-centric models can be located on the taxonomy by asking which downstream tasks they support, making comparisons and gap analysis more systematic.
  • The unified perception-and-generation family indicates that LLMs and multimodal LLMs are becoming the default hub, with human-centric signals treated as foreign languages.
  • The agentic family marks embodiment, interaction, and humanoid control as a frontier distinct from perception and generation.
  • The survey's challenges section implies that progress hinges on solving human-data scarcity, holistic body-face-hand representation, interactivity, and privacy ethics.
  • If the taxonomy is accepted, it can serve as a shared reference for structuring future human-centric foundation-model research and evaluation.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial extension: the four families are likely to blur, since a capable agentic model may soon include its own perceptual and generative components, so future taxonomies may need fewer or overlapping categories rather than four disjoint buckets.
  • Editorial extension: the fixed-versus-extended-vocabulary distinction predicts a testable trade-off—frozen-vocabulary tool use is cheaper and safer to extend, while grown vocabularies give finer control over human-centric outputs at higher training cost.
  • Editorial extension: a concrete way to stress-test the taxonomy is to classify the full set of recent humanoid robotics models; many may straddle the vision-language and vision-language-action split, suggesting the agentic category needs its own finer-grained subcategories.
  • Editorial extension: if the 'first survey' claim becomes influential, later papers should increasingly cite this taxonomy as their organizing frame, which is a checkable bibliographic prediction.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper presents a survey of human-centric foundation models (HcFMs), proposing a taxonomy that divides the field into four categories: (1) human-centric perception foundation models, (2) human-centric AIGC foundation models, (3) unified perception and generation models, and (4) human-centric agentic foundation models. For each category, the authors review representative methods, describe their underlying learning frameworks (e.g., contrastive learning, masked image modeling, diffusion transformers, LLM-integrated vocabularies), and discuss challenges and future directions in data, representation, interactivity, and ethics. The paper claims to be the first survey on HcFMs with a novel taxonomy and positions itself as a roadmap for researchers and practitioners working on digital humans and humanoid embodiments.

Significance. If the proposed taxonomy is accepted and its coverage is representative, the survey would provide a useful organizational structure for a rapidly growing but fragmented field. The paper includes several strengths: it gives clear conceptual distinctions among perception, generation, unified, and agentic models; it illustrates each category with framework diagrams (Figs. 2–5); it discusses both self-supervised and supervised paradigms within perception and generation; and it explicitly enumerates open challenges in data, representation, interactivity, and ethics. The survey also draws attention to the emerging area of humanoid agentic models. However, the significance of the survey as a 'roadmap' depends on two unverified premises: that no prior survey covers the same scope, and that the four-category taxonomy and the selected representative works fairly capture the field. These premises are not established in the manuscript, which limits the current contribution to a potentially useful but incomplete organizational proposal.

major comments (4)
  1. [Section 1 and References] The claim that 'this is the first survey about human-centric foundation models with a novel taxonomy' is not supported by evidence. The paper cites [Feng and others, 2024] ('Foundation Models for 3D Humans') in Section 1 but never states how the present taxonomy differs from or improves upon that prior survey. Without an explicit comparison to earlier surveys—including [Feng and others, 2024] and any other relevant overviews—the novelty claim remains an assertion. The authors should either provide such a comparison or soften the claim to avoid an unverifiable first-survey statement.
  2. [Section 2 and Fig. 1] The taxonomy is introduced with no inclusion criteria or search protocol. The text states that models are grouped 'according to their supported downstream tasks,' but it does not specify how works were identified, screened, or selected for inclusion. Fig. 1 lists only about 29 representative examples, which is far from the full landscape of human-centric foundation models. The absence of a completeness argument means the four categories cannot be verified as exhaustive, and the roadmap promise in the abstract is not justified. The authors should add a methodology subsection describing the literature search and selection process, and clearly state which works are representative rather than exhaustive.
  3. [Section 6 and Fig. 1] The agentic category is disproportionately thin compared to the other three categories. Section 6 discusses only HumanVLA, SuperPADL, and GR00T, and the text itself admits that applying VLA models to humanoid robotics 'remains largely unexplored.' If agentic modeling is meant to be a major fourth pillar of the taxonomy, the representative choice needs explicit justification, and the survey should explain why this category is included at the same level as perception, generation, and unified models despite its limited maturity. This is not a fatal flaw, but it weakens the claim of a balanced and exhaustive taxonomy.
  4. [Fig. 1 and Sections 3, 5] The selection of representative examples appears to be skewed toward works with which the authors are affiliated. For example, PATH (Tang et al., 2023), UniHCP (Ci et al., 2023), Hulk (Wang et al., 2023), MotionGPT (Jiang et al., 2023), and MotionGPT-2 (Wang et al., 2024b) are all associated with the authors of this survey, and several are given prominent placement in Fig. 1 and the main text. This is not inherently improper, but the survey provides no statement of how representatives were chosen or whether conflicts of interest influenced selection. A brief note on selection neutrality or a statement of author contributions to surveyed works would address the concern.
minor comments (5)
  1. [Throughout] There are numerous typographical and formatting errors: 'NeruIPS' for NeurIPS (reference [Jiang et al., 2023]), 'labled' instead of 'labeled' (Section 3.2), 'Yoshikawaet al.' missing space (reference [Yoshikawa et al., 2023]), and 'V ocabulary' with an extra space in the Fig. 4 caption. A careful proofreading pass is needed.
  2. [Section 2] The sentence 'The four categories of foundation models are interconnected rather than mutually exclusive' is in tension with the earlier statement that models are classified 'according to their supported downstream tasks.' The authors should clarify how a model that spans multiple categories (e.g., a unified perception-generation model that also has agentic capabilities) would be classified.
  3. [Section 3.1] In the text introducing contrastive learning methods, the phrase 'instead of commonly used momentum encoders, multiple encoders were used' is unclear. It is not obvious what is being contrasted with what; consider rewriting for clarity.
  4. [Fig. 2 caption] The caption says 'Parameters in modules with are used in downstream tasks,' which appears to be missing a symbol or a word. This obscures the intended meaning and should be fixed.
  5. [Section 7] The 'Ethics' paragraph mentions that 'anonymization methods should be applied to all training data,' but it does not discuss the trade-offs between anonymization and data utility, nor does it reference specific existing work on privacy-preserving human-centric learning. Adding a pointer to relevant literature would strengthen this discussion.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the survey's taxonomy is defined by supported downstream tasks, an external criterion, and the self-citations in the representative list are illustrative, not load-bearing.

full rationale

This is a survey paper, not a derivation or prediction paper. Its central claim is a taxonomy of human-centric foundation models into four categories, defined in Section 2 by 'supported downstream tasks' — an external, task-oriented criterion independent of the authors' own works. The survey does not fit parameters, derive equations, or make first-principles predictions that could reduce to its inputs. The self-citations (PATH, UniHCP, Hulk, MotionGPT-2, etc.) appear only as representative examples in Figure 1 and in the text; they are not used to justify the taxonomy's structure or to force any conclusion. The 'first survey with a novel taxonomy' claim is a novelty assertion that would require external comparison with prior surveys (e.g., Feng et al. 2024) to be verified, but a missing comparison is a completeness or correctness issue, not circularity. The paper even acknowledges the agentic category is 'largely unexplored,' disclosing its thinness. No step exhibits a definitional equivalence, a fitted parameter renamed as a prediction, or a load-bearing self-citation chain. Therefore, no circularity is present.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

No free parameters or invented physical entities appear because the paper is a survey. Its central contribution rests on taxonomy choices and on the representativeness of selected papers, both assumed rather than derived.

assumptions (3)
  • ad hoc to paper The four-category taxonomy (perception, AIGC, unified, agentic) is exhaustive and partitions all relevant human-centric foundation models.
    Introduced by the authors in Section 2 as the organizing principle of the survey; no systematic search or formal coverage proof is given.
  • domain assumption A 'foundation model' in this context must satisfy generalization, broad applicability, and high fidelity.
    Stated in Section 1 as the criteria for inclusion; other definitions based on pretraining scale or task coverage would shift the taxonomy.
  • domain assumption The cited representative examples are sufficiently characteristic of each category.
    Used throughout Sections 3-6; coverage and representativeness are asserted, not demonstrated by a selection protocol.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Human-Centric Foundation Models: Perception, Generation and Agentic Modeling." pith.science (2026). https://pith.science/paper/DRQF42MA

@misc{pith2026250208556,
  author       = {Pith},
  title        = {Pith review of: Human-Centric Foundation Models: Perception, Generation and Agentic Modeling},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/DRQF42MA}},
  note         = {Machine review of arXiv:2502.08556}
}
read the original abstract

Human understanding and generation are critical for modeling digital humans and humanoid embodiments. Recently, Human-centric Foundation Models (HcFMs) inspired by the success of generalist models, such as large language and vision models, have emerged to unify diverse human-centric tasks into a single framework, surpassing traditional task-specific approaches. In this survey, we present a comprehensive overview of HcFMs by proposing a taxonomy that categorizes current approaches into four groups: (1) Human-centric Perception Foundation Models that capture fine-grained features for multi-modal 2D and 3D understanding. (2) Human-centric AIGC Foundation Models that generate high-fidelity, diverse human-related content. (3) Unified Perception and Generation Models that integrate these capabilities to enhance both human understanding and synthesis. (4) Human-centric Agentic Foundation Models that extend beyond perception and generation to learn human-like intelligence and interactive behaviors for humanoid embodied tasks. We review state-of-the-art techniques, discuss emerging challenges and future research directions. This survey aims to serve as a roadmap for researchers and practitioners working towards more robust, versatile, and intelligent digital human and embodiments modeling.

Figures

Figures reproduced from arXiv: 2502.08556 by the authors.

Figure 1
Figure 1. A taxonomy of human-centric foundation models with representative examples. [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Different Frameworks of human-centric perception foundation models. Parameters in modules with [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Different frameworks of human-centric AIGC foundation models. Unsupervised learning methods: (a) A sampled noise is [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: Human-centric foundation models for unified perception [PITH_FULL_IMAGE:figures/full_fig_p006_4.png]
Figure 5
Figure 5. Figure 5: Frameworks of human-centric agentic foundation models [PITH_FULL_IMAGE:figures/full_fig_p007_5.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

49 extracted references · 37 canonical work pages

  1. [1]

    Gpt-4 technical report

    [Achiam et al., 2023] Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Alt- man, Shyamal Anadkat, et al. Gpt-4 technical report. arXiv:2303.08774,

  2. [4]

    A morphable model for the synthesis of 3d faces

    [Blanz and Vetter, 2023] V olker Blanz and Thomas Vetter. A morphable model for the synthesis of 3d faces. In Seminal Graphics Papers: Pushing the Boundaries

  3. [5]

    Instructpix2pix: Learning to follow image editing instructions

    [Brooks et al., 2023] Tim Brooks, Aleksander Holynski, and Alexei A Efros. Instructpix2pix: Learning to follow image editing instructions. In CVPR,

  4. [6]

    Smpler-x: Scaling up expressive human pose and shape estimation

    [Cai et al., 2024] Zhongang Cai, Wanqi Yin, Ailing Zeng, Chen Wei, Qingping Sun, Wang Yanjun, Hui En Pang, Haiyi Mei, Mingyuan Zhang, Lei Zhang, et al. Smpler-x: Scaling up expressive human pose and shape estimation. NeurIPS,

  5. [7]

    Sofgan: A portrait image generator with dynamic styling

    [Chen et al., 2022] Anpei Chen, Ruiyang Liu, Ling Xie, Zhang Chen, Hao Su, and Jingyi Yu. Sofgan: A portrait image generator with dynamic styling. TOG,

  6. [8]

    The language of motion: Uni- fying verbal and non-verbal language of 3d human motion

    [Chen et al., 2024] Changan Chen, Juze Zhang, Shrinidhi K Lakshmikanth, Yusu Fang, Ruizhi Shao, Gordon Wetzstein, Li Fei-Fei, and Ehsan Adeli. The language of motion: Uni- fying verbal and non-verbal language of 3d human motion. arXiv:2412.10523,

  7. [9]

    Unihcp: A unified model for human-centric perceptions

    [Ci et al., 2023] Yuanzheng Ci, Yizhou Wang, Meilin Chen, Shixiang Tang, Lei Bai, Feng Zhu, Rui Zhao, Fengwei Yu, Donglian Qi, and Wanli Ouyang. Unihcp: A unified model for human-centric perceptions. In CVPR,

  8. [10]

    Ag3d: Learning to generate 3d avatars from 2d image col- lections

    [Dong et al., 2023] Zijian Dong, Xu Chen, Jinlong Yang, Michael J Black, Otmar Hilliges, and Andreas Geiger. Ag3d: Learning to generate 3d avatars from 2d image col- lections. In CVPR,

Show all 49 references
  1. [11]

    Bringing robots home: The rise of ai robots in consumer electronics

    [Dong et al., 2024] Haiwei Dong, Yang Liu, Ted Chu, and Abdulmotaleb El Saddik. Bringing robots home: The rise of ai robots in consumer electronics. arXiv:2403.14449,

  2. [12]

    Foundation Models for 3D Humans,

    [Feng and others, 2024] Yao Feng et al. Foundation Models for 3D Humans,

  3. [13]

    Chatpose: Chatting about 3d human pose

    [Feng et al., 2024] Yao Feng, Jing Lin, Sai Kumar Dwivedi, Yu Sun, Priyanka Patel, and Michael J Black. Chatpose: Chatting about 3d human pose. In CVPR,

  4. [14]

    Stylegan-human: A data-centric odyssey of human generation

    [Fu et al., 2022] Jianglin Fu, Shikai Li, Yuming Jiang, Kwan- Yee Lin, Chen Qian, Chen Change Loy, Wayne Wu, and Ziwei Liu. Stylegan-human: A data-centric odyssey of human generation. In ECCV,

  5. [16]

    Instruct-reid: A multi- purpose person re-identification task with instructions

    [He et al., 2024] Weizhen He, Yiheng Deng, Shixiang Tang, Qihao Chen, Qingsong Xie, Yizhou Wang, Lei Bai, Feng Zhu, Rui Zhao, Wanli Ouyang, et al. Instruct-reid: A multi- purpose person re-identification task with instructions. In CVPR,

  6. [17]

    Versatile multi-modal pre-training for human-centric perception

    [Hong et al., 2022] Fangzhou Hong, Liang Pan, Zhongang Cai, and Ziwei Liu. Versatile multi-modal pre-training for human-centric perception. In CVPR,

  7. [18]

    Animate anyone: Consistent and control- lable image-to-video synthesis for character animation

    [Hu, 2024] Li Hu. Animate anyone: Consistent and control- lable image-to-video synthesis for character animation. In CVPR,

  8. [19]

    Refhcm: A unified model for referring perceptions in human-centric scenarios

    [Huang et al., 2024a] Jie Huang, Ruibing Hou, Jiahe Zhao, Hong Chang, and Shiguang Shan. Refhcm: A unified model for referring perceptions in human-centric scenarios. arXiv:2412.14643,

  9. [20]

    Motiongpt: Human motion as a foreign language

    [Jiang et al., 2023] Biao Jiang, Xin Chen, Wen Liu, Jingyi Yu, Gang Yu, and Tao Chen. Motiongpt: Human motion as a foreign language. NeruIPS,

  10. [21]

    You only learn one query: learning unified human query for single-stage multi-person multi-task human-centric perception

    [Jin et al., 2024] Sheng Jin, Shuhuai Li, Tong Li, Wentao Liu, Chen Qian, and Ping Luo. You only learn one query: learning unified human query for single-stage multi-person multi-task human-centric perception. In ECCV,

  11. [22]

    Humansd: A native skeleton-guided diffusion model for human image generation

    [Ju et al., 2023] Xuan Ju, Ailing Zeng, Chenchen Zhao, Jianan Wang, Lei Zhang, and Qiang Xu. Humansd: A native skeleton-guided diffusion model for human image generation. In CVPR,

  12. [23]

    Superpadl: Scaling language- directed physics-based control with progressive supervised distillation

    [Juravsky et al., 2024] Jordan Juravsky, Yunrong Guo, Sanja Fidler, and Xue Bin Peng. Superpadl: Scaling language- directed physics-based control with progressive supervised distillation. In SIGGRAPH,

  13. [24]

    Fashion- vdm: Video diffusion model for virtual try-on

    [Karras et al., 2024] Johanna Karras, Yingwei Li, Nan Liu, Luyang Zhu, Innfarn Yoo, Andreas Lugmayr, et al. Fashion- vdm: Video diffusion model for virtual try-on. In SIG- GRAPH Asia,

  14. [25]

    Sapiens: Foundation for human vision models

    [Khirodkar et al., 2024] Rawal Khirodkar, Timur Bagautdi- nov, Julieta Martinez, Su Zhaoen, Austin James, Peter Selednik, Stuart Anderson, and Shunsuke Saito. Sapiens: Foundation for human vision models. In ECCV,

  15. [26]

    Dreamhuman: Animatable 3d avatars from text

    [Kolotouros et al., 2023] Nikos Kolotouros, Thiemo Alldieck, Andrei Zanfir, Eduard Gabriel Bazavan, Mihai Fieraru, and Cristian Sminchisescu. Dreamhuman: Animatable 3d avatars from text. arXiv:2306.09329,

  16. [27]

    Blip-2: Bootstrapping language-image pre- training with frozen image encoders and large language models

    [Li et al., 2023] Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. Blip-2: Bootstrapping language-image pre- training with frozen image encoders and large language models. In ICML,

  17. [28]

    Unipose: A unified multi- modal framework for human pose comprehension, genera- tion and editing

    [Li et al., 2024c] Yiheng Li, Ruibing Hou, Hong Chang, Shiguang Shan, and Xilin Chen. Unipose: A unified multi- modal framework for human pose comprehension, genera- tion and editing. arXiv:2411.16781,

  18. [29]

    Motion-x: A large-scale 3d expressive whole-body human motion dataset

    [Lin et al., 2023] Jing Lin, Ailing Zeng, Shunlin Lu, Yuan- hao Cai, Ruimao Zhang, Haoqian Wang, and Lei Zhang. Motion-x: A large-scale 3d expressive whole-body human motion dataset. NeurIPS,

  19. [30]

    Chathuman: Language-driven 3d hu- man understanding with retrieval-augmented tool reasoning

    [Lin et al., 2024] Jing Lin, Yao Feng, Weiyang Liu, and Michael J Black. Chathuman: Language-driven 3d hu- man understanding with retrieval-augmented tool reasoning. arXiv:2405.04533,

  20. [31]

    Omnihuman-1: Rethinking the scaling-up of one- stage conditioned human animation models,

    [Lin et al., 2025] Gaojie Lin, Jianwen Jiang, Jiaqi Yang, and otherss. Omnihuman-1: Rethinking the scaling-up of one- stage conditioned human animation models,

  21. [33]

    Visual instruction tuning

    [Liu et al., 2024] Haotian Liu, Chunyuan Li, et al. Visual instruction tuning. NeurIPS,

  22. [34]

    M3-gpt: An advanced multimodal, multitask framework for motion comprehension and gener- ation

    [Luo et al., 2024] Mingshuang Luo, Ruibing Hou, Hong Chang, Zimo Liu, et al. M3-gpt: An advanced multimodal, multitask framework for motion comprehension and gener- ation. arXiv:2405.16273,

  23. [35]

    Efficient multi-modal human-centric contrastive pre-training with a pseudo body-structured prior

    [Meng et al., 2024] Yihang Meng, Hao Cheng, Zihua Wang, Hongyuan Zhu, Xiuxian Lao, and Yu Zhang. Efficient multi-modal human-centric contrastive pre-training with a pseudo body-structured prior. In PRCV,

  24. [36]

    Drag your gan: Interactive point-based manipulation on the generative image manifold

    [Pan et al., 2023] Xingang Pan, Ayush Tewari, Thomas Leimk¨uhler, et al. Drag your gan: Interactive point-based manipulation on the generative image manifold. In SIG- GRAPH,

  25. [37]

    360-degree human video generation with 4d diffusion transformer

    [Shao et al., 2024] Ruizhi Shao, Youxin Pang, Zerong Zheng, Jingxiang Sun, and Yebin Liu. 360-degree human video generation with 4d diffusion transformer. ACM Transac- tions on Graphics (TOG),

  26. [38]

    Humanbench: To- wards general human-centric perception with projector as- sisted pretraining

    [Tang et al., 2023] Shixiang Tang, Cheng Chen, Qingsong Xie, Meilin Chen, Yizhou Wang, Yuanzheng Ci, Lei Bai, Feng Zhu, Haiyang Yang, Li Yi, et al. Humanbench: To- wards general human-centric perception with projector as- sisted pretraining. In CVPR,

  27. [39]

    Llama: Open and efficient foundation language models

    [Touvron et al., 2023] Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timoth´ee Lacroix, Baptiste Rozi `ere, Naman Goyal, Eric Hambro, Faisal Azhar, et al. Llama: Open and efficient foundation language models. arXiv:2302.13971,

  28. [40]

    Hulk: A universal knowledge trans- lator for human-centric tasks

    [Wang et al., 2023] Yizhou Wang, Yixuan Wu, Shixiang Tang, Weizhen He, Xun Guo, Feng Zhu, Lei Bai, Rui Zhao, Jian Wu, Tong He, et al. Hulk: A universal knowledge trans- lator for human-centric tasks. arXiv:2312.01697,

  29. [41]

    Facegpt: Self- supervised learning to chat about 3d human faces

    [Wang et al., 2024a] Haoran Wang, Mohit Mendiratta, Chris- tian Theobalt, and Adam Kortylewski. Facegpt: Self- supervised learning to chat about 3d human faces. arXiv:2406.07163,

  30. [42]

    Motiongpt-2: A general- purpose motion-language model for motion generation and understanding

    [Wang et al., 2024b] Yuan Wang, Di Huang, Yaqi Zhang, Wanli Ouyang, Jile Jiao, Xuetao Feng, Yan Zhou, Pengfei Wan, Shixiang Tang, and Dan Xu. Motiongpt-2: A general- purpose motion-language model for motion generation and understanding. arXiv:2410.21747,

  31. [43]

    Aniportraitgan: animatable 3d portrait generation from 2d image collections

    [Wu et al., 2023] Yue Wu, Sicheng Xu, Jianfeng Xiang, Fangyun Wei, Qifeng Chen, Jiaolong Yang, and Xin Tong. Aniportraitgan: animatable 3d portrait generation from 2d image collections. In SIGGRAPH Asia,

  32. [44]

    Motionllm: Multimodal motion-language learning with large language models

    [Wu et al., 2024] Qi Wu, Yubo Zhao, Yifan Wang, Yu-Wing Tai, and Chi-Keung Tang. Motionllm: Multimodal motion-language learning with large language models. arXiv:2405.17013,

  33. [45]

    Get3dhuman: Lifting stylegan-human into a 3d generative model using pixel- aligned reconstruction priors

    [Xiong et al., 2023] Zhangyang Xiong, Di Kang, Derong Jin, Weikai Chen, Linchao Bao, et al. Get3dhuman: Lifting stylegan-human into a 3d generative model using pixel- aligned reconstruction priors. In ICCV,

  34. [46]

    Humanvla: Towards vision- language directed object rearrangement by physical hu- manoid

    [Xu et al., 2024] Xinyu Xu, Yizheng Zhang, Yong-Lu Li, Lei Han, and Cewu Lu. Humanvla: Towards vision- language directed object rearrangement by physical hu- manoid. arXiv:2406.19972,

  35. [47]

    Stylehumanclip: Text-guided garment manipulation for stylegan-human

    [Yoshikawaet al., 2023] Takato Yoshikawa, Yuki Endo, and Yoshihiro Kanamori. Stylehumanclip: Text-guided garment manipulation for stylegan-human. arXiv:2305.16759,

  36. [48]

    Hap: Structure-aware masked image modeling for human-centric perception

    [Yuan et al., 2024] Junkun Yuan, Xinyu Zhang, Hao Zhou, Jian Wang, et al. Hap: Structure-aware masked image modeling for human-centric perception. NeurIPS,

  37. [49]

    Avatargpt: All-in-one framework for motion un- derstanding planning generation and beyond

    [Zhou et al., 2024] Zixiang Zhou, Yu Wan, and Baoyuan Wang. Avatargpt: All-in-one framework for motion un- derstanding planning generation and beyond. In CVPR, 2024

  38. [2022]

    Unitedhuman: Har- nessing multi-source data for high-resolution human gener- ation

    [Fu et al., 2023] Jianglin Fu, Shikai Li, Yuming Jiang, Kwan- Yee Lin, Wayne Wu, and Ziwei Liu. Unitedhuman: Har- nessing multi-source data for high-resolution human gener- ation. In ICCV,

  39. [2023]

    Cross- view and cross-pose completion for 3d human understand- ing

    [Armando et al., 2024] Matthieu Armando, Salma Galaaoui, Fabien Baradel, Thomas Lucas, Vincent Leroy, Romain Br´egier, Philippe Weinzaepfel, and Gr´egory Rogez. Cross- view and cross-pose completion for 3d human understand- ing. In CVPR,

  40. [2024]

    Chatgarment: Garment estimation, generation and editing via large language models

    [Bian et al., 2024] Siyuan Bian, Chenghao Xu, Yuliang Xiu, Artur Grigorev, Zhen Liu, Cewu Lu, Michael J Black, and Yao Feng. Chatgarment: Garment estimation, generation and editing via large language models. arXiv:2412.17811,

  41. [2025]

    Hyperhuman: Hyper- realistic human generation with latent structural diffusion

    [Liu et al., 2023] Xian Liu, Jian Ren, Aliaksandr Siarohin, Ivan Skorokhodov, Yanyu Li, Dahua Lin, Xihui Liu, Zi- wei Liu, and Sergey Tulyakov. Hyperhuman: Hyper- realistic human generation with latent structural diffusion. arXiv:2310.08579,

Pith tools

Reviewed August 8, 2026 · model on record in the stance chip above.