Pith. sign in

REVIEW 4 major objections 5 minor 5 cited by

A parametric body model built from artist-defined shape prototypes rather than 3D scans can replace scan-trained models in human mesh recovery, the paper claims, achieving 2.4 mm scan-fitting accuracy and state-of-the-art HMR results.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

A scan-free, interpretable body model built from MakeHuman artist assets matches scan-trained SMPL-X models for human mesh recovery and scan fitting.

T0 review reviewed 2026-08-03 challenge →

load-bearing objection An open, scan-free body model that actually works for HMR, with a shape-coverage caveat the authors acknowledge but then overclaim past. the 4 major comments →

arxiv 2511.03589 v3 pith:BWWKZOWW submitted 2025-11-05 cs.CV

Human Mesh Modeling for Anny Body

classification cs.CV
keywords human mesh recoveryparametric body modelscan-free 3D human modelingphenotype parametersblendshape interpolationsynthetic training databody shape diversityWHO anthropometric calibration
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Anny is a fully differentiable, open-source body model whose shape space is defined by interpretable phenotype parameters (age, gender, height, weight, muscle, and local traits) via piecewise-multilinear blendshape interpolation, with distributions calibrated to WHO growth and BMI statistics. The paper argues that this scan-free, procedural design is sufficient for real tasks: it fits adult scans to 2.4 mm mean point-to-mesh error, fits child scans, and when used to generate an 800k-image synthetic dataset (Anny-One), trains HMR models that reach state-of-the-art accuracy on benchmarks including children—matching or beating SMPL-X, a scan-trained baseline. If true, this collapses the assumption that high-fidelity body models require proprietary, demographically narrow scan collections, and makes a single unified model covering the full lifespan feasible.

Core claim

The paper's central claim is that a body model constructed entirely from open, artist-authored anthropometric knowledge—with no 3D scan data—is sufficient for modern human mesh recovery. Anny encodes shape through continuous phenotype parameters in [0,1] that interpolate between prototype meshes; WHO calibration makes the parameter distribution reflect global population statistics. The authors show that Anny registers to adult scans at 2.4 mm mean point-to-mesh error (excluding head/hands), that it fits children's scans, and that HMR models trained with Anny match or outperform the same networks trained with SMPL-X on 3DPW and EHF, improve on AGORA especially for children, and reach state-of

What carries the argument

The load-bearing mechanism is piecewise-multilinear interpolation between a set of prototypical blendshape meshes, each corresponding to a phenotype corner (e.g., 'female baby, small muscle, average weight, large height'). Scalar phenotype parameters in [0,1] weight these prototypes, producing topologically consistent meshes while keeping the shape space interpretable and human-readable; a statistical layer calibrates the parameter distributions to WHO age-gender-height-weight/BMI data. Mesh deformation is completed by forward kinematics and blend skinning (implemented for autodiff), and linear regressors map Anny meshes to existing topologies (SMPL-X among others) with mean cyclic error 3.2

Load-bearing premise

The load-bearing premise is that the artist-designed prototype shapes from the open community character-modeling project truly cover real human morphology across ages and body types—WHO calibration adjusts the statistical distribution of the parameters but cannot create shape variation the prototypes lack.

What would settle it

Fit Anny to a diverse set of 3D scans spanning populations and body types well outside the design range of the community project (e.g., global regions, elderly, extremely tall or short) and compare the residuals; if the mean error substantially exceeds the 2.4 mm reported on 3DBodyTex, or if the residual variance after projecting real scans onto Anny's shape space is large, the phenotype space is missing real morphologies and the sufficiency claim fails.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • Training HMR models no longer requires collecting or licensing 3D body scans; open synthetic data suffices for state-of-the-art accuracy.
  • One model covers the full human lifespan, removing the need for separate child-specific body models in the reported benchmarks.
  • The semantic phenotype parameters give users direct control over height, weight, age, and local traits, enabling controlled synthetic data generation and shape editing.
  • The 800k-image Anny-One dataset, paired with Anny, yields state-of-the-art multi-person HMR results on standard benchmarks.
  • The 2.4 mm scan-fitting result indicates the procedural shape space is also precise enough for geometry-oriented applications, not just recognition.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If the sufficiency claim holds, the bottleneck in body modeling shifts from scan collection to the quality and coverage of the procedural shape space; a testable prediction is that HMR accuracy will scale with the number and diversity of phenotype prototypes.
  • WHO calibration only fixes the first two moments of the parameter distribution; an extension would validate the full joint distribution against real anthropometric surveys and add explicit covariates (e.g., ethnicity, disability) that the current phenotypes do not encode.
  • Because Anny's parameters are meaningful, the same model could support controllable aging simulation, virtual try-on, or ergonomic analysis—applications where abstract latent spaces are hard to steer.
  • Since the paper reports a ~3.2 mm cyclic mapping error to existing mesh topologies, benchmark numbers mix model error with mapping error; evaluating on native Anny ground truth would sharpen conclusions.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper introduces Anny, a differentiable, scan-free parametric human body model built from MakeHuman's artist-authored phenotype blendshapes, with a piecewise-multilinear shape parameterization and a distribution calibrated to WHO anthropometric statistics. The authors also introduce Anny-One, a synthetic dataset of 800k photorealistic images with corresponding Anny annotations. They evaluate Anny by registering it to adult and child scans and by training HMR2.0 and Multi-HMR with Anny versus SMPL-X under several training-data regimes. The paper reports that Anny is comparable or superior to SMPL-X in some settings and that a Multi-HMR model trained on Anny-One plus standard data achieves state-of-the-art results on 3DPW, EMDB, Hi4D, and CMU-Toddler. It concludes that Anny can serve as a drop-in replacement for SMPL-X and that scan-free body models are sufficient for HMR.

Significance. If the claims hold, this is a significant contribution: an open, interpretable body model that avoids proprietary scan data, a large synthetic dataset for HMR, and an empirical demonstration that a non-scan-based shape space can compete with SMPL-X in downstream tasks. The release of the model and code under Apache 2.0 is a clear strength, as is the direct comparison with SMPL-X under the same training data in Table 1. However, the current validation is incomplete: the 'without real images' claim is contradicted by the SOTA comparison setup, the body-model and dataset effects are confounded in Table 2, and the 'globally representative' claim rests on a narrow set of validation scans. These gaps must be addressed before the paper's central conclusions can be accepted.

major comments (4)
  1. [Section 2 (Related Work) and Section 6.3 (Table 3)] The paper states: 'we ... obtain SOTA performance without using real images for training' (Section 2, p.3). However, the SOTA comparison in Table 3 is for Multi-HMR trained on a mixture that explicitly includes MS-COCO and MPII, which are real-image datasets (Section 6.3: 'Finally we train using both Anny-One and standards HMR training data [28] including BEDLAM [10], MS-COCO [33], and MPII [7]'). Thus the claim that Anny-One alone enables SOTA without real images is not supported by the provided evidence. Please report results for Anny-One-only training on the Table 3 benchmarks, and remove or carefully qualify the 'without real images' statement.
  2. [Section 6.3, Table 2] The comparison that supports the 'drop-in replacement' claim is confounded. In the last three rows of Table 2, pre-training data (BEDLAM vs Anny-One) and body model (SMPL-X, SMPL-X+A, Anny) change simultaneously. The observed gains attributed to 'Anny-One+Anny' could come from the more child-diverse synthetic dataset rather than from Anny itself. To isolate the effect of the body model, the authors should include cross-ablation rows, e.g., Anny trained on BEDLAM and SMPL-X trained on Anny-One, both with and without AGORA fine-tuning. Without this, the conclusion that Anny is a sufficient drop-in replacement for SMPL-X is not established.
  3. [Sections 3, 4 and 5] The paper claims that Anny is 'globally representative' (Introduction) and provides 'demographically grounded' shape variation (Abstract). Yet the direct shape-coverage validation is limited to 3DBodyTex (400 Western adults) and three child scans. The WHO calibration in Section 4 only matches means and standard deviations of height/weight/BMI; it cannot create new shape dimensions absent from the MakeHuman prototype blendshapes. The authors themselves caution that phenotypes encode artists' preconceptions and 'should not be expected to faithfully encode any identity-related characteristics.' On the evidence presented, the coverage of non-Western, elderly, or extreme body types is untested. Please provide a direct evaluation on a more diverse set of scans or an analysis of which phenotypes are activated when fitting to such scans, or clearly scope the representativeness claims to the popu
  4. [Section 6.3, Tables 1 and 3] No error bars or multiple-seed results are reported. Several key differences are small (e.g., 86.5 vs 86.0 mm MPJPE for HMR2.0 in Table 1) and could be within run-to-run noise. The SOTA claims in Table 3 would be considerably more convincing with at least three seeds and mean±std reporting. This is particularly load-bearing because the paper argues that Anny matches or outperforms SMPL-X; without uncertainty measures, the equivalence claim is not statistically grounded.
minor comments (5)
  1. [Section 4] The 'empirically defined bijective mapping between the age parameter of Anny and some morphological age in years' is not described. Please provide the mapping, its derivation, or a reference.
  2. [Section 6.1] The text says body shapes are sampled from a distribution 'derived from WHO population data (Section 5)', but the statistical modeling appears in Section 4. Please correct the cross-reference.
  3. [Section 3 (Interoperability)] The Anny-to-SMPL-X regressor has a mean cyclic error of 3.2 mm. The paper does not discuss how this regressor error affects evaluation on benchmarks with SMPL-X ground truth, particularly PVE metrics. A brief analysis would be helpful.
  4. [Table 3] The CMU-Toddler results show very large errors (MPJPE 102.1 vs 153.6 for Multi-HMR) without discussion. Consider adding a note on why errors are substantially higher on this benchmark.
  5. [Figure 8] The runtime plot uses logarithmic axes without labeling them as such. Please clarify the axes and ensure the reader can interpret the scaling.

Circularity Check

0 steps flagged

No significant circularity: Anny is validated on external benchmarks and WHO statistics; the mesh regressor and fitting residuals are not reused as predictions.

full rationale

The derivation chain is externally grounded rather than self-referential. The Anny shape space is built from MakeHuman's open, artist-authored blendshapes and piecewise-multilinear interpolation, not from Anny's own outputs or from the HMR benchmarks. The WHO calibration in Section 4 fits only means and standard deviations of height/weight/BMI by age and gender, which are external statistics and do not encode the 3DPW, EHF, AGORA, EMDB, Hi4D, or CMU-Toddler metrics that the paper later reports. The HMR evaluations therefore compare against independent ground truth. The Anny-to-SMPL-X regressor (3.2 mm cyclic error) is learned by mesh-to-mesh fitting for topology conversion only; it is not trained on any benchmark outcome, so it does not force the reported accuracy. The 2.4 mm scan-fit error on 3DBodyTex is a registration residual, presented as fitting capacity rather than as a held-out prediction, so it is not a fitted value disguised as a prediction. Section 3's 'Word of caution' openly states that phenotype labels encode artist preconceptions and should not be read as faithful identity characteristics; this is a coverage limitation that bears on the 'globally representative' claim, but it is not a circularity. The self-citations to Multi-HMR [9] and Condimen [50] are used as published, code-released tools and metric references, not as the justification for Anny's shape space or performance; under the review rules they count as real evidence. No equation in the paper makes a claimed output equal to its own input by construction. The overall circularity score is therefore low, at most reflecting minor non-load-bearing self-citation rather than any substantive circular step.

Axiom & Free-Parameter Ledger

4 free parameters · 5 axioms · 0 invented entities

Anny and Anny-One are engineered artifacts, not new physical or conceptual entities. The model's freedom comes from inherited MakeHuman blendshapes and hand-fitted WHO calibration, not from a new force, particle, dimension, or conserved quantity.

free parameters (4)
  • Age-to-years mapping = not specified
    Section 4: 'We empirically define a bijective mapping between the age parameter of Anny and some morphological age in years.' This hand-defined mapping is needed for WHO calibration; its exact form is not reported.
  • Beta distribution parameters for phenotype prior = not specified
    Section 4: distributions are 'jointly calibrated to match the mean and standard deviation of reference growth standards and body mass index from WHO'. Parameters are fitted to WHO summary statistics but not listed in the paper.
  • MakeHuman prototype blendshape meshes and 267 phenotype controls = artist-defined, not given
    Section 3: Anny's shape space is built from MakeHuman's procedural prototypes. These inherited artist choices are free parameters of the model, neither learned nor independently validated in this paper.
  • Sparse linear regressors Anny-SMPL-X and Anny-HumGen3D = 3.2 mm cyclic error (SMPL-X), 1.7 mm (HumGen3D)
    Section 3, Interoperability: regression matrices R are optimized by mesh fitting; the 3.2 mm cyclic error is part of the evaluation chain and is not propagated into HMR benchmark errors.
axioms (5)
  • domain assumption MakeHuman artist-defined morphology is an adequate proxy for real human shape variation.
    Sections 2-3: central premise of scan-free modeling; explicitly acknowledged as artistic stereotypes in the 'Word of caution'. If artist prototypes miss real morphologies, scan fitting and HMR fail.
  • domain assumption Piecewise-multilinear interpolation between prototypes yields topologically consistent and plausible human meshes for all parameter combinations.
    Section 3: 'This use of interpolation constraints the structure of the shape space and helps producing topologically consistent meshes'; no proof that off-grid combinations remain plausible.
  • domain assumption WHO mean and standard deviation statistics are sufficient to characterize a global population shape distribution.
    Section 4: calibration only matches first two moments of WHO curves and BMI; higher-order shape correlations are not constrained.
  • domain assumption The learned Anny-to-SMPL-X regressor error (3.2 mm) does not materially change HMR benchmark comparisons.
    Section 3, Interoperability, used in Section 6.2: evaluation of Anny HMR is done after converting to SMPL-X topology; the 3.2 mm cyclic error is not propagated or corrected in benchmark numbers.
  • standard math Standard linear blend skinning and forward kinematics, implemented in PyTorch/Warp, are correct and differentiable.
    Section 3, 'Differentiable deformation': relies on well-established skinning and inverse-kinematics machinery.

reviewed 2026-08-03 · how reviews work

0 comments
Cite this review

Pith. "Pith review of Human Mesh Modeling for Anny Body." pith.science (2026). https://pith.science/paper/BWWKZOWW

@misc{pith2026251103589,
  author       = {Pith},
  title        = {Pith review of: Human Mesh Modeling for Anny Body},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/BWWKZOWW}},
  note         = {Machine review of arXiv:2511.03589}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Parametric body models provide the structural basis for many human-centric tasks, yet existing models often rely on costly 3D scans and learned shape spaces that are proprietary and demographically narrow. We introduce Anny, a simple, fully differentiable, and scan-free human body model grounded in anthropometric knowledge from the MakeHuman community. Anny defines a continuous, interpretable shape space, where phenotype parameters (e.g. gender, age, height, weight) control blendshapes spanning a wide range of human forms---across ages (from infants to elders), body types, and proportions. Calibrated using WHO population statistics, Anny provides realistic and demographically grounded human shape variation within a single unified model. We release the Anny body model and its code under the Apache 2.0 license. Thanks to its openness and semantic control, Anny serves as a versatile foundation for 3D human modeling---supporting millimeter-accurate scan fitting, controlled synthetic data generation, and Human Mesh Recovery (HMR). We further introduce Anny-One, a collection of 780k photorealistic images generated with Anny, showing that despite its simplicity, HMR models trained with Anny can match the performance of those trained with scan-based body models.

Figures

Figures reproduced from arXiv: 2511.03589 by Fabien Baradel, Gr\'egory Rogez, Gu\'enol\'e Fiche, Laura Bravo-S\'anchez, Matthieu Armando, Philippe Weinzaepfel, Romain Br\'egier, Thomas Lucas.

Figure 1
Figure 1. Figure 1: Anny is a unified, open and interpretable human parametric body model aiming to capture the diversity of human shapes and ages, from infants to elders. Abstract Parametric body models provide the structural basis for many human-centric tasks, yet existing models often rely on costly 3D scans and learned shape spaces that are proprietary and demographically narrow. We introduce Anny, a simple, fully differe… view at source ↗
Figure 3
Figure 3. Figure 3: Example of local morphological variations covered by the model. large-scale synthetic dataset of 800k photorealistic hu￾mans generated with Anny, featuring expressive full￾body poses, hands, and faces across diverse environ￾ments. HMR models trained with Anny-One achieve competitive accuracy on standard benchmarks and out￾perform existing approaches when body-shape diversity is high — using a single parame… view at source ↗
Figure 4
Figure 4. Figure 4: Mesh and skeleton of Anny. Left: default mesh and skeleton, modeling two different morpholo￾gies. Middle: coarser mesh (1,229 vertices) using only a skeleton subset. Right: use of the same skeleton as a Mixamo [6] character for pose re-targeting. obtain SOTA performance without using real images for training. 3. Modeling Anny Body We propose Anny, a differentiable parametric mesh model aiming at modeling a… view at source ↗
Figure 5
Figure 5. Figure 5: Model calibration. Calibration of Anny morphological shape distribution with WHO Child Growth standards [39] (boys). that purpose we define a mapping to regress SMPL-X vertices from Anny meshes and vice versa. The sec￾ond one is to generate synthetic 3D scenes with Anny annotations. Our synthetic data generation pipeline is built on Humgen3D [5], thus we also learn regres￾sors for this body model. More spe… view at source ↗
Figure 7
Figure 7. Figure 7: Random samples from the Anny-One dataset. between 30° and 130° to capture a wide range of spa￾tial compositions. Body poses are randomly sampled from AMASS [36], while hand poses are independently drawn from GRAB [57]. Body shapes are sampled from a statistical distribution of human phenotypes derived from WHO population data (Section 5). To ensure phys￾ical plausibility, we apply a self-collision check to… view at source ↗
Figure 8
Figure 8. Figure 8: Runtime performance of a body model for￾ward pass, measured on a NVIDIA A100 GPU. images with human scans from 3DPEOPLE constitute the new validation set composed of 2k images, ensuring an equal percentage of children in both training and validation splits. 6.3. HMR Results To isolate the effect of our body model from that of our synthetic dataset, we first evaluate Anny by re-training existing HMR methods… view at source ↗
Figure 9
Figure 9. Figure 9: Qualitative examples on real-world images sourced from Pexels [3] that contain both adults and children. We observe that our model outputs childlike body shapes for children, in terms of height and build. Model 3DPW EMDB Hi4D CMU-Toddler PA-MPJPE MPJPE PVE PA-MPJPE MPJPE PVE PA-MPJPE MPJPE Pair-PA-MPJPE MPJPE Pair-PA-MPJPE AiOS [54] 45.0 68.8 90.9 63.3 90.6 108.1 49.9 71.4 234.2 162.4 723.1 SAT-HMR [53] 52… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. ResiHMR: Residual-Limb Aware Single-Image 3D Human Mesh Recovery for Individuals with Limb Loss

    cs.CV 2026-04 unverdicted novelty 8.0

    ResiHMR is the first single-image system to explicitly reconstruct residual-limb surfaces and perform topology-adaptive optimization for people with limb loss.

  2. Anny-Fit: All-Age Human Mesh Recovery

    cs.CV 2026-05 unverdicted novelty 7.0

    Anny-Fit jointly optimizes all-age multi-person 3D human meshes in camera coordinates using complementary signals from off-the-shelf depth, segmentation, keypoint, and VLM networks, yielding better reprojection, depth...

  3. Direct Clinical Joint Angle Extraction from Parametric Body Model Rotation Matrices

    cs.CV 2026-07 conditional novelty 6.0

    A swing-twist decomposition plus a per-body-model calibration table converts body-model rotation matrices into clinical joint angles at 4.50° MAE on OpenCap LabValidation.

  4. PRISM: A 3D Probabilistic Neural Representation for Interpretable Shape Modeling

    cs.LG 2026-02 reject novelty 5.0

    PRISM learns a neural Gaussian field of shape displacements conditioned on age and uses Fisher information to estimate local temporal uncertainty, but the Cramér–Rao justification confuses latent age with an estimator.

  5. Comparing Commercial Depth Sensor Accuracy for Medical Applications

    cs.RO 2026-06 unverdicted novelty 3.0

    Zivid 2M+ 60 outperformed Intel RealSense D405, PMD Flexx2, and Stereolabs ZED 2i on all tested specimens and metrics; ZED ranked second on real tissue but last on the phantom.

Reference graph

Works this paper leans on

69 extracted references · 3 linked inside Pith · cited by 5 Pith papers

  1. [1]

    http://www.makehumancommunity

    Makehuman. http://www.makehumancommunity. org/. 2, 3, 4

  2. [2]

    https://github.com/ makehumancommunity/mpfb2

    MPFB2. https://github.com/ makehumancommunity/mpfb2. 4

  3. [3]

    Pexels.https://www.pexels.com. 8

  4. [4]

    https://www.tc2.com/size-usa.html, 2017

    SizeUSA. https://www.tc2.com/size-usa.html, 2017. 5

  5. [5]

    https://www.humgen3d.com/, 2025

    HumGen3D. https://www.humgen3d.com/, 2025. 3, 5

  6. [6]

    Adobe Inc. Mixamo. https://www.mixamo.com/,

  7. [7]

    2d human pose estimation: New benchmark and state of the art analysis

    Mykhaylo Andriluka, Leonid Pishchulin, Peter Gehler, and Bernt Schiele. 2d human pose estimation: New benchmark and state of the art analysis. InCVPR, 2014. 3, 7

  8. [8]

    Cross-view and cross- pose completion for 3d human understanding

    Matthieu Armando, Salma Galaaoui, Fabien Baradel, Thomas Lucas, Vincent Leroy, Romain Brégier, Philippe Weinzaepfel, and Grégory Rogez. Cross-view and cross- pose completion for 3d human understanding. InCVPR,

  9. [9]

    Multi-hmr: Multi-person whole- body human mesh recovery in a single shot

    Fabien Baradel, Matthieu Armando, Salma Galaaoui, Romain Brégier, Philippe Weinzaepfel, Grégory Rogez, and Thomas Lucas. Multi-hmr: Multi-person whole- body human mesh recovery in a single shot. InECCV,

  10. [10]

    Black, Priyanka Patel, Joachim Tesch, and Jinlong Yang

    Michael J. Black, Priyanka Patel, Joachim Tesch, and Jinlong Yang. BEDLAM: A synthetic dataset of bodies exhibiting detailed lifelike animated motion. InCVPR,

  11. [11]

    Blender Foundation. Blender. https://www. blender.org/. 6

  12. [12]

    Federica Bogo, Javier Romero, Matthew Loper, and Michael J. Black. FAUST: Dataset and evaluation for 3D mesh registration. InCVPR, 2014. 1

  13. [13]

    Keep it SMPL: Automatic estimation of 3d human pose and shape from a single image

    Federica Bogo, Angjoo Kanazawa, Christoph Lassner, Peter Gehler, Javier Romero, and Michael J Black. Keep it SMPL: Automatic estimation of 3d human pose and shape from a single image. InECCV, 2016. 3

  14. [14]

    How far are we from solving the 2d & 3d face alignment problem? (and a dataset of 230,000 3d facial landmarks)

    Adrian Bulat and Georgios Tzimiropoulos. How far are we from solving the 2d & 3d face alignment problem? (and a dataset of 230,000 3d facial landmarks). InICCV,

  15. [15]

    Smpler-x: Scaling up expressive human pose and shape estimation

    Zhongang Cai, Wanqi Yin, Ailing Zeng, Chen Wei, Qingping Sun, Wang Yanjun, Hui En Pang, Haiyi Mei, Mingyuan Zhang, Lei Zhang, et al. Smpler-x: Scaling up expressive human pose and shape estimation. In NeurIPS, 2023. 3

  16. [16]

    Collaborative regression of expressive bodies using moderation

    Yao Feng, Vasileios Choutas, Timo Bolkart, Dimitrios Tzionas, and Michael J Black. Collaborative regression of expressive bodies using moderation. In3DV, 2021. 3

  17. [17]

    Hu- mans in 4D: Reconstructing and tracking humans with transformers

    Shubham Goel, Georgios Pavlakos, Jathushan Ra- jasegaran, Angjoo Kanazawa, and Jitendra Malik. Hu- mans in 4D: Reconstructing and tracking humans with transformers. InICCV, 2023. 1, 3, 6, 7

  18. [18]

    Densepose: Dense human pose estimation in the wild

    Rıza Alp Güler, Natalia Neverova, and Iasonas Kokkinos. Densepose: Dense human pose estimation in the wild. InCVPR, 2018. 3

  19. [19]

    Black, Christoph Bodensteiner, Michael Arens, Ulrich G

    Nikolas Hesse, Sergi Pujades, Javier Romero, Michael J. Black, Christoph Bodensteiner, Michael Arens, Ulrich G. Hofmann, Uta Tacke, Mijna Hadders-Algra, Raphael Weinberger, Wolfgang Müller-Felber, and A. Sebas- tian Schroeder. Learning an Infant Body Model from RGB-D Data for Accurate Full Body Motion Analysis. In MICCAI, 2018. 2, 3

  20. [20]

    Towards accurate marker- less human shape and pose estimation over time

    Yinghao Huang, Federica Bogo, Christoph Lassner, Angjoo Kanazawa, Peter V Gehler, Javier Romero, Ijaz Akhter, and Michael J Black. Towards accurate marker- less human shape and pose estimation over time. In 3DV, 2017. 3

  21. [21]

    Human3.6M: Large scale datasets and predictive methods for 3d human sensing in natural environments.IEEE Trans

    Catalin Ionescu, Dragos Papava, Vlad Olaru, and Cris- tian Sminchisescu. Human3.6M: Large scale datasets and predictive methods for 3d human sensing in natural environments.IEEE Trans. PAMI, 2013. 3 9 Anny

  22. [22]

    Clustered pose and nonlinear appearance models for human pose esti- mation

    Sam Johnson and Mark Everingham. Clustered pose and nonlinear appearance models for human pose esti- mation. InBMVC, 2010. 3

  23. [23]

    Panoptic studio: A massively multi- view system for social interaction capture.IEEE trans

    Hanbyul Joo, Tomas Simon, Xulong Li, Hao Liu, Lei Tan, Lin Gui, Sean Banerjee, Timothy Scott Godisart, Bart Nabbe, IainMatthews, TakeoKanade, ShoheiNobuhara, and Yaser Sheikh. Panoptic studio: A massively multi- view system for social interaction capture.IEEE trans. PAMI, 2017. 7

  24. [24]

    Exemplar fine-tuning for 3d human model fitting to- wards in-the-wild 3d human pose estimation

    Hanbyul Joo, Natalia Neverova, and Andrea Vedaldi. Exemplar fine-tuning for 3d human model fitting to- wards in-the-wild 3d human pose estimation. In3DV,

  25. [25]

    End-to-end recovery of human shape and pose

    Angjoo Kanazawa, Michael J Black, David W Jacobs, and Jitendra Malik. End-to-end recovery of human shape and pose. InCVPR, 2018. 3

  26. [26]

    EMDB: The Electromagnetic Database of Global 3D Human Pose and Shape in the Wild

    Manuel Kaufmann, Jie Song, Chen Guo, Kaiyue Shen, TianjianJiang,ChengchengTang,JuanJoséZárate,and Otmar Hilliges. EMDB: The Electromagnetic Database of Global 3D Human Pose and Shape in the Wild. In ICCV, 2023. 7

  27. [27]

    Vibe: Video inference for human body pose and shape estimation

    Muhammed Kocabas, Nikos Athanasiou, and Michael J Black. Vibe: Video inference for human body pose and shape estimation. InCVPR, 2020. 1

  28. [28]

    Learningtoreconstruct3dhuman pose and shape via model-fitting in the loop

    Nikos Kolotouros, Georgios Pavlakos, Michael J Black, andKostasDaniilidis. Learningtoreconstruct3dhuman pose and shape via model-fitting in the loop. InICCV,

  29. [29]

    Unite the people: Closing the loop between 3D and 2D human representations

    Christoph Lassner, Javier Romero, Martin Kiefel, Feder- ica Bogo, Michael J Black, and Peter V Gehler. Unite the people: Closing the loop between 3D and 2D human representations. InCVPR, 2017. 3

  30. [30]

    Pose Space Deformation: A Unified Approach to Shape Interpola- tion and Skeleton-Driven Deformation

    J P Lewis, Matt Cordner, and Nickson Fong. Pose Space Deformation: A Unified Approach to Shape Interpola- tion and Skeleton-Driven Deformation. InSIGGRAPH,

  31. [31]

    CLIFF: Carrying location in- formation in full frames into human pose and shape estimation

    Zhihao Li, Jianzhuang Liu, Zhensong Zhang, Songcen Xu, and Youliang Yan. CLIFF: Carrying location in- formation in full frames into human pose and shape estimation. InECCV, 2022. 3

  32. [32]

    One-stage 3d whole-body mesh recovery with component aware transformer

    Jing Lin, Ailing Zeng, Haoqian Wang, Lei Zhang, and Yu Li. One-stage 3d whole-body mesh recovery with component aware transformer. InCVPR, 2023. 3

  33. [33]

    Microsoft coco: Common objects in context

    Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. Microsoft coco: Common objects in context. InECCV, 2014. 3, 7

  34. [34]

    Smpl: askinned multi-person linear model

    Matthew Loper, Naureen Mahmood, Javier Romero, GerardPons-Moll, andMichaelJBlack. Smpl: askinned multi-person linear model. InACM Trans. Graphics,

  35. [35]

    Warp: A high-performance python framework for gpu simulation and graphics.https: //github.com/nvidia/warp, 2022

    Miles Macklin. Warp: A high-performance python framework for gpu simulation and graphics.https: //github.com/nvidia/warp, 2022. NVIDIA GPU Technology Conference (GTC). 4

  36. [36]

    Troje, Gerard Pons-Moll, and Michael Black

    Naureen Mahmood, Nima Ghorbani, Nikolaus F. Troje, Gerard Pons-Moll, and Michael Black. AMASS: Archive of Motion Capture As Surface Shapes. InICCV, 2019. 6

  37. [37]

    Monocular 3d human pose estimation in the wild using improved cnn supervision

    Dushyant Mehta, Helge Rhodin, Dan Casas, Pascal Fua, Oleksandr Sotnychenko, Weipeng Xu, and Christian Theobalt. Monocular 3d human pose estimation in the wild using improved cnn supervision. In3DV, 2017. 3

  38. [38]

    Neuralannot: Neural annotator for 3d human mesh training sets

    Gyeongsik Moon, Hongsuk Choi, and Kyoung Mu Lee. Neuralannot: Neural annotator for 3d human mesh training sets. InCVPR, 2022. 3

  39. [39]

    WHO child growth standards: length/height-for-age, weight-for-age, weight-for-length, weight-for-height and body mass index-for-age: methods and develop- ment

    Department of Nutrition for Health and Development. WHO child growth standards: length/height-for-age, weight-for-age, weight-for-length, weight-for-height and body mass index-for-age: methods and develop- ment. Technical report, World Health Organisation,

  40. [40]

    Dinov2: Learning robust visual features without supervision.TMLR, 2024

    Maxime Oquab, Timothée Darcet, Théo Moutakanni, Huy Vo, Marc Szafraniec, Vasil Khalidov, Pierre Fer- nandez, Daniel Haziza, Francisco Massa, Alaaeldin El- Nouby, et al. Dinov2: Learning robust visual features without supervision.TMLR, 2024. 7

  41. [41]

    STAR: Sparse trained articulated human body regressor

    Ahmed AA Osman, Timo Bolkart, and Michael J Black. STAR: Sparse trained articulated human body regressor. InECCV, 2020. 1, 3, 4

  42. [42]

    Ahmed A A Osman, Timo Bolkart, Dimitrios Tzionas, and Michael J. Black. SUPR: A sparse unified part-based human body model. InECCV, 2022. 1, 3, 4, 5

  43. [43]

    Atlas: De- coupling skeletal and shape parameters for expressive parametric human modeling

    Jinhyung Park, Javier Romero, Shunsuke Saito, Fabian Prada, Takaaki Shiratori, Yichen Xu, Federica Bogo, Shoou-I Yu, Kris Kitani, and Rawal Khirodkar. Atlas: De- coupling skeletal and shape parameters for expressive parametric human modeling. InICCV, 2025. 3, 4, 5

  44. [44]

    Auto- matic differentiation in pytorch

    Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer. Auto- matic differentiation in pytorch. InNeurIPS workshop,

  45. [45]

    CameraHMR:Align- ing people with perspective

    PriyankaPatelandMichaelJBlack. CameraHMR:Align- ing people with perspective. In3DV, 2025. 3, 5

  46. [46]

    AGORA: Avatars in geography optimized for regression analysis

    Priyanka Patel, Chun-Hao P Huang, Joachim Tesch, David T Hoffmann, Shashank Tripathi, and Michael J Black. AGORA: Avatars in geography optimized for regression analysis. InCVPR, 2021. 2, 3, 7

  47. [47]

    Expressive body capture: 3d hands, face, and body from a single image

    Georgios Pavlakos, Vasileios Choutas, Nima Ghorbani, TimoBolkart,AhmedAAOsman,DimitriosTzionas,and Michael J Black. Expressive body capture: 3d hands, face, and body from a single image. InCVPR, 2019. 1, 3, 4, 7, 9

  48. [48]

    Infinigen indoors: Photorealistic indoor scenes using procedural generation

    Alexander Raistrick, Lingjie Mei, Karhan Kayan, David Yan, Yiming Zuo, Beining Han, Hongyu Wen, Meenal Parakh, Stamatis Alexandropoulos, Lahav Lipson, Zeyu Ma, and Jia Deng. Infinigen indoors: Photorealistic indoor scenes using procedural generation. InCVPR,

  49. [49]

    Robinette, Sherri Blackwell, Hein Daanen, Mark Boehmer, and Scott Fleming

    Kathleen M. Robinette, Sherri Blackwell, Hein Daanen, Mark Boehmer, and Scott Fleming. Civilian American and European Surface Anthropometry Resource (CAE- SAR), Final Report. Volume 1. Summary:. Technical report, Defense Technical Information Center, 2002. 1, 3, 5 10 Anny

  50. [50]

    Condimen: Condi- tional multi-person mesh recovery.arXiv preprint arXiv:2412.13058, 2024

    Brégier Romain, Baradel Fabien, Lucas Thomas, Galaaoui Salma, Armando Matthieu, Weinzaepfel Philippe, and Rogez Grégory. Condimen: Condi- tional multi-person mesh recovery.arXiv preprint arXiv:2412.13058, 2024. 7

  51. [51]

    3DBodyTex: Textured 3D Body Dataset

    Alexandre Saint, Eman Ahmed, Abd El Rahman Shabayek, Kseniya Cherenkova, Gleb Gusev, Djamila Aouada, and Bjorn Ottersten. 3DBodyTex: Textured 3D Body Dataset. In3DV, 2018. 5

  52. [52]

    Wham: Reconstructing world-grounded humans with accurate 3d motion

    Soyong Shin, Juyong Kim, Eni Halilaj, and Michael J Black. Wham: Reconstructing world-grounded humans with accurate 3d motion. InCVPR, 2024. 1

  53. [53]

    Sat- hmr: Real-time multi-person 3d mesh estimation via scale-adaptive tokens

    Chi Su, Xiaoxuan Ma, Jiajun Su, and Yizhou Wang. Sat- hmr: Real-time multi-person 3d mesh estimation via scale-adaptive tokens. InCVPR, 2025. 8

  54. [54]

    Aios: All-in-one-stage expres- sive human pose and shape estimation

    Qingping Sun, Yanjun Wang, Ailing Zeng, Wanqi Yin, Chen Wei, Wenjia Wang, Haiyi Mei, Chi-Sing Leung, Ziwei Liu, Lei Yang, et al. Aios: All-in-one-stage expres- sive human pose and shape estimation. InCVPR, 2024. 3, 8

  55. [55]

    Monocular, one-stage, regression of multiple 3d people

    Yu Sun, Qian Bao, Wu Liu, Yili Fu, Michael J Black, and Tao Mei. Monocular, one-stage, regression of multiple 3d people. InICCV, 2021. 3

  56. [56]

    Putting people in their place: Monoc- ular regression of 3d people in depth

    Yu Sun, Wu Liu, Qian Bao, Yili Fu, Tao Mei, and Michael J Black. Putting people in their place: Monoc- ular regression of 3d people in depth. InCVPR, 2022. 3

  57. [57]

    GRAB: A Dataset of Whole-Body Human Grasping of Objects

    Omid Taheri, Nima Ghorbani, Michael J Black, and Dimitrios Tzionas. GRAB: A Dataset of Whole-Body Human Grasping of Objects. InECCV, 2020. 6

  58. [58]

    Joachim Tesch, Giorgio Becherini, Prerana Achar, Anas- tasios Yiannakidis, Muhammed Kocabas, Priyanka Patel, and Michael J. Black. BEDLAM2.0: Synthetic humans and cameras in motion. InNeurIPS, 2025. 5

  59. [59]

    MetaHuman

    Unreal Engine. MetaHuman. https://www. unrealengine.com/fr/metahuman. 3

  60. [60]

    Learning from synthetic humans

    Gul Varol, Javier Romero, Xavier Martin, Naureen Mah- mood, Michael J Black, Ivan Laptev, and Cordelia Schmid. Learning from synthetic humans. InCVPR,

  61. [61]

    Recovering accurate 3d human pose in the wild using imus and a moving camera

    Timo von Marcard, Roberto Henschel, Michael Black, Bodo Rosenhahn, and Gerard Pons-Moll. Recovering accurate 3d human pose in the wild using imus and a moving camera. InECCV, 2018. 7

  62. [62]

    BLADE: Single-view Body Mesh Learning through Accurate Depth Estimation.arXiv preprint arXiv:2412.08640, 2024

    Shengze Wang, Jiefeng Li, Tianye Li, Ye Yuan, Henry Fuchs, Shalini De Mello, Koki Nagano, and Michael Stengel. BLADE: Single-view Body Mesh Learning through Accurate Depth Estimation.arXiv preprint arXiv:2412.08640, 2024. 3

  63. [63]

    Zolly: Zoom focal length correctly for perspective-distorted human mesh reconstruction

    Wenjia Wang, Yongtao Ge, Haiyi Mei, Zhongang Cai, Qingping Sun, Yanjun Wang, Chunhua Shen, Lei Yang, and Taku Komura. Zolly: Zoom focal length correctly for perspective-distorted human mesh reconstruction. InICCV, 2023. 3

  64. [64]

    GHUM & GHUML: Generative 3d human shape and articulated pose models

    Hongyi Xu, Eduard Gabriel Bazavan, Andrei Zanfir, William T Freeman, Rahul Sukthankar, and Cristian Sminchisescu. GHUM & GHUML: Generative 3d human shape and articulated pose models. InCVPR, 2020. 3, 4

  65. [65]

    WHAC: World- grounded humans and cameras

    Wanqi Yin, Zhongang Cai, Ruisi Wang, Fanzhou Wang, Chen Wei, Haiyi Mei, Weiye Xiao, Zhitao Yang, Qing- ping Sun, Atsushi Yamashita, et al. WHAC: World- grounded humans and cameras. InECCV, 2024. 3

  66. [66]

    SMPLest-X: Ultimate scaling for expres- sive human pose and shape estimation.arXiv preprint arXiv:2501.09782, 2025

    Wanqi Yin, Zhongang Cai, Ruisi Wang, Ailing Zeng, Chen Wei, Qingping Sun, Haiyi Mei, Yanjun Wang, Hui En Pang, Mingyuan Zhang, Lei Zhang, Chen Change Loy, Atsushi Yamashita, Lei Yang, and Ziwei Liu. SMPLest-X: Ultimate scaling for expres- sive human pose and shape estimation.arXiv preprint arXiv:2501.09782, 2025. 1

  67. [67]

    Hi4d: 4d instance seg- mentation of close human interaction

    Yifei Yin, Chen Guo, Manuel Kaufmann, Juan Zarate, Jie Song, and Otmar Hilliges. Hi4d: 4d instance seg- mentation of close human interaction. InCVPR, 2023. 7

  68. [68]

    Py- MAF: 3d human pose and shape regression with pyra- midal mesh alignment feedback loop

    Hongwen Zhang, Yating Tian, Xinchi Zhou, Wanli Ouyang, Yebin Liu, Limin Wang, and Zhenan Sun. Py- MAF: 3d human pose and shape regression with pyra- midal mesh alignment feedback loop. InICCV, 2021. 1 11

  69. [2025]

    Online 3D character animation service. 4

This paper was first reviewed by deepseek-v4-flash on August 3, 2026.