Pith. sign in

REVIEW 5 major objections 3 minor 56 references

DyCrowd: Towards Dynamic Crowd Reconstruction from a Large-scene Video

T0 review · 5 major / 3 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read A single large-scene video can yield consistent 3D poses, positions, and shapes for hundreds of people over time.

desk verdict A plausible video-based crowd reconstruction idea with a genuinely new task framing, but the SOTA claim currently rests on an authors-made synthetic benchmark and the supplied text is unreadable, so the evidence cannot yet be checked. read the letter →

arxiv 2508.12644 v1 pith:NTLYO3TL submitted 2025-08-18 cs.CV

classification cs.CV
keywords dynamiccrowdreconstructionlarge-scenevideo3Dhumanposeestimationtemporalconsistencyocclusionreasoninggroup-guidedoptimizationmotionpriorVirtual
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper introduces DyCrowd, a framework that takes one large-scene video and reconstructs the 3D pose, ground position, and body shape of hundreds of individuals frame by frame while keeping each person's motion consistent over time. This matters because existing methods reconstruct crowds from static images, so they have no temporal consistency and are vulnerable to occlusions. DyCrowd's central idea is that crowd motion is collective: people with similar motion segments are grouped, and clearly visible members of a group are used to guide the recovery of members who are occluded. The paper also contributes VirtualCrowd, a virtual benchmark dataset for evaluating large-scene dynamic crowd reconstruction, and reports state-of-the-art results on it.

What carries the argument

The load-bearing mechanism is segment-level group-guided optimization. Motion sequences are divided into segments, individuals with similar motion segments are clustered, and each cluster's motion is optimized together in a coarse-to-fine manner; within a cluster, visible unoccluded segments act as the temporal reference for reconstructing occluded segments. The VAE-based human motion prior constrains optimizations to plausible body motions, while the AMC loss enforces consistency between group members despite asynchrony and rhythmic variation. This turns occlusion from a per-person missing-data problem into a group inference problem.

What would settle it

Take a video where a person walks behind a long wall while everyone else in view either stands still or moves differently. If DyCrowd reconstructs the hidden walker's motion as following the standing crowd rather than continuing along the observed pre- and post-occlusion path, the group-guidance premise is falsified; a multi-camera ground-truth capture of the same scene would settle it.

Watch

Extended reading notes

Core claim

The central claim is that DyCrowd is the first framework to reconstruct hundreds of individuals' poses, positions and shapes with spatio-temporal consistency from a single large-scene video. The argument rests on a coarse-to-fine, group-guided motion optimization: motion sequences are split into segments, similar segments are grouped, and the group's motions are optimized jointly so that reliable, unoccluded segments transfer their temporal information to occluded ones. A VAE-based human motion prior keeps reconstructed motions plausible, and the Asynchronous Motion Consistency (AMC) loss makes the transfer robust to people moving out of sync or at different rhythms. The paper reports that this method achieves state-of-the-art performance on the large-scene dynamic crowd reconstruction task, evaluated on its new VirtualCrowd benchmark.

Load-bearing premise

The method bets that people with similar motion can be grouped and that seeing some group members clearly is enough to reconstruct the hidden members; if a hidden person's motion resembles no one visible, or the entire group is occluded at once, the temporal guidance has nothing to draw on.

Editorial extensions

If this is right

  • Large-scene video becomes a viable input for crowd reconstruction, giving temporal consistency that static-image methods cannot offer.
  • Long-term occlusions can be resolved by borrowing movement patterns from similarly behaving people who are visible, rather than relying only on the occluded person's own observed past.
  • Reconstruction quality should scale with crowd regularity: a stream of pedestrians with similar gaits should recover occluded members better than a crowd of independent, dissimilar actors.
  • The VirtualCrowd benchmark provides a quantitative way to compare methods on temporal consistency and occlusion recovery in large scenes.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The method's advantage should grow with crowd density, since dense crowds provide more similar visible segments to guide each occluded person, until everyone is simultaneously hidden.
  • Group-guided temporal transfer of the same kind could apply to other correlated multi-object scenes, such as vehicle traffic or animal herds, wherever motion is shared enough to form reliable groups.
  • A concrete stress test is to move the VAE motion prior outside its training distribution: unusual gaits or choreographed actions may remain unrecoverable even when visible peers move similarly, because the prior will pull reconstructions back to familiar motions.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 3 minor

Summary. The manuscript introduces DyCrowd, a framework for reconstructing the 3D poses, positions, and shapes of hundreds of people over time from a single large-scene video. The proposed method combines coarse-to-fine group-guided motion optimization with a VAE-based human motion prior and an Asynchronous Motion Consistency (AMC) loss, targeting long-term dynamic occlusions in large scenes. The authors also present VirtualCrowd, a synthetic benchmark dataset for evaluating dynamic crowd reconstruction. The abstract claims state-of-the-art performance, but the submitted full text is encoding-corrupted and largely unreadable, so the methodology, equations, experimental setup, and quantitative comparisons cannot be verified from the available material.

Significance. The task addressed by DyCrowd is timely and practically relevant for city surveillance and crowd analysis, and the goal of spatio-temporally consistent reconstruction of hundreds of individuals from a single video is a genuine gap in the current literature. The contribution of a benchmark dataset, even a synthetic one, is useful if it is carefully validated and released. However, the significance of the work cannot currently be assessed: the central state-of-the-art claim rests on the abstract alone, no quantitative metrics are reported in readable form, and the only named evaluation artifact is a synthetic dataset constructed by the same authors. If the method and dataset are made fully readable and are validated against real-world large-scene data, the contribution could be substantial; at present it remains unsubstantiated.

major comments (5)
  1. [Abstract and Full Text] The central claim of state-of-the-art performance is not supported by any readable quantitative result. The abstract provides no metrics, ablations, or error bars, and the full text is encoding-corrupted, making the comparison tables and experimental sections inaccessible. The authors should provide a readable manuscript with concrete numbers on the VirtualCrowd benchmark and, ideally, on at least one real-world large-scene video with annotations or proxy metrics.
  2. [VirtualCrowd evaluation and circularity] The only evaluation benchmark named in the paper, VirtualCrowd, is introduced by the authors, and the abstract justifies its creation by the absence of existing well-annotated large-scene video datasets. If the synthetic data are generated using the same group-motion assumptions and occlusion patterns that DyCrowd explicitly encodes, then the reported state-of-the-art result would be at least partly self-confirming. The paper should describe the data generation process in detail, provide statistics on occlusion rates and motion diversity, and include a transfer experiment to real video to demonstrate that the method generalizes beyond the synthetic distribution.
  3. [Group-guided occlusion mechanism] The core mechanism assumes that visible, unoccluded members of a group with similar motion segments can reliably guide the reconstruction of occluded members. The abstract does not address failure cases where an occluded individual has a unique motion not shared by any visible group member, or where an entire group becomes occluded simultaneously. Without an analysis of these failure modes or an ablation that varies occlusion severity and group size, the claim of 'robust and plausible motion recovery' remains unsupported.
  4. [Provenance of the submitted file] The full text contains the line 'arXiv:2508.12645v5 [cs.IR] 18 Jan 2026', which does not match the claimed identifier '2508.12644' nor the subject area cs.CV. This mismatch raises a provenance concern: the authors should confirm that the correct version of the paper was uploaded, and the inserted line should be removed from the camera-ready version.
  5. [Readability of the full text] The submitted full text is not readable due to encoding corruption; the equations, including the definitions of the AMC loss and the VAE motion prior, along with all tables and algorithm pseudocode, are unreadable. This prevents the verification of every technical contribution and every experimental result, and it must be fixed before any further assessment can occur.
minor comments (3)
  1. [Abstract] The abstract promises that 'code and dataset will be available for research purposes' but gives no details on licenses or access conditions; please specify the intended release terms.
  2. [Abstract] The acronym 'AMC' is used without a definition in the abstract; please add a brief clarification when the loss is first mentioned.
  3. [General] Once the full text is readable, please ensure that all tables report error bars or statistical significance, and that ablations isolate the contributions of the group-guided optimization, the VAE prior, and the AMC loss.

Circularity Check

0 steps flagged · score 0.0 of 10

No circular derivation is established from the available text; the SOTA claim rests on an author-provided synthetic benchmark, which is a generalization concern rather than a definitional circularity.

full rationale

The supplied full text is encoding-corrupted, so the equations, losses, ablations, and comparison tables cannot be checked. From the abstract alone, DyCrowd's components are not defined in terms of the target reconstruction (no self-definition), no fitted parameter is relabeled as a prediction, and no load-bearing claim reduces to a self-citation. The only identified concern is that the state-of-the-art claim is supported by VirtualCrowd, a benchmark contributed by the same paper rather than an external benchmark. That makes the evaluation self-referential in provenance, but it is not, on the quoted evidence, circular by construction: the abstract does not state that the benchmark generator encodes the method's group-motion assumptions, and no equation-level reduction is available to exhibit. Under the hard rule requiring a quoted reduction, the appropriate circularity finding is no significant circularity; the missing external validation and the provenance mismatch of the inserted arXiv identifier affect verifiability and correctness risk, not the circularity score.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The central claims rest on several domain assumptions stated in the abstract but not derivable from it. No free parameters are visible at the abstract level. VirtualCrowd is a contributed artifact, not a physical entity, so it is not listed as an invented entity.

assumptions (3)
  • domain assumption A VAE trained on human motion provides a realistic prior for occluded poses over time.
    The abstract introduces a 'VAE-based human motion prior' to stabilize temporally occluded reconstruction; its adequacy is assumed.
  • domain assumption People with similar motion segments can be clustered, and visible members can guide invisible members.
    The core 'group-guided motion optimization' depends on this collective-behavior assumption for occlusion recovery.
  • domain assumption VirtualCrowd, a synthetic benchmark, is representative of real large-scene crowd videos.
    The abstract states the dataset fills the gap of no annotated real data; transferability to real scenes is assumed.

how reviews work

0 comments
Cite this review

Pith. "Pith review of DyCrowd: Towards Dynamic Crowd Reconstruction from a Large-scene Video." pith.science (2026). https://pith.science/paper/NTLYO3TL

@misc{pith2026250812644,
  author       = {Pith},
  title        = {Pith review of: DyCrowd: Towards Dynamic Crowd Reconstruction from a Large-scene Video},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/NTLYO3TL}},
  note         = {Machine review of arXiv:2508.12644}
}
read the original abstract

3D reconstruction of dynamic crowds in large scenes has become increasingly important for applications such as city surveillance and crowd analysis. However, current works attempt to reconstruct 3D crowds from a static image, causing a lack of temporal consistency and inability to alleviate the typical impact caused by occlusions. In this paper, we propose DyCrowd, the first framework for spatio-temporally consistent 3D reconstruction of hundreds of individuals' poses, positions and shapes from a large-scene video. We design a coarse-to-fine group-guided motion optimization strategy for occlusion-robust crowd reconstruction in large scenes. To address temporal instability and severe occlusions, we further incorporate a VAE (Variational Autoencoder)-based human motion prior along with a segment-level group-guided optimization. The core of our strategy leverages collective crowd behavior to address long-term dynamic occlusions. By jointly optimizing the motion sequences of individuals with similar motion segments and combining this with the proposed Asynchronous Motion Consistency (AMC) loss, we enable high-quality unoccluded motion segments to guide the motion recovery of occluded ones, ensuring robust and plausible motion recovery even in the presence of temporal desynchronization and rhythmic inconsistencies. Additionally, in order to fill the gap of no existing well-annotated large-scene video dataset, we contribute a virtual benchmark dataset, VirtualCrowd, for evaluating dynamic crowd reconstruction from large-scene videos. Experimental results demonstrate that the proposed method achieves state-of-the-art performance in the large-scene dynamic crowd reconstruction task. The code and dataset will be available for research purposes.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

56 extracted references · 54 canonical work pages

  1. [1]

    H. Wen, J. Huang, H. Cui, H. Lin, Y.-K. Lai, L. Fang, and K. Li, ``Crowd3 D : Towards hundreds of people reconstruction from a single image,'' in Proc. IEEE Conference on Computer Vision and Pattern Recognition, 2023, pp. 8937--8946

  2. [2]

    Huang, J

    B. Huang, J. Ju, Z. Li, and Y. Wang, ``Reconstructing groups of people with hypergraph relational reasoning,'' in Proc. IEEE International Conference on Computer Vision, 2023, pp. 14\,873--14\,883

  3. [3]

    Kocabas, N

    M. Kocabas, N. Athanasiou, and M. J. Black, `` VIBE : Video inference for human body pose and shape estimation,'' in Proc. IEEE Conference on Computer Vision and Pattern Recognition, 2020, pp. 5253--5263

  4. [4]

    H. Choi, G. Moon, J. Y. Chang, and K. M. Lee, ``Beyond static features for temporally consistent 3 D human pose and shape from a video,'' in Proc. IEEE Conference on Computer Vision and Pattern Recognition, 2021, pp. 1964--1973

  5. [5]

    Wei, J.-C

    W.-L. Wei, J.-C. Lin, T.-L. Liu, and H.-Y. M. Liao, ``Capturing humans in motion: Temporal-attentive 3 D human pose and shape estimation from monocular video,'' in Proc. IEEE Conference on Computer Vision and Pattern Recognition, 2022, pp. 13\,211--13\,220

  6. [6]

    X. Shen, Z. Yang, X. Wang, J. Ma, C. Zhou, and Y. Yang, ``Global-to-local modeling for video-based 3 D human pose and shape estimation,'' in Proc. IEEE Conference on Computer Vision and Pattern Recognition, 2023, pp. 8887--8896

  7. [7]

    Y. Yuan, U. Iqbal, P. Molchanov, K. Kitani, and J. Kautz, `` GLAMR : Global occlusion-aware human mesh recovery with dynamic cameras,'' in Proc. IEEE Conference on Computer Vision and Pattern Recognition, 2022, pp. 11\,038--11\,049

  8. [8]

    Y. Wang, Z. Wang, L. Liu, and K. Daniilidis, `` TRAM : Global trajectory and motion of 3 D humans from in-the-wild videos,'' arXiv preprint arXiv:2403.17346, 2024

Show all 56 references
  1. [9]

    K. Li, Y. Liu, Y.-K. Lai, and J. Yang, `` MILI : Multi-person inference from a low-resolution image,'' Fundamental Research, vol. 3, no. 3, pp. 434--441, 2023

  2. [10]

    Y. Sun, Q. Bao, W. Liu, T. Mei, and M. J. Black, `` TRACE : 5 D temporal regression of avatars with dynamic cameras in 3 D environments,'' in Proc. IEEE Conference on Computer Vision and Pattern Recognition, 2023, pp. 8856--8866

  3. [11]

    Z. Qiu, Q. Yang, J. Wang, H. Feng, J. Han, E. Ding, C. Xu, D. Fu, and J. Wang, `` PSVT : End-to-end multi-person 3 D pose and shape estimation with progressive video transformers,'' in Proc. IEEE Conference on Computer Vision and Pattern Recognition, 2023, pp. 21\,254--21\,263

  4. [12]

    V. Ye, G. Pavlakos, J. Malik, and A. Kanazawa, ``Decoupling human and camera motion from videos in the wild,'' in Proc. IEEE Conference on Computer Vision and Pattern Recognition, 2023, pp. 21\,222--21\,232

  5. [13]

    Kocabas, Y

    M. Kocabas, Y. Yuan, P. Molchanov, Y. Guo, M. J. Black, O. Hilliges, J. Kautz, and U. Iqbal, `` PACE : Human and camera motion estimation from in-the-wild videos,'' in Proc. IEEE Int. Conf. 3D Vis. 1em plus 0.5em minus 0.4em IEEE, 2024, pp. 397--408

  6. [14]

    B. Zhou, X. Tang, and X. Wang, ``Measuring crowd collectiveness,'' in Proc. IEEE Conference on Computer Vision and Pattern Recognition, 2013, pp. 3049--3056

  7. [15]

    Y. Wu, Y. Ye, C. Zhao, and Z. Shi, ``Collective density clustering for coherent motion detection,'' IEEE Trans. Multimedia, vol. 20, no. 6, pp. 1418--1431, 2017

  8. [16]

    L. Mei, J. Lai, Z. Chen, and X. Xie, ``Measuring crowd collectiveness via global motion correlation,'' in Proc. IEEE International Conference on Computer Vision, 2019, pp. 0--0

  9. [17]

    Zhang and S

    Y. Zhang and S. Tang, ``The wanderings of odysseus in 3 D scenes,'' in Proc. IEEE Conference on Computer Vision and Pattern Recognition, 2022, pp. 20\,481--20\,491

  10. [18]

    K. Zhao, Y. Zhang, S. Wang, T. Beeler, and S. Tang, ``Synthesizing diverse human motions in 3 D indoor scenes,'' in Proc. IEEE International Conference on Computer Vision, 2023, pp. 14\,738--14\,749

  11. [19]

    [Online]

    iCity3D , `` iCity3D ,'' 2024. [Online]. Available: https://icity3d.com/

  12. [20]

    Pavlakos, V

    G. Pavlakos, V. Choutas, N. Ghorbani, T. Bolkart, A. A. Osman, D. Tzionas, and M. J. Black, ``Expressive body capture: 3 D hands, face, and body from a single image,'' in Proc. IEEE Conference on Computer Vision and Pattern Recognition, 2019, pp. 10\,975--10\,985

  13. [21]

    D. J. Brady, M. E. Gehm, R. A. Stack, D. L. Marks, D. S. Kittle, D. R. Golish, E. Vera, and S. D. Feller, ``Multiscale gigapixel photography,'' Nature, vol. 486, no. 7403, pp. 386--389, 2012

  14. [22]

    X. Yuan, L. Fang, Q. Dai, D. J. Brady, and Y. Liu, ``Multiscale gigapixel video: A cross resolution image matching and warping approach,'' in 2017 IEEE International Conference on Computational Photography (ICCP). 1em plus 0.5em minus 0.4em IEEE, 2017, pp. 1--9

  15. [23]

    X. Wang, X. Zhang, Y. Zhu, Y. Guo, X. Yuan, L. Xiang, Z. Wang, G. Ding, D. Brady, Q. Dai et al., `` PANDA : A gigapixel-level human-centric video dataset,'' in Proc. IEEE Conference on Computer Vision and Pattern Recognition, 2020, pp. 3268--3278

  16. [24]

    C. Liu, H. Wei, J. Yang, J. Liu, W. Li, Y. Guo, and L. Fang, `` GigaHumanDet : Exploring full-body detection on gigapixel-level images,'' in Proc. AAAI Conference on Artificial Intelligence, vol. 38, no. 9, 2024, pp. 10\,092--10\,100

  17. [25]

    W. Mo, W. Zhang, H. Wei, R. Cao, Y. Ke, and Y. Luo, `` PVDet : Towards pedestrian and vehicle detection on gigapixel-level images,'' Engineering Applications of Artificial Intelligence, vol. 118, p. 105705, 2023

  18. [26]

    W. Li, R. Zhang, H. Lin, Y. Guo, C. Ma, and X. Yang, `` SaccadeDet : A novel dual-stage architecture for rapid and accurate detection in gigapixel images,'' in Joint European Conference on Machine Learning and Knowledge Discovery in Databases. 1em plus 0.5em minus 0.4em Spring...

  19. [27]

    Zhang, L

    J. Zhang, L. Gu, Y.-K. Lai, X. Wang, and K. Li, ``Towards grouping in large scenes with occlusion-aware spatio-temporal transformers,'' IEEE Transactions on Circuits and Systems for Video Technology, 2023

  20. [28]

    X. Xu, H. Chen, F. Moreno-Noguer, L. A. Jeni, and F. De la Torre, ``3 D human pose, shape and texture from low-resolution images and videos,'' IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 44, no. 9, pp. 4490--4504, 2021

  21. [29]

    S. Shin, J. Kim, E. Halilaj, and M. J. Black, `` WHAM : Reconstructing world-grounded humans with accurate 3 D motion,'' in Proc. IEEE Conference on Computer Vision and Pattern Recognition, 2024, pp. 2070--2080

  22. [30]

    Zhang, D

    J. Zhang, D. Yu, J. H. Liew, X. Nie, and J. Feng, ``Body meshes as points,'' in Proc. IEEE Conference on Computer Vision and Pattern Recognition, 2021, pp. 546--556

  23. [31]

    Y. Sun, Q. Bao, W. Liu, Y. Fu, M. J. Black, and T. Mei, ``Monocular, one-stage, regression of multiple 3 D people,'' in Proc. IEEE International Conference on Computer Vision, 2021, pp. 11\,179--11\,188

  24. [32]

    Y. Sun, W. Liu, Q. Bao, Y. Fu, T. Mei, and M. J. Black, ``Putting people in their place: Monocular regression of 3 D people in depth,'' in Proc. IEEE Conference on Computer Vision and Pattern Recognition, 2022, pp. 13\,243--13\,252

  25. [33]

    Rempe, T

    D. Rempe, T. Birdal, A. Hertzmann, J. Yang, S. Sridhar, and L. J. Guibas, `` HuMoR : 3 D human motion model for robust pose estimation,'' in Proc. IEEE International Conference on Computer Vision, 2021, pp. 11\,488--11\,499

  26. [34]

    C. He, J. Saito, J. Zachary, H. Rushmeier, and Y. Zhou, `` NeMF : Neural motion fields for kinematic animation,'' Proc. Adv. Neural Inform. Process. Syst., vol. 35, pp. 4244--4256, 2022

  27. [35]

    Loper, N

    M. Loper, N. Mahmood, J. Romero, G. Pons-Moll, and M. J. Black, `` SMPL : A skinned multi-person linear model,'' ACM Trans. Graphics (Proc. SIGGRAPH Asia), vol. 34, no. 6, pp. 248:1--248:16, Oct. 2015

  28. [36]

    Y. Li, H. Mao, R. Girshick, and K. He, ``Exploring plain vision transformer backbones for object detection,'' in Proc. European Conference on Computer Vision. 1em plus 0.5em minus 0.4em Springer, 2022, pp. 280--296

  29. [37]

    Z. Yang, A. Zeng, C. Yuan, and Y. Li, ``Effective whole-body pose estimation with two-stages distillation,'' in Proc. IEEE International Conference on Computer Vision, 2023, pp. 4210--4220

  30. [38]

    S. Goel, G. Pavlakos, J. Rajasegaran, A. Kanazawa, and J. Malik, ``Humans in 4 D : Reconstructing and tracking humans with transformers,'' in Proc. IEEE International Conference on Computer Vision, 2023, pp. 14\,783--14\,794

  31. [39]

    Rajasegaran, G

    J. Rajasegaran, G. Pavlakos, A. Kanazawa, and J. Malik, ``Tracking people by predicting 3 D appearance, location and pose,'' in Proc. IEEE Conference on Computer Vision and Pattern Recognition, 2022, pp. 2740--2749

  32. [40]

    Pavlakos, V

    G. Pavlakos, V. Choutas, N. Ghorbani, T. Bolkart, A. A. A. Osman, D. Tzionas, and M. J. Black, ``Expressive body capture: 3 D hands, face, and body from a single image,'' in Proc. IEEE Conference on Computer Vision and Pattern Recognition, 2019

  33. [41]

    X. Chen, B. Jiang, W. Liu, Z. Huang, B. Fu, T. Chen, and G. Yu, ``Executing your commands via motion diffusion in latent space,'' in Proc. IEEE Conference on Computer Vision and Pattern Recognition, 2023, pp. 18\,000--18\,010

  34. [42]

    Rempe, L

    D. Rempe, L. J. Guibas, A. Hertzmann, B. Russell, R. Villegas, and J. Yang, ``Contact and human dynamics from monocular video,'' in Proc. European Conference on Computer Vision. 1em plus 0.5em minus 0.4em Springer, 2020, pp. 71--87

  35. [43]

    Zhang, Y

    S. Zhang, Y. Zhang, F. Bogo, M. Pollefeys, and S. Tang, ``Learning motion priors for 4d human body capture in 3 D scenes,'' in Proc. IEEE International Conference on Computer Vision, 2021, pp. 11\,343--11\,353

  36. [44]

    C. Guo, S. Zou, X. Zuo, S. Wang, W. Ji, X. Li, and L. Cheng, ``Generating diverse and natural 3 D human motions from text,'' in Proc. IEEE Conference on Computer Vision and Pattern Recognition, 2022, pp. 5152--5161

  37. [45]

    Zhang, B

    S. Zhang, B. L. Bhatnagar, Y. Xu, A. Winkler, P. Kadlecek, S. Tang, and F. Bogo, `` RoHM : Robust human motion reconstruction via diffusion,'' in Proc. IEEE Conference on Computer Vision and Pattern Recognition, 2024, pp. 14\,606--14\,617

  38. [46]

    Mahmood, N

    N. Mahmood, N. Ghorbani, N. F. Troje, G. Pons-Moll, and M. J. Black, `` AMASS : Archive of motion capture as surface shapes,'' in Proc. IEEE International Conference on Computer Vision, 2019, pp. 5442--5451

  39. [47]

    Besse, B

    P. Besse, B. Guillouet, J.-M. Loubes, and R. Fran c ois, ``Review and perspective for distance based trajectory clustering,'' arXiv preprint arXiv:1508.04904, 2015

  40. [48]

    B. J. Frey and D. Dueck, ``Clustering by passing messages between data points,'' science, vol. 315, no. 5814, pp. 972--976, 2007

  41. [49]

    Cuturi and M

    M. Cuturi and M. Blondel, `` Soft-DTW : a differentiable loss function for time-series,'' in International conference on machine learning. 1em plus 0.5em minus 0.4em PMLR, 2017, pp. 894--903

  42. [50]

    Sakoe and S

    H. Sakoe and S. Chiba, ``Dynamic programming algorithm optimization for spoken word recognition,'' IEEE transactions on acoustics, speech, and signal processing, vol. 26, no. 1, pp. 43--49, 1978

  43. [51]

    Z. Yang, Z. Cai, H. Mei, S. Liu, Z. Chen, W. Xiao, Y. Wei, Z. Qing, C. Wei, B. Dai et al., `` SynBody : Synthetic dataset with layered human models for 3 D human perception and modeling,'' in Proc. IEEE International Conference on Computer Vision, 2023, pp. 20\,282--20\,292

  44. [52]

    X. Liu, T. Zhou, H. Kang, J. Ma, Z. Wang, J. Huang, W. Weng, Y.-K. Lai, and K. Li, `` RESCUE : Crowd evacuation simulation via controlling sdm-united characters,'' in 2025 International Conference on Computer Vision (ICCV), 2025

  45. [53]

    J. Zhen, Q. Fang, J. Sun, W. Liu, W. Jiang, H. Bao, and X. Zhou, `` SMAP : Single-shot multi-person absolute 3 D pose estimation,'' in Proc. European Conference on Computer Vision. 1em plus 0.5em minus 0.4em Springer, 2020, pp. 550--566

  46. [54]

    Paszke, S

    A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga et al., `` PyTorch : An imperative style, high-performance deep learning library,'' Proc. Adv. Neural Inform. Process. Syst., vol. 32, 2019

  47. [55]

    Hinton, N

    G. Hinton, N. Srivastava, and K. Swersky, ``Neural networks for machine learning lecture 6a overview of mini-batch gradient descent,'' Cited on, vol. 14, no. 8, p. 2, 2012

  48. [56]

    Huang, J

    B. Huang, J. Ju, Y. Shu, and Y. Wang, ``Simultaneously recovering multi-person meshes and multi-view cameras with human semantics,'' IEEE Transactions on Circuits and Systems for Video Technology, vol. 34, no. 6, pp. 4229--4242, 2023

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.