REVIEW 5 major objections 3 minor 56 references
DyCrowd: Towards Dynamic Crowd Reconstruction from a Large-scene Video
T0 review · 5 major / 3 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read A single large-scene video can yield consistent 3D poses, positions, and shapes for hundreds of people over time.
desk verdict A plausible video-based crowd reconstruction idea with a genuinely new task framing, but the SOTA claim currently rests on an authors-made synthetic benchmark and the supplied text is unreadable, so the evidence cannot yet be checked. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is segment-level group-guided optimization. Motion sequences are divided into segments, individuals with similar motion segments are clustered, and each cluster's motion is optimized together in a coarse-to-fine manner; within a cluster, visible unoccluded segments act as the temporal reference for reconstructing occluded segments. The VAE-based human motion prior constrains optimizations to plausible body motions, while the AMC loss enforces consistency between group members despite asynchrony and rhythmic variation. This turns occlusion from a per-person missing-data problem into a group inference problem.
What would settle it
Take a video where a person walks behind a long wall while everyone else in view either stands still or moves differently. If DyCrowd reconstructs the hidden walker's motion as following the standing crowd rather than continuing along the observed pre- and post-occlusion path, the group-guidance premise is falsified; a multi-camera ground-truth capture of the same scene would settle it.
Extended reading notes
Core claim
The central claim is that DyCrowd is the first framework to reconstruct hundreds of individuals' poses, positions and shapes with spatio-temporal consistency from a single large-scene video. The argument rests on a coarse-to-fine, group-guided motion optimization: motion sequences are split into segments, similar segments are grouped, and the group's motions are optimized jointly so that reliable, unoccluded segments transfer their temporal information to occluded ones. A VAE-based human motion prior keeps reconstructed motions plausible, and the Asynchronous Motion Consistency (AMC) loss makes the transfer robust to people moving out of sync or at different rhythms. The paper reports that this method achieves state-of-the-art performance on the large-scene dynamic crowd reconstruction task, evaluated on its new VirtualCrowd benchmark.
Load-bearing premise
The method bets that people with similar motion can be grouped and that seeing some group members clearly is enough to reconstruct the hidden members; if a hidden person's motion resembles no one visible, or the entire group is occluded at once, the temporal guidance has nothing to draw on.
Editorial extensions
If this is right
- Large-scene video becomes a viable input for crowd reconstruction, giving temporal consistency that static-image methods cannot offer.
- Long-term occlusions can be resolved by borrowing movement patterns from similarly behaving people who are visible, rather than relying only on the occluded person's own observed past.
- Reconstruction quality should scale with crowd regularity: a stream of pedestrians with similar gaits should recover occluded members better than a crowd of independent, dissimilar actors.
- The VirtualCrowd benchmark provides a quantitative way to compare methods on temporal consistency and occlusion recovery in large scenes.
Reading between the lines
- The method's advantage should grow with crowd density, since dense crowds provide more similar visible segments to guide each occluded person, until everyone is simultaneously hidden.
- Group-guided temporal transfer of the same kind could apply to other correlated multi-object scenes, such as vehicle traffic or animal herds, wherever motion is shared enough to form reliable groups.
- A concrete stress test is to move the VAE motion prior outside its training distribution: unusual gaits or choreographed actions may remain unrecoverable even when visible peers move similarly, because the prior will pull reconstructions back to familiar motions.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript introduces DyCrowd, a framework for reconstructing the 3D poses, positions, and shapes of hundreds of people over time from a single large-scene video. The proposed method combines coarse-to-fine group-guided motion optimization with a VAE-based human motion prior and an Asynchronous Motion Consistency (AMC) loss, targeting long-term dynamic occlusions in large scenes. The authors also present VirtualCrowd, a synthetic benchmark dataset for evaluating dynamic crowd reconstruction. The abstract claims state-of-the-art performance, but the submitted full text is encoding-corrupted and largely unreadable, so the methodology, equations, experimental setup, and quantitative comparisons cannot be verified from the available material.
Significance. The task addressed by DyCrowd is timely and practically relevant for city surveillance and crowd analysis, and the goal of spatio-temporally consistent reconstruction of hundreds of individuals from a single video is a genuine gap in the current literature. The contribution of a benchmark dataset, even a synthetic one, is useful if it is carefully validated and released. However, the significance of the work cannot currently be assessed: the central state-of-the-art claim rests on the abstract alone, no quantitative metrics are reported in readable form, and the only named evaluation artifact is a synthetic dataset constructed by the same authors. If the method and dataset are made fully readable and are validated against real-world large-scene data, the contribution could be substantial; at present it remains unsubstantiated.
major comments (5)
- [Abstract and Full Text] The central claim of state-of-the-art performance is not supported by any readable quantitative result. The abstract provides no metrics, ablations, or error bars, and the full text is encoding-corrupted, making the comparison tables and experimental sections inaccessible. The authors should provide a readable manuscript with concrete numbers on the VirtualCrowd benchmark and, ideally, on at least one real-world large-scene video with annotations or proxy metrics.
- [VirtualCrowd evaluation and circularity] The only evaluation benchmark named in the paper, VirtualCrowd, is introduced by the authors, and the abstract justifies its creation by the absence of existing well-annotated large-scene video datasets. If the synthetic data are generated using the same group-motion assumptions and occlusion patterns that DyCrowd explicitly encodes, then the reported state-of-the-art result would be at least partly self-confirming. The paper should describe the data generation process in detail, provide statistics on occlusion rates and motion diversity, and include a transfer experiment to real video to demonstrate that the method generalizes beyond the synthetic distribution.
- [Group-guided occlusion mechanism] The core mechanism assumes that visible, unoccluded members of a group with similar motion segments can reliably guide the reconstruction of occluded members. The abstract does not address failure cases where an occluded individual has a unique motion not shared by any visible group member, or where an entire group becomes occluded simultaneously. Without an analysis of these failure modes or an ablation that varies occlusion severity and group size, the claim of 'robust and plausible motion recovery' remains unsupported.
- [Provenance of the submitted file] The full text contains the line 'arXiv:2508.12645v5 [cs.IR] 18 Jan 2026', which does not match the claimed identifier '2508.12644' nor the subject area cs.CV. This mismatch raises a provenance concern: the authors should confirm that the correct version of the paper was uploaded, and the inserted line should be removed from the camera-ready version.
- [Readability of the full text] The submitted full text is not readable due to encoding corruption; the equations, including the definitions of the AMC loss and the VAE motion prior, along with all tables and algorithm pseudocode, are unreadable. This prevents the verification of every technical contribution and every experimental result, and it must be fixed before any further assessment can occur.
minor comments (3)
- [Abstract] The abstract promises that 'code and dataset will be available for research purposes' but gives no details on licenses or access conditions; please specify the intended release terms.
- [Abstract] The acronym 'AMC' is used without a definition in the abstract; please add a brief clarification when the loss is first mentioned.
- [General] Once the full text is readable, please ensure that all tables report error bars or statistical significance, and that ablations isolate the contributions of the group-guided optimization, the VAE prior, and the AMC loss.
Circularity Check
No circular derivation is established from the available text; the SOTA claim rests on an author-provided synthetic benchmark, which is a generalization concern rather than a definitional circularity.
full rationale
The supplied full text is encoding-corrupted, so the equations, losses, ablations, and comparison tables cannot be checked. From the abstract alone, DyCrowd's components are not defined in terms of the target reconstruction (no self-definition), no fitted parameter is relabeled as a prediction, and no load-bearing claim reduces to a self-citation. The only identified concern is that the state-of-the-art claim is supported by VirtualCrowd, a benchmark contributed by the same paper rather than an external benchmark. That makes the evaluation self-referential in provenance, but it is not, on the quoted evidence, circular by construction: the abstract does not state that the benchmark generator encodes the method's group-motion assumptions, and no equation-level reduction is available to exhibit. Under the hard rule requiring a quoted reduction, the appropriate circularity finding is no significant circularity; the missing external validation and the provenance mismatch of the inserted arXiv identifier affect verifiability and correctness risk, not the circularity score.
Assumptions & free parameters
assumptions (3)
- domain assumption A VAE trained on human motion provides a realistic prior for occluded poses over time.
- domain assumption People with similar motion segments can be clustered, and visible members can guide invisible members.
- domain assumption VirtualCrowd, a synthetic benchmark, is representative of real large-scene crowd videos.
Cite this review
Pith. "Pith review of DyCrowd: Towards Dynamic Crowd Reconstruction from a Large-scene Video." pith.science (2026). https://pith.science/paper/NTLYO3TL
@misc{pith2026250812644,
author = {Pith},
title = {Pith review of: DyCrowd: Towards Dynamic Crowd Reconstruction from a Large-scene Video},
year = {2026},
howpublished = {\url{https://pith.science/paper/NTLYO3TL}},
note = {Machine review of arXiv:2508.12644}
}
read the original abstract
3D reconstruction of dynamic crowds in large scenes has become increasingly important for applications such as city surveillance and crowd analysis. However, current works attempt to reconstruct 3D crowds from a static image, causing a lack of temporal consistency and inability to alleviate the typical impact caused by occlusions. In this paper, we propose DyCrowd, the first framework for spatio-temporally consistent 3D reconstruction of hundreds of individuals' poses, positions and shapes from a large-scene video. We design a coarse-to-fine group-guided motion optimization strategy for occlusion-robust crowd reconstruction in large scenes. To address temporal instability and severe occlusions, we further incorporate a VAE (Variational Autoencoder)-based human motion prior along with a segment-level group-guided optimization. The core of our strategy leverages collective crowd behavior to address long-term dynamic occlusions. By jointly optimizing the motion sequences of individuals with similar motion segments and combining this with the proposed Asynchronous Motion Consistency (AMC) loss, we enable high-quality unoccluded motion segments to guide the motion recovery of occluded ones, ensuring robust and plausible motion recovery even in the presence of temporal desynchronization and rhythmic inconsistencies. Additionally, in order to fill the gap of no existing well-annotated large-scene video dataset, we contribute a virtual benchmark dataset, VirtualCrowd, for evaluating dynamic crowd reconstruction from large-scene videos. Experimental results demonstrate that the proposed method achieves state-of-the-art performance in the large-scene dynamic crowd reconstruction task. The code and dataset will be available for research purposes.
Reference graph
Works this paper leans on
-
[1]
H. Wen, J. Huang, H. Cui, H. Lin, Y.-K. Lai, L. Fang, and K. Li, ``Crowd3 D : Towards hundreds of people reconstruction from a single image,'' in Proc. IEEE Conference on Computer Vision and Pattern Recognition, 2023, pp. 8937--8946
work page 2023
- [2]
-
[3]
M. Kocabas, N. Athanasiou, and M. J. Black, `` VIBE : Video inference for human body pose and shape estimation,'' in Proc. IEEE Conference on Computer Vision and Pattern Recognition, 2020, pp. 5253--5263
work page 2020
-
[4]
H. Choi, G. Moon, J. Y. Chang, and K. M. Lee, ``Beyond static features for temporally consistent 3 D human pose and shape from a video,'' in Proc. IEEE Conference on Computer Vision and Pattern Recognition, 2021, pp. 1964--1973
work page 2021
- [5]
-
[6]
X. Shen, Z. Yang, X. Wang, J. Ma, C. Zhou, and Y. Yang, ``Global-to-local modeling for video-based 3 D human pose and shape estimation,'' in Proc. IEEE Conference on Computer Vision and Pattern Recognition, 2023, pp. 8887--8896
work page 2023
-
[7]
Y. Yuan, U. Iqbal, P. Molchanov, K. Kitani, and J. Kautz, `` GLAMR : Global occlusion-aware human mesh recovery with dynamic cameras,'' in Proc. IEEE Conference on Computer Vision and Pattern Recognition, 2022, pp. 11\,038--11\,049
work page 2022
-
[8]
Y. Wang, Z. Wang, L. Liu, and K. Daniilidis, `` TRAM : Global trajectory and motion of 3 D humans from in-the-wild videos,'' arXiv preprint arXiv:2403.17346, 2024
arXiv 2024
Show all 56 references
-
[9]
K. Li, Y. Liu, Y.-K. Lai, and J. Yang, `` MILI : Multi-person inference from a low-resolution image,'' Fundamental Research, vol. 3, no. 3, pp. 434--441, 2023
2023
-
[10]
Y. Sun, Q. Bao, W. Liu, T. Mei, and M. J. Black, `` TRACE : 5 D temporal regression of avatars with dynamic cameras in 3 D environments,'' in Proc. IEEE Conference on Computer Vision and Pattern Recognition, 2023, pp. 8856--8866
2023
-
[11]
Z. Qiu, Q. Yang, J. Wang, H. Feng, J. Han, E. Ding, C. Xu, D. Fu, and J. Wang, `` PSVT : End-to-end multi-person 3 D pose and shape estimation with progressive video transformers,'' in Proc. IEEE Conference on Computer Vision and Pattern Recognition, 2023, pp. 21\,254--21\,263
2023
-
[12]
V. Ye, G. Pavlakos, J. Malik, and A. Kanazawa, ``Decoupling human and camera motion from videos in the wild,'' in Proc. IEEE Conference on Computer Vision and Pattern Recognition, 2023, pp. 21\,222--21\,232
2023
-
[13]
Kocabas, Y
M. Kocabas, Y. Yuan, P. Molchanov, Y. Guo, M. J. Black, O. Hilliges, J. Kautz, and U. Iqbal, `` PACE : Human and camera motion estimation from in-the-wild videos,'' in Proc. IEEE Int. Conf. 3D Vis. 1em plus 0.5em minus 0.4em IEEE, 2024, pp. 397--408
2024
-
[14]
B. Zhou, X. Tang, and X. Wang, ``Measuring crowd collectiveness,'' in Proc. IEEE Conference on Computer Vision and Pattern Recognition, 2013, pp. 3049--3056
2013
-
[15]
Y. Wu, Y. Ye, C. Zhao, and Z. Shi, ``Collective density clustering for coherent motion detection,'' IEEE Trans. Multimedia, vol. 20, no. 6, pp. 1418--1431, 2017
2017
-
[16]
L. Mei, J. Lai, Z. Chen, and X. Xie, ``Measuring crowd collectiveness via global motion correlation,'' in Proc. IEEE International Conference on Computer Vision, 2019, pp. 0--0
2019
-
[17]
Zhang and S
Y. Zhang and S. Tang, ``The wanderings of odysseus in 3 D scenes,'' in Proc. IEEE Conference on Computer Vision and Pattern Recognition, 2022, pp. 20\,481--20\,491
2022
-
[18]
K. Zhao, Y. Zhang, S. Wang, T. Beeler, and S. Tang, ``Synthesizing diverse human motions in 3 D indoor scenes,'' in Proc. IEEE International Conference on Computer Vision, 2023, pp. 14\,738--14\,749
2023
-
[19]
[Online]
iCity3D , `` iCity3D ,'' 2024. [Online]. Available: https://icity3d.com/
2024
-
[20]
Pavlakos, V
G. Pavlakos, V. Choutas, N. Ghorbani, T. Bolkart, A. A. Osman, D. Tzionas, and M. J. Black, ``Expressive body capture: 3 D hands, face, and body from a single image,'' in Proc. IEEE Conference on Computer Vision and Pattern Recognition, 2019, pp. 10\,975--10\,985
2019
-
[21]
D. J. Brady, M. E. Gehm, R. A. Stack, D. L. Marks, D. S. Kittle, D. R. Golish, E. Vera, and S. D. Feller, ``Multiscale gigapixel photography,'' Nature, vol. 486, no. 7403, pp. 386--389, 2012
2012
-
[22]
X. Yuan, L. Fang, Q. Dai, D. J. Brady, and Y. Liu, ``Multiscale gigapixel video: A cross resolution image matching and warping approach,'' in 2017 IEEE International Conference on Computational Photography (ICCP). 1em plus 0.5em minus 0.4em IEEE, 2017, pp. 1--9
2017
-
[23]
X. Wang, X. Zhang, Y. Zhu, Y. Guo, X. Yuan, L. Xiang, Z. Wang, G. Ding, D. Brady, Q. Dai et al., `` PANDA : A gigapixel-level human-centric video dataset,'' in Proc. IEEE Conference on Computer Vision and Pattern Recognition, 2020, pp. 3268--3278
2020
-
[24]
C. Liu, H. Wei, J. Yang, J. Liu, W. Li, Y. Guo, and L. Fang, `` GigaHumanDet : Exploring full-body detection on gigapixel-level images,'' in Proc. AAAI Conference on Artificial Intelligence, vol. 38, no. 9, 2024, pp. 10\,092--10\,100
2024
-
[25]
W. Mo, W. Zhang, H. Wei, R. Cao, Y. Ke, and Y. Luo, `` PVDet : Towards pedestrian and vehicle detection on gigapixel-level images,'' Engineering Applications of Artificial Intelligence, vol. 118, p. 105705, 2023
2023
-
[26]
W. Li, R. Zhang, H. Lin, Y. Guo, C. Ma, and X. Yang, `` SaccadeDet : A novel dual-stage architecture for rapid and accurate detection in gigapixel images,'' in Joint European Conference on Machine Learning and Knowledge Discovery in Databases. 1em plus 0.5em minus 0.4em Spring...
2024
-
[27]
Zhang, L
J. Zhang, L. Gu, Y.-K. Lai, X. Wang, and K. Li, ``Towards grouping in large scenes with occlusion-aware spatio-temporal transformers,'' IEEE Transactions on Circuits and Systems for Video Technology, 2023
2023
-
[28]
X. Xu, H. Chen, F. Moreno-Noguer, L. A. Jeni, and F. De la Torre, ``3 D human pose, shape and texture from low-resolution images and videos,'' IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 44, no. 9, pp. 4490--4504, 2021
2021
-
[29]
S. Shin, J. Kim, E. Halilaj, and M. J. Black, `` WHAM : Reconstructing world-grounded humans with accurate 3 D motion,'' in Proc. IEEE Conference on Computer Vision and Pattern Recognition, 2024, pp. 2070--2080
2024
-
[30]
Zhang, D
J. Zhang, D. Yu, J. H. Liew, X. Nie, and J. Feng, ``Body meshes as points,'' in Proc. IEEE Conference on Computer Vision and Pattern Recognition, 2021, pp. 546--556
2021
-
[31]
Y. Sun, Q. Bao, W. Liu, Y. Fu, M. J. Black, and T. Mei, ``Monocular, one-stage, regression of multiple 3 D people,'' in Proc. IEEE International Conference on Computer Vision, 2021, pp. 11\,179--11\,188
2021
-
[32]
Y. Sun, W. Liu, Q. Bao, Y. Fu, T. Mei, and M. J. Black, ``Putting people in their place: Monocular regression of 3 D people in depth,'' in Proc. IEEE Conference on Computer Vision and Pattern Recognition, 2022, pp. 13\,243--13\,252
2022
-
[33]
Rempe, T
D. Rempe, T. Birdal, A. Hertzmann, J. Yang, S. Sridhar, and L. J. Guibas, `` HuMoR : 3 D human motion model for robust pose estimation,'' in Proc. IEEE International Conference on Computer Vision, 2021, pp. 11\,488--11\,499
2021
-
[34]
C. He, J. Saito, J. Zachary, H. Rushmeier, and Y. Zhou, `` NeMF : Neural motion fields for kinematic animation,'' Proc. Adv. Neural Inform. Process. Syst., vol. 35, pp. 4244--4256, 2022
2022
-
[35]
Loper, N
M. Loper, N. Mahmood, J. Romero, G. Pons-Moll, and M. J. Black, `` SMPL : A skinned multi-person linear model,'' ACM Trans. Graphics (Proc. SIGGRAPH Asia), vol. 34, no. 6, pp. 248:1--248:16, Oct. 2015
2015
-
[36]
Y. Li, H. Mao, R. Girshick, and K. He, ``Exploring plain vision transformer backbones for object detection,'' in Proc. European Conference on Computer Vision. 1em plus 0.5em minus 0.4em Springer, 2022, pp. 280--296
2022
-
[37]
Z. Yang, A. Zeng, C. Yuan, and Y. Li, ``Effective whole-body pose estimation with two-stages distillation,'' in Proc. IEEE International Conference on Computer Vision, 2023, pp. 4210--4220
2023
-
[38]
S. Goel, G. Pavlakos, J. Rajasegaran, A. Kanazawa, and J. Malik, ``Humans in 4 D : Reconstructing and tracking humans with transformers,'' in Proc. IEEE International Conference on Computer Vision, 2023, pp. 14\,783--14\,794
2023
-
[39]
Rajasegaran, G
J. Rajasegaran, G. Pavlakos, A. Kanazawa, and J. Malik, ``Tracking people by predicting 3 D appearance, location and pose,'' in Proc. IEEE Conference on Computer Vision and Pattern Recognition, 2022, pp. 2740--2749
2022
-
[40]
Pavlakos, V
G. Pavlakos, V. Choutas, N. Ghorbani, T. Bolkart, A. A. A. Osman, D. Tzionas, and M. J. Black, ``Expressive body capture: 3 D hands, face, and body from a single image,'' in Proc. IEEE Conference on Computer Vision and Pattern Recognition, 2019
2019
-
[41]
X. Chen, B. Jiang, W. Liu, Z. Huang, B. Fu, T. Chen, and G. Yu, ``Executing your commands via motion diffusion in latent space,'' in Proc. IEEE Conference on Computer Vision and Pattern Recognition, 2023, pp. 18\,000--18\,010
2023
-
[42]
Rempe, L
D. Rempe, L. J. Guibas, A. Hertzmann, B. Russell, R. Villegas, and J. Yang, ``Contact and human dynamics from monocular video,'' in Proc. European Conference on Computer Vision. 1em plus 0.5em minus 0.4em Springer, 2020, pp. 71--87
2020
-
[43]
Zhang, Y
S. Zhang, Y. Zhang, F. Bogo, M. Pollefeys, and S. Tang, ``Learning motion priors for 4d human body capture in 3 D scenes,'' in Proc. IEEE International Conference on Computer Vision, 2021, pp. 11\,343--11\,353
2021
-
[44]
C. Guo, S. Zou, X. Zuo, S. Wang, W. Ji, X. Li, and L. Cheng, ``Generating diverse and natural 3 D human motions from text,'' in Proc. IEEE Conference on Computer Vision and Pattern Recognition, 2022, pp. 5152--5161
2022
-
[45]
Zhang, B
S. Zhang, B. L. Bhatnagar, Y. Xu, A. Winkler, P. Kadlecek, S. Tang, and F. Bogo, `` RoHM : Robust human motion reconstruction via diffusion,'' in Proc. IEEE Conference on Computer Vision and Pattern Recognition, 2024, pp. 14\,606--14\,617
2024
-
[46]
Mahmood, N
N. Mahmood, N. Ghorbani, N. F. Troje, G. Pons-Moll, and M. J. Black, `` AMASS : Archive of motion capture as surface shapes,'' in Proc. IEEE International Conference on Computer Vision, 2019, pp. 5442--5451
2019
-
[47]
Besse, B
P. Besse, B. Guillouet, J.-M. Loubes, and R. Fran c ois, ``Review and perspective for distance based trajectory clustering,'' arXiv preprint arXiv:1508.04904, 2015
2015 arXiv
-
[48]
B. J. Frey and D. Dueck, ``Clustering by passing messages between data points,'' science, vol. 315, no. 5814, pp. 972--976, 2007
2007
-
[49]
Cuturi and M
M. Cuturi and M. Blondel, `` Soft-DTW : a differentiable loss function for time-series,'' in International conference on machine learning. 1em plus 0.5em minus 0.4em PMLR, 2017, pp. 894--903
2017
-
[50]
Sakoe and S
H. Sakoe and S. Chiba, ``Dynamic programming algorithm optimization for spoken word recognition,'' IEEE transactions on acoustics, speech, and signal processing, vol. 26, no. 1, pp. 43--49, 1978
1978
-
[51]
Z. Yang, Z. Cai, H. Mei, S. Liu, Z. Chen, W. Xiao, Y. Wei, Z. Qing, C. Wei, B. Dai et al., `` SynBody : Synthetic dataset with layered human models for 3 D human perception and modeling,'' in Proc. IEEE International Conference on Computer Vision, 2023, pp. 20\,282--20\,292
2023
-
[52]
X. Liu, T. Zhou, H. Kang, J. Ma, Z. Wang, J. Huang, W. Weng, Y.-K. Lai, and K. Li, `` RESCUE : Crowd evacuation simulation via controlling sdm-united characters,'' in 2025 International Conference on Computer Vision (ICCV), 2025
2025
-
[53]
J. Zhen, Q. Fang, J. Sun, W. Liu, W. Jiang, H. Bao, and X. Zhou, `` SMAP : Single-shot multi-person absolute 3 D pose estimation,'' in Proc. European Conference on Computer Vision. 1em plus 0.5em minus 0.4em Springer, 2020, pp. 550--566
2020
-
[54]
Paszke, S
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga et al., `` PyTorch : An imperative style, high-performance deep learning library,'' Proc. Adv. Neural Inform. Process. Syst., vol. 32, 2019
2019
-
[55]
Hinton, N
G. Hinton, N. Srivastava, and K. Swersky, ``Neural networks for machine learning lecture 6a overview of mini-batch gradient descent,'' Cited on, vol. 14, no. 8, p. 2, 2012
2012
-
[56]
Huang, J
B. Huang, J. Ju, Y. Shu, and Y. Wang, ``Simultaneously recovering multi-person meshes and multi-view cameras with human semantics,'' IEEE Transactions on Circuits and Systems for Video Technology, vol. 34, no. 6, pp. 4229--4242, 2023
2023
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.