Pith. sign in

REVIEW 3 cited by

Playing for 3D Human Recovery

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2110.07588 v3 pith:NRT4NSCO submitted 2021-10-14 cs.CV

classification cs.CV
keywords datagta-humanhumanrealrecoverydatasetdatasetsgame
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Image- and video-based 3D human recovery (i.e., pose and shape estimation) have achieved substantial progress. However, due to the prohibitive cost of motion capture, existing datasets are often limited in scale and diversity. In this work, we obtain massive human sequences by playing the video game with automatically annotated 3D ground truths. Specifically, we contribute GTA-Human, a large-scale 3D human dataset generated with the GTA-V game engine, featuring a highly diverse set of subjects, actions, and scenarios. More importantly, we study the use of game-playing data and obtain five major insights. First, game-playing data is surprisingly effective. A simple frame-based baseline trained on GTA-Human outperforms more sophisticated methods by a large margin. For video-based methods, GTA-Human is even on par with the in-domain training set. Second, we discover that synthetic data provides critical complements to the real data that is typically collected indoor. Our investigation into domain gap provides explanations for our data mixture strategies that are simple yet useful. Third, the scale of the dataset matters. The performance boost is closely related to the additional data available. A systematic study reveals the model sensitivity to data density from multiple key aspects. Fourth, the effectiveness of GTA-Human is also attributed to the rich collection of strong supervision labels (SMPL parameters), which are otherwise expensive to acquire in real datasets. Fifth, the benefits of synthetic data extend to larger models such as deeper convolutional neural networks (CNNs) and Transformers, for which a significant impact is also observed. We hope our work could pave the way for scaling up 3D human recovery to the real world. Homepage: https://caizhongang.github.io/projects/GTA-Human/

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. BlanketGen2-Fit3D: Synthetic Blanket Augmentation Towards Improving Real-World In-Bed Blanket Occluded Human Pose Estimation

    cs.CV 2025-01 conditional novelty 6.0 of 10

    Synthetic blanket augmentation of Fit3D improves ViTPose-B pose estimation on real blanket-occluded in-bed images by an absolute 2.3% PCK.

  2. SMPLest-X: Ultimate Scaling for Expressive Human Pose and Shape Estimation

    cs.CV 2025-01 conditional novelty 6.0 of 10

    Training on 40 datasets and larger vision transformers steadily improves whole-body and hand pose estimation, achieving state-of-the-art results on multiple benchmarks.

  3. Human4K: A Large-Scale 4K Multi-View Mocap Dataset for Whole-Body 3D Human Reconstruction

    cs.CV 2026-07 conditional novelty 5.0 of 10

    A new dataset of six million 4K multi-view frames with Vicon-mocap-derived SMPL-X annotations improves whole-body 3D human reconstruction when added to public training data.

Pith tools