Pith. sign in

REVIEW 15 cited by

Sapiens: Foundation for Human Vision Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.12569 v3 pith:QZEFOG4O submitted 2024-08-22 cs.CV

classification cs.CV
keywords modelssapienstaskshumanhuman-centricacrossdatadepth
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present Sapiens, a family of models for four fundamental human-centric vision tasks -- 2D pose estimation, body-part segmentation, depth estimation, and surface normal prediction. Our models natively support 1K high-resolution inference and are extremely easy to adapt for individual tasks by simply fine-tuning models pretrained on over 300 million in-the-wild human images. We observe that, given the same computational budget, self-supervised pretraining on a curated dataset of human images significantly boosts the performance for a diverse set of human-centric tasks. The resulting models exhibit remarkable generalization to in-the-wild data, even when labeled data is scarce or entirely synthetic. Our simple model design also brings scalability -- model performance across tasks improves as we scale the number of parameters from 0.3 to 2 billion. Sapiens consistently surpasses existing baselines across various human-centric benchmarks. We achieve significant improvements over the prior state-of-the-art on Humans-5K (pose) by 7.6 mAP, Humans-2K (part-seg) by 17.1 mIoU, Hi4D (depth) by 22.4% relative RMSE, and THuman2 (normal) by 53.5% relative angular error. Project page: https://about.meta.com/realitylabs/codecavatars/sapiens.

Discussion (0). Sign in to comment.

Forward citations

Cited by 15 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. STROKEVISION-BENCH: A Multimodal Video And 2D Pose Benchmark For Tracking Stroke Recovery

    eess.IV 2025-09 conditional novelty 6.0 of 10

    A new 1,000-video benchmark of stroke patients performing box-and-block sub-actions, with raw frames and 2D skeletons, establishes baseline action classification accuracy for seven models.

  2. GRMM: Real-Time High-Fidelity Gaussian Morphable Head Model with Learned Residuals

    cs.GR 2025-09 conditional novelty 6.0 of 10

    GRMM combines a classic 3D face template with learned fine detail residuals to render controllable full-head avatars in real time.

  3. Generative Video Matting

    cs.CV 2025-08 conditional novelty 6.0 of 10

    Fine-tuning Stable Video Diffusion with a flow-matching schedule, hybrid losses, and synthetic plus pseudo-labeled data yields a video matting model that beats regression-based baselines on humans and animals, includi...

  4. Delay-constrained re-entry governs large-scale brain seizures and other network pathologies

    q-bio.NC 2025-08 unverdicted novelty 6.0 of 10

    An epilepsy modeling preprint claims delay-constrained re-entry of traveling excitation drives seizures and predicts 184 recorded seizures, but the submitted full text is an unrelated computer vision paper, so the cla...

  5. Diffuman4D: 4D Consistent Human View Synthesis from Sparse-View Videos with Spatio-Temporal Diffusion Models

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A sliding iterative denoising scheme that alternates spatial and temporal passes, combined with skeleton conditioning, lets a diffusion model create spatio-temporally consistent multi-view human videos from sparse inp...

  6. Video Virtual Try-on with Conditional Diffusion Transformer Inpainter

    cs.CV 2025-06 conditional novelty 6.0 of 10

    ViTI reformulates video virtual try-on as conditional video inpainting with a full 3D attention diffusion transformer, and reports the best VFID score on VVT (2.121).

  7. DreamCube: 3D Panorama Generation via Multi-plane Synchronization

    cs.GR 2025-06 conditional novelty 6.0 of 10

    A synchronized multi-plane adaptation of 2D diffusion operators enables seam-consistent cubemap generation, and DreamCube extends this to joint RGB-D panorama generation and 3D scene lifting.

  8. Total-Editing: Head Avatar with Editable Appearance, Motion, and Lighting

    cs.CV 2025-05 conditional novelty 6.0 of 10

    Total-Editing is a unified 3D head avatar framework that separately controls appearance, motion, and lighting through an intrinsically decomposed neural radiance field, and reports stronger identity, expression, pose,...

  9. PoseBH: Prototypical Multi-Dataset Training Beyond Human Pose Estimation

    cs.CV 2025-05 conditional novelty 6.0 of 10

    PoseBH unifies pose estimation across human, whole-body, and animal skeletons using nonparametric keypoint prototypes and cross-type self-supervision, improving animal-dataset accuracy while preserving human benchmark...

  10. InfinityHuman: Towards Long-Term Audio-Driven Human

    cs.CV 2025-08 conditional novelty 5.0 of 10

    A coarse-to-fine audio-driven animation framework that uses pose-guided refinement and hand-specific reward learning to generate long, identity-stable talking videos.

  11. HuGeDiff: 3D Human Generation via Diffusion with Gaussian Splatting

    cs.CV 2025-06 conditional novelty 5.0 of 10

    HuGeDiff generates 3D human avatars from text by training a diffusion model on 3D Gaussian parameters lifted from FLUX-generated synthetic images.

  12. Part Segmentation of Human Meshes via Multi-View Human Parsing

    cs.CV 2025-07 conditional novelty 4.0 of 10

    A multi-view 2D human parsing backprojection pipeline generates pseudo-ground-truth labels for THuman2.1 meshes, and a PointTransformer trained on geometry alone reaches up to 74.4 mIoU when measured against those pse...

  13. EgoAnimate: Generating Human Animations from Egocentric top-down Views

    cs.CV 2025-07 conditional novelty 4.0 of 10

    EgoAnimate synthesizes a frontal T-pose image from an egocentric top-down photo using a fine-tuned Stable Diffusion model, then animates it with off-the-shelf image-to-motion methods to produce an animatable avatar.

  14. CuriosAI Submission to the EgoExo4D Proficiency Estimation Challenge 2025

    cs.CV 2025-07 conditional novelty 4.0 of 10

    On EgoExo4D proficiency estimation, a two-stage pipeline with zero-shot scenario recognition and per-scenario, per-view VideoMAE classifiers (47.8% validation) outperforms a Sapiens-2B multi-task model (43.6%).

  15. Controllable and Expressive One-Shot Video Head Swapping

    cs.CV 2025-06

Pith tools