REVIEW 15 cited by
Sapiens: Foundation for Human Vision Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We present Sapiens, a family of models for four fundamental human-centric vision tasks -- 2D pose estimation, body-part segmentation, depth estimation, and surface normal prediction. Our models natively support 1K high-resolution inference and are extremely easy to adapt for individual tasks by simply fine-tuning models pretrained on over 300 million in-the-wild human images. We observe that, given the same computational budget, self-supervised pretraining on a curated dataset of human images significantly boosts the performance for a diverse set of human-centric tasks. The resulting models exhibit remarkable generalization to in-the-wild data, even when labeled data is scarce or entirely synthetic. Our simple model design also brings scalability -- model performance across tasks improves as we scale the number of parameters from 0.3 to 2 billion. Sapiens consistently surpasses existing baselines across various human-centric benchmarks. We achieve significant improvements over the prior state-of-the-art on Humans-5K (pose) by 7.6 mAP, Humans-2K (part-seg) by 17.1 mIoU, Hi4D (depth) by 22.4% relative RMSE, and THuman2 (normal) by 53.5% relative angular error. Project page: https://about.meta.com/realitylabs/codecavatars/sapiens.
Forward citations
Cited by 15 Pith papers
-
STROKEVISION-BENCH: A Multimodal Video And 2D Pose Benchmark For Tracking Stroke Recovery
A new 1,000-video benchmark of stroke patients performing box-and-block sub-actions, with raw frames and 2D skeletons, establishes baseline action classification accuracy for seven models.
-
GRMM: Real-Time High-Fidelity Gaussian Morphable Head Model with Learned Residuals
GRMM combines a classic 3D face template with learned fine detail residuals to render controllable full-head avatars in real time.
-
Generative Video Matting
Fine-tuning Stable Video Diffusion with a flow-matching schedule, hybrid losses, and synthetic plus pseudo-labeled data yields a video matting model that beats regression-based baselines on humans and animals, includi...
-
Delay-constrained re-entry governs large-scale brain seizures and other network pathologies
An epilepsy modeling preprint claims delay-constrained re-entry of traveling excitation drives seizures and predicts 184 recorded seizures, but the submitted full text is an unrelated computer vision paper, so the cla...
-
Diffuman4D: 4D Consistent Human View Synthesis from Sparse-View Videos with Spatio-Temporal Diffusion Models
A sliding iterative denoising scheme that alternates spatial and temporal passes, combined with skeleton conditioning, lets a diffusion model create spatio-temporally consistent multi-view human videos from sparse inp...
-
Video Virtual Try-on with Conditional Diffusion Transformer Inpainter
ViTI reformulates video virtual try-on as conditional video inpainting with a full 3D attention diffusion transformer, and reports the best VFID score on VVT (2.121).
-
DreamCube: 3D Panorama Generation via Multi-plane Synchronization
A synchronized multi-plane adaptation of 2D diffusion operators enables seam-consistent cubemap generation, and DreamCube extends this to joint RGB-D panorama generation and 3D scene lifting.
-
Total-Editing: Head Avatar with Editable Appearance, Motion, and Lighting
Total-Editing is a unified 3D head avatar framework that separately controls appearance, motion, and lighting through an intrinsically decomposed neural radiance field, and reports stronger identity, expression, pose,...
-
PoseBH: Prototypical Multi-Dataset Training Beyond Human Pose Estimation
PoseBH unifies pose estimation across human, whole-body, and animal skeletons using nonparametric keypoint prototypes and cross-type self-supervision, improving animal-dataset accuracy while preserving human benchmark...
-
InfinityHuman: Towards Long-Term Audio-Driven Human
A coarse-to-fine audio-driven animation framework that uses pose-guided refinement and hand-specific reward learning to generate long, identity-stable talking videos.
-
HuGeDiff: 3D Human Generation via Diffusion with Gaussian Splatting
HuGeDiff generates 3D human avatars from text by training a diffusion model on 3D Gaussian parameters lifted from FLUX-generated synthetic images.
-
Part Segmentation of Human Meshes via Multi-View Human Parsing
A multi-view 2D human parsing backprojection pipeline generates pseudo-ground-truth labels for THuman2.1 meshes, and a PointTransformer trained on geometry alone reaches up to 74.4 mIoU when measured against those pse...
-
EgoAnimate: Generating Human Animations from Egocentric top-down Views
EgoAnimate synthesizes a frontal T-pose image from an egocentric top-down photo using a fine-tuned Stable Diffusion model, then animates it with off-the-shelf image-to-motion methods to produce an animatable avatar.
-
CuriosAI Submission to the EgoExo4D Proficiency Estimation Challenge 2025
On EgoExo4D proficiency estimation, a two-stage pipeline with zero-shot scenario recognition and per-scenario, per-view VideoMAE classifiers (47.8% validation) outperforms a Sapiens-2B multi-task model (43.6%).
- Controllable and Expressive One-Shot Video Head Swapping
Discussion (0). Sign in to comment.