Pith. sign in

REVIEW 5 cited by

ViTPose: Simple Vision Transformer Baselines for Human Pose Estimation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2204.12484 v3 pith:IB3QJV5X submitted 2022-04-26 cs.CV

classification cs.CV
keywords vitposemodelposeestimationknowledgesimpletransformersvision
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Although no specific domain knowledge is considered in the design, plain vision transformers have shown excellent performance in visual recognition tasks. However, little effort has been made to reveal the potential of such simple structures for pose estimation tasks. In this paper, we show the surprisingly good capabilities of plain vision transformers for pose estimation from various aspects, namely simplicity in model structure, scalability in model size, flexibility in training paradigm, and transferability of knowledge between models, through a simple baseline model called ViTPose. Specifically, ViTPose employs plain and non-hierarchical vision transformers as backbones to extract features for a given person instance and a lightweight decoder for pose estimation. It can be scaled up from 100M to 1B parameters by taking the advantages of the scalable model capacity and high parallelism of transformers, setting a new Pareto front between throughput and performance. Besides, ViTPose is very flexible regarding the attention type, input resolution, pre-training and finetuning strategy, as well as dealing with multiple pose tasks. We also empirically demonstrate that the knowledge of large ViTPose models can be easily transferred to small ones via a simple knowledge token. Experimental results show that our basic ViTPose model outperforms representative methods on the challenging MS COCO Keypoint Detection benchmark, while the largest model sets a new state-of-the-art. The code and models are available at https://github.com/ViTAE-Transformer/ViTPose.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. IMASHRIMP: Automatic White Shrimp (Penaeus vannamei) Biometrical Analysis from Laboratory Images Using Computer Vision and Deep Learning

    cs.CV 2025-07 conditional novelty 6.0 of 10

    IMASHRIMP combines ViTPose keypoint estimation, ResNet-50 discriminators, and SVM regression to measure 23 white-shrimp traits from RGBD images, reporting a mean error of 0.07 cm.

  2. Voice-guided Orchestrated Intelligence for Clinical Evaluation (VOICE): A Voice AI Agent System for Prehospital Stroke Assessment

    cs.HC 2025-06 conditional novelty 6.0 of 10

    A voice AI agent guided lay users through stroke assessments in a 10-scenario simulation, achieving 86% stroke sensitivity and 33% specificity, with a physician confident in only 40% of reports.

  3. Joint angle based learning to refine kinematic human pose estimation

    cs.CV 2025-07 conditional novelty 5.0 of 10

    Joint angle refinement (JAR), trained on Fourier-synthesized angle sequences and applied with a BiGRU-Attention network, smooths keypoint trajectories and corrects outliers in human pose estimates, outperforming Smoot...

  4. Wan-S2V: Audio-Driven Cinematic Video Generation

    cs.CV 2025-08 reject novelty 4.0 of 10

    Wan-S2V is an audio-driven video generator built on Wan, claiming better cinematic character animation than prior systems, though the evaluation is limited.

  5. Prompt Engineering in Segment Anything Model: Methodologies, Applications, and Emerging Challenges

    cs.CV 2025-07 conditional novelty 2.0 of 10

    A structured survey of prompt engineering methods for the Segment Anything Model, covering geometric, textual, and multimodal prompts and their applications.

Pith tools