Pith. sign in

REVIEW 3 cited by

DirectPose: Direct End-to-End Multi-Person Pose Estimation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1911.07451 v2 pith:NVEOLCTT submitted 2019-11-18 cs.CV

classification cs.CV
keywords end-to-endframeworkestimationmulti-personposealignmentbottom-upbounding-boxes
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We propose the first direct end-to-end multi-person pose estimation framework, termed DirectPose. Inspired by recent anchor-free object detectors, which directly regress the two corners of target bounding-boxes, the proposed framework directly predicts instance-aware keypoints for all the instances from a raw input image, eliminating the need for heuristic grouping in bottom-up methods or bounding-box detection and RoI operations in top-down ones. We also propose a novel Keypoint Alignment (KPAlign) mechanism, which overcomes the main difficulty: lack of the alignment between the convolutional features and predictions in this end-to-end framework. KPAlign improves the framework's performance by a large margin while still keeping the framework end-to-end trainable. With the only postprocessing non-maximum suppression (NMS), our proposed framework can detect multi-person keypoints with or without bounding-boxes in a single shot. Experiments demonstrate that the end-to-end paradigm can achieve competitive or better performance than previous strong baselines, in both bottom-up and top-down methods. We hope that our end-to-end approach can provide a new perspective for the human pose estimation task.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. From Sharp to Blur: Unsupervised Domain Adaptation for 2D Human Pose Estimation Under Extreme Motion Blur Using Event Cameras

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A domain adaptation method uses event-derived motion to synthesize blur and iteratively cleans pseudo-labels, improving multi-person 2D pose estimation on blurry images without annotations.

  2. DETRPose: Real-Time End-to-End Multi-Person Pose Estimation via Modified Transformer Decoder and Novel Denoising Keypoints

    cs.CV 2025-06 conditional novelty 6.0 of 10

    DETRPose is a real-time end-to-end transformer model family for multi-person pose estimation that trains 5 to 10 times faster than RTMO.

  3. A New Teacher-Reviewer-Student Framework for Semi-supervised 2D Human Pose Estimation

    cs.CV 2025-01 conditional novelty 4.0 of 10

    A semi-supervised 2D human pose estimation method that adds EMA-based reviewer networks, multi-level feature supervision, and a keypoint-mix augmentation to improve accuracy with few labels.

Pith tools