REVIEW 4 cited by
Rethinking on Multi-Stage Networks for Human Pose Estimation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Existing pose estimation approaches fall into two categories: single-stage and multi-stage methods. While multi-stage methods are seemingly more suited for the task, their performance in current practice is not as good as single-stage methods. This work studies this issue. We argue that the current multi-stage methods' unsatisfactory performance comes from the insufficiency in various design choices. We propose several improvements, including the single-stage module design, cross stage feature aggregation, and coarse-to-fine supervision. The resulting method establishes the new state-of-the-art on both MS COCO and MPII Human Pose dataset, justifying the effectiveness of a multi-stage architecture. The source code is publicly available for further research.
Forward citations
Cited by 4 Pith papers
-
SPGNet: Semantic Prediction Guidance for Scene Parsing
SPGNet improves semantic segmentation by using a first stage's per-pixel predictions to re-weight features entering a second encoder-decoder stage.
-
Waterfall Transformer for Multi-person Pose Estimation
WTPose combines a Swin transformer backbone with a multi-scale waterfall module using dilated neighborhood attention and reports +1.2 AP over Swin-B on COCO val.
-
SPARK: Low Latency Single-Camera 3D Pose Estimation for Autonomous Racing using Keypoints
SPARK applies keypoint detection with YOLO models to monocular images for low-latency 3D pose estimation of racing opponents, claiming better accuracy and speed than prior camera methods on real racing data.
-
VISUALCENT: Visual Human Analysis using Dynamic Centroid Representation
VISUALCENT reports new state-of-the-art bottom-up human pose and instance segmentation numbers on COCO and OCHuman using disk-based keypoint heatmaps, per-keypoint offset fields, and dynamic keypoint-anchored mask clustering.
Discussion (0). Continue with ORCID to comment.