Pith. sign in

REVIEW 4 cited by

Rethinking on Multi-Stage Networks for Human Pose Estimation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1901.00148 v4 pith:7Z3UUXN4 submitted 2019-01-01 cs.CV

classification cs.CV
keywords multi-stagemethodsposesingle-stagecurrentdesignestimationhuman
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Existing pose estimation approaches fall into two categories: single-stage and multi-stage methods. While multi-stage methods are seemingly more suited for the task, their performance in current practice is not as good as single-stage methods. This work studies this issue. We argue that the current multi-stage methods' unsatisfactory performance comes from the insufficiency in various design choices. We propose several improvements, including the single-stage module design, cross stage feature aggregation, and coarse-to-fine supervision. The resulting method establishes the new state-of-the-art on both MS COCO and MPII Human Pose dataset, justifying the effectiveness of a multi-stage architecture. The source code is publicly available for further research.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SPGNet: Semantic Prediction Guidance for Scene Parsing

    cs.CV 2019-08 conditional novelty 6.0 of 10

    SPGNet improves semantic segmentation by using a first stage's per-pixel predictions to re-weight features entering a second encoder-decoder stage.

  2. Waterfall Transformer for Multi-person Pose Estimation

    cs.CV 2024-11 conditional novelty 5.0 of 10

    WTPose combines a Swin transformer backbone with a multi-scale waterfall module using dilated neighborhood attention and reports +1.2 AP over Swin-B on COCO val.

  3. SPARK: Low Latency Single-Camera 3D Pose Estimation for Autonomous Racing using Keypoints

    cs.RO 2026-06 unverdicted novelty 4.0 of 10

    SPARK applies keypoint detection with YOLO models to monocular images for low-latency 3D pose estimation of racing opponents, claiming better accuracy and speed than prior camera methods on real racing data.

  4. VISUALCENT: Visual Human Analysis using Dynamic Centroid Representation

    cs.CV 2025-04 conditional novelty 4.0 of 10

    VISUALCENT reports new state-of-the-art bottom-up human pose and instance segmentation numbers on COCO and OCHuman using disk-based keypoint heatmaps, per-keypoint offset fields, and dynamic keypoint-anchored mask clustering.

Pith tools