REVIEW 4 cited by
Pose Flow: Efficient Online Pose Tracking
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Multi-person articulated pose tracking in unconstrained videos is an important while challenging problem. In this paper, going along the road of top-down approaches, we propose a decent and efficient pose tracker based on pose flows. First, we design an online optimization framework to build the association of cross-frame poses and form pose flows (PF-Builder). Second, a novel pose flow non-maximum suppression (PF-NMS) is designed to robustly reduce redundant pose flows and re-link temporal disjoint ones. Extensive experiments show that our method significantly outperforms best-reported results on two standard Pose Tracking datasets by 13 mAP 25 MOTA and 6 mAP 3 MOTA respectively. Moreover, in the case of working on detected poses in individual frames, the extra computation of pose tracker is very minor, guaranteeing online 10FPS tracking. Our source codes are made publicly available(https://github.com/YuliangXiu/PoseFlow).
Forward citations
Cited by 4 Pith papers
-
SpatioTemporal Learning for Human Pose Estimation in Sparsely-Labeled Videos
STDPose combines a Dynamic-Aware Mask, temporal feature and heatmap fusion, and a mutual information loss to improve video pose propagation and pose estimation on sparsely-labeled videos, achieving new state-of-the-ar...
-
An End-to-End Framework for Video Multi-Person Pose Estimation
An end-to-end video pose transformer built on PETR with spatio-temporal encoders and an instance consistency loss reaches 83.0 mAP on PoseTrack2017 and appears around 4x faster than DCPose.
-
FashionPose: Unified Text-Driven Fashion Synthesis with Joint Geometric and Photometric Control
A single caption can drive pose generation, person-image synthesis, and relighting through a three-stage FashionPose pipeline, with reported text-to-pose gains on DF-PASS that are undermined by inconsistent tables.
-
Optimizing Human Pose Estimation Through Focused Human and Joint Regions
VREMD combines human and keypoint masks with bidirectional deformable cross-attention to reach state-of-the-art mAP on three PoseTrack benchmarks.
Discussion (0). Continue with ORCID to comment.