REVIEW 3 cited by
A Graph-Based Approach for Category-Agnostic Pose Estimation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Traditional 2D pose estimation models are limited by their category-specific design, making them suitable only for predefined object categories. This restriction becomes particularly challenging when dealing with novel objects due to the lack of relevant training data. To address this limitation, category-agnostic pose estimation (CAPE) was introduced. CAPE aims to enable keypoint localization for arbitrary object categories using a few-shot single model, requiring minimal support images with annotated keypoints. We present a significant departure from conventional CAPE techniques, which treat keypoints as isolated entities, by treating the input pose data as a graph. We leverage the inherent geometrical relations between keypoints through a graph-based network to break symmetry, preserve structure, and better handle occlusions. We validate our approach on the MP-100 benchmark, a comprehensive dataset comprising over 20,000 images spanning over 100 categories. Our solution boosts performance by 0.98% under a 1-shot setting, achieving a new state-of-the-art for CAPE. Additionally, we enhance the dataset with skeleton annotations. Our code and data are publicly available.
Forward citations
Cited by 3 Pith papers
-
ProtoSnap: Prototype Alignment for Cuneiform Signs
ProtoSnap recovers the internal stroke structure of photographed cuneiform signs by snapping font-derived skeleton prototypes onto target images, and uses the resulting alignments to generate synthetic data that impro...
-
Unified Understanding of Environment, Task, and Human for Human-Robot Interaction in Real-World Environments
A service robot using a multi-layer indoor map and GPT-4 task representations served ordered items correctly in 37 of 41 trials in a simulated restaurant, with some operator and customer help.
-
Skel3D: Skeleton Guided Novel View Synthesis
Skel3D conditions a Free3D-style diffusion model on target-view skeleton images, showing small but significant metric gains on Objaverse, yet the evaluation relies on oracle skeletons rather than skeletons estimated f...
Discussion (0). Continue with ORCID to comment.