REVIEW 4 cited by
HO-Cap: A Capture System and Dataset for 3D Reconstruction and Pose Tracking of Hand-Object Interaction
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We introduce a data capture system and a new dataset, HO-Cap, for 3D reconstruction and pose tracking of hands and objects in videos. The system leverages multiple RGBD cameras and a HoloLens headset for data collection, avoiding the use of expensive 3D scanners or mocap systems. We propose a semi-automatic method for annotating the shape and pose of hands and objects in the collected videos, significantly reducing the annotation time compared to manual labeling. With this system, we captured a video dataset of humans interacting with objects to perform various tasks, including simple pick-and-place actions, handovers between hands, and using objects according to their affordance, which can serve as human demonstrations for research in embodied AI and robot manipulation. Our data capture setup and annotation framework will be available for the community to use in reconstructing 3D shapes of objects and human hands and tracking their poses in videos.
Forward citations
Cited by 4 Pith papers
-
AnyHand: A Large-Scale Synthetic Dataset for RGB(-D) Hand Pose Estimation
Co-training HaMeR and WiLoR on AnyHand (2.5M single-hand + 4.1M hand-object RGB-D images) improves FreiHAND/HO-3D metrics and a lightweight depth-fusion model beats prior RGB-D methods.
-
DemoBridge: A Simulation-in-the-Loop Toolkit for Single-View Human Demonstration Retargeting
DemoBridge retargets single-view human hand demonstrations into physics-validated, collision-aware robot trajectories via whole-trajectory optimization and simulation-in-the-loop re-planning.
-
Glove2Hand: Synthesizing Natural Hand-Object Interaction from Multi-Modal Sensing Gloves
A 3D-Gaussian-plus-diffusion pipeline translates multi-modal glove HOI videos into photorealistic bare-hand videos, yielding the HandSense dataset that improves contact estimation and occluded tracking.
-
OpenEgo: A Large-Scale Multimodal Egocentric Dataset for Dexterous Manipulation
OpenEgo is a 1,107-hour unified egocentric manipulation dataset with standardized 21-joint hand poses and timestamped action language, plus a small validation showing a language-conditioned policy learns short-horizon...
Discussion (0). Continue with ORCID to comment.