REVIEW 14 cited by
Introducing HOT3D: An Egocentric Dataset for 3D Hand and Object Tracking
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
We introduce HOT3D, a publicly available dataset for egocentric hand and object tracking in 3D. The dataset offers over 833 minutes (more than 3.7M images) of multi-view RGB/monochrome image streams showing 19 subjects interacting with 33 diverse rigid objects, multi-modal signals such as eye gaze or scene point clouds, as well as comprehensive ground truth annotations including 3D poses of objects, hands, and cameras, and 3D models of hands and objects. In addition to simple pick-up/observe/put-down actions, HOT3D contains scenarios resembling typical actions in a kitchen, office, and living room environment. The dataset is recorded by two head-mounted devices from Meta: Project Aria, a research prototype of light-weight AR/AI glasses, and Quest 3, a production VR headset sold in millions of units. Ground-truth poses were obtained by a professional motion-capture system using small optical markers attached to hands and objects. Hand annotations are provided in the UmeTrack and MANO formats and objects are represented by 3D meshes with PBR materials obtained by an in-house scanner. We aim to accelerate research on egocentric hand-object interaction by making the HOT3D dataset publicly available and by co-organizing public challenges on the dataset at ECCV 2024. The dataset can be downloaded from the project website: https://facebookresearch.github.io/hot3d/.
Forward citations
Cited by 14 Pith papers
-
EgoPolice: A Benchmark for Egocentric Video Understanding in High-Stakes Police Body-Worn Camera Footage
EgoPolice introduces a 185-hour annotated police body-worn camera benchmark showing state-of-the-art video models fail on high-stakes actions due to motion, occlusion, and low inter-class visual separability.
-
EgoVerse: An Egocentric Human Dataset for Robot Learning from Around the World
EgoVerse releases 1,362 hours of standardized egocentric human data across 1,965 tasks and shows via multi-lab experiments that robot policy performance scales with human data volume when the data aligns with robot ob...
-
Hierarchical and Holistic Open-Vocabulary Functional 3D Scene Graphs for Indoor Spaces
An open-vocabulary pipeline anchors functional edges via 2D visual grounding then uses temporal 3D graph optimization with evidence accumulation and entropy regularization to build hierarchical scene graphs for dense ...
-
ECHO: Ego-Centric modeling of Human-Object interactions
ECHO jointly predicts human pose, object trajectory, and contact from sparse head-and-wrist tracking using a tri-variate diffusion transformer, and reports the best egocentric human-object interaction reconstruction r...
-
MR6D: Benchmarking 6D Pose Estimation for Mobile Robots
MR6D is a new benchmark of 92 real industrial scenes for 6D pose estimation on mobile robots, and current unseen-object pipelines achieve only 0.35 average recall with ground-truth masks.
-
Perceiving and Acting in First-Person: A Dataset and Benchmark for Egocentric Human-Object-Human Interactions
Claimed first large-scale egocentric and multi-view dataset of human-object-human assistance (11.4 hours, 1.2M frames) with three benchmarks; only the abstract was assessable because the submitted body text is a diffe...
-
EgoM2P: Egocentric Multimodal Multitask Pretraining
A masked pretraining model over RGB, depth, gaze, and camera-pose tokens matches specialist egocentric vision systems on four tasks while running at 300+ frames per second.
-
GoTrack: Generic 6DoF Object Pose Refinement and Tracking
GoTrack uses optical flow between a synthetic object render and the input image to refine and track 6D poses of unseen objects, improving accuracy and speed over prior methods.
-
Generating 6DoF Object Manipulation Trajectories from Action Description in Egocentric Vision
A new 28,497-sample dataset of 6DoF object manipulation trajectories is automatically extracted from egocentric video, and vision-language models are trained to generate these trajectories from action descriptions.
-
How Do I Do That? Synthesizing 3D Hand Motion and Contacts for Everyday Interactions
A codebook-based model predicts future 3D hand poses and contact maps from a single image, action text, and a 3D contact point, outperforming baselines on a new large-scale benchmark.
-
GigaHands: A Massive Annotated Dataset of Bimanual Hand Activities
GigaHands provides 14,000 bimanual hand clips, 84,000 text annotations, and 183 million frames from 51 cameras, outperforming smaller datasets in text-to-motion and captioning tasks.
-
OpenEgo: A Large-Scale Multimodal Egocentric Dataset for Dexterous Manipulation
OpenEgo is a 1,107-hour unified egocentric manipulation dataset with standardized 21-joint hand poses and timestamped action language, plus a small validation showing a language-conditioned policy learns short-horizon...
-
HaWoR: World-Space Hand Motion Reconstruction from Egocentric Videos
HaWoR estimates metric world-space hand trajectories from egocentric video by masking hands from SLAM bundle adjustment, aligning SLAM scale with Metric3D depth, and infilling missing hand frames with a transformer.
-
Data Pyramid for Embodied Manipulation: A Survey
Embodied training data form a five-layer pyramid—real-robot, UMI, ego/exo, simulation, general V–L—ordered by the trade-off between scale and robot alignment, and model capabilities track how those layers are mixed.
Discussion (0). Continue with ORCID to comment.