Pith. sign in

REVIEW 14 cited by

Introducing HOT3D: An Egocentric Dataset for 3D Hand and Object Tracking

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.09598 v1 pith:JCWMPMMU submitted 2024-06-13 cs.CV

classification cs.CV
keywords datasethot3dobjectsegocentrichandhandsactionsannotations
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

We introduce HOT3D, a publicly available dataset for egocentric hand and object tracking in 3D. The dataset offers over 833 minutes (more than 3.7M images) of multi-view RGB/monochrome image streams showing 19 subjects interacting with 33 diverse rigid objects, multi-modal signals such as eye gaze or scene point clouds, as well as comprehensive ground truth annotations including 3D poses of objects, hands, and cameras, and 3D models of hands and objects. In addition to simple pick-up/observe/put-down actions, HOT3D contains scenarios resembling typical actions in a kitchen, office, and living room environment. The dataset is recorded by two head-mounted devices from Meta: Project Aria, a research prototype of light-weight AR/AI glasses, and Quest 3, a production VR headset sold in millions of units. Ground-truth poses were obtained by a professional motion-capture system using small optical markers attached to hands and objects. Hand annotations are provided in the UmeTrack and MANO formats and objects are represented by 3D meshes with PBR materials obtained by an in-house scanner. We aim to accelerate research on egocentric hand-object interaction by making the HOT3D dataset publicly available and by co-organizing public challenges on the dataset at ECCV 2024. The dataset can be downloaded from the project website: https://facebookresearch.github.io/hot3d/.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 14 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. EgoPolice: A Benchmark for Egocentric Video Understanding in High-Stakes Police Body-Worn Camera Footage

    cs.CV 2026-07 conditional novelty 7.0 of 10

    EgoPolice introduces a 185-hour annotated police body-worn camera benchmark showing state-of-the-art video models fail on high-stakes actions due to motion, occlusion, and low inter-class visual separability.

  2. EgoVerse: An Egocentric Human Dataset for Robot Learning from Around the World

    cs.RO 2026-04 unverdicted novelty 6.5 of 10

    EgoVerse releases 1,362 hours of standardized egocentric human data across 1,965 tasks and shows via multi-lab experiments that robot policy performance scales with human data volume when the data aligns with robot ob...

  3. Hierarchical and Holistic Open-Vocabulary Functional 3D Scene Graphs for Indoor Spaces

    cs.RO 2026-05 unverdicted novelty 6.0 of 10

    An open-vocabulary pipeline anchors functional edges via 2D visual grounding then uses temporal 3D graph optimization with evidence accumulation and entropy regularization to build hierarchical scene graphs for dense ...

  4. ECHO: Ego-Centric modeling of Human-Object interactions

    cs.CV 2025-08 conditional novelty 6.0 of 10

    ECHO jointly predicts human pose, object trajectory, and contact from sparse head-and-wrist tracking using a tri-variate diffusion transformer, and reports the best egocentric human-object interaction reconstruction r...

  5. MR6D: Benchmarking 6D Pose Estimation for Mobile Robots

    cs.CV 2025-08 conditional novelty 6.0 of 10

    MR6D is a new benchmark of 92 real industrial scenes for 6D pose estimation on mobile robots, and current unseen-object pipelines achieve only 0.35 average recall with ground-truth masks.

  6. Perceiving and Acting in First-Person: A Dataset and Benchmark for Egocentric Human-Object-Human Interactions

    cs.CV 2025-08 unverdicted novelty 6.0 of 10

    Claimed first large-scale egocentric and multi-view dataset of human-object-human assistance (11.4 hours, 1.2M frames) with three benchmarks; only the abstract was assessable because the submitted body text is a diffe...

  7. EgoM2P: Egocentric Multimodal Multitask Pretraining

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A masked pretraining model over RGB, depth, gaze, and camera-pose tokens matches specialist egocentric vision systems on four tasks while running at 300+ frames per second.

  8. GoTrack: Generic 6DoF Object Pose Refinement and Tracking

    cs.CV 2025-06 conditional novelty 6.0 of 10

    GoTrack uses optical flow between a synthetic object render and the input image to refine and track 6D poses of unseen objects, improving accuracy and speed over prior methods.

  9. Generating 6DoF Object Manipulation Trajectories from Action Description in Egocentric Vision

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A new 28,497-sample dataset of 6DoF object manipulation trajectories is automatically extracted from egocentric video, and vision-language models are trained to generate these trajectories from action descriptions.

  10. How Do I Do That? Synthesizing 3D Hand Motion and Contacts for Everyday Interactions

    cs.CV 2025-04 conditional novelty 6.0 of 10

    A codebook-based model predicts future 3D hand poses and contact maps from a single image, action text, and a 3D contact point, outperforming baselines on a new large-scale benchmark.

  11. GigaHands: A Massive Annotated Dataset of Bimanual Hand Activities

    cs.CV 2024-12 conditional novelty 6.0 of 10

    GigaHands provides 14,000 bimanual hand clips, 84,000 text annotations, and 183 million frames from 51 cameras, outperforming smaller datasets in text-to-motion and captioning tasks.

  12. OpenEgo: A Large-Scale Multimodal Egocentric Dataset for Dexterous Manipulation

    cs.CV 2025-09 conditional novelty 5.0 of 10

    OpenEgo is a 1,107-hour unified egocentric manipulation dataset with standardized 21-joint hand poses and timestamped action language, plus a small validation showing a language-conditioned policy learns short-horizon...

  13. HaWoR: World-Space Hand Motion Reconstruction from Egocentric Videos

    cs.CV 2025-01 conditional novelty 5.0 of 10

    HaWoR estimates metric world-space hand trajectories from egocentric video by masking hands from SLAM bundle adjustment, aligning SLAM scale with Metric3D depth, and infilling missing hand frames with a transformer.

  14. Data Pyramid for Embodied Manipulation: A Survey

    cs.RO 2026-07 conditional novelty 3.0 of 10

    Embodied training data form a five-layer pyramid—real-robot, UMI, ego/exo, simulation, general V–L—ordered by the trade-off between scale and robot alignment, and model capabilities track how those layers are mixed.

Pith tools