Pith. sign in

REVIEW 2 cited by

EFM3D: A Benchmark for Measuring Progress Towards 3D Egocentric Foundation Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.10224 v1 pith:DSZ7ICA5 submitted 2024-06-14 cs.CV

classification cs.CV
keywords egocentricbenchmarkefm3dfoundationmodelsdataefmsprogress
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The advent of wearable computers enables a new source of context for AI that is embedded in egocentric sensor data. This new egocentric data comes equipped with fine-grained 3D location information and thus presents the opportunity for a novel class of spatial foundation models that are rooted in 3D space. To measure progress on what we term Egocentric Foundation Models (EFMs) we establish EFM3D, a benchmark with two core 3D egocentric perception tasks. EFM3D is the first benchmark for 3D object detection and surface regression on high quality annotated egocentric data of Project Aria. We propose Egocentric Voxel Lifting (EVL), a baseline for 3D EFMs. EVL leverages all available egocentric modalities and inherits foundational capabilities from 2D foundation models. This model, trained on a large simulated dataset, outperforms existing methods on the EFM3D benchmark.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing

    cs.CV 2025-06 conditional novelty 7.0 of 10

    SAVVY-Bench tests audio-visual LLMs on dynamic 3D spatial questions, and the SAVVY pipeline, combining visual tracks with spatial audio and global mapping, lifts Gemini-2.5-pro accuracy from 50.9% to 58.0%.

  2. HD-EPIC: A Highly-Detailed Egocentric Video Dataset

    cs.CV 2025-02 conditional novelty 7.0 of 10

    A new densely annotated, 3D-grounded egocentric kitchen dataset with a 26K-question VQA benchmark that current video-language models mostly fail.

Pith tools