Pith. sign in

REVIEW 8 cited by

Aria Everyday Activities Dataset

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.13349 v2 pith:2VHAUBIR submitted 2024-02-20 cs.CV cs.AIcs.HC

classification cs.CVcs.AIcs.HC
keywords datasetariaprojectrecordedactivitiesalignedcontainsdata
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present Aria Everyday Activities (AEA) Dataset, an egocentric multimodal open dataset recorded using Project Aria glasses. AEA contains 143 daily activity sequences recorded by multiple wearers in five geographically diverse indoor locations. Each of the recording contains multimodal sensor data recorded through the Project Aria glasses. In addition, AEA provides machine perception data including high frequency globally aligned 3D trajectories, scene point cloud, per-frame 3D eye gaze vector and time aligned speech transcription. In this paper, we demonstrate a few exemplar research applications enabled by this dataset, including neural scene reconstruction and prompted segmentation. AEA is an open source dataset that can be downloaded from https://www.projectaria.com/datasets/aea/. We are also providing open-source implementations and examples of how to use the dataset in Project Aria Tools https://github.com/facebookresearch/projectaria_tools.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing

    cs.CV 2025-06 conditional novelty 7.0 of 10

    SAVVY-Bench tests audio-visual LLMs on dynamic 3D spatial questions, and the SAVVY pipeline, combining visual tracks with spatial audio and global mapping, lifts Gemini-2.5-pro accuracy from 50.9% to 58.0%.

  2. HeadRoom: Lightweight, Edge-deployable Pipeline for Adaptive Notification Routing

    cs.HC 2026-07 conditional novelty 6.0 of 10

    Under high perceptual load, routing brief probes to the channel with higher HeadRoom-estimated availability cuts response time versus the less available channel.

  3. EgoEverything: A Benchmark for Human Behavior Inspired Long Context Egocentric Video Understanding in AR Environment

    cs.LG 2026-04 unverdicted novelty 6.0 of 10

    EgoEverything is a new benchmark for long-context egocentric video understanding that uses human gaze-based attention signals to generate questions reflecting natural behavior.

  4. EgoCogNav: Cognition-aware Human Egocentric Navigation

    cs.LG 2025-11 conditional novelty 6.0 of 10

    EgoCogNav jointly predicts walking path, head motion, and perceived route uncertainty from egocentric sensors, and the CEN dataset makes such joint forecasting possible.

  5. EgoM2P: Egocentric Multimodal Multitask Pretraining

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A masked pretraining model over RGB, depth, gaze, and camera-pose tokens matches specialist egocentric vision systems on four tasks while running at 300+ frames per second.

  6. Seeing in the Dark: Benchmarking Egocentric 3D Vision with the Oxford Day-and-Night Dataset

    cs.CV 2025-06 conditional novelty 6.0 of 10

    This paper releases and benchmarks a 30 km egocentric day-and-night dataset with SLAM poses and TLS ground truth, and shows current NVS and relocalization methods degrade sharply at night.

  7. Limits and Trade-Offs of Shift-Invariant Meta-Optical Encoders for Image Compression

    physics.optics 2026-02 conditional novelty 5.0 of 10

    At equal compression ratios and low noise, lens-based spatial binning reconstructs images at least as well as random or orthogonal multi-channel optical encoders, while multi-channel designs tolerate more noise.

  8. EgoAdapt: Adaptive Multisensory Distillation and Policy Learning for Efficient Egocentric Perception

    cs.CV 2025-06 conditional novelty 5.0 of 10

    A joint distillation and policy-learning framework claims near-teacher accuracy on egocentric action recognition, active speaker localization, and behavior anticipation at a fraction of the compute.

Pith tools