Pith. sign in

REVIEW 2 cited by

The BabyView dataset: High-resolution egocentric videos of infants' and young children's everyday experiences

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.10447 v2 pith:6B4GIF3W submitted 2024-06-14 cs.CV

classification cs.CV
keywords datachildrenmodelsdatasetegocentrichumanperformancesystems
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Human children far exceed modern machine learning algorithms in their sample efficiency, achieving high performance in key domains with much less data than current models. This ''data gap'' is a key challenge both for building intelligent artificial systems and for understanding human development. Egocentric video capturing children's experience--their ''training data''--is a key ingredient for comparison of humans and models and for the development of algorithmic innovations to bridge this gap. Yet there are few such datasets available, and extant data are low-resolution, have limited metadata, and importantly, represent only a small set of children's experiences. Here, we provide the first release of a large developmental egocentric video dataset--the BabyView dataset--recorded using a high-resolution camera with a large vertical field-of-view and gyroscope/accelerometer data. This 868 hour dataset includes egocentric videos from children spanning 6 months to 3 years of age in longitudinal, at-home contexts. We provide gold-standard annotations for the evaluation of speech transcription, speaker diarization, and human pose estimation, and evaluate models in each of these domains. We train self-supervised language and vision models and evaluate their transfer to out-of-distribution tasks, including syntactic structure learning, object recognition, depth estimation, and image segmentation. Although performance in each domain scales with dataset size, overall performance is relatively lower than when models are trained on curated datasets, especially in the visual domain. Our dataset stands as an open challenge for robust, human-like AI systems: how can such systems achieve human-levels of success on the same scale and distribution of training data as humans?

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Precise Action-to-Video Generation Through Visual Action Prompts

    cs.CV 2025-08 conditional novelty 6.0 of 10

    Skeleton-based visual action prompts give precise, cross-domain action control for video generation of human and robot interactions.

  2. Challenges in Automated Processing of Speech from Child Wearables: The Case of Voice Type Classifier

    eess.AS 2025-06 conditional novelty 6.0 of 10

    Adding more training data improves voice-type classification more than architecture or representation changes, which plateau near 50% average F-score, still below human annotator agreement of about 70%.

Pith tools