Pith. sign in

REVIEW 14 cited by

Ego4D: Around the World in 3,000 Hours of Egocentric Video

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2110.07058 v3 pith:WZOPS6BK submitted 2021-10-13 cs.CV cs.AI

classification cs.CVcs.AI
keywords videoegocentricbenchmarkego4darounddatasetfirst-personhours
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We introduce Ego4D, a massive-scale egocentric video dataset and benchmark suite. It offers 3,670 hours of daily-life activity video spanning hundreds of scenarios (household, outdoor, workplace, leisure, etc.) captured by 931 unique camera wearers from 74 worldwide locations and 9 different countries. The approach to collection is designed to uphold rigorous privacy and ethics standards with consenting participants and robust de-identification procedures where relevant. Ego4D dramatically expands the volume of diverse egocentric video footage publicly available to the research community. Portions of the video are accompanied by audio, 3D meshes of the environment, eye gaze, stereo, and/or synchronized videos from multiple egocentric cameras at the same event. Furthermore, we present a host of new benchmark challenges centered around understanding the first-person visual experience in the past (querying an episodic memory), present (analyzing hand-object manipulation, audio-visual conversation, and social interactions), and future (forecasting activities). By publicly sharing this massive annotated dataset and benchmark suite, we aim to push the frontier of first-person perception. Project page: https://ego4d-data.org/

Discussion (0). Sign in to comment.

Forward citations

Cited by 14 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. How Can AI Augment Access to Justice? Public Defenders' Perspectives on AI Adoption

    cs.CY 2025-10 conditional novelty 7.0 of 10

    Public defenders view AI as most useful for evidence investigation but limited in courtroom work and strategy, with adoption blocked by costs, confidentiality risks, and norms, requiring human oversight and open development.

  2. EgoSim: Egocentric World Simulator for Embodied Interaction Generation

    cs.CV 2026-04 conditional novelty 6.5 of 10

    EgoSim generates spatially consistent egocentric interaction videos by conditioning a video diffusion model on updatable 3D point-cloud states and action keypoints extracted at scale from monocular videos.

  3. Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills

    cs.RO 2026-08 conditional novelty 6.0 of 10

    A taxonomy of robot learning on a weights-versus-skills axis, with a five-rung self-improvement ladder whose top cell (feedback plus memory plus search) holds only a few recent systems.

  4. HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data Alone

    cs.RO 2026-07 conditional novelty 6.0 of 10

    Robot-free HiFi-UMI demonstrations can replace teleoperated real-robot data in post-training: three policy backbones matched in-domain teleoperation within 3.1 percentage points, including 85% success on a precision i...

  5. MEMORA: Embodied Action Memory from Egocentric Videos for Reasoning and Planning

    cs.RO 2026-07 conditional novelty 6.0 of 10

    A typed, editable memory built from egocentric video improves memory-grounded question answering and out-of-distribution robot planning over flat-text and graph baselines.

  6. Towards Real-World Wearable Motion Reconstruction

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A consumer-wearable MoCap dataset plus WHIP, a flow-matching model that reconstructs full-body motion from arbitrary sensor subsets and quantifies sensor complementarity.

  7. Spotlighting Task-Relevant Features: Object-Centric Representations for Better Generalization in Robotic Manipulation

    cs.RO 2026-01 conditional novelty 6.0 of 10

    Slot-based object-centric visual representations, especially with robot-video pretraining, improve out-of-distribution generalization of robotic manipulation policies compared to global and dense pre-trained features.

  8. CausalVQA: A Physically Grounded Causal Reasoning Benchmark for Video Models

    cs.CV 2025-06 conditional novelty 6.0 of 10

    CausalVQA provides 793 paired real-video causal reasoning questions on which the best multimodal model scores 61.66% versus 84.78% for humans, with the largest gaps on anticipation and hypothetical questions.

  9. VideoConviction: A Multimodal Benchmark for Human Conviction and Stock Market Recommendations

    cs.MM 2025-06 conditional novelty 6.0 of 10

    VideoConviction provides the first expert-annotated multimodal benchmark of financial influencer video recommendations, showing MLLMs extract tickers better but struggle with actions and conviction, and that an invers...

  10. iFLYTEK-Embodied-Omni Technical Report

    cs.AI 2026-06 conditional novelty 5.0 of 10

    A three-branch Omni model (VLM+VGM brain, AGM cerebellum) with shared multimodal attention and four-stage training reaches 89.6% zero-shot on LIBERO-Plus and ~93% on RoboTwin 2.0.

  11. Enhancing Scene Transition Awareness in Video Generation via Post-Training

    cs.CV 2025-07 conditional novelty 5.0 of 10

    A post-training dataset of transition-centered clips increases the number of scenes an open-source video generator produces for multi-scene prompts, with mixed effects on quality.

  12. How Far Can Off-the-Shelf Multimodal Large Language Models Go in Online Episodic Memory Question Answering?

    cs.CV 2025-06 conditional novelty 5.0 of 10

    A training-free pipeline that summarizes egocentric video clips into a few kilobytes of text per minute and answers multiple-choice episodic memory questions with an LLM reasoner reaches 56.0% accuracy on QAEgo4D-Clos...

  13. Comparing Learning Paradigms for Egocentric Video Summarization

    cs.CV 2025-06 reject novelty 4.0 of 10

    A prompt-engineered GPT-4o (quality score 64.95) outperformed Shotluck Holmes (61.19) and TAC-SUM (58.43) on a 21-video egocentric summary evaluation, though all scores were modest.

  14. Data Pyramid for Embodied Manipulation

    cs.RO 2026-07 conditional novelty 3.0 of 10

    Embodied training data form a five-layer pyramid—real-robot, UMI, ego/exo, simulation, general V–L—ordered by the trade-off between scale and robot alignment, and model capabilities track how those layers are mixed.

Pith tools