Pith. sign in

REVIEW 9 cited by

EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.08380 v1 pith:MYBQHHR7 submitted 2024-11-13 cs.CV

classification cs.CV
keywords egocentricgenerationvideoactiondatasetegovid-5mdataannotations
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Video generation has emerged as a promising tool for world simulation, leveraging visual data to replicate real-world environments. Within this context, egocentric video generation, which centers on the human perspective, holds significant potential for enhancing applications in virtual reality, augmented reality, and gaming. However, the generation of egocentric videos presents substantial challenges due to the dynamic nature of egocentric viewpoints, the intricate diversity of actions, and the complex variety of scenes encountered. Existing datasets are inadequate for addressing these challenges effectively. To bridge this gap, we present EgoVid-5M, the first high-quality dataset specifically curated for egocentric video generation. EgoVid-5M encompasses 5 million egocentric video clips and is enriched with detailed action annotations, including fine-grained kinematic control and high-level textual descriptions. To ensure the integrity and usability of the dataset, we implement a sophisticated data cleaning pipeline designed to maintain frame consistency, action coherence, and motion smoothness under egocentric conditions. Furthermore, we introduce EgoDreamer, which is capable of generating egocentric videos driven simultaneously by action descriptions and kinematic control signals. The EgoVid-5M dataset, associated action annotations, and all data cleansing metadata will be released for the advancement of research in egocentric video generation.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. EgoSim: Egocentric World Simulator for Embodied Interaction Generation

    cs.CV 2026-04 conditional novelty 6.5 of 10

    EgoSim generates spatially consistent egocentric interaction videos by conditioning a video diffusion model on updatable 3D point-cloud states and action keypoints extracted at scale from monocular videos.

  2. Precise Action-to-Video Generation Through Visual Action Prompts

    cs.CV 2025-08 conditional novelty 6.0 of 10

    Skeleton-based visual action prompts give precise, cross-domain action control for video generation of human and robot interactions.

  3. FreeLong++: Training-Free Long Video Generation via Multi-band SpectralFusion

    cs.CV 2025-06 conditional novelty 6.0 of 10

    FreeLong++ extends short-video diffusion models to 4x to 8x longer clips, without retraining, by fusing multiple windowed attention branches through frequency-domain filters and a spectral noise initialization.

  4. EgoM2P: Egocentric Multimodal Multitask Pretraining

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A masked pretraining model over RGB, depth, gaze, and camera-pose tokens matches specialist egocentric vision systems on four tasks while running at 300+ frames per second.

  5. Seeing in the Dark: Benchmarking Egocentric 3D Vision with the Oxford Day-and-Night Dataset

    cs.CV 2025-06 conditional novelty 6.0 of 10

    This paper releases and benchmarks a 30 km egocentric day-and-night dataset with SLAM poses and TLS ground truth, and shows current NVS and relocalization methods degrade sharply at night.

  6. CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning

    cs.CV 2025-05 conditional novelty 6.0 of 10

    CoMo learns continuous latent motion self-supervised from internet videos and uses it as pseudo action labels to improve robot policy co-training.

  7. EgoSafe: A First-Person Mobile-Captured Benchmark for Visual Safety Understanding

    cs.CV 2026-07 conditional novelty 5.0 of 10

    A new egocentric safety benchmark shows current video-language models can describe scenes well but fail at multi-step causal reasoning about blind spots and covert actions.

  8. iFLYTEK-Embodied-Omni Technical Report

    cs.AI 2026-06 conditional novelty 5.0 of 10

    A three-branch Omni model (VLM+VGM brain, AGM cerebellum) with shared multimodal attention and four-stage training reaches 89.6% zero-shot on LIBERO-Plus and ~93% on RoboTwin 2.0.

  9. VISTA: A Controllable Platform for Generating and Auditing Egocentric Assistance Scenarios

    cs.CL 2026-05 unverdicted novelty 5.0 of 10

    VISTA is a video synthesis framework that creates controllable egocentric videos of daily tasks with reactive and proactive agent intervention modes via causal reverse reasoning scripts.

Pith tools