REVIEW 9 cited by
EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Video generation has emerged as a promising tool for world simulation, leveraging visual data to replicate real-world environments. Within this context, egocentric video generation, which centers on the human perspective, holds significant potential for enhancing applications in virtual reality, augmented reality, and gaming. However, the generation of egocentric videos presents substantial challenges due to the dynamic nature of egocentric viewpoints, the intricate diversity of actions, and the complex variety of scenes encountered. Existing datasets are inadequate for addressing these challenges effectively. To bridge this gap, we present EgoVid-5M, the first high-quality dataset specifically curated for egocentric video generation. EgoVid-5M encompasses 5 million egocentric video clips and is enriched with detailed action annotations, including fine-grained kinematic control and high-level textual descriptions. To ensure the integrity and usability of the dataset, we implement a sophisticated data cleaning pipeline designed to maintain frame consistency, action coherence, and motion smoothness under egocentric conditions. Furthermore, we introduce EgoDreamer, which is capable of generating egocentric videos driven simultaneously by action descriptions and kinematic control signals. The EgoVid-5M dataset, associated action annotations, and all data cleansing metadata will be released for the advancement of research in egocentric video generation.
Forward citations
Cited by 9 Pith papers
-
EgoSim: Egocentric World Simulator for Embodied Interaction Generation
EgoSim generates spatially consistent egocentric interaction videos by conditioning a video diffusion model on updatable 3D point-cloud states and action keypoints extracted at scale from monocular videos.
-
Precise Action-to-Video Generation Through Visual Action Prompts
Skeleton-based visual action prompts give precise, cross-domain action control for video generation of human and robot interactions.
-
FreeLong++: Training-Free Long Video Generation via Multi-band SpectralFusion
FreeLong++ extends short-video diffusion models to 4x to 8x longer clips, without retraining, by fusing multiple windowed attention branches through frequency-domain filters and a spectral noise initialization.
-
EgoM2P: Egocentric Multimodal Multitask Pretraining
A masked pretraining model over RGB, depth, gaze, and camera-pose tokens matches specialist egocentric vision systems on four tasks while running at 300+ frames per second.
-
Seeing in the Dark: Benchmarking Egocentric 3D Vision with the Oxford Day-and-Night Dataset
This paper releases and benchmarks a 30 km egocentric day-and-night dataset with SLAM poses and TLS ground truth, and shows current NVS and relocalization methods degrade sharply at night.
-
CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning
CoMo learns continuous latent motion self-supervised from internet videos and uses it as pseudo action labels to improve robot policy co-training.
-
EgoSafe: A First-Person Mobile-Captured Benchmark for Visual Safety Understanding
A new egocentric safety benchmark shows current video-language models can describe scenes well but fail at multi-step causal reasoning about blind spots and covert actions.
-
iFLYTEK-Embodied-Omni Technical Report
A three-branch Omni model (VLM+VGM brain, AGM cerebellum) with shared multimodal attention and four-stage training reaches 89.6% zero-shot on LIBERO-Plus and ~93% on RoboTwin 2.0.
-
VISTA: A Controllable Platform for Generating and Auditing Egocentric Assistance Scenarios
VISTA is a video synthesis framework that creates controllable egocentric videos of daily tasks with reactive and proactive agent intervention modes via causal reverse reasoning scripts.
Discussion (0). Continue with ORCID to comment.