Pith. sign in

Paper Citation Record · LEDGER

EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation

As of 20 August 2026, this Paper Citation Record lists 89 of 89 outbound references and 20 inbound Pith citation observations for arXiv:2411.08380.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.08380 v1

Coverage vector

measured 89 of 89 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T21:41:59.785064Z

measured 109 of 109 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 20 of 20 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T22:30:42.570671Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T04:09:35.541046Z

Reference resolution

89 of 89 outbound references displayed

  • verified exact0
  • verified fuzzy37
  • unresolved51
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ade03dd6-589d-4841-93ba-acd5d7f4c032 · outbound

This paper cites Self-supervised object detection from egocentric videos.

EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation Self-supervised object detection from egocentric videos

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T21:41:59.378302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:41:59.378302Z digest=sha256:3eb91534bd824867f3e4def1c5dd4446b8dafdc1048bc0f8f211f481298fd496

Observation 1fbcac96-1f07-4511-a0ad-251ac8bc8c6e · outbound

This paper cites Frozen in time: A joint video and image encoder for end-to-end retrieval.

EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation Frozen in time: A joint video and image encoder for end-to-end retrieval

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T21:41:59.384225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:41:59.384225Z digest=sha256:f7e6a0f34f2fffa19c16fcae86a3614911afb98bf8754c3e84888ff11e901bc1

Observation 84e4c6b2-ee04-44c5-9037-1d6f79041a67 · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T21:41:59.389473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:41:59.389473Z digest=sha256:0cdd2f6101e402ece880c0e45a9d587f7b04e0038fc6fda0530f682aa10cc9ad

Observation 4a4716cb-4e37-4f7b-911b-08734f530412 · outbound

This paper cites Video generation models as world simulators.

EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation Video generation models as world simulators

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T21:41:59.394717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:41:59.394717Z digest=sha256:693ecad24cf965b176a044fdeb8ff62e0cd23fd4abec449f821851f9aa906d9e

Observation 36cfff75-c6d4-4b27-8297-afa889c274b0 · outbound

This paper cites Genie: Generative Interactive Environments.

EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation Genie: Generative Interactive Environments

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T21:41:59.399286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:41:59.399286Z digest=sha256:4d9f2fa0eed92fa372e11b50c99b5e7e90434913995373cdbcbd3da80eaab708

Observation 05cd4cc0-50fc-422d-96c5-b9fb3bdc9608 · outbound

This paper cites Videocrafter1: Open diffusion models for high-quality video generation, 2023.

EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation Videocrafter1: Open diffusion models for high-quality video generation, 2023

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T21:41:59.404250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:41:59.404250Z digest=sha256:f50bbd306fd9bf079d9c2ddc1b0358ef337e28f451d5dad9c7fbd3625115d071

Observation 123aa0ef-6344-476d-af7d-a4b7d42cc2e3 · outbound

This paper cites Videocrafter2: Overcoming data limitations for high-quality video diffusion models, 2024.

EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation Videocrafter2: Overcoming data limitations for high-quality video diffusion models, 2024

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T21:41:59.409263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:41:59.409263Z digest=sha256:0156b25dab72d6707144be7928152627b568e66e2aeed898428f13efb06d5c5e

Observation e4b3653b-3b9c-4bc2-8806-031e7a5238df · outbound

This paper cites Panda-70m: Captioning 70m videos with multiple cross-modality teachers.

EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation Panda-70m: Captioning 70m videos with multiple cross-modality teachers

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T21:41:59.414184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:41:59.414184Z digest=sha256:7e050c4d8cdfa6c2ddc514d6e36ce3e81235ed24759918778a11cf65a492cdbb

Observation ccedbfd6-f54b-4d3e-af78-37f014cc304d · outbound

This paper cites Improved-aesthetic-predictor, 2022.

EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation Improved-aesthetic-predictor, 2022

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T21:41:59.418013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:41:59.418013Z digest=sha256:fa448e83cabdd94b2a4ddd7efdf0ce4cfaf2955e92ab541b78e825ce10782d7f

Observation 8413f425-2ee0-4fa5-984f-99e9fc524e72 · outbound

This paper cites Lillicrap and Mohammad Norouzi and Jimmy Ba.

EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation Lillicrap and Mohammad Norouzi and Jimmy Ba

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T21:41:59.422044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:41:59.422044Z digest=sha256:9b2c6de0754c500aa5a02aa67c83ffad83ebc88d42cf631faa1dc15bf3cfcbe0

Observation ac0927ea-dfe4-4f37-bd0f-5fffbce5b5a1 · outbound

This paper cites Aigcbench: Comprehensive evaluation of image-to-video content generated by ai.

EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation Aigcbench: Comprehensive evaluation of image-to-video content generated by ai

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T21:41:59.425825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:41:59.425825Z digest=sha256:539aeb94aef79648ffbb11c43bd70ca8d96d0afa536144548f9ee95bb8144224

Observation c07df424-656f-44f6-b37a-2cad4fc5c67c · outbound

This paper cites On the content bias in fr ´echet video distance.

EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation On the content bias in fr ´echet video distance

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:42:01.082838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T21:41:59.429463Z digest=sha256:291d10151a98b6bb2bee89032508c18c89ca5565efac3b1044b1efb97f59bb07

Observation fc0951f0-d238-4965-9c6f-d3e473a06871 · outbound

This paper cites Introducing general world models.

EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation Introducing general world models

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:42:01.069093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T21:41:59.433433Z digest=sha256:e25f489c72e99a65778ffdfd7188264e7665beaf9f3fbdce347f86c7e8b245ef

Observation c8c93335-39cd-4581-bf2d-ea0dc6e54544 · outbound

This paper cites Introducing gen-3 alpha.

EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation Introducing gen-3 alpha

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:42:01.053640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T21:41:59.437282Z digest=sha256:635b5e0b4a8a2a1da4ec60c19e3eb7411839723ef3f7f8d746290665d59853f9

Observation 183e3ab4-627a-4d8d-acfd-64859e3cc218 · outbound

This paper cites Mmg-ego4d: Multimodal generalization in egocentric action recognition.

EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation Mmg-ego4d: Multimodal generalization in egocentric action recognition

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:42:01.038444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T21:41:59.440913Z digest=sha256:ba2f4582a63adefad39566f954e2cac9605a89f94026094c550ee4fb7dd1d34d

Observation d29a241b-1037-499b-b1fa-4c8f34aa4ad3 · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video.

EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation Ego4d: Around the world in 3,000 hours of egocentric video

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:42:01.022455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T21:41:59.444397Z digest=sha256:c988bd2ff0a0eda4e125ec9fa548add1517519bc0a25e05c182c1aac611e47e5

Observation 1c6268db-3cf8-4873-a733-39c88a8ff679 · outbound

This paper cites Jawahar, Richard Newcombe, Hyun Soo Park, James M.

EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation Jawahar, Richard Newcombe, Hyun Soo Park, James M

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:42:01.006220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T21:41:59.448690Z digest=sha256:7608041d7b39e7901aec565f652960d0685c5c95e1935ba9511cd85146e5565a

Observation dacf5b83-bf9d-4037-9716-bca80310ecbb · outbound

This paper cites Lillicrap, Jimmy Ba, and Mo- hammad Norouzi.

EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation Lillicrap, Jimmy Ba, and Mo- hammad Norouzi

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:42:00.991474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T21:41:59.453600Z digest=sha256:8a6d80660d9c60ecf592a1d631b4b5adc5e10a58f24068e0599d32641c1412cf

Observation eb6a4b20-239b-44ea-b928-c8311a6533f0 · outbound

This paper cites Mastering Diverse Domains through World Models.

EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation Mastering Diverse Domains through World Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T21:41:59.458026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:41:59.458026Z digest=sha256:9687e9ab751f20ffaad8d76ee8cd3974e79f90b725c782bf4cb95c0b8133bb47

Observation 8fb2c236-5511-46e3-b6a8-31a99feddbc1 · outbound

This paper cites CameraCtrl: Enabling Camera Control for Text-to-Video Generation.

EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation CameraCtrl: Enabling Camera Control for Text-to-Video Generation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T21:41:59.462862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:41:59.462862Z digest=sha256:e814a6308080fce2a035d72994b4fbbddbc2d812fd727043b95d7a6df3ff9421

Observation 1d1d25a0-074c-4190-b504-2bd867c214a5 · outbound

This paper cites Imagen Video: High Definition Video Generation with Diffusion Models.

EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation Imagen Video: High Definition Video Generation with Diffusion Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T21:41:59.467532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:41:59.467532Z digest=sha256:64af8b908466b7bcf428bc2cbc3a280f5e707201a8df94f18591a222e6824880

Observation 910d8d6e-f0a7-4124-b177-87e34b881a04 · outbound

This paper cites Denoising diffu- sion probabilistic models.

EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation Denoising diffu- sion probabilistic models

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:42:00.976388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T21:41:59.472384Z digest=sha256:ebec45742a3cc6b4c65a6293735ae3a450b5e97eb30f7f1cb7934d5d064e8831

Observation a5e1147f-8905-47aa-89e1-2b43ba144cf1 · outbound

This paper cites CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers.

EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T21:41:59.476861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:41:59.476861Z digest=sha256:35a6830ea7240ef52d79ab0d984c1fc2f089eea6f4121a6b8c401eb0500cba6d

Observation 27443513-f220-4fa7-a8c5-ca5d190e8a70 · outbound

This paper cites Training-free Camera Control for Video Generation.

EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation Training-free Camera Control for Video Generation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T21:41:59.481577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:41:59.481577Z digest=sha256:54e840451073d3ee0591147e0037d01439ed215b60a63ca72ef569ffe737d94d

Observation 0c75d9d3-a3f2-4cee-b3cf-b0492833142e · outbound

This paper cites GAIA-1: A Generative World Model for Autonomous Driving.

EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation GAIA-1: A Generative World Model for Autonomous Driving

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T21:41:59.486124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:41:59.486124Z digest=sha256:a78ebaa731d69f11897e497ee70c1f6972cb0ac2538ac1dfa959582e9fa63723

Observation e31e0c5f-7777-4b70-9dc5-0fff524309aa · outbound

This paper cites MotionMaster: Training-free Camera Motion Transfer For Video Generation.

EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation MotionMaster: Training-free Camera Motion Transfer For Video Generation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T21:41:59.490941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:41:59.490941Z digest=sha256:ee693902241860eb0287cb565f568e0c39eb320c9f2919414147d0a3b4f6e45d

Observation 719501f7-84b7-415b-832c-be6a1ecf5eb7 · outbound

This paper cites VideoControlNet: A Motion-Guided Video-to-Video Translation Framework by Using Diffusion Model with ControlNet.

EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation VideoControlNet: A Motion-Guided Video-to-Video Translation Framework by Using Diffusion Model with ControlNet

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T21:41:59.495588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:41:59.495588Z digest=sha256:7540571639f01169e16a1f64361a865686c0392c16acdae982cab3dfe2ae1f1b

Observation bfb2d091-9eda-4778-a294-1e41161649b1 · outbound

This paper cites Vbench: Comprehensive bench- mark suite for video generative models.

EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation Vbench: Comprehensive bench- mark suite for video generative models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:42:00.961046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T21:41:59.500458Z digest=sha256:cb3aa58326813e75e306f99b645fecb85cbd15e87f31865be77ef592ef531857

Observation b0485c69-e397-4d95-a64e-cf6017a16aa4 · outbound

This paper cites Elucidating the design space of diffusion-based generative models.

EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation Elucidating the design space of diffusion-based generative models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:42:00.946079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T21:41:59.504909Z digest=sha256:0459e47511b972ed51c15913779b3fce6a8b62ceb235d33e867368f598036a7e

Observation f9c34f5c-1692-48ed-b8d0-9b708ec2b10c · outbound

This paper cites Open-sora-plan, Apr.

EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation Open-sora-plan, Apr

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T21:41:59.509265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:41:59.509265Z digest=sha256:9f3b20c6b28af324980cbe8d4a1faf9f93ffde0abd159738214c9c837547d9cd

Observation 57e1d6c8-5d99-47e4-acd6-92b480a9afb3 · outbound

This paper cites 11 EgoGen: An Egocentric Synthetic Data Generator.

EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation 11 EgoGen: An Egocentric Synthetic Data Generator

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:42:00.920695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T21:41:59.513761Z digest=sha256:1df0183c7c3cd2c6ca8e7d08b9116e5ee87d9da2118f81c5372de9cddc870bc6

Observation 320e9458-4d98-422c-9e83-10693c84488f · outbound

This paper cites Ego-exo: Transferring visual representations from third-person to first-person videos.

EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation Ego-exo: Transferring visual representations from third-person to first-person videos

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:42:00.905826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T21:41:59.518969Z digest=sha256:e758ed96cb53b79420a3e470ee591afc4126b99a9f4713852cef3ca93b1a5daa

Observation 46819068-aaff-4a7e-9cce-5c4e8549afb4 · outbound

This paper cites Amt: All-pairs multi-field transforms for efficient frame interpolation.

EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation Amt: All-pairs multi-field transforms for efficient frame interpolation

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:42:00.892507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T21:41:59.523831Z digest=sha256:684b629a22d92319ef7e34c285e5df9735fbc0db7271b4da77663c761c5354e9

Observation 54fa81b0-b3a8-48be-b171-a6c860ae9f41 · outbound

This paper cites Egocentric video-language pretraining.

EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation Egocentric video-language pretraining

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:42:00.879016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T21:41:59.528528Z digest=sha256:af71f3a8d2ad55ed5ac0f7516b60d67b182d7208a3a60c19c8d6f38b57f2bb6b

Observation f9b67bc0-73f0-4b3d-8406-ef7f45e9e842 · outbound

This paper cites Flow Matching for Generative Modeling.

EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation Flow Matching for Generative Modeling

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T21:41:59.532830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:41:59.532830Z digest=sha256:01833029c5ac1dcd129d2a04d8fffadc9a4bc29e2330e7647500312a0d859091

Observation 623918a3-585b-466f-8193-1f19b5926de8 · outbound

This paper cites Cross-view exocentric to egocentric video synthe- sis.

EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation Cross-view exocentric to egocentric video synthe- sis

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:42:00.864579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T21:41:59.536606Z digest=sha256:7e831b83ebbc989e3bc68e99330fb1c59d4a4d424204b7fa2079d60c0ae3b057

Observation 0ffe1a77-c4c8-428b-a7d8-3e5f3f53d9c9 · outbound

This paper cites A hybrid egocentric activ- ity anticipation framework via memory-augmented recurrent and one-shot representation forecasting.

EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation A hybrid egocentric activ- ity anticipation framework via memory-augmented recurrent and one-shot representation forecasting

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:42:00.849577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T21:41:59.541337Z digest=sha256:2f45dc26845b1fc0baf6b256c621d21143a805c80f3347ef6d6a34847687e873

Observation 92840d04-770b-4ae9-882c-4d70f0df9046 · outbound

This paper cites Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow.

EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T21:41:59.545797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:41:59.545797Z digest=sha256:48a7a9d990f67290f30910fc9016bb7bc3de9ffc91e5a83d9eada4d151a26d2a

Observation c3860403-d1fd-4fc6-8bdc-c3776ec7ba94 · outbound

This paper cites Intention-driven Ego-to-Exo Video Generation.

EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation Intention-driven Ego-to-Exo Video Generation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T21:41:59.551366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:41:59.551366Z digest=sha256:85637ecec0e85040fb826c163b8a3913f0a6b4440bd5fd95a6e47ab7314423be

Observation 7c301cfa-db54-4652-bab8-2a0e570abf70 · outbound

This paper cites TrailBlazer: Trajectory Control for Diffusion-Based Video Generation.

EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation TrailBlazer: Trajectory Control for Diffusion-Based Video Generation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T21:41:59.556092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:41:59.556092Z digest=sha256:3fca057562c6355227edb11020ba9042f4012c0dc6e5d5f27a532b79e326d119

Observation 42418197-1469-432b-8cfd-b4d8cd92c076 · outbound

This paper cites Latte: Latent Diffusion Transformer for Video Generation.

EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation Latte: Latent Diffusion Transformer for Video Generation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T21:41:59.561766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:41:59.561766Z digest=sha256:00b6e0af32322e5df6dd579f3f6eca4ad0c3e97647173594d9e21c9655d12be2

Observation 57033659-e889-403c-82ce-e2dfabf40dfe · outbound

This paper cites Egoloc: Revisiting 3d object local- ization from egocentric videos with visual queries.

EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation Egoloc: Revisiting 3d object local- ization from egocentric videos with visual queries

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T21:41:59.567046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:41:59.567046Z digest=sha256:46abd40f477653a185de5b194fa9f35d267a9b14007e42b8dd8860f3765f5f22

Observation 8cb6329e-c68b-4e37-add0-c0a56f8c9314 · outbound

This paper cites Howto100m: Learning a text-video embedding by watching hundred million narrated video clips.

EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation Howto100m: Learning a text-video embedding by watching hundred million narrated video clips

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:42:00.825333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T21:41:59.571520Z digest=sha256:c8f2e6ac27a3ff4723bc2742a8d4e3ba5a3d8eefe0d9998ed07ab7e984f7d1f4

Observation 7a6661c6-1f97-465f-9f55-8a19caa9b4b9 · outbound

This paper cites OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation.

EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T21:41:59.576134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:41:59.576134Z digest=sha256:83c484a36aa834aeb8c63a734b2d85ba8877d0e14af0855c46bbe9b3925d7c91

Observation 872f3483-2661-4ef0-b60b-8471c6c9549c · outbound

This paper cites EgoVideo: Exploring Egocentric Foundation Model and Downstream Adaptation.

EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation EgoVideo: Exploring Egocentric Foundation Model and Downstream Adaptation

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T21:41:59.581876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:41:59.581876Z digest=sha256:f389a15f5e8edeca8a0d08e0b4bc2feaf840d9fd0b0687d5dd5a2e45b0f780ca

Observation 4021565f-fab1-414a-abf4-bbfa8952e7e5 · outbound

This paper cites ControlNeXt: Powerful and Efficient Control for Image and Video Generation.

EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation ControlNeXt: Powerful and Efficient Control for Image and Video Generation

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T21:41:59.586749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:41:59.586749Z digest=sha256:ff959e2c4d1437cba09e84c5c0ed621b9e5d9cfe6754f8157d08cfac97833318

Observation ac3526ff-b076-4a63-9513-2f4dc5142331 · outbound

This paper cites What can a cook in italy teach a mechanic in in- dia? action recognition generalisation over scenarios and lo- cations.

EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation What can a cook in italy teach a mechanic in in- dia? action recognition generalisation over scenarios and lo- cations

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:42:00.810193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T21:41:59.591400Z digest=sha256:6137f240c0ac6fbf6868d854488383de64e19b0cf18bcfea752343b304b5f981

Observation 5f8ac5b4-7494-411a-9e10-01e0716c6b3f · outbound

This paper cites Multimodal distillation for egocentric action recognition.

EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation Multimodal distillation for egocentric action recognition

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:42:00.795666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T21:41:59.595211Z digest=sha256:09a482a8f74ffba1d0acbc5a72e36865cac857ba16e820ba91a5d927cb3e7134

Observation 5732c4e1-772c-4555-acb6-b4e04cc4cc55 · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation Learn- ing transferable visual models from natural language super- vision

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:42:00.781177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T21:41:59.599578Z digest=sha256:10cc1e9678ba1ba3e8f0fbeaf40e6593a53828e16a45a429a27ca994b1d8b497

Observation faa03250-f214-463b-a9e4-9f59b0323ed2 · outbound

This paper cites A dataset for movie description.

EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation A dataset for movie description

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:42:00.766236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T21:41:59.603830Z digest=sha256:7baa22cccdce3b1b40fdde43829c8e5fdd4ab3357b0f1d10a91817bad8c82df4

Observation 18334e94-1057-4092-ba8c-79b76be18b43 · outbound

This paper cites First order motion model for image animation.

EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation First order motion model for image animation

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:42:00.751554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T21:41:59.608221Z digest=sha256:84edc1a549e4d3506568afa38bab9fc4a5d41734002eae7df5dbb0d6ec366af1

Observation e733c57b-0751-45e0-a041-78141c6389b4 · outbound

This paper cites Light field networks: Neu- ral scene representations with single-evaluation rendering.

EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation Light field networks: Neu- ral scene representations with single-evaluation rendering

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:42:00.735854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T21:41:59.612979Z digest=sha256:8fcd2bee1db5c8a69f27e18032c5c8c26777e0ce25bd1f5acd11e7c4e40792ce

Observation f2625ca3-5d7a-48d8-b08d-73909418bac7 · outbound

This paper cites UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild.

EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T21:41:59.617554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:41:59.617554Z digest=sha256:f84984fc7b78c5ef2e08bffaa66334e37e68a52aa5f14bd7d9fed3c8a8e6e949

Observation 66ee06a8-3fa8-41df-86a8-25744596c3b0 · outbound

This paper cites Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation.

EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-12T21:41:59.622397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:41:59.622397Z digest=sha256:4427c64c5c6961abc274f57c0bd18f5ae2ca553af52d42ba2aa2e0169f35252a

Observation d3c18b2b-2818-4fc1-9ad7-c5af5c13b606 · outbound

This paper cites VidGen-1M: A Large-Scale Dataset for Text-to-video Generation.

EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation VidGen-1M: A Large-Scale Dataset for Text-to-video Generation

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-12T21:41:59.627308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:41:59.627308Z digest=sha256:852ad9a55fc9cc3d5deb4b0330534545a51e8bc339dda2878704930ebdddd18c

Observation 9daa2966-bb37-4ece-b5f3-a1c1ee307771 · outbound

This paper cites Raft: Recurrent all-pairs field transforms for optical flow.

EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation Raft: Recurrent all-pairs field transforms for optical flow

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:42:00.721210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T21:41:59.632680Z digest=sha256:6327c7a2efbe5d2529807b764f4c6cc3e4534e6f0a1cf40ee84801803c332e1e

Observation e819fe0c-9a0a-4103-ad1a-0568953f6d81 · outbound

This paper cites Towards Accurate Generative Models of Video: A New Metric & Challenges.

EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation Towards Accurate Generative Models of Video: A New Metric & Challenges

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-12T21:41:59.637155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:41:59.637155Z digest=sha256:0f6c2a784b8e6df6a34f3202c456a6d8d916b1835a2f0bd834e0971c3cd47ac8

Observation 87282600-96f5-4e13-a76c-78c11be100f9 · outbound

This paper cites DriveDreamer: Towards Real-world-driven World Models for Autonomous Driving.

EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation DriveDreamer: Towards Real-world-driven World Models for Autonomous Driving

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-12T21:41:59.642313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:41:59.642313Z digest=sha256:adaec3a8d078645a9eab80e06ebd561c0bce2b839f84b69d04cc6404f4839df8

Observation f5a9d4a4-8530-45b9-abf1-1edf48125159 · outbound

This paper cites WorldDreamer: Towards General World Models for Video Generation via Predicting Masked Tokens.

EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation WorldDreamer: Towards General World Models for Video Generation via Predicting Masked Tokens

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-12T21:41:59.648070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:41:59.648070Z digest=sha256:38bfbb3f5acbd3bc6576c44575f82d92e24941efe84dbf86b1f8048c579e8edb

Observation 96f50292-a182-42d3-872f-949ad6711854 · outbound

This paper cites Driving into the Future: Multiview Visual Forecasting and Planning with World Model for Autonomous Driving.

EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation Driving into the Future: Multiview Visual Forecasting and Planning with World Model for Autonomous Driving

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-12T21:41:59.651962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:41:59.651962Z digest=sha256:a1489f2b5734ecd81396446d66909c68218c3d1c8e66db773cb6d0c83191dab3

Observation bb62d593-214e-43ab-b791-82a51b4ba356 · outbound

This paper cites InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation.

EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-12T21:41:59.655925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:41:59.655925Z digest=sha256:77ece31da80f3330bc2ad246dca53d8866b6f7243aeab9012e25a290a14efbcb

Observation c926186f-4ff1-4e77-8587-0340bf85850e · outbound

This paper cites Motionctrl: A unified and flexible motion controller for video generation.

EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation Motionctrl: A unified and flexible motion controller for video generation

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:42:00.706112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T21:41:59.659934Z digest=sha256:37673eb93f32813988abd7d9983604e18f309fe273400a143f2ca78364020733

Observation 781232c7-fa68-47de-a26d-92673ca7d1fc · outbound

This paper cites Exploring video quality assessment on user generated contents from aesthetic and technical perspectives.

EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation Exploring video quality assessment on user generated contents from aesthetic and technical perspectives

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:42:00.691076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T21:41:59.664268Z digest=sha256:cef851231cae93bcecb95b57958e96ba3ac6750a775821a1672e58a505486fb1

Observation 35eb08b1-90d0-47cc-9424-3d45f01780a8 · outbound

This paper cites Daydreamer: World models for physical robot learning.

EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation Daydreamer: World models for physical robot learning

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:42:00.676667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T21:41:59.669034Z digest=sha256:228875f83d647d4902d9366ceea12e1cb294f264f42742049faa9e9f85625051

Observation f356bd31-2fda-4f24-b69e-dbf695baa145 · outbound

This paper cites Dynamicrafter: Animating open-domain images with video diffusion priors.

EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation Dynamicrafter: Animating open-domain images with video diffusion priors

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-12T21:41:59.674276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:41:59.674276Z digest=sha256:f7a4637305bacaf6723629a574f8a7ce5823d18c350112a48792afc08e8d3692

Observation 4cd277c3-c34e-41be-b3f4-83dcb4496ca9 · outbound

This paper cites CamCo: Camera-Controllable 3D-Consistent Image-to-Video Generation.

EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation CamCo: Camera-Controllable 3D-Consistent Image-to-Video Generation

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-12T21:41:59.683353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:41:59.683353Z digest=sha256:7cc7cd2617daf4cbf3e16c56b5a6ac9ef17115dee72fcee26a0d66afdbdf0de1

Observation 6a0df8b5-548a-40f7-8ebf-8e79720ee127 · outbound

This paper cites Egopca: A new framework for egocentric hand-object interaction under- standing.

EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation Egopca: A new framework for egocentric hand-object interaction under- standing

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:42:00.638485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T21:41:59.688070Z digest=sha256:4c4802d1d70ff130c3c45a2fbe948b90e50c2c172281fb3e34bc40607ecd4039

Observation 24eebc82-7aa2-4325-a416-6734eb8969fa · outbound

This paper cites Ad- vancing high-resolution video-language representation with large-scale video transcriptions.

EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation Ad- vancing high-resolution video-language representation with large-scale video transcriptions

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-12T21:41:59.692480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:41:59.692480Z digest=sha256:c1e9daf4ff4e18ff3f94ead6c482aa14a5c4e601d7ece65f5cddee06fb63a1ea

Observation d6e83ca7-25a2-4bcb-9ffd-f219405cf2e6 · outbound

This paper cites VideoGPT: Video Generation using VQ-VAE and Transformers.

EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation VideoGPT: Video Generation using VQ-VAE and Transformers

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-12T21:41:59.696931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:41:59.696931Z digest=sha256:6b6d857a70638d112f2d5ac467c57d2010e30792ea284b715a2e5074bb8b2d09

Observation f07a7966-7fb3-4260-ac2b-5b0abbc2471f · outbound

This paper cites Qwen2 Technical Report.

EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation Qwen2 Technical Report

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-12T21:41:59.701806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:41:59.701806Z digest=sha256:1b05d47608b88c62c9fbcb603a1118c5df581e31e17f37a58d553c15e082aa8d

Observation a103f838-3238-4796-9f92-8e624ba607a7 · outbound

This paper cites Generalized predictive model for autonomous driving.

EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation Generalized predictive model for autonomous driving

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:42:00.614000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T21:41:59.706587Z digest=sha256:1c06fe9c82a8d9fe423af81dcdb99bd188a2c4faab8c3166217b061112cbeedf

Observation 1d0e4ce7-9a03-472f-a774-a975cfc89d1e · outbound

This paper cites Learning interactive real-world simulators.

EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation Learning interactive real-world simulators

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:42:00.598987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T21:41:59.711373Z digest=sha256:a20d97dd978d2cb34180c268fc0ed6d52de64d01af6e74bc799e4608e681559b

Observation 64b15505-9498-4089-b6ed-831b50ab1613 · outbound

This paper cites Direct-a-video: Customized video generation with user- directed camera movement and object motion.

EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation Direct-a-video: Customized video generation with user- directed camera movement and object motion

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:42:00.583132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T21:41:59.716103Z digest=sha256:64af92b05920b9bf255395f237c058485a31fde9b534adf5418c2e4f4992a1ab

Observation 11e9ebcb-76c3-4628-acfe-df96849c5ce2 · outbound

This paper cites Magvit: Masked generative video transformer.

EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation Magvit: Masked generative video transformer

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:42:00.567936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T21:41:59.720571Z digest=sha256:b0fd4998ad325de536a4d4ae9ca633248f85f094d1b1e5ff3a64ffb97ef2fea1

Observation 8e8dd9a6-fccd-4359-b451-d3ccfdec33b6 · outbound

This paper cites Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation.

EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-12T21:41:59.725143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:41:59.725143Z digest=sha256:d54352579794febcbb78679e1949c610974655e4a3dfb9225592169c53d7c66c

Observation 0ca08549-2058-478e-b5f1-b14b3b459368 · outbound

This paper cites Adding conditional control to text-to-image diffusion models.

EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation Adding conditional control to text-to-image diffusion models

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-12T21:41:59.730832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:41:59.730832Z digest=sha256:2a53188728df4580072d2daccae2bce816d7d54ab16d70ef4151668b3b0366be

Observation ce36ee1b-fab0-494e-a6ef-a8b747ba24e3 · outbound

This paper cites Llava- next: A strong zero-shot video understanding model, 2024.

EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation Llava- next: A strong zero-shot video understanding model, 2024

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:42:00.543823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T21:41:59.735082Z digest=sha256:e26e4295d4d0e77617850da93f038d00e9115aa18e694781853b3fe99b523866

Observation 83226228-51d0-480c-a068-0a527c51e990 · outbound

This paper cites DriveDreamer4D: World Models Are Effective Data Machines for 4D Driving Scene Representation.

EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation DriveDreamer4D: World Models Are Effective Data Machines for 4D Driving Scene Representation

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-12T21:41:59.739668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:41:59.739668Z digest=sha256:16b10240ef64f3069a590d440a818ace0f6b59a72cb613bdd53529f801c2aec2

Observation 65d23f81-0508-43a2-948f-b971f4c16946 · outbound

This paper cites DriveDreamer-2: LLM-Enhanced World Models for Diverse Driving Video Generation.

EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation DriveDreamer-2: LLM-Enhanced World Models for Diverse Driving Video Generation

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-12T21:41:59.744181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:41:59.744181Z digest=sha256:d22f02466fa191c545b3f4641055721efa107d3b17e92dc5d51c0ec1427ef274

Observation 01f957f4-392d-4161-92a3-6d3224185e36 · outbound

This paper cites Particlesfm: Exploiting dense point trajecto- ries for localizing moving cameras in the wild.

EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation Particlesfm: Exploiting dense point trajecto- ries for localizing moving cameras in the wild

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-12T21:41:59.748069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:41:59.748069Z digest=sha256:a16da035a642a96f4e6c339607fa6753de8740271e9cdf2675260e5cdb1e9ec8

Observation 013c6124-70d7-4abf-af8c-c117d933b248 · outbound

This paper cites Open-sora: Democratizing efficient video production for all, March 2024.

EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation Open-sora: Democratizing efficient video production for all, March 2024

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:42:00.520079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T21:41:59.752458Z digest=sha256:ed51a43b1e49797cd82becd7dd4437317873b0eda9a07ec93382988c28717037

Observation 0de8bfae-f254-4e5a-a28e-40276a54ea21 · outbound

This paper cites RoboDreamer: Learning Compositional World Models for Robot Imagination.

EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation RoboDreamer: Learning Compositional World Models for Robot Imagination

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-12T21:41:59.757007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:41:59.757007Z digest=sha256:bc7c7283df7097a0561c9d614e1820230683b5e1638e75318ecb6cc44f83b0ba

Observation f50bfe22-6945-4882-b889-a14139ef7987 · outbound

This paper cites Is sora a world simulator? a comprehensive survey on general world models and beyond.

EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation Is sora a world simulator? a comprehensive survey on general world models and beyond

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-12T21:41:59.761872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:41:59.761872Z digest=sha256:d0e542936e79563969176f4ed752a7dca1c6c39bbf25baf0bf36c13a7d43ba93

Observation dd89e2f8-0b6f-43c3-af69-480c6f6ce5d4 · outbound

This paper cites Kinematic Annotation Details To enhance kinematic annotation accuracy, we fuse cam- era poses from IMU and ParticleSfM [80], utilizing the Kalman filter.

EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation Kinematic Annotation Details To enhance kinematic annotation accuracy, we fuse cam- era poses from IMU and ParticleSfM [80], utilizing the Kalman filter

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:42:00.503662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T21:41:59.766300Z digest=sha256:90dd453bef850cc8a6a73f31164e28f42bdb40f1decf1d01913d87fdeb8298b7

Observation 5d259be6-9489-4875-bbdb-1672d24aa502 · outbound

This paper cites Videos cleaned from the five-point optical flow strategy (average optical flow below 3, and the proportion of optical flow (≥12 pixels) is greater than 3%).

EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation Videos cleaned from the five-point optical flow strategy (average optical flow below 3, and the proportion of optical flow (≥12 pixels) is greater than 3%)

Reference 86

Resolution
malformed identifier
raw_fallback, observed 2026-08-12T21:42:00.488053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T21:41:59.771109Z digest=sha256:ec905635fa63849a4ff257f2b16d7d680dd72f1ad68431a78bd377c67e97fc0b

Observation 96d271da-81ed-456a-a36b-98d58c82f283 · outbound

This paper cites an unresolved cited work.

EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation Unresolved cited work

Reference 87

Resolution
unresolved
raw_fallback, observed 2026-08-12T21:42:00.472007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T21:41:59.775784Z digest=sha256:dffdd9d832ca2fd7bce2f1065e4d6105b4d5e8f5c1b04714e3ba5af96b9d3b36

Observation c15cfd60-00ce-4e2c-813b-49e7c21f3cee · outbound

This paper cites These met- rics are as follows: Overall Quality.

EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation These met- rics are as follows: Overall Quality

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:42:00.457058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T21:41:59.780245Z digest=sha256:c400718747a4236aa113d0f3a734c68dcadb29b2f0235bb7b6fbee24814a364b

Observation 2748e6ca-4b9b-427d-8b0f-02089335ae08 · outbound

This paper cites As shown in Fig.

EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation As shown in Fig

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:42:00.441201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T21:41:59.785064Z digest=sha256:f7cd55a54510b0c18a864d966de70bbee0cb602d9a46528a29009acc52496ee6

Observation 627e89fb-d820-4ead-9c7f-7c84e23ec34b · outbound

This paper cites an unresolved cited work.

EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation Unresolved cited work

Reference 2023

Resolution
unresolved
raw_fallback, observed 2026-08-12T21:42:00.652797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T21:41:59.678863Z digest=sha256:626d8a8b4d0ab3b84636c5e5eafd062c4575ecc1eed55d842b48179eb1c646e7

Pith citing papers

Observation f7aaf33b-9987-4e4e-aaa1-a13eb1713491 · inbound

ReconDreamer: Crafting World Models for Driving Scene Reconstruction via Online Restoration cites this paper.

ReconDreamer: Crafting World Models for Driving Scene Reconstruction via Online Restoration EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T10:09:30.704015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:09:30.704015Z digest=sha256:208fadcb92d6640a20c4f057dc21666d60e21c9d1c0bb8dfa8fbc39f6528e5b1

Observation 6f5f1db1-2755-4779-a441-58553b189d1e · inbound

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation cites this paper.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-11T21:45:10.136640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:45:10.136640Z digest=sha256:379cf92938b9cab4817bd81c2b08a4884cd9344b2efe4ae78db44a638819ac27

Observation 1b605e66-2194-46b2-b065-0ed2a0f6d052 · inbound

DAPE: Dual-Stage Parameter-Efficient Fine-Tuning for Consistent Video Editing with Diffusion Models cites this paper.

DAPE: Dual-Stage Parameter-Efficient Fine-Tuning for Consistent Video Editing with Diffusion Models EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T22:30:42.570671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:30:42.570671Z digest=sha256:9b52bc3bc4bbf7c649e535e77a52d1c3901a8696f315c472c0780a113e9632a1

Observation eba3de72-db0f-477b-ac34-c9d316dc5fd1 · inbound

CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning cites this paper.

CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:40.227298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:56:40.227298Z digest=sha256:33d3b69684687dedfe18c2455386fc620f6b06b09d0eae5a61ecbe95199506b7

Observation 259f6f08-efe0-4d4b-ac2c-913566405df4 · inbound

Seeing in the Dark: Benchmarking Egocentric 3D Vision with the Oxford Day-and-Night Dataset cites this paper.

Seeing in the Dark: Benchmarking Egocentric 3D Vision with the Oxford Day-and-Night Dataset EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:51.168830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:51.168830Z digest=sha256:5fa16998a94fe0dcca34a527f851cfd7ce26b4fee7aea7878d35b14164988d7b

Observation d96b73dd-335a-4a97-8689-3d79f91c0368 · inbound

EgoM2P: Egocentric Multimodal Multitask Pretraining cites this paper.

EgoM2P: Egocentric Multimodal Multitask Pretraining EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation

Reference 108

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.617889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.617889Z digest=sha256:9b6defda7137fb0cfaebff93fcf726702ea698e1c9fe3938740ecc7afcb73f85

Observation e2e62075-90b9-442d-8a0f-54f627e3b7a1 · inbound

FreeLong++: Training-Free Long Video Generation via Multi-band SpectralFusion cites this paper.

FreeLong++: Training-Free Long Video Generation via Multi-band SpectralFusion EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T21:28:17.840081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:28:17.840081Z digest=sha256:69aa880b58e01410c7a20f6755edc04d7349bc3e29b1e52b68951e3232d045c7

Observation ec597a67-6a35-4e64-a58c-56dfcbe03a3f · inbound

Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey cites this paper.

Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation

Reference 228

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:28:16.373183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-17T20:28:15.818016Z digest=sha256:be14f8e04df86b614c4b8613b719a87418c9be8f3991c786c5dbfcc9eca5f650

Observation 8d4ee067-2a01-470c-a0e8-d8cbe31222ec · inbound

Precise Action-to-Video Generation Through Visual Action Prompts cites this paper.

Precise Action-to-Video Generation Through Visual Action Prompts EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:34.292314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:34.292314Z digest=sha256:bfcad79a65b3d3949c25b72e5e2d1cb9eb88ec1d5f3007999e3a207562215702

Observation be883004-0e67-4bd2-a90c-745b6dfc6a8e · inbound

SFHand: Learning Embodied Manipulation by Streaming Egocentric 3D Hand Forecasting cites this paper.

SFHand: Learning Embodied Manipulation by Streaming Egocentric 3D Hand Forecasting EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-21T17:54:18.270974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-21T17:53:19.002603Z digest=sha256:eaea61fb37c3b6dc8741c1f555726e66618ca5c7dbc25535290f2bc9a28ff096

Observation 67625f16-802b-41f8-acc7-3a7e45e13f3c · inbound

EgoSim: Egocentric World Simulator for Embodied Interaction Generation cites this paper.

EgoSim: Egocentric World Simulator for Embodied Interaction Generation EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-13T14:40:24.818842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T14:40:24.818842Z digest=sha256:0b5d31eb7c84d97c06fce251cd1c6a4039d3e756d0742ed3beea710b9117a313

Observation ff58e91a-20bb-474b-8274-95c618a83ef0 · inbound

EgoTL: Egocentric Think-Aloud Chains for Long-Horizon Tasks cites this paper.

EgoTL: Egocentric Think-Aloud Chains for Long-Horizon Tasks EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:26:02.978634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T17:10:36.209152Z digest=sha256:d1e9be915aeeaafc4589759df88fc16780b15d0387c56fbb60b5b3e4affc2206

Observation 57fcf62c-8edd-4aee-8b9b-a8a04ea19581 · inbound

Robot Learning from Human Videos: A Survey cites this paper.

Robot Learning from Human Videos: A Survey EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T10:41:29.720260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-07T04:55:44.273643Z digest=sha256:22d7307c14efeee9bd76ebed0fb2c7a2eb5ae1e696763d1b9dc47d11d206854e

Observation f066cf9c-ec7e-4459-8a22-426517529e66 · inbound

VISTA: A Controllable Platform for Generating and Auditing Egocentric Assistance Scenarios cites this paper.

VISTA: A Controllable Platform for Generating and Auditing Egocentric Assistance Scenarios EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T14:24:57.896385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:24:57.896385Z digest=sha256:c19aee6b2ca024eec636ec1833e880e8bc4358b97c145e060b529302b63358ae

Observation 9a614153-e28a-419a-b3dc-dbd1de4d3215 · inbound

World Action Models: The Next Frontier in Embodied AI cites this paper.

World Action Models: The Next Frontier in Embodied AI EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation

Reference 183

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:07:18.085461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-13T05:01:16.802019Z digest=sha256:580adb8da8c2e5b65b5c5f12d64ffc4d44340400ed79532328ce31b6bbb196bd

Observation 03223971-dea8-4844-936a-34fb0f43d6ea · inbound

E$^3$C: Video Generation with 3D Environmental Memory and Ego-Exo Human Pose Control cites this paper.

E$^3$C: Video Generation with 3D Environmental Memory and Ego-Exo Human Pose Control EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation

Reference 60

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T22:34:02.491593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-29T22:24:38.389294Z digest=sha256:e4e024937e04e63e69e64d20876455c2bd0441d14e839e2d5e30b893498530f1

Observation f7401e87-89f8-4008-8fab-a9f396fbf08a · inbound

World Action Models: A Survey cites this paper.

World Action Models: A Survey EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation

Reference 163

Resolution
verified exact
arxiv_id, observed 2026-07-04T04:09:35.543085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-26T17:11:12.686936Z digest=sha256:ba22dec1845345a278537682a15329ca8a1016ae907ee424ca5613ab8cd11ecd

Observation 2d0fbbc3-546e-4783-98f4-cc931057e2bc · inbound

HandsOnWorld: Unconstrained Egocentric Video Generation with Camera-Disentangled Hand Control cites this paper.

HandsOnWorld: Unconstrained Egocentric Video Generation with Camera-Disentangled Hand Control EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-07-03T15:58:37.380383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-03T15:52:18.371702Z digest=sha256:827d59d4736b0369c832464270db9d3a5859460b5a0db70ea35857a77273c3af

Observation 58d28c96-4fad-4bb4-b6a7-adabeebe6cf0 · inbound

iFLYTEK-Embodied-Omni Technical Report cites this paper.

iFLYTEK-Embodied-Omni Technical Report EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-12T12:21:16.963000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T12:21:16.963000Z digest=sha256:716e0938771b11c56b9d4e799aff0c2eebe1c408336fb8789c748042e46679c1

Observation 3f261905-45cb-4fb0-b2d3-10ff1f52fe30 · inbound

EgoSafe: A First-Person Mobile-Captured Benchmark for Visual Safety Understanding cites this paper.

EgoSafe: A First-Person Mobile-Captured Benchmark for Visual Safety Understanding EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T14:08:29.408596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:08:29.408596Z digest=sha256:b4d804878f6c2fae08c9c28a7cc03671e04aefb52e8a0f85a96463869b8a12c6