Pith. sign in

Paper Citation Record · LEDGER

EgoM2P: Egocentric Multimodal Multitask Pretraining

As of 7 August 2026, this Paper Citation Record lists 100 of 138 outbound references and 0 inbound Pith citation observations for arXiv:2506.07886.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.07886 v3

Coverage vector

measured 100 of 138 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:31:42.599141Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 138 outbound references displayed

  • verified exact0
  • verified fuzzy26
  • unresolved74
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 15894f6c-633e-4832-806d-3e43a6a26cda · outbound

This paper cites an unresolved cited work.

EgoM2P: Egocentric Multimodal Multitask Pretraining Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.346223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.346223Z digest=sha256:64c4b01de7c988e0dbf3a96ba87bdd27cbbd71d9818c412ee6208fc4a016e93e

Observation 0a8f9399-7488-4ef2-a179-2678ae232f49 · outbound

This paper cites Phi-3 technical report: A highly capable language model locally on your phone.

EgoM2P: Egocentric Multimodal Multitask Pretraining Phi-3 technical report: A highly capable language model locally on your phone

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.349741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.349741Z digest=sha256:b4da55e4dbb8c92baf176cb9b22d70b72bf05894cc0233c796024fac3a3e606b

Observation cf5137c5-86e3-4333-9e18-23dcab4e0c56 · outbound

This paper cites Cosmos World Foundation Model Platform for Physical AI.

EgoM2P: Egocentric Multimodal Multitask Pretraining Cosmos World Foundation Model Platform for Physical AI

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.352925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.352925Z digest=sha256:f8f89234e18a761fb8605f1c8c0fc5739699cc94a92ec9e78999389d9a4b67f9

Observation f63069d3-06bb-4d13-8e40-4a4c0cd3b781 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

EgoM2P: Egocentric Multimodal Multitask Pretraining Flamingo: a visual language model for few-shot learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.356074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.356074Z digest=sha256:8c1bc7369b57c5561911ed855ddb4e60e410df126b98737750657c074a2fec9f

Observation 7fad7658-0c0b-409e-886d-56e0c4d08177 · outbound

This paper cites Scenescript: Reconstructing scenes with an autoregressive structured language model.

EgoM2P: Egocentric Multimodal Multitask Pretraining Scenescript: Reconstructing scenes with an autoregressive structured language model

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.358916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.358916Z digest=sha256:89ed44f5fb7f8c9435516dd1d597e6f629ac11eb19e112ab98774c0bb4cc75e8

Observation 0617dc0d-0e34-4703-96c7-c1db19e69ccb · outbound

This paper cites Newcombe, and Vasileios Balntas.

EgoM2P: Egocentric Multimodal Multitask Pretraining Newcombe, and Vasileios Balntas

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.361669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.361669Z digest=sha256:beaa45c19993fc018ee9a157db60b3125400418df360be140a54805f537b069c

Observation bc8ce495-3b8b-4b19-bc39-a6519066f386 · outbound

This paper cites MultiMAE: Multi-modal multi-task masked autoencoders.

EgoM2P: Egocentric Multimodal Multitask Pretraining MultiMAE: Multi-modal multi-task masked autoencoders

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.364505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.364505Z digest=sha256:1cbe010e2f79f0f47fd12304d617b65c3f414a16c42a792f5c3812211868c4b9

Observation 449cfd9e-580e-4616-901f-ab5e8a91cdd1 · outbound

This paper cites 4M-21: An any-to-any vision model for tens of tasks and modalities.

EgoM2P: Egocentric Multimodal Multitask Pretraining 4M-21: An any-to-any vision model for tens of tasks and modalities

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.366966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.366966Z digest=sha256:1ae02373a32e17431121f93fdde016f156e0077c61bb2f0b0f9950133f6f6b80

Observation 13c506a7-8635-4d99-9918-5431ab0404ce · outbound

This paper cites Lending a hand: Detecting hands and recognizing activities in complex egocentric interactions.

EgoM2P: Egocentric Multimodal Multitask Pretraining Lending a hand: Detecting hands and recognizing activities in complex egocentric interactions

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.369371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.369371Z digest=sha256:f55ab6a6b976fca0e493f83bab4e0a30f35d72f6de0db6b5008eb8b7781cab32

Observation f68b6e4d-f5d9-4ffe-b548-1bef9b959f7b · outbound

This paper cites Introducing HOT3D: An Egocentric Dataset for 3D Hand and Object Tracking.

EgoM2P: Egocentric Multimodal Multitask Pretraining Introducing HOT3D: An Egocentric Dataset for 3D Hand and Object Tracking

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.372012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.372012Z digest=sha256:5be687768c31dca3b2c725626fc2836e490b7b72d4cf82e466b6e510d57ffafd

Observation b8f5cdf0-83c4-44f1-8530-eb19bc33e420 · outbound

This paper cites Stable video diffusion: Scaling latent video diffusion models to large datasets, 2023.

EgoM2P: Egocentric Multimodal Multitask Pretraining Stable video diffusion: Scaling latent video diffusion models to large datasets, 2023

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.375059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.375059Z digest=sha256:bb36e37503fd66aee3156712679ddb1ab64169c086912bbf43fc6739f2b87649

Observation 49ff6499-ee5d-430f-8f77-31886228d558 · outbound

This paper cites Align your latents: High-resolution video synthe- sis with latent diffusion models.

EgoM2P: Egocentric Multimodal Multitask Pretraining Align your latents: High-resolution video synthe- sis with latent diffusion models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.377428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.377428Z digest=sha256:77904fd6264fb0ba5f776553504c83abe3ec528feb8c9c8e4ced6c763cd2cfa7

Observation 541dd731-2077-4ef5-b5a6-00c8efb30f53 · outbound

This paper cites Scene coordinate reconstruction: Posing of image collections via incremental learning of a relocalizer.

EgoM2P: Egocentric Multimodal Multitask Pretraining Scene coordinate reconstruction: Posing of image collections via incremental learning of a relocalizer

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.379761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.379761Z digest=sha256:360b075db0fcde194a1d50e0a2d698ddb0b2a52279446cb9c70cc62aaea4c420

Observation a852ea0f-f485-4732-8c3c-2aa14db15dee · outbound

This paper cites Genie: Generative interactive environments.

EgoM2P: Egocentric Multimodal Multitask Pretraining Genie: Generative interactive environments

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.382148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.382148Z digest=sha256:f7ec8d1b76667e2bf473333a8622ebd0b60ddb600313f97c15f38f6d5d1633f9

Observation 5c1d5162-9dbf-4d9a-b756-670d26569b51 · outbound

This paper cites End-to-end object detection with transformers.

EgoM2P: Egocentric Multimodal Multitask Pretraining End-to-end object detection with transformers

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.384474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.384474Z digest=sha256:128b6a904313c60e311bbda0b9c312c902a3cc8cd77fb2e2ed4d3aacbb07675c

Observation 3f9630f1-da1a-4b4d-9abd-2f82a4159002 · outbound

This paper cites Emerging properties in self-supervised vision transformers.

EgoM2P: Egocentric Multimodal Multitask Pretraining Emerging properties in self-supervised vision transformers

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.386991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.386991Z digest=sha256:34ae2aad929a68b7fe6a4dd9aa29e8cdf6f8132af4c3abb7490b67128c8a8a75

Observation 15fb6131-0fbc-44c4-a196-1853f71245d8 · outbound

This paper cites Videocrafter2: Overcoming data limitations for high-quality video diffu- sion models, 2024.

EgoM2P: Egocentric Multimodal Multitask Pretraining Videocrafter2: Overcoming data limitations for high-quality video diffu- sion models, 2024

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.389243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.389243Z digest=sha256:ccb99a28b9c7339da46bf98cefdaf004fa09e8cbe86310cfce2bdd6f8f5b0b13

Observation f6e90f27-0f2f-4cdd-98fd-c557c5db6b1b · outbound

This paper cites Fleet, and Geoffrey Hinton.

EgoM2P: Egocentric Multimodal Multitask Pretraining Fleet, and Geoffrey Hinton

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.391946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.391946Z digest=sha256:75da5d48db240680373315b262f237f0af20da1eb64439c12204f873de34bee8

Observation 8afff46f-d8e4-473d-8a61-84d53bd87593 · outbound

This paper cites Control- a-video: Controllable text-to-video diffusion models with motion prior and reward feedback learning, 2024.

EgoM2P: Egocentric Multimodal Multitask Pretraining Control- a-video: Controllable text-to-video diffusion models with motion prior and reward feedback learning, 2024

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.394269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.394269Z digest=sha256:4e1fd8f3dac47b2511684932d91f13065ec449d1f48ced14b43dcc1c8c936954

Observation 2b910072-2163-4309-89dc-77bd1f85f945 · outbound

This paper cites Scaling egocentric vision: The epic- kitchens dataset.

EgoM2P: Egocentric Multimodal Multitask Pretraining Scaling egocentric vision: The epic- kitchens dataset

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.396601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.396601Z digest=sha256:68fdfc059490673c4b3e71617bf0398e5d2cb98a5265f11a5cc736f793d53d63

Observation 00a2618b-7fce-4b42-9bec-9b401e93b5e4 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

EgoM2P: Egocentric Multimodal Multitask Pretraining An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.399010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.399010Z digest=sha256:560e14f2a59a9edbdb4f6427bdada967866f1e73a484ba33c6f83bb7c0bdfb17

Observation 6d95ac89-690b-435f-af8c-9526af2a8c0b · outbound

This paper cites Structure and Content-Guided Video Synthesis with Diffusion Mod- els.

EgoM2P: Egocentric Multimodal Multitask Pretraining Structure and Content-Guided Video Synthesis with Diffusion Mod- els

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.401827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.401827Z digest=sha256:7a12ceb7863b6f81545a56462f4f83e2a371421c69c3c979681d129c52f2b378

Observation 34ffb50e-9f66-460b-8ce2-7ad3e4a53694 · outbound

This paper cites Black, and Otmar Hilliges.

EgoM2P: Egocentric Multimodal Multitask Pretraining Black, and Otmar Hilliges

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.404240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.404240Z digest=sha256:96cf6844af8307288be6fcfe6729c3b41c036f2eb315312d58db051bdaeea10c

Observation 3dfc6897-3adf-4384-893a-18b9fd95211d · outbound

This paper cites HOLD: Category-agnostic 3d reconstruction of interacting hands and objects from video.

EgoM2P: Egocentric Multimodal Multitask Pretraining HOLD: Category-agnostic 3d reconstruction of interacting hands and objects from video

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.406558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.406558Z digest=sha256:67af445685351c2c44c944cc0dea9cabf50efff79a4f57d720415d3ae9a97ba6

Observation 505417fb-eae6-4fa0-b2de-e1c0b6e9f9d2 · outbound

This paper cites VIOLET : End-to-End Video-Language Transformers with Masked Visual-token Modeling.

EgoM2P: Egocentric Multimodal Multitask Pretraining VIOLET : End-to-End Video-Language Transformers with Masked Visual-token Modeling

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.408928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.408928Z digest=sha256:72fb5eaf45dc9392525cf03ae2930876a54022dd639ec67bfa0a19db1dcc4f0c

Observation 55a2aeec-0259-4acf-8ae5-04c17ce47536 · outbound

This paper cites First-person hand action bench- mark with rgb-d videos and 3d hand pose annotations.

EgoM2P: Egocentric Multimodal Multitask Pretraining First-person hand action bench- mark with rgb-d videos and 3d hand pose annotations

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.411701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.411701Z digest=sha256:04f5dd678d557e3328c66e716232b3fbdcb12bb6a73be4412bac911267a32d85

Observation 5fd4f5f4-7520-4d04-a576-e2df90ba89ff · outbound

This paper cites Imagebind: One embedding space to bind them all.

EgoM2P: Egocentric Multimodal Multitask Pretraining Imagebind: One embedding space to bind them all

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.414302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.414302Z digest=sha256:2f0b5a1224ef8d814168ce4e5b603cddc2790210962107e4ee2053cd2c9db380

Observation 9e5a3b5b-3cb4-4816-a31b-6ad45b96c684 · outbound

This paper cites Accurate, large minibatch sgd: Training imagenet in 1 hour, 2018.

EgoM2P: Egocentric Multimodal Multitask Pretraining Accurate, large minibatch sgd: Training imagenet in 1 hour, 2018

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.416754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.416754Z digest=sha256:82b6ed7123f581b3f4e36c53d9d829c94db0a6803f3f69b85ebf3ce033e95c0b

Observation 0801895d-a503-4cb6-a12f-0c36348f3d69 · outbound

This paper cites Ego4d: Around the World in 3,000 Hours of Egocentric Video.

EgoM2P: Egocentric Multimodal Multitask Pretraining Ego4d: Around the World in 3,000 Hours of Egocentric Video

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.419200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.419200Z digest=sha256:e504c23b906fce497a059907c86f0d562ab57e9a3b810e8ed8498262b706f2eb

Observation bcf61676-1751-4004-80ab-d9232c3a7592 · outbound

This paper cites Ego-exo4d: Understanding skilled human activity from first- and third-person perspec- tives.

EgoM2P: Egocentric Multimodal Multitask Pretraining Ego-exo4d: Understanding skilled human activity from first- and third-person perspec- tives

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.421469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.421469Z digest=sha256:03ad17061dc2eb4c6b317b431dadaad5cea99bb55b8002e42d23192842c5b4c6

Observation 59fcde70-c33c-4575-9deb-69eab36e6483 · outbound

This paper cites Animatediff: Animate your personalized text- to-image diffusion models without specific tuning.

EgoM2P: Egocentric Multimodal Multitask Pretraining Animatediff: Animate your personalized text- to-image diffusion models without specific tuning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.423750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.423750Z digest=sha256:7a1fd4b078cf5e51fe955946c6b65d355e2784e13fa85745f98d5ba19cbb3284

Observation 9af1b819-83cf-4904-80fc-840d4704c4aa · outbound

This paper cites World Models.

EgoM2P: Egocentric Multimodal Multitask Pretraining World Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.426124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.426124Z digest=sha256:a9865c58b0a0ab762c8e9b147c51f9b5f381d085c74b8efd7dda897d58d21aa1

Observation 7c1575ea-1ee2-4ed0-8020-9e970c0fd46a · outbound

This paper cites Girshick.

EgoM2P: Egocentric Multimodal Multitask Pretraining Girshick

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.428661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.428661Z digest=sha256:68832a26217d6d9d0404a5efd7528cb9ebe8ccb8da1f1cbb79c76bee96b10162

Observation cfecc510-d9c4-4fa7-8696-949fad01d9f0 · outbound

This paper cites Classifier-free diffusion guidance.

EgoM2P: Egocentric Multimodal Multitask Pretraining Classifier-free diffusion guidance

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.430965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.430965Z digest=sha256:607a16d8c535580a262705b9007660f7ff4d89e8dd2d350b23f46ce4a9b857ec

Observation c087ee39-de85-40d5-ab00-c371152241bb · outbound

This paper cites Video diffu- sion models.

EgoM2P: Egocentric Multimodal Multitask Pretraining Video diffu- sion models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.433709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.433709Z digest=sha256:05bcda8ab6a7a876ec4260b13990140a30d7e2c54f239a15937f7f8556567eca

Observation 4a4af980-0b14-40ca-abf4-c16d7b4263c4 · outbound

This paper cites The curious case of neural text degeneration.

EgoM2P: Egocentric Multimodal Multitask Pretraining The curious case of neural text degeneration

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.436271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.436271Z digest=sha256:37ababd49e5509ab1ac935814abc7807acf03e9d8bc58a31ce220fb56565f1fc

Observation 7c459429-91fc-4ea3-a3cc-f48d486d3cf5 · outbound

This paper cites Cogvideo: Large-scale pretraining for text-to- video generation via transformers.

EgoM2P: Egocentric Multimodal Multitask Pretraining Cogvideo: Large-scale pretraining for text-to- video generation via transformers

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.438690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.438690Z digest=sha256:c1c2ddd8bfe99fb4c2cb01924d925d77b07629cc1550efa420f4be64ebeb49a6

Observation 1e6a5851-42d7-4baf-be4c-4db19fd1e816 · outbound

This paper cites Gritsenko, Jasmijn Bastings, Ben Poole, Rianne van den Berg, and Tim Salimans.

EgoM2P: Egocentric Multimodal Multitask Pretraining Gritsenko, Jasmijn Bastings, Ben Poole, Rianne van den Berg, and Tim Salimans

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.441013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.441013Z digest=sha256:f56fc05456839be7be0c25b4946eec1d1a44b8585b5986c91c21a864070f813c

Observation 0ff59b2a-d994-4462-9651-ffcf2ec61f47 · outbound

This paper cites Ross, and Alireza Fathi.

EgoM2P: Egocentric Multimodal Multitask Pretraining Ross, and Alireza Fathi

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.443348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.443348Z digest=sha256:04f5bab51ebdb31a306b5897bf0b369282fd47c7110bc867b20b3d6f49d0adf5

Observation d63a9227-59ab-40b8-b469-aae67ab99370 · outbound

This paper cites Predicting gaze in egocentric video by learning task- dependent attention transition.

EgoM2P: Egocentric Multimodal Multitask Pretraining Predicting gaze in egocentric video by learning task- dependent attention transition

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.445693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.445693Z digest=sha256:87dd22c1e1a67b3ab0e5e0261c025c1eb1fab7791674f9740b72a637462d19b1

Observation 2a3c13d0-fb71-46ce-b967-3ad57c9ea5f2 · outbound

This paper cites GPT-4o System Card.

EgoM2P: Egocentric Multimodal Multitask Pretraining GPT-4o System Card

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.447923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.447923Z digest=sha256:281ea6a1191004764b1e3c677024f9797a8abe9a2dc4c45b84fa2cfa2d29c448

Observation 400f5c80-0591-4a7f-a8ab-ffd96cbb14b8 · outbound

This paper cites Epic-fusion: Audio-visual temporal binding for egocentric action recognition.

EgoM2P: Egocentric Multimodal Multitask Pretraining Epic-fusion: Audio-visual temporal binding for egocentric action recognition

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.450515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.450515Z digest=sha256:a2e1c60c7ea469f5c5e613d7a0e9bf969e01ad1f765fa8b97ca418e1dbbc4865

Observation 04454c8a-dc38-4169-a86b-a8039f528166 · outbound

This paper cites Re- purposing diffusion-based image generators for monocular depth estimation.

EgoM2P: Egocentric Multimodal Multitask Pretraining Re- purposing diffusion-based image generators for monocular depth estimation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.453671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.453671Z digest=sha256:6db32e9d8de32f0b3d01abd286274ae4aaddca9deb1103608062eeea4ef75f84

Observation c185bded-22cc-4cc7-bf12-3d5c9ceec635 · outbound

This paper cites Video depth without video models.

EgoM2P: Egocentric Multimodal Multitask Pretraining Video depth without video models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.456344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.456344Z digest=sha256:764f1e7b37824c26a458617d52dab527a9adb1bc76f9d333a7e520b451e84e08

Observation 36c03888-f351-40d8-834e-23e0d4fd5506 · outbound

This paper cites Text2Video-Zero: Text- to-Image Diffusion Models are Zero-Shot Video Generators.

EgoM2P: Egocentric Multimodal Multitask Pretraining Text2Video-Zero: Text- to-Image Diffusion Models are Zero-Shot Video Generators

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.458614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.458614Z digest=sha256:9a7356e8a05604cd2f1a73a027588247a03cb7151fd836fa69e2f51759915f9a

Observation 306b8dd8-70c4-4016-8e9a-b90231c1cee4 · outbound

This paper cites Segment anything.

EgoM2P: Egocentric Multimodal Multitask Pretraining Segment anything

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.460971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.460971Z digest=sha256:3d50ce93342d3888bc1f6c1c06b7ab9436c63360220a9a9199d75420c23f6b6c

Observation 85cfafe2-386a-4e0b-a8b3-b6856a206452 · outbound

This paper cites Harmsen, and Neil Houlsby.

EgoM2P: Egocentric Multimodal Multitask Pretraining Harmsen, and Neil Houlsby

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.463435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.463435Z digest=sha256:9c49d18606d0bb50d52800940120965f89e273b3c721f835879fb58e75e41f84

Observation 56202ac3-0576-4504-b221-2e677d77f744 · outbound

This paper cites VideoPoet: A large language model for zero-shot video generation.

EgoM2P: Egocentric Multimodal Multitask Pretraining VideoPoet: A large language model for zero-shot video generation

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.465709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.465709Z digest=sha256:8bc26e94611be2b7baa598bb7ebb31b31bb1e96ad772a78858b65fd5ac7ab469

Observation 8aaa8fa9-d631-4558-9fcf-3e60bf0ae668 · outbound

This paper cites H2o: Two hands manipulating objects for first person interaction recognition.

EgoM2P: Egocentric Multimodal Multitask Pretraining H2o: Two hands manipulating objects for first person interaction recognition

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.468298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.468298Z digest=sha256:684d361f06c7128877085419665600b5defcf0dfbb2494dff75afd21cdefe4ee

Observation 4caa288a-432f-4f65-affa-3db175476b05 · outbound

This paper cites an unresolved cited work.

EgoM2P: Egocentric Multimodal Multitask Pretraining Unresolved cited work

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.470764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.470764Z digest=sha256:8586bb18e6ab166870cbb0cf0beccc69a5cd98c8b2690aa4ad194fe8e77de5fe

Observation a82b8a1d-6305-4cfa-93db-c0ffa7b3405a · outbound

This paper cites Lisa: Reasoning segmenta- tion via large language model.

EgoM2P: Egocentric Multimodal Multitask Pretraining Lisa: Reasoning segmenta- tion via large language model

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.473203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.473203Z digest=sha256:321c934b46bb89a8dbdfdb816139d15a2f0fbf7c6ea265ba6753de82bb5c455f

Observation 31d3e8ba-650b-4fb4-992f-ca9afb0f6ad1 · outbound

This paper cites Egogen: An egocentric synthetic data generator.

EgoM2P: Egocentric Multimodal Multitask Pretraining Egogen: An egocentric synthetic data generator

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.475810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.475810Z digest=sha256:ef1b21ae9a8e307f969c8605df4bc8b0e1a50c848ccc984cff76d5473a11bbd4

Observation f4bd1b8c-c516-46d8-a784-fb266f7e1a4f · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

EgoM2P: Egocentric Multimodal Multitask Pretraining VideoChat: Chat-Centric Video Understanding

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.478451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.478451Z digest=sha256:4ee82ad093c7023a1ff672eb7b4d57f53b3733e41f9d1d4a0e01a7ac1ab96432

Observation aec46962-59cc-45c6-acbc-8457da029965 · outbound

This paper cites Megasam: Accurate, fast and robust structure and motion from casual dynamic videos.

EgoM2P: Egocentric Multimodal Multitask Pretraining Megasam: Accurate, fast and robust structure and motion from casual dynamic videos

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.481256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.481256Z digest=sha256:944cd3f6b373b5bf075c856dbdc66c5be4522ee0f7582b338f3972fdf7783bda

Observation 45d097af-9a90-49d3-91e0-60c25371e42e · outbound

This paper cites Video-LLaV A: Learning united visual representation by alignment before projection.

EgoM2P: Egocentric Multimodal Multitask Pretraining Video-LLaV A: Learning united visual representation by alignment before projection

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.483820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.483820Z digest=sha256:49c968f5f89806541a839edf526917aaf7f0bf7f5aac7249d421a3812e634daf

Observation 3e263881-3e9a-4794-86af-d218551e21ae · outbound

This paper cites Cross-view exocentric to egocentric video synthesis.

EgoM2P: Egocentric Multimodal Multitask Pretraining Cross-view exocentric to egocentric video synthesis

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.486760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.486760Z digest=sha256:2fa3113b9ef42731b82d95409cd553ed9962347ba18cc15e07e7b4329c19c7e3

Observation 01a27f75-5b00-4edf-8daa-383a5ad7a7c2 · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024.

EgoM2P: Egocentric Multimodal Multitask Pretraining Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.489037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.489037Z digest=sha256:567e0d2ca793f58955add2d34e05afa8f75544ba17dc0fe796f0cbd4d5a30182

Observation 6b3c69ab-e05a-4fbb-b93b-29013773e9f4 · outbound

This paper cites Exocentric-to-egocentric video gener- ation.

EgoM2P: Egocentric Multimodal Multitask Pretraining Exocentric-to-egocentric video gener- ation

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.491572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.491572Z digest=sha256:c9b8760b53e903ba0c961bcf8831ff17ec5493b539eabdd4b4db3703be32e4f4

Observation 6f7ff5ef-cb26-4e70-a7fc-dc3429d9c4d3 · outbound

This paper cites Li, Ying Shan, and Ge Li.

EgoM2P: Egocentric Multimodal Multitask Pretraining Li, Ying Shan, and Ge Li

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.494202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.494202Z digest=sha256:f29383e45712d646d117dd4bdc300508f74082f2d381435ac6674bb828ebec49

Observation fee22a39-4c72-4ac1-8728-4142f468e909 · outbound

This paper cites Hoi4d: A 4d egocentric dataset for category-level human-object interaction.

EgoM2P: Egocentric Multimodal Multitask Pretraining Hoi4d: A 4d egocentric dataset for category-level human-object interaction

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.499922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.499922Z digest=sha256:2b5880f75f589ece0b04b6555c3b1934d958f1949a81851081dd45c048cc6d0e

Observation baa0c2cf-52c5-41ae-99fe-26b4bc392390 · outbound

This paper cites an unresolved cited work.

EgoM2P: Egocentric Multimodal Multitask Pretraining Unresolved cited work

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.502461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.502461Z digest=sha256:84fa6f5a5e3b1ddd57923cbef2255b45198ce2dafabd48d6cdad6cbef2be55d8

Observation 9064648b-65c1-4b38-a48f-fdd5d5acbc60 · outbound

This paper cites A convnet for the 2020s.

EgoM2P: Egocentric Multimodal Multitask Pretraining A convnet for the 2020s

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.505098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.505098Z digest=sha256:93989c02d46f7a49b55a02513cbf4bf25ac9611307b333ca2af6997411d98df8

Observation b95d2052-f35f-4041-8469-e7085267b8e4 · outbound

This paper cites Decoupled weight de- cay regularization.

EgoM2P: Egocentric Multimodal Multitask Pretraining Decoupled weight de- cay regularization

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.507569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.507569Z digest=sha256:c149970dc830c626bfb434c2fac00f3324c9dfbb417c4de017c3583b4bc94b0b

Observation d35445cd-14d2-407e-a5e3-3c65dc60ab61 · outbound

This paper cites Unified-io 2: Scaling autoregressive mul- timodal models with vision, language, audio, and action.

EgoM2P: Egocentric Multimodal Multitask Pretraining Unified-io 2: Scaling autoregressive mul- timodal models with vision, language, audio, and action

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.509938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.509938Z digest=sha256:034fdee7c097f644818130405bbe1e07d2b16ac04a870f3537754d02636dd6d7

Observation 4ddd5a67-06cf-42cf-968e-3ddfee63a802 · outbound

This paper cites UNIFIED-IO: A uni- fied model for vision, language, and multi-modal tasks.

EgoM2P: Egocentric Multimodal Multitask Pretraining UNIFIED-IO: A uni- fied model for vision, language, and multi-modal tasks

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:31:43.345152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:31:42.512343Z digest=sha256:1b930e986c060dd145564881b18cb1e662f2f88ff527c3f5a3cbb708c29c400b

Observation a87fcc94-61c6-4fb1-a7f3-9790ca147d68 · outbound

This paper cites Align3r: Aligned monocular depth estimation for dynamic videos.

EgoM2P: Egocentric Multimodal Multitask Pretraining Align3r: Aligned monocular depth estimation for dynamic videos

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:31:43.337663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:31:42.515421Z digest=sha256:35cf0587601d24a0a100493cd82f34c08186c7e0b3f25ef7d0e3a6aeab700a61

Observation e196fdd4-630d-480a-8b97-1abbf6d199b6 · outbound

This paper cites Dream Machine.

EgoM2P: Egocentric Multimodal Multitask Pretraining Dream Machine

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:31:43.329825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:31:42.518003Z digest=sha256:035193c94f355b715a4f36f37525c368fde12fecb57e2a9831cb385e39df6bf2

Observation 85040bb0-72f9-4587-923d-e6ba6321b337 · outbound

This paper cites Videofusion: Decomposed diffusion models for high-quality video generation.

EgoM2P: Egocentric Multimodal Multitask Pretraining Videofusion: Decomposed diffusion models for high-quality video generation

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:31:43.322374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:31:42.520485Z digest=sha256:5fe2d354919645977f67ea6de8a64d279b7ce4bd26168d41ad5db9442e2210ba

Observation b49d3759-826d-4058-912b-f96bc2c8be84 · outbound

This paper cites Aria Everyday Activities Dataset.

EgoM2P: Egocentric Multimodal Multitask Pretraining Aria Everyday Activities Dataset

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.522928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.522928Z digest=sha256:c9f35584500eee0cc3c4fb8f33ce711fcaa0842816891916874f982651d8b1dc

Observation b6e9ab21-5033-431c-80fe-6db7083e494f · outbound

This paper cites Nymeria: A Massive Collection of Multimodal Egocentric Daily Motion in the Wild.

EgoM2P: Egocentric Multimodal Multitask Pretraining Nymeria: A Massive Collection of Multimodal Egocentric Daily Motion in the Wild

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.525404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.525404Z digest=sha256:da46d7e407835b2acd02bd7697e458f1da22db6253ef3a514ba860cd45807fc0

Observation 1fd07221-c02f-4ec3-a438-49767fe5dfa2 · outbound

This paper cites Video-chatgpt: Towards detailed video understanding via large vision and language models.

EgoM2P: Egocentric Multimodal Multitask Pretraining Video-chatgpt: Towards detailed video understanding via large vision and language models

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:31:43.315038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:31:42.527939Z digest=sha256:a2ce09ea4f869c9f7ec0bbf1447da20a46ecc9b097b36c0d32bea3a53c514c23

Observation 9977c314-da17-4ddf-b118-483416f1e3f9 · outbound

This paper cites Mm1: Methods, analysis & insights from multimodal llm pre-training, 2024.

EgoM2P: Egocentric Multimodal Multitask Pretraining Mm1: Methods, analysis & insights from multimodal llm pre-training, 2024

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:31:43.306917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:31:42.530377Z digest=sha256:83c8efdc831de2d2ed2d6f6a91066b321ecd2fa71d531dad36d6a29b00852849

Observation d9deb7a6-65ed-4c3a-b7db-84c72c36390f · outbound

This paper cites Project Aria Glasses.

EgoM2P: Egocentric Multimodal Multitask Pretraining Project Aria Glasses

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:31:43.299588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:31:42.532866Z digest=sha256:080354fcd144fbed587ced8bced44b8dab951ec379cf11d79bef59de3b56dfdd

Observation 8ba322c7-646c-418d-9622-261ee4a5c97b · outbound

This paper cites Transformers are Sample-Efficient World Models.

EgoM2P: Egocentric Multimodal Multitask Pretraining Transformers are Sample-Efficient World Models

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.535714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.535714Z digest=sha256:24a01eae1c9dd7cc811bae4590b313a169ebe839024f727d9f8419a34b48a869

Observation 5dc57102-cd1d-463d-b21a-8dcf0ada7d48 · outbound

This paper cites HoloLens 2.

EgoM2P: Egocentric Multimodal Multitask Pretraining HoloLens 2

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:31:43.292402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:31:42.538301Z digest=sha256:0c321b3328501b2b3390515b5157b608dea053c1ec8272a9e165a25f76874490

Observation c137608e-1b3b-4edf-bd0f-d879c2a3dbf5 · outbound

This paper cites 4M: Massively multimodal masked modeling.

EgoM2P: Egocentric Multimodal Multitask Pretraining 4M: Massively multimodal masked modeling

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:31:43.285124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:31:42.540832Z digest=sha256:940d5eb20ba0a7de84a2f4e3ef0da52a1c4617359ebda5b67a8dc5d62ce8d9b8

Observation b4420ba5-ac62-4b89-b5cb-6046902b1220 · outbound

This paper cites AssemblyHands: towards egocentric activity understanding via 3d hand pose esti- mation.

EgoM2P: Egocentric Multimodal Multitask Pretraining AssemblyHands: towards egocentric activity understanding via 3d hand pose esti- mation

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:31:43.277532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:31:42.543229Z digest=sha256:9e3245c8aa170d049bdf3e5c365d612ded82b8d34a98486ae5af43ffd6a1b5a9

Observation c79c1e4d-4e9a-4db0-aae0-0eff5aeb866a · outbound

This paper cites Video generation models as world simula- tors.

EgoM2P: Egocentric Multimodal Multitask Pretraining Video generation models as world simula- tors

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:31:43.269716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:31:42.545537Z digest=sha256:0ddfd0a2a69864a43dc771e0277df52dda70fe65ad56ca9ddfef0b6138815e21

Observation 374a7076-4d4a-4983-9a51-fd3c31220660 · outbound

This paper cites an unresolved cited work.

EgoM2P: Egocentric Multimodal Multitask Pretraining Unresolved cited work

Reference 79

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:31:43.262258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:31:42.548214Z digest=sha256:3a6453f9b0445ceba50f1c1e14401277fd24d7894d9f85ffa0b7e57d9a2b5937

Observation 3fb761d3-f991-4650-a924-fac389397dcf · outbound

This paper cites Aria digital twin: A new benchmark dataset for egocentric 3d machine perception.

EgoM2P: Egocentric Multimodal Multitask Pretraining Aria digital twin: A new benchmark dataset for egocentric 3d machine perception

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:31:43.254661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:31:42.550637Z digest=sha256:245bf08a1de32c64b0293bee9fb9bff2d3452d7e8c96b42347d8488360b9a734

Observation b0563564-1aa4-491d-a10b-6305dffedad5 · outbound

This paper cites Movie Gen: A Cast of Media Foundation Models.

EgoM2P: Egocentric Multimodal Multitask Pretraining Movie Gen: A Cast of Media Foundation Models

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.552979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.552979Z digest=sha256:22ddb76fcc91c0bcae7c6bab9c4c53be6cc483e4910fcb14df31d54a188c817d

Observation 2a104cc6-49f3-486a-a341-1b1e0f6f3a57 · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

EgoM2P: Egocentric Multimodal Multitask Pretraining Learn- ing transferable visual models from natural language super- vision

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:31:43.245929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:31:42.555752Z digest=sha256:2a125b1ec11c252a2cdb27cc855432c203df4a2dd476b1df2964b79e0881a95d

Observation 43e157ec-730b-452b-8b52-c5821f2d7f52 · outbound

This paper cites an unresolved cited work.

EgoM2P: Egocentric Multimodal Multitask Pretraining Unresolved cited work

Reference 83

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:31:43.237643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:31:42.558132Z digest=sha256:71ffe618501d6c93d1d8b80bbd5210ae74bd63fc927381fbe4d629ef2a183721

Observation 64ae6b91-6d2e-4268-bbab-dff2f7364b2d · outbound

This paper cites High-Resolution Im- age Synthesis with Latent Diffusion Models.

EgoM2P: Egocentric Multimodal Multitask Pretraining High-Resolution Im- age Synthesis with Latent Diffusion Models

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:31:43.229469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:31:42.560589Z digest=sha256:d2b3b1c5d8502e485704a78fd18f64444867c138d501303b261b21b8eacd8ccb

Observation b3f70a23-1293-4c39-9ff0-f360abc10974 · outbound

This paper cites Gen-3 Alpha.

EgoM2P: Egocentric Multimodal Multitask Pretraining Gen-3 Alpha

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:31:43.221152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:31:42.562991Z digest=sha256:ab5c584c350122488b959d278a9ed5415d53c5e155c88db54c0716125a618a7d

Observation 66bddc3a-7e9c-4f03-bd41-b88c485106d3 · outbound

This paper cites Lamar: Bench- marking localization and mapping for augmented reality.

EgoM2P: Egocentric Multimodal Multitask Pretraining Lamar: Bench- marking localization and mapping for augmented reality

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:31:43.212813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:31:42.565470Z digest=sha256:5ff4793043e0b6f577decb6df954fa310890fd05284e085c4d5e8e8a79f013e3

Observation 07d1e775-b7d6-4c92-94dc-a1b1a51ed7b1 · outbound

This paper cites Sener, D.

EgoM2P: Egocentric Multimodal Multitask Pretraining Sener, D

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:31:43.203556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:31:42.567826Z digest=sha256:3a83e111190ebba71d7415e7ae9149429a5dddf261f3be4be2d442f297ee02b9

Observation ee2d64bc-3bc5-478e-af14-5ff57d27f4ce · outbound

This paper cites Make-a-video: Text-to-video generation without text-video data.

EgoM2P: Egocentric Multimodal Multitask Pretraining Make-a-video: Text-to-video generation without text-video data

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.570245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.570245Z digest=sha256:b2f1177f8a2c16696e7accd1ed6133aabbf1938a7d9d9fa1e31561930c2ee5ed

Observation 2706b4d7-8665-4f5e-9dc2-06dd67743634 · outbound

This paper cites The Replica Dataset: A Digital Replica of Indoor Spaces.

EgoM2P: Egocentric Multimodal Multitask Pretraining The Replica Dataset: A Digital Replica of Indoor Spaces

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.572653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.572653Z digest=sha256:96521b6fe79a8f4c9678bfc992ba30d68ea17f76c872c5d1fdc2a0d83eb5123d

Observation c84926b5-1768-451c-b056-ae70ede2ec88 · outbound

This paper cites Part, Ioannis Papaioannou, Arash Eshghi, Ioannis Konstas, and Oliver Lemon.

EgoM2P: Egocentric Multimodal Multitask Pretraining Part, Ioannis Papaioannou, Arash Eshghi, Ioannis Konstas, and Oliver Lemon

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:31:43.191261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:31:42.575522Z digest=sha256:d99fb0c0081915c1d6ae7d73454ef55d2b068e64ec245a8d0403492a1159d13c

Observation c5e558a0-b2ed-4d3d-bf55-4fb269ab28b3 · outbound

This paper cites Emu: Generative pretraining in multimodality.

EgoM2P: Egocentric Multimodal Multitask Pretraining Emu: Generative pretraining in multimodality

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:31:43.183503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:31:42.577766Z digest=sha256:cd8a0d8c430651b5bcf71d151cf229ba8c62729fa2513a745484e6bae4e015a5

Observation 9ba77c0d-e496-42e5-a6c2-5147d5e6eb40 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

EgoM2P: Egocentric Multimodal Multitask Pretraining Gemini: A Family of Highly Capable Multimodal Models

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.580076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.580076Z digest=sha256:7d26cf0740aa27db77406bfa752cbacb976e0a09bfcfd83cfab0efa729583d5a

Observation 522fd6c4-b46a-4fc0-b37e-a072ab425b44 · outbound

This paper cites Kling ai video generator.

EgoM2P: Egocentric Multimodal Multitask Pretraining Kling ai video generator

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:31:43.176365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:31:42.582501Z digest=sha256:9ebad23f0b6af2267f512fea9f0c6f9484a10cddd72e4efa59f6b50ff0995841

Observation e8559f55-245d-4006-a0d8-af30007aedce · outbound

This paper cites DROID-SLAM: Deep vi- sual SLAM for monocular, stereo, and RGB-d cameras.

EgoM2P: Egocentric Multimodal Multitask Pretraining DROID-SLAM: Deep vi- sual SLAM for monocular, stereo, and RGB-d cameras

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:31:43.168267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:31:42.584808Z digest=sha256:ee8761be833d0222c5cdf0999ebaf6224579b2f377444e86cc9a3d7859491a73

Observation 2fa6a850-a295-4b97-a6bc-d0ddc6ddd58f · outbound

This paper cites VideoMAE: Masked autoencoders are data-efficient learn- ers for self-supervised video pre-training.

EgoM2P: Egocentric Multimodal Multitask Pretraining VideoMAE: Masked autoencoders are data-efficient learn- ers for self-supervised video pre-training

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:31:43.160061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:31:42.587075Z digest=sha256:5aff050c25b1cd7e75357114ba0c4177bfab6a6f5917066f7357ba5b20cc2c35

Observation 2a8dbe85-a2d5-4cb7-a498-9d7c352ce84d · outbound

This paper cites Towards accurate generative models of video: A new metric & challenges, 2019.

EgoM2P: Egocentric Multimodal Multitask Pretraining Towards accurate generative models of video: A new metric & challenges, 2019

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:31:43.151538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:31:42.589235Z digest=sha256:5b698f52d01e29e341ce15ff9a50f5287baa5e20c95c22dbc148f3f12b1c492d

Observation 1e2f2197-f1ed-4ec3-9ad4-e1979f876097 · outbound

This paper cites Diffusion Models Are Real-Time Game Engines.

EgoM2P: Egocentric Multimodal Multitask Pretraining Diffusion Models Are Real-Time Game Engines

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.591660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.591660Z digest=sha256:da3be5e8591ece8de128d04516fb7fd954d55f429e25f7ff985b2428af80b732

Observation 7e673542-0d8e-4206-bd37-c9d6f93d6215 · outbound

This paper cites Neural discrete representation learn- ing.

EgoM2P: Egocentric Multimodal Multitask Pretraining Neural discrete representation learn- ing

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:31:43.142981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:31:42.594555Z digest=sha256:5a3937fd453a8bc24ffcae93d196db77229948d44e67d5f3cf9c877e5c0c0052

Observation 3fc44bf5-10be-4363-8f68-d9a1423b43b9 · outbound

This paper cites Attention is all you need.

EgoM2P: Egocentric Multimodal Multitask Pretraining Attention is all you need

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:31:43.134665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:31:42.596902Z digest=sha256:96f57f64e6c33d31bef08c2c0c4fa253009553cab8c53ed5f7af02ff2a0e6aae

Observation f383cfda-5eff-479b-b21c-46daadaed112 · outbound

This paper cites Phenaki: Variable length video generation from open do- main textual descriptions.

EgoM2P: Egocentric Multimodal Multitask Pretraining Phenaki: Variable length video generation from open do- main textual descriptions

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:31:43.126418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:31:42.599141Z digest=sha256:00b7edfa35f9d98744ae830bb1a3996d3b608666ef5fd5e6384281b133b2d354

Pith citing papers

No inbound Pith citation observations are available.