Pith. sign in

Paper Citation Record · LEDGER

EgoM2P: Egocentric Multimodal Multitask Pretraining

As of 20 August 2026, this Paper Citation Record lists 100 of 138 outbound references and 0 inbound Pith citation observations for arXiv:2506.07886.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.07886 v3

Coverage vector

measured 100 of 138 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:31:42.599141Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 138 outbound references displayed

  • verified exact0
  • verified fuzzy26
  • unresolved74
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 15894f6c-633e-4832-806d-3e43a6a26cda · outbound

This paper cites an unresolved cited work.

EgoM2P: Egocentric Multimodal Multitask Pretraining Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.346223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.346223Z digest=sha256:800c49599e1fa5571ce59ef4b7caf73ec13216ad5bd8dffdaeb73045174a43ca

Observation 0a8f9399-7488-4ef2-a179-2678ae232f49 · outbound

This paper cites Phi-3 technical report: A highly capable language model locally on your phone.

EgoM2P: Egocentric Multimodal Multitask Pretraining Phi-3 technical report: A highly capable language model locally on your phone

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.349741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.349741Z digest=sha256:01fa4e2aaa41a7225d99145daa89c730a677b14d5a6e477cfe0d9f0d7a67220f

Observation cf5137c5-86e3-4333-9e18-23dcab4e0c56 · outbound

This paper cites Cosmos World Foundation Model Platform for Physical AI.

EgoM2P: Egocentric Multimodal Multitask Pretraining Cosmos World Foundation Model Platform for Physical AI

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.352925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.352925Z digest=sha256:05f3b2c885cac61fe858e197067d5b72e0c91677a92cd0809b58198b15d336fc

Observation f63069d3-06bb-4d13-8e40-4a4c0cd3b781 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

EgoM2P: Egocentric Multimodal Multitask Pretraining Flamingo: a visual language model for few-shot learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.356074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.356074Z digest=sha256:79fd50b484d8bbbdf070ba7da1611e043821d220e1c6555f8073678d1a57c63a

Observation 7fad7658-0c0b-409e-886d-56e0c4d08177 · outbound

This paper cites Scenescript: Reconstructing scenes with an autoregressive structured language model.

EgoM2P: Egocentric Multimodal Multitask Pretraining Scenescript: Reconstructing scenes with an autoregressive structured language model

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.358916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.358916Z digest=sha256:f3d9a8fc858c5a5ca3b7a82483898c82c69529ab74a8f6f74ffab19394e755a5

Observation 0617dc0d-0e34-4703-96c7-c1db19e69ccb · outbound

This paper cites Newcombe, and Vasileios Balntas.

EgoM2P: Egocentric Multimodal Multitask Pretraining Newcombe, and Vasileios Balntas

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.361669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.361669Z digest=sha256:6d19d57a4e1ea30ce8ea4412c5517490f95ff2a57d9de8aa299f7c2ef0066afc

Observation bc8ce495-3b8b-4b19-bc39-a6519066f386 · outbound

This paper cites MultiMAE: Multi-modal multi-task masked autoencoders.

EgoM2P: Egocentric Multimodal Multitask Pretraining MultiMAE: Multi-modal multi-task masked autoencoders

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.364505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.364505Z digest=sha256:e630862fffb19d30206147762b3583a4d5f95a2c3e1309e9be4e75e9cb4f6bd0

Observation 449cfd9e-580e-4616-901f-ab5e8a91cdd1 · outbound

This paper cites 4M-21: An any-to-any vision model for tens of tasks and modalities.

EgoM2P: Egocentric Multimodal Multitask Pretraining 4M-21: An any-to-any vision model for tens of tasks and modalities

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.366966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.366966Z digest=sha256:3cc5c4757a5024a696b023ba8933440afc469f3b67f847184a7857df4b346d62

Observation 13c506a7-8635-4d99-9918-5431ab0404ce · outbound

This paper cites Lending a hand: Detecting hands and recognizing activities in complex egocentric interactions.

EgoM2P: Egocentric Multimodal Multitask Pretraining Lending a hand: Detecting hands and recognizing activities in complex egocentric interactions

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.369371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.369371Z digest=sha256:079200c43556f0423b892ae99544bb3d5fc3acf7597d391d62016ec5f92aa356

Observation f68b6e4d-f5d9-4ffe-b548-1bef9b959f7b · outbound

This paper cites Introducing HOT3D: An Egocentric Dataset for 3D Hand and Object Tracking.

EgoM2P: Egocentric Multimodal Multitask Pretraining Introducing HOT3D: An Egocentric Dataset for 3D Hand and Object Tracking

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.372012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.372012Z digest=sha256:a80e49a4e7aef54ade45dee08d01eb1694b7b8802240546d990174f00c580c4f

Observation b8f5cdf0-83c4-44f1-8530-eb19bc33e420 · outbound

This paper cites Stable video diffusion: Scaling latent video diffusion models to large datasets, 2023.

EgoM2P: Egocentric Multimodal Multitask Pretraining Stable video diffusion: Scaling latent video diffusion models to large datasets, 2023

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.375059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.375059Z digest=sha256:22703779e270ba9c9a72f433c18732a84f64179e97880202f90dfcb3a1204722

Observation 49ff6499-ee5d-430f-8f77-31886228d558 · outbound

This paper cites Align your latents: High-resolution video synthe- sis with latent diffusion models.

EgoM2P: Egocentric Multimodal Multitask Pretraining Align your latents: High-resolution video synthe- sis with latent diffusion models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.377428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.377428Z digest=sha256:10caeb1ceffee527bd81c7257566aea3f2e85782a3e2cc054323aee0245f5752

Observation 541dd731-2077-4ef5-b5a6-00c8efb30f53 · outbound

This paper cites Scene coordinate reconstruction: Posing of image collections via incremental learning of a relocalizer.

EgoM2P: Egocentric Multimodal Multitask Pretraining Scene coordinate reconstruction: Posing of image collections via incremental learning of a relocalizer

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.379761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.379761Z digest=sha256:ab3e3a4e441a58835da66610e96a5b478e38c2f42d9265a17f944093ed497093

Observation a852ea0f-f485-4732-8c3c-2aa14db15dee · outbound

This paper cites Genie: Generative interactive environments.

EgoM2P: Egocentric Multimodal Multitask Pretraining Genie: Generative interactive environments

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.382148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.382148Z digest=sha256:c4e1cd443b676ad2f50113015074bd0ff8e76335b85f50145aac4577ddddfd8d

Observation 5c1d5162-9dbf-4d9a-b756-670d26569b51 · outbound

This paper cites End-to-end object detection with transformers.

EgoM2P: Egocentric Multimodal Multitask Pretraining End-to-end object detection with transformers

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.384474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.384474Z digest=sha256:17bd83c7744b86ec3eab3c5afe23853ed9365f87e6abb6755ebc49726780a5ec

Observation 3f9630f1-da1a-4b4d-9abd-2f82a4159002 · outbound

This paper cites Emerging properties in self-supervised vision transformers.

EgoM2P: Egocentric Multimodal Multitask Pretraining Emerging properties in self-supervised vision transformers

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.386991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.386991Z digest=sha256:d2ce954d6cdf3e256eeb8a92b49ab2d0606828fb9fefc376841032902754c6e0

Observation 15fb6131-0fbc-44c4-a196-1853f71245d8 · outbound

This paper cites Videocrafter2: Overcoming data limitations for high-quality video diffu- sion models, 2024.

EgoM2P: Egocentric Multimodal Multitask Pretraining Videocrafter2: Overcoming data limitations for high-quality video diffu- sion models, 2024

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.389243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.389243Z digest=sha256:0e3f452d1322c12820a787d303e9b6963277adfce6ea7111058d4ffa65f99bf0

Observation f6e90f27-0f2f-4cdd-98fd-c557c5db6b1b · outbound

This paper cites Fleet, and Geoffrey Hinton.

EgoM2P: Egocentric Multimodal Multitask Pretraining Fleet, and Geoffrey Hinton

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.391946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.391946Z digest=sha256:abff33599f6783805bcef1d2e745c7e98d387eb94eb29fe1f1385d2a9ee818e3

Observation 8afff46f-d8e4-473d-8a61-84d53bd87593 · outbound

This paper cites Control- a-video: Controllable text-to-video diffusion models with motion prior and reward feedback learning, 2024.

EgoM2P: Egocentric Multimodal Multitask Pretraining Control- a-video: Controllable text-to-video diffusion models with motion prior and reward feedback learning, 2024

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.394269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.394269Z digest=sha256:5a10c9d33448fd095255ab9423d97a64e4a1f67c220acacdee71322533d54dd2

Observation 2b910072-2163-4309-89dc-77bd1f85f945 · outbound

This paper cites Scaling egocentric vision: The epic- kitchens dataset.

EgoM2P: Egocentric Multimodal Multitask Pretraining Scaling egocentric vision: The epic- kitchens dataset

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.396601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.396601Z digest=sha256:b5073257b76dc2f56d5510d7f459d81b46843f5fda4c284a21b3772e4eb63668

Observation 00a2618b-7fce-4b42-9bec-9b401e93b5e4 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

EgoM2P: Egocentric Multimodal Multitask Pretraining An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.399010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.399010Z digest=sha256:7ba0ce9637f90c607e6d8007790c6af13610fd4d4e0759ed00fa49ceac4bd3e2

Observation 6d95ac89-690b-435f-af8c-9526af2a8c0b · outbound

This paper cites Structure and Content-Guided Video Synthesis with Diffusion Mod- els.

EgoM2P: Egocentric Multimodal Multitask Pretraining Structure and Content-Guided Video Synthesis with Diffusion Mod- els

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.401827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.401827Z digest=sha256:0a1f4c3aaefdf2c4a6b8c1a84e99b09000d8ae90de7087f47b29f62ec0217352

Observation 34ffb50e-9f66-460b-8ce2-7ad3e4a53694 · outbound

This paper cites Black, and Otmar Hilliges.

EgoM2P: Egocentric Multimodal Multitask Pretraining Black, and Otmar Hilliges

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.404240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.404240Z digest=sha256:c3468f96c3598d7a51a4d2e1c21edfbbaa837fe04a9eb834c2cba623f8dca7c6

Observation 3dfc6897-3adf-4384-893a-18b9fd95211d · outbound

This paper cites HOLD: Category-agnostic 3d reconstruction of interacting hands and objects from video.

EgoM2P: Egocentric Multimodal Multitask Pretraining HOLD: Category-agnostic 3d reconstruction of interacting hands and objects from video

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.406558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.406558Z digest=sha256:747806fd6a22d540104f4ee1fdefb56c0443c1e8c1d14f318fb6af4c56856a5a

Observation 505417fb-eae6-4fa0-b2de-e1c0b6e9f9d2 · outbound

This paper cites VIOLET : End-to-End Video-Language Transformers with Masked Visual-token Modeling.

EgoM2P: Egocentric Multimodal Multitask Pretraining VIOLET : End-to-End Video-Language Transformers with Masked Visual-token Modeling

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.408928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.408928Z digest=sha256:157e331edb96519a40311e29b5b45d82dc4a38e840304ea4ae7995cc78bdb3f9

Observation 55a2aeec-0259-4acf-8ae5-04c17ce47536 · outbound

This paper cites First-person hand action bench- mark with rgb-d videos and 3d hand pose annotations.

EgoM2P: Egocentric Multimodal Multitask Pretraining First-person hand action bench- mark with rgb-d videos and 3d hand pose annotations

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.411701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.411701Z digest=sha256:d641c49712dda3f2029c49d2dad6021a49ba907dfe4e6c0325d5c5132a71dcfe

Observation 5fd4f5f4-7520-4d04-a576-e2df90ba89ff · outbound

This paper cites Imagebind: One embedding space to bind them all.

EgoM2P: Egocentric Multimodal Multitask Pretraining Imagebind: One embedding space to bind them all

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.414302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.414302Z digest=sha256:d0dccbcc077a1a91d73635a714ea1830f6aa3fc6904bf92b873205c94585df7d

Observation 9e5a3b5b-3cb4-4816-a31b-6ad45b96c684 · outbound

This paper cites Accurate, large minibatch sgd: Training imagenet in 1 hour, 2018.

EgoM2P: Egocentric Multimodal Multitask Pretraining Accurate, large minibatch sgd: Training imagenet in 1 hour, 2018

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.416754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.416754Z digest=sha256:af95d947c16126322f134c7bd4820b539eb7fff85480e51870a28c0f5667d9f2

Observation 0801895d-a503-4cb6-a12f-0c36348f3d69 · outbound

This paper cites Ego4d: Around the World in 3,000 Hours of Egocentric Video.

EgoM2P: Egocentric Multimodal Multitask Pretraining Ego4d: Around the World in 3,000 Hours of Egocentric Video

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.419200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.419200Z digest=sha256:be30992b7238420df6117d6638750a148d70f2eefd8ebbd0540dd2af96e6f652

Observation bcf61676-1751-4004-80ab-d9232c3a7592 · outbound

This paper cites Ego-exo4d: Understanding skilled human activity from first- and third-person perspec- tives.

EgoM2P: Egocentric Multimodal Multitask Pretraining Ego-exo4d: Understanding skilled human activity from first- and third-person perspec- tives

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.421469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.421469Z digest=sha256:0ad901b181686b37d009057954468433152f123c0766d6cf4d86e8ab84aa095a

Observation 59fcde70-c33c-4575-9deb-69eab36e6483 · outbound

This paper cites Animatediff: Animate your personalized text- to-image diffusion models without specific tuning.

EgoM2P: Egocentric Multimodal Multitask Pretraining Animatediff: Animate your personalized text- to-image diffusion models without specific tuning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.423750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.423750Z digest=sha256:c1399f0580056a278914abac9c23a6edb5bb0bd49bb16efa8b882fe876c6d4e0

Observation 9af1b819-83cf-4904-80fc-840d4704c4aa · outbound

This paper cites World Models.

EgoM2P: Egocentric Multimodal Multitask Pretraining World Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.426124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.426124Z digest=sha256:7a553a50362377c2301f9bd49b19ca5427c660b4ca6ce3b81a979a75e539417e

Observation 7c1575ea-1ee2-4ed0-8020-9e970c0fd46a · outbound

This paper cites Girshick.

EgoM2P: Egocentric Multimodal Multitask Pretraining Girshick

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.428661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.428661Z digest=sha256:6ef0781ee3c3ec92e4ff4eae2712414653764a655207f58614c04ca49e912f47

Observation cfecc510-d9c4-4fa7-8696-949fad01d9f0 · outbound

This paper cites Classifier-free diffusion guidance.

EgoM2P: Egocentric Multimodal Multitask Pretraining Classifier-free diffusion guidance

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.430965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.430965Z digest=sha256:7ed5d52c1ee9ea8da0a0521d45a1c2da64c5db0c05ac42517745c4f4e42f3a18

Observation c087ee39-de85-40d5-ab00-c371152241bb · outbound

This paper cites Video diffu- sion models.

EgoM2P: Egocentric Multimodal Multitask Pretraining Video diffu- sion models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.433709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.433709Z digest=sha256:278dbd57faa30cc516c405a9c549b6c6bcf364c618823de6ee63ccdf5e517798

Observation 4a4af980-0b14-40ca-abf4-c16d7b4263c4 · outbound

This paper cites The curious case of neural text degeneration.

EgoM2P: Egocentric Multimodal Multitask Pretraining The curious case of neural text degeneration

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.436271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.436271Z digest=sha256:a28cd9a1f0277547f1cb1f86696ce8781c7aa59314abff42c66c3d71988c8ff6

Observation 7c459429-91fc-4ea3-a3cc-f48d486d3cf5 · outbound

This paper cites Cogvideo: Large-scale pretraining for text-to- video generation via transformers.

EgoM2P: Egocentric Multimodal Multitask Pretraining Cogvideo: Large-scale pretraining for text-to- video generation via transformers

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.438690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.438690Z digest=sha256:0a399b7d12b5a9e638dcea12a0059326cd14996a35d43ce5c71ac1e3bd97c29c

Observation 1e6a5851-42d7-4baf-be4c-4db19fd1e816 · outbound

This paper cites Gritsenko, Jasmijn Bastings, Ben Poole, Rianne van den Berg, and Tim Salimans.

EgoM2P: Egocentric Multimodal Multitask Pretraining Gritsenko, Jasmijn Bastings, Ben Poole, Rianne van den Berg, and Tim Salimans

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.441013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.441013Z digest=sha256:bb8b9b7ae964ba1870f95710310791135b0b4fbace08a32d745ccdca40c2bc65

Observation 0ff59b2a-d994-4462-9651-ffcf2ec61f47 · outbound

This paper cites Ross, and Alireza Fathi.

EgoM2P: Egocentric Multimodal Multitask Pretraining Ross, and Alireza Fathi

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.443348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.443348Z digest=sha256:8be466e729e46d70292972fc17e93165a880fa31bab0fcc9e90fed6f40cc0cb6

Observation d63a9227-59ab-40b8-b469-aae67ab99370 · outbound

This paper cites Predicting gaze in egocentric video by learning task- dependent attention transition.

EgoM2P: Egocentric Multimodal Multitask Pretraining Predicting gaze in egocentric video by learning task- dependent attention transition

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.445693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.445693Z digest=sha256:d9638f733930a0c7dd310b9dd2ff4c57e2d456ae6e0841acf9ceb914e228b2a0

Observation 2a3c13d0-fb71-46ce-b967-3ad57c9ea5f2 · outbound

This paper cites GPT-4o System Card.

EgoM2P: Egocentric Multimodal Multitask Pretraining GPT-4o System Card

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.447923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.447923Z digest=sha256:8a748cfa9efcdcf63616d45cef3f06b9814ab03f102bb1ede829dd29974c768c

Observation 400f5c80-0591-4a7f-a8ab-ffd96cbb14b8 · outbound

This paper cites Epic-fusion: Audio-visual temporal binding for egocentric action recognition.

EgoM2P: Egocentric Multimodal Multitask Pretraining Epic-fusion: Audio-visual temporal binding for egocentric action recognition

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.450515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.450515Z digest=sha256:56117e0586dbfb255439f3ba5771fbb9b0cfb1b2de6c8a4d2bfffc091f2c87ee

Observation 04454c8a-dc38-4169-a86b-a8039f528166 · outbound

This paper cites Re- purposing diffusion-based image generators for monocular depth estimation.

EgoM2P: Egocentric Multimodal Multitask Pretraining Re- purposing diffusion-based image generators for monocular depth estimation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.453671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.453671Z digest=sha256:b615b2c4653f918d3dc4549a8e3985e7d3c4d319612ec2d0deb21192a41e6f24

Observation c185bded-22cc-4cc7-bf12-3d5c9ceec635 · outbound

This paper cites Video depth without video models.

EgoM2P: Egocentric Multimodal Multitask Pretraining Video depth without video models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.456344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.456344Z digest=sha256:e60c75dbbf2b5896da1e8e88b0c47051dbcc3ea34ff5d24450a8856218964d0d

Observation 36c03888-f351-40d8-834e-23e0d4fd5506 · outbound

This paper cites Text2Video-Zero: Text- to-Image Diffusion Models are Zero-Shot Video Generators.

EgoM2P: Egocentric Multimodal Multitask Pretraining Text2Video-Zero: Text- to-Image Diffusion Models are Zero-Shot Video Generators

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.458614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.458614Z digest=sha256:74593f0de277a9e5c1edefc812b4b7db80a296989f58c74f421b8efeed49f2bf

Observation 306b8dd8-70c4-4016-8e9a-b90231c1cee4 · outbound

This paper cites Segment anything.

EgoM2P: Egocentric Multimodal Multitask Pretraining Segment anything

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.460971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.460971Z digest=sha256:f56343d7df1a0ffee2ec17f5b99eb3877dcc9082243b6aab45691828c175799b

Observation 85cfafe2-386a-4e0b-a8b3-b6856a206452 · outbound

This paper cites Harmsen, and Neil Houlsby.

EgoM2P: Egocentric Multimodal Multitask Pretraining Harmsen, and Neil Houlsby

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.463435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.463435Z digest=sha256:a2ed23e5a12e913073588d8e1aaca1167246436c476bdbe9756474c29545562a

Observation 56202ac3-0576-4504-b221-2e677d77f744 · outbound

This paper cites VideoPoet: A large language model for zero-shot video generation.

EgoM2P: Egocentric Multimodal Multitask Pretraining VideoPoet: A large language model for zero-shot video generation

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.465709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.465709Z digest=sha256:f8456f9488fd349e886a615926c7f1d4bd418545daa4b125f5bb12d43c6da8d0

Observation 8aaa8fa9-d631-4558-9fcf-3e60bf0ae668 · outbound

This paper cites H2o: Two hands manipulating objects for first person interaction recognition.

EgoM2P: Egocentric Multimodal Multitask Pretraining H2o: Two hands manipulating objects for first person interaction recognition

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.468298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.468298Z digest=sha256:0f32820aac197a641e84fc893b5098397ab56c24152988afde4d2aec14076936

Observation 4caa288a-432f-4f65-affa-3db175476b05 · outbound

This paper cites an unresolved cited work.

EgoM2P: Egocentric Multimodal Multitask Pretraining Unresolved cited work

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.470764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.470764Z digest=sha256:bad46006b05ce9d1ba28ca7cf83f17d43e3f31cb8d09a9e8080a152f3f2fea2d

Observation a82b8a1d-6305-4cfa-93db-c0ffa7b3405a · outbound

This paper cites Lisa: Reasoning segmenta- tion via large language model.

EgoM2P: Egocentric Multimodal Multitask Pretraining Lisa: Reasoning segmenta- tion via large language model

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.473203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.473203Z digest=sha256:fb66d38032bce386bcb2bb6122158e7579768564e6dd6ace3f2cc4107782adbf

Observation 31d3e8ba-650b-4fb4-992f-ca9afb0f6ad1 · outbound

This paper cites Egogen: An egocentric synthetic data generator.

EgoM2P: Egocentric Multimodal Multitask Pretraining Egogen: An egocentric synthetic data generator

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.475810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.475810Z digest=sha256:9dbc32f0137028ca884c51671ce71b4dde4425579eba0aa7b19100b510e72106

Observation f4bd1b8c-c516-46d8-a784-fb266f7e1a4f · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

EgoM2P: Egocentric Multimodal Multitask Pretraining VideoChat: Chat-Centric Video Understanding

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.478451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.478451Z digest=sha256:214eb1f36b2c25cdb454f19ae4ee8370d0cd2e24051ba4b50cf366644658c731

Observation aec46962-59cc-45c6-acbc-8457da029965 · outbound

This paper cites Megasam: Accurate, fast and robust structure and motion from casual dynamic videos.

EgoM2P: Egocentric Multimodal Multitask Pretraining Megasam: Accurate, fast and robust structure and motion from casual dynamic videos

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.481256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.481256Z digest=sha256:e917a80c9e4fed12facfdc00494896638ef9d19766b260d88dec630b5a9d22e8

Observation 45d097af-9a90-49d3-91e0-60c25371e42e · outbound

This paper cites Video-LLaV A: Learning united visual representation by alignment before projection.

EgoM2P: Egocentric Multimodal Multitask Pretraining Video-LLaV A: Learning united visual representation by alignment before projection

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.483820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.483820Z digest=sha256:ecc896a6078ec24dd9a4ffb38a663fea6a9d271ca982d75a5cc6e501b9a12c6a

Observation 3e263881-3e9a-4794-86af-d218551e21ae · outbound

This paper cites Cross-view exocentric to egocentric video synthesis.

EgoM2P: Egocentric Multimodal Multitask Pretraining Cross-view exocentric to egocentric video synthesis

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.486760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.486760Z digest=sha256:4fe681fa8a7a92bee275999adc0397d115ed40f5b584f287c8a4a31744ffbce8

Observation 01a27f75-5b00-4edf-8daa-383a5ad7a7c2 · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024.

EgoM2P: Egocentric Multimodal Multitask Pretraining Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.489037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.489037Z digest=sha256:8b3b0ff649ff5afdc1509e6230b4a6f468bedef7a3822e0f3e097ea63e6ff83d

Observation 6b3c69ab-e05a-4fbb-b93b-29013773e9f4 · outbound

This paper cites Exocentric-to-egocentric video gener- ation.

EgoM2P: Egocentric Multimodal Multitask Pretraining Exocentric-to-egocentric video gener- ation

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.491572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.491572Z digest=sha256:723505f4cfd9dfc1ec5db87b256915c91a22af936b7efee834384ee00ccf9122

Observation 6f7ff5ef-cb26-4e70-a7fc-dc3429d9c4d3 · outbound

This paper cites Li, Ying Shan, and Ge Li.

EgoM2P: Egocentric Multimodal Multitask Pretraining Li, Ying Shan, and Ge Li

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.494202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.494202Z digest=sha256:c634de493637bd83f38e1679174a149599be3d03152207f2435c79d4e3acc8a5

Observation fee22a39-4c72-4ac1-8728-4142f468e909 · outbound

This paper cites Hoi4d: A 4d egocentric dataset for category-level human-object interaction.

EgoM2P: Egocentric Multimodal Multitask Pretraining Hoi4d: A 4d egocentric dataset for category-level human-object interaction

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.499922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.499922Z digest=sha256:688a67f4a4fc6c4685a21e553aa4cfec3695b2495b11d7aaa4b45d0c5b5f79c2

Observation baa0c2cf-52c5-41ae-99fe-26b4bc392390 · outbound

This paper cites an unresolved cited work.

EgoM2P: Egocentric Multimodal Multitask Pretraining Unresolved cited work

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.502461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.502461Z digest=sha256:dd015846170fc8148d8e8d22d8737da3376ea40c438ff749f837e22a2c674028

Observation 9064648b-65c1-4b38-a48f-fdd5d5acbc60 · outbound

This paper cites A convnet for the 2020s.

EgoM2P: Egocentric Multimodal Multitask Pretraining A convnet for the 2020s

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.505098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.505098Z digest=sha256:45134802109b378c525f51409d33556e48f8f83529f82466772de5cd4c06d152

Observation b95d2052-f35f-4041-8469-e7085267b8e4 · outbound

This paper cites Decoupled weight de- cay regularization.

EgoM2P: Egocentric Multimodal Multitask Pretraining Decoupled weight de- cay regularization

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.507569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.507569Z digest=sha256:4b29dad4630ba8e9012a48a6e02987fc39ff3fd881d411e3d750eaab088f2dc3

Observation d35445cd-14d2-407e-a5e3-3c65dc60ab61 · outbound

This paper cites Unified-io 2: Scaling autoregressive mul- timodal models with vision, language, audio, and action.

EgoM2P: Egocentric Multimodal Multitask Pretraining Unified-io 2: Scaling autoregressive mul- timodal models with vision, language, audio, and action

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.509938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.509938Z digest=sha256:b65875e84b5808f3870e536a8b6b5200a9143963610e9c609463eb795126b1f2

Observation 4ddd5a67-06cf-42cf-968e-3ddfee63a802 · outbound

This paper cites UNIFIED-IO: A uni- fied model for vision, language, and multi-modal tasks.

EgoM2P: Egocentric Multimodal Multitask Pretraining UNIFIED-IO: A uni- fied model for vision, language, and multi-modal tasks

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:31:43.345152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:31:42.512343Z digest=sha256:25a46940c13f2e18f4201a6df25330eccce386b1b1ddea4c6b8125714115766d

Observation a87fcc94-61c6-4fb1-a7f3-9790ca147d68 · outbound

This paper cites Align3r: Aligned monocular depth estimation for dynamic videos.

EgoM2P: Egocentric Multimodal Multitask Pretraining Align3r: Aligned monocular depth estimation for dynamic videos

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:31:43.337663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:31:42.515421Z digest=sha256:07482f9205cc4d128781c524f3998bba211596db6b157b6cd77cc8e05c6667c1

Observation e196fdd4-630d-480a-8b97-1abbf6d199b6 · outbound

This paper cites Dream Machine.

EgoM2P: Egocentric Multimodal Multitask Pretraining Dream Machine

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:31:43.329825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:31:42.518003Z digest=sha256:0a52732917e88e6435de472f9bac17d2502bad09345c6fd964f7c5498534cfb1

Observation 85040bb0-72f9-4587-923d-e6ba6321b337 · outbound

This paper cites Videofusion: Decomposed diffusion models for high-quality video generation.

EgoM2P: Egocentric Multimodal Multitask Pretraining Videofusion: Decomposed diffusion models for high-quality video generation

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:31:43.322374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:31:42.520485Z digest=sha256:deeb5f70a2a7a40d5464c23cbb413b020784bbb45c29ec91fcfc4f937f6b0fa3

Observation b49d3759-826d-4058-912b-f96bc2c8be84 · outbound

This paper cites Aria Everyday Activities Dataset.

EgoM2P: Egocentric Multimodal Multitask Pretraining Aria Everyday Activities Dataset

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.522928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.522928Z digest=sha256:42372e3a79100df91a41d354d06e809cc94e96411ceed690b049a8260e9222d9

Observation b6e9ab21-5033-431c-80fe-6db7083e494f · outbound

This paper cites Nymeria: A Massive Collection of Multimodal Egocentric Daily Motion in the Wild.

EgoM2P: Egocentric Multimodal Multitask Pretraining Nymeria: A Massive Collection of Multimodal Egocentric Daily Motion in the Wild

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.525404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.525404Z digest=sha256:c3bc28599cede187f1ca42b03a8f7f280b713d3c8f6ea8f198fd32f4f4ea6b27

Observation 1fd07221-c02f-4ec3-a438-49767fe5dfa2 · outbound

This paper cites Video-chatgpt: Towards detailed video understanding via large vision and language models.

EgoM2P: Egocentric Multimodal Multitask Pretraining Video-chatgpt: Towards detailed video understanding via large vision and language models

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:31:43.315038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:31:42.527939Z digest=sha256:ef87d5d51fba0a0468cc73e5c25ceaaca85e988d0e1a3bd57dce9a212f5c9449

Observation 9977c314-da17-4ddf-b118-483416f1e3f9 · outbound

This paper cites Mm1: Methods, analysis & insights from multimodal llm pre-training, 2024.

EgoM2P: Egocentric Multimodal Multitask Pretraining Mm1: Methods, analysis & insights from multimodal llm pre-training, 2024

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:31:43.306917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:31:42.530377Z digest=sha256:e4b1371025835e6e42071b0557323d29917cfbe450380eaadcbdd51e020e48d9

Observation d9deb7a6-65ed-4c3a-b7db-84c72c36390f · outbound

This paper cites Project Aria Glasses.

EgoM2P: Egocentric Multimodal Multitask Pretraining Project Aria Glasses

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:31:43.299588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:31:42.532866Z digest=sha256:892bdaa8e59d0b026a94242db3bc16eee41eb930c3e5d7bcae6cd99d82e0d40b

Observation 8ba322c7-646c-418d-9622-261ee4a5c97b · outbound

This paper cites Transformers are Sample-Efficient World Models.

EgoM2P: Egocentric Multimodal Multitask Pretraining Transformers are Sample-Efficient World Models

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.535714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.535714Z digest=sha256:f6908831be3c2456086b46a2b581e8ed7f1eed0620138b814f57ea53c31d08e1

Observation 5dc57102-cd1d-463d-b21a-8dcf0ada7d48 · outbound

This paper cites HoloLens 2.

EgoM2P: Egocentric Multimodal Multitask Pretraining HoloLens 2

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:31:43.292402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:31:42.538301Z digest=sha256:b218aa5e34d6c4f0d1d2f3cd0562cc38bf0b1e0c8934094911c4c874cf0dc2b8

Observation c137608e-1b3b-4edf-bd0f-d879c2a3dbf5 · outbound

This paper cites 4M: Massively multimodal masked modeling.

EgoM2P: Egocentric Multimodal Multitask Pretraining 4M: Massively multimodal masked modeling

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:31:43.285124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:31:42.540832Z digest=sha256:95f953fd8338bd5ca76fb4831a4c86bbcc105ff09a3654eb43e143de6c52334a

Observation b4420ba5-ac62-4b89-b5cb-6046902b1220 · outbound

This paper cites AssemblyHands: towards egocentric activity understanding via 3d hand pose esti- mation.

EgoM2P: Egocentric Multimodal Multitask Pretraining AssemblyHands: towards egocentric activity understanding via 3d hand pose esti- mation

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:31:43.277532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:31:42.543229Z digest=sha256:b240a1df82499ec0646980538ba1bb6d7bad706500097c67a012c166226f4ee3

Observation c79c1e4d-4e9a-4db0-aae0-0eff5aeb866a · outbound

This paper cites Video generation models as world simula- tors.

EgoM2P: Egocentric Multimodal Multitask Pretraining Video generation models as world simula- tors

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:31:43.269716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:31:42.545537Z digest=sha256:89d684d1b6c2fc40635030b11536856ff4c454738c9cdfaad9ef7401bc480642

Observation 374a7076-4d4a-4983-9a51-fd3c31220660 · outbound

This paper cites an unresolved cited work.

EgoM2P: Egocentric Multimodal Multitask Pretraining Unresolved cited work

Reference 79

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:31:43.262258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:31:42.548214Z digest=sha256:08dda621c8ab66ecd03d62cbd0600ca854efa0f6f9b1a584d6a66eb1969d6235

Observation 3fb761d3-f991-4650-a924-fac389397dcf · outbound

This paper cites Aria digital twin: A new benchmark dataset for egocentric 3d machine perception.

EgoM2P: Egocentric Multimodal Multitask Pretraining Aria digital twin: A new benchmark dataset for egocentric 3d machine perception

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:31:43.254661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:31:42.550637Z digest=sha256:824c900c1c76b9f65efbdfea0968c2219e4f8dcb16bee0bfed739ccba724ca56

Observation b0563564-1aa4-491d-a10b-6305dffedad5 · outbound

This paper cites Movie Gen: A Cast of Media Foundation Models.

EgoM2P: Egocentric Multimodal Multitask Pretraining Movie Gen: A Cast of Media Foundation Models

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.552979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.552979Z digest=sha256:a064f1275d8dcc5942f830e304cdff17f21edccba39649aaf428c1ffcf559923

Observation 2a104cc6-49f3-486a-a341-1b1e0f6f3a57 · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

EgoM2P: Egocentric Multimodal Multitask Pretraining Learn- ing transferable visual models from natural language super- vision

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:31:43.245929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:31:42.555752Z digest=sha256:d57c9bd2655281e7ac1d5fde5ed1dca343a09351f02cec4da79dc68d9ea83f5a

Observation 43e157ec-730b-452b-8b52-c5821f2d7f52 · outbound

This paper cites an unresolved cited work.

EgoM2P: Egocentric Multimodal Multitask Pretraining Unresolved cited work

Reference 83

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:31:43.237643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:31:42.558132Z digest=sha256:ab4ae5df241616986f6a90e5a8f4df5c67d6280656d3e014ad9edd1a7c699775

Observation 64ae6b91-6d2e-4268-bbab-dff2f7364b2d · outbound

This paper cites High-Resolution Im- age Synthesis with Latent Diffusion Models.

EgoM2P: Egocentric Multimodal Multitask Pretraining High-Resolution Im- age Synthesis with Latent Diffusion Models

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:31:43.229469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:31:42.560589Z digest=sha256:b48d693941951bdc415caa740ff9b4ebf0337b0a2c395300c7bd6b8a031dc538

Observation b3f70a23-1293-4c39-9ff0-f360abc10974 · outbound

This paper cites Gen-3 Alpha.

EgoM2P: Egocentric Multimodal Multitask Pretraining Gen-3 Alpha

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:31:43.221152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:31:42.562991Z digest=sha256:9e56ff45ce0be5bd03745e1bf7dc1ede445dce828911a2c03602a5bbac9233af

Observation 66bddc3a-7e9c-4f03-bd41-b88c485106d3 · outbound

This paper cites Lamar: Bench- marking localization and mapping for augmented reality.

EgoM2P: Egocentric Multimodal Multitask Pretraining Lamar: Bench- marking localization and mapping for augmented reality

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:31:43.212813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:31:42.565470Z digest=sha256:1d27a3a36bc22a7861f5435f7d4195f62cbacd00ccbdbb5599f471b64bdab9bd

Observation 07d1e775-b7d6-4c92-94dc-a1b1a51ed7b1 · outbound

This paper cites Sener, D.

EgoM2P: Egocentric Multimodal Multitask Pretraining Sener, D

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:31:43.203556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:31:42.567826Z digest=sha256:7044a11eca54d99afe5afe20a6164804e510c666c02507d8e01519df13490b5d

Observation ee2d64bc-3bc5-478e-af14-5ff57d27f4ce · outbound

This paper cites Make-a-video: Text-to-video generation without text-video data.

EgoM2P: Egocentric Multimodal Multitask Pretraining Make-a-video: Text-to-video generation without text-video data

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.570245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.570245Z digest=sha256:e245be0154f9de95cfbe8bfbaf8f5c3d97aae1c52bff4f3a5b309f1bbb95deff

Observation 2706b4d7-8665-4f5e-9dc2-06dd67743634 · outbound

This paper cites The Replica Dataset: A Digital Replica of Indoor Spaces.

EgoM2P: Egocentric Multimodal Multitask Pretraining The Replica Dataset: A Digital Replica of Indoor Spaces

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.572653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.572653Z digest=sha256:a0ecb45561c0da17cc96b25786399917cb6cdb9592b7708644c21ffff2aa460b

Observation c84926b5-1768-451c-b056-ae70ede2ec88 · outbound

This paper cites Part, Ioannis Papaioannou, Arash Eshghi, Ioannis Konstas, and Oliver Lemon.

EgoM2P: Egocentric Multimodal Multitask Pretraining Part, Ioannis Papaioannou, Arash Eshghi, Ioannis Konstas, and Oliver Lemon

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:31:43.191261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:31:42.575522Z digest=sha256:394a54aae0cd3c4562031871f4907c2bc5ce93aebcd1031eb9461f80545f87ae

Observation c5e558a0-b2ed-4d3d-bf55-4fb269ab28b3 · outbound

This paper cites Emu: Generative pretraining in multimodality.

EgoM2P: Egocentric Multimodal Multitask Pretraining Emu: Generative pretraining in multimodality

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:31:43.183503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:31:42.577766Z digest=sha256:5032b590ddebf57725f4e7d2a191128d08f7b9e1a1ba2070fa2709e5e0e58180

Observation 9ba77c0d-e496-42e5-a6c2-5147d5e6eb40 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

EgoM2P: Egocentric Multimodal Multitask Pretraining Gemini: A Family of Highly Capable Multimodal Models

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.580076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.580076Z digest=sha256:706876c5f5772fb1097664a35ea0e070b031318a838d1409a2bdc24f25658e94

Observation 522fd6c4-b46a-4fc0-b37e-a072ab425b44 · outbound

This paper cites Kling ai video generator.

EgoM2P: Egocentric Multimodal Multitask Pretraining Kling ai video generator

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:31:43.176365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:31:42.582501Z digest=sha256:04af58e55c262e8938f92dd06206c889fcf512deadd73af4aa5025a08890a79a

Observation e8559f55-245d-4006-a0d8-af30007aedce · outbound

This paper cites DROID-SLAM: Deep vi- sual SLAM for monocular, stereo, and RGB-d cameras.

EgoM2P: Egocentric Multimodal Multitask Pretraining DROID-SLAM: Deep vi- sual SLAM for monocular, stereo, and RGB-d cameras

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:31:43.168267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:31:42.584808Z digest=sha256:71c27ebcfff104d7e53cfdbb17755316ddbcdca7761557e74c40d3058215030e

Observation 2fa6a850-a295-4b97-a6bc-d0ddc6ddd58f · outbound

This paper cites VideoMAE: Masked autoencoders are data-efficient learn- ers for self-supervised video pre-training.

EgoM2P: Egocentric Multimodal Multitask Pretraining VideoMAE: Masked autoencoders are data-efficient learn- ers for self-supervised video pre-training

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:31:43.160061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:31:42.587075Z digest=sha256:2e4d8d024c6359726953b55fd846312e32e0b7ffb5fd5d960c18a700b61e55f0

Observation 2a8dbe85-a2d5-4cb7-a498-9d7c352ce84d · outbound

This paper cites Towards accurate generative models of video: A new metric & challenges, 2019.

EgoM2P: Egocentric Multimodal Multitask Pretraining Towards accurate generative models of video: A new metric & challenges, 2019

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:31:43.151538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:31:42.589235Z digest=sha256:4eb901fd28e4216b11e3c409debda5f810bfea9137c1f3d4b395ffe95518f60b

Observation 1e2f2197-f1ed-4ec3-9ad4-e1979f876097 · outbound

This paper cites Diffusion Models Are Real-Time Game Engines.

EgoM2P: Egocentric Multimodal Multitask Pretraining Diffusion Models Are Real-Time Game Engines

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.591660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.591660Z digest=sha256:a81f411042ae18895b777073eb097d5dab1096bc521b6533bc06c59cf3e93baa

Observation 7e673542-0d8e-4206-bd37-c9d6f93d6215 · outbound

This paper cites Neural discrete representation learn- ing.

EgoM2P: Egocentric Multimodal Multitask Pretraining Neural discrete representation learn- ing

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:31:43.142981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:31:42.594555Z digest=sha256:7fa601808e310f31c49f2ba53ff73d2f22237665b8643c2e3374666269855936

Observation 3fc44bf5-10be-4363-8f68-d9a1423b43b9 · outbound

This paper cites Attention is all you need.

EgoM2P: Egocentric Multimodal Multitask Pretraining Attention is all you need

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:31:43.134665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:31:42.596902Z digest=sha256:04c1e61046c06070938384655681e175695139f1c1e424f33ff388200ae82dcb

Observation f383cfda-5eff-479b-b21c-46daadaed112 · outbound

This paper cites Phenaki: Variable length video generation from open do- main textual descriptions.

EgoM2P: Egocentric Multimodal Multitask Pretraining Phenaki: Variable length video generation from open do- main textual descriptions

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:31:43.126418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:31:42.599141Z digest=sha256:40aca56ef03daad3f07783678e2667bb84d887429308320da4742406fc652a43

Pith citing papers

No inbound Pith citation observations are available.