Pith. sign in

Paper Citation Record · LEDGER

Extending Video Masked Autoencoders to 128 frames

As of 18 August 2026, this Paper Citation Record lists 82 of 82 outbound references and 0 inbound Pith citation observations for arXiv:2411.13683.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.13683 v1

Coverage vector

measured 82 of 82 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T16:22:25.355326Z

measured 82 of 82 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

82 of 82 outbound references displayed

  • verified exact0
  • verified fuzzy71
  • unresolved10
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 84cf1d58-7cf6-4e59-b24a-18365b0499be · outbound

This paper cites Towards long-form video understanding.

Extending Video Masked Autoencoders to 128 frames Towards long-form video understanding

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T16:22:25.014117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:22:25.014117Z digest=sha256:df4b9cc82fdc4a581019a0717f456cc000c94eb7d535897179dcf0b3d87b9db9

Observation 57d954e0-e57e-4b91-834e-9520d8a00794 · outbound

This paper cites Egoschema: A diagnostic benchmark for very long-form video language understanding.

Extending Video Masked Autoencoders to 128 frames Egoschema: A diagnostic benchmark for very long-form video language understanding

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T16:22:25.018939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:22:25.018939Z digest=sha256:e78b7f74fcd78aa4bfd49ad33e99fc1fcf48604401e6f624be7e21db6bd3c463

Observation 1e2090b2-dcc2-4824-91ac-5108d39886e7 · outbound

This paper cites Long movie clip classification with state-space video models.

Extending Video Masked Autoencoders to 128 frames Long movie clip classification with state-space video models

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.442059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T16:22:25.023182Z digest=sha256:f04602b5c181f16fea02d8ecb3da36f45bccc7fdcfd06f962308c398e4cd6c72

Observation 3abf5ab4-b5ec-4dfe-a7e7-e0e7e9f6434f · outbound

This paper cites Selective structured state-spaces for long-form video understanding.

Extending Video Masked Autoencoders to 128 frames Selective structured state-spaces for long-form video understanding

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.428871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T16:22:25.027369Z digest=sha256:f521dd31d1ee47a442559bea178b170b081887c57e45204aa59e5a9f6b3aed6d

Observation e1d2eb4b-a9ef-4679-8eb6-ccfb70df90db · outbound

This paper cites Memory consolidation enables long-context video understanding.

Extending Video Masked Autoencoders to 128 frames Memory consolidation enables long-context video understanding

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.415408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T16:22:25.031425Z digest=sha256:014a003d88790cc4a91a7d5e6f19253bcdfee6e1e85d97d45e06f70ff90298a1

Observation fad0160e-a622-46f0-8cbe-1ba5c085f687 · outbound

This paper cites Memvit: Memory-augmented multiscale vision transformer for efficient long-term video recognition.

Extending Video Masked Autoencoders to 128 frames Memvit: Memory-augmented multiscale vision transformer for efficient long-term video recognition

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.401384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T16:22:25.036030Z digest=sha256:91ab08e7689cbd2acb720edbe1ede0ea6ec933828d06cd7b812be025572f925c

Observation 9136e534-9098-4a59-9d9b-2b43c8578db5 · outbound

This paper cites Token turing machines.

Extending Video Masked Autoencoders to 128 frames Token turing machines

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.386697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T16:22:25.040506Z digest=sha256:882f5402e80d98cf027e74e453344977a5d60bf416e08b73c1f21b14d17607e1

Observation 1ee8d817-2f90-4da5-b5ff-b16c8255302b · outbound

This paper cites Video recap: Recursive captioning of hour-long videos.

Extending Video Masked Autoencoders to 128 frames Video recap: Recursive captioning of hour-long videos

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.374304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T16:22:25.045808Z digest=sha256:251d7036417bf6adfc336b554c6156030a49c057e0011c34b8be14312f5afd49

Observation 15282ec6-ac03-48a8-bb04-1f898e7a3bc4 · outbound

This paper cites A simple recipe for contrastively pre-training video-first encoders beyond 16 frames.

Extending Video Masked Autoencoders to 128 frames A simple recipe for contrastively pre-training video-first encoders beyond 16 frames

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.360794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T16:22:25.050401Z digest=sha256:b4f80a814855982a6e55731da33ed8d726fa2dbbce5fb2b019643048ad531a47

Observation 96c04c42-cb02-4ce3-a8e5-aba75807fbeb · outbound

This paper cites A simple llm framework for long-range video question-answering.

Extending Video Masked Autoencoders to 128 frames A simple llm framework for long-range video question-answering

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T16:22:25.054724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:22:25.054724Z digest=sha256:c93345087037783725f166de49c0ef2a251d4e28c7d83ff670223c770decac52

Observation b28fee82-238e-451f-a888-05e92652f15a · outbound

This paper cites Long-form video- language pre-training with multimodal temporal contrastive learning.

Extending Video Masked Autoencoders to 128 frames Long-form video- language pre-training with multimodal temporal contrastive learning

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.341094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T16:22:25.058702Z digest=sha256:5f8d0458b0a98a1c417b0ceabf4474074f641ce3875328bee47b59f62cfc5099

Observation 4dc4e05a-4e81-45f9-ba36-ef5bdde63ccb · outbound

This paper cites Koala: Key frame-conditioned long video-llm.

Extending Video Masked Autoencoders to 128 frames Koala: Key frame-conditioned long video-llm

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.328284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T16:22:25.062400Z digest=sha256:6942779bdd5957623c0f0b93abe05f84ae6ff7c9330cd3139c5b4ba342638cc7

Observation 9104b8a1-ea81-4e5d-b003-51f7c116eb30 · outbound

This paper cites Masked autoencoders as spatiotemporal learners.

Extending Video Masked Autoencoders to 128 frames Masked autoencoders as spatiotemporal learners

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.316949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T16:22:25.066391Z digest=sha256:2a6c8df0cbccf8b0f855047f35b0fd530654320e56c9ffe73787a98b7eb88d41

Observation 30a10d9c-6a53-456b-9d89-fc6543a31d37 · outbound

This paper cites Videomae v2: Scaling video masked autoencoders with dual masking.

Extending Video Masked Autoencoders to 128 frames Videomae v2: Scaling video masked autoencoders with dual masking

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.304662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T16:22:25.071077Z digest=sha256:889d096fabaa2118927f73fac9508fa7fcf976999b45b436d5cb168feff691e2

Observation 61aa62bd-3850-4120-ade0-88283e9a3003 · outbound

This paper cites VideoMAE: Masked autoencoders are data-efficient learners for self-supervised video pre-training.

Extending Video Masked Autoencoders to 128 frames VideoMAE: Masked autoencoders are data-efficient learners for self-supervised video pre-training

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.290950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T16:22:25.074927Z digest=sha256:bcb1acefef95b95ab977f769459ebdcc0a5997548b9bbd4a45698feb8e21aec0

Observation cdd82006-37e8-4192-a1fe-5725f23d39cc · outbound

This paper cites How can objects help action recognition? In CVPR, 2023.

Extending Video Masked Autoencoders to 128 frames How can objects help action recognition? In CVPR, 2023

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.278784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T16:22:25.078492Z digest=sha256:ae2bd87fd998e69320fe9165793dba0b36f708f3b487f61adfe9f1673475fc91

Observation 2d832d9b-485e-4aca-8091-c506b8885877 · outbound

This paper cites Everest: Efficient masked video autoencoder by removing redundant spatiotemporal tokens.

Extending Video Masked Autoencoders to 128 frames Everest: Efficient masked video autoencoder by removing redundant spatiotemporal tokens

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.266111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T16:22:25.082930Z digest=sha256:5b52acfe81897b18ead4a2adaff70c43705c3fdd06153792e2f03299400046b5

Observation 9d1350a1-3813-4f80-bbf4-6a55f4baba20 · outbound

This paper cites Rescaling egocentric vision: Collection, pipeline and challenges for epic-kitchens-100.

Extending Video Masked Autoencoders to 128 frames Rescaling egocentric vision: Collection, pipeline and challenges for epic-kitchens-100

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.253869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T16:22:25.086594Z digest=sha256:d4e1e24e1a9b85e4a5437a2b0d38fe2bcf20b748f364ea9d0073d74911a1d94d

Observation 6c0d1bea-e94f-46af-ab04-b4c051bd18c1 · outbound

This paper cites Resound: Towards action recognition without representation bias.

Extending Video Masked Autoencoders to 128 frames Resound: Towards action recognition without representation bias

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.241287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T16:22:25.090422Z digest=sha256:1c44fddf0cf84ca8a7dfdfceb6e62180fc0f29e4723fa4f8d44f9cb380e4034b

Observation 21f1223f-011f-4832-ade6-2ee718f0afb4 · outbound

This paper cites an unresolved cited work.

Extending Video Masked Autoencoders to 128 frames Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-12T16:22:26.229238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T16:22:25.094637Z digest=sha256:e2ad632708123833c9b740635c122f77732c3055edf752ea245497e3f28dbb9f

Observation 05497a92-60ff-443e-9f04-71f253a54044 · outbound

This paper cites Bevt: Bert pretraining of video transformers.

Extending Video Masked Autoencoders to 128 frames Bevt: Bert pretraining of video transformers

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.218024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T16:22:25.098703Z digest=sha256:cd6d20b543d40ed467d227f69641cd1c3d1b423c245882e5eabae9440d0727c5

Observation e5e7babb-2187-497a-af9d-76469a1373b0 · outbound

This paper cites Girdhar, A.

Extending Video Masked Autoencoders to 128 frames Girdhar, A

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.205639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T16:22:25.103922Z digest=sha256:8c63ed7a9159c722659b74d31b314f7ce57d760e8ebb0eb66725a58f065fbf0d

Observation 713e89ef-4263-417a-b095-74a19c92e633 · outbound

This paper cites Zero-shot text-to-image generation.

Extending Video Masked Autoencoders to 128 frames Zero-shot text-to-image generation

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.193574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T16:22:25.107960Z digest=sha256:82998d8778fd1202afcfce4d5592683bf29e9dff9aa5293e8b77b4b0aa3777d4

Observation bc7e4fd2-2979-4311-a58b-b08529e687ad · outbound

This paper cites Masked video distillation: Rethinking masked feature modeling for self-supervised video representation learning.

Extending Video Masked Autoencoders to 128 frames Masked video distillation: Rethinking masked feature modeling for self-supervised video representation learning

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.182211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T16:22:25.111942Z digest=sha256:a5ed21d4c019c882b0644ce5a235377cb39211098179608bc0a3826863ac8918

Observation e5511110-75f1-465e-bd26-2f2d846c7286 · outbound

This paper cites Magvit: Masked generative video transformer.

Extending Video Masked Autoencoders to 128 frames Magvit: Masked generative video transformer

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.170158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T16:22:25.115504Z digest=sha256:8fc3a7947fc4a6324783fe5e4c28626f172029fbaf975f5bb33d325a7b38ed20

Observation 929db5b8-04b5-4e3d-a6cc-91d48a12431c · outbound

This paper cites Tokenlearner: What can 8 learned tokens do for images and videos? In NeurIPS, 2021.

Extending Video Masked Autoencoders to 128 frames Tokenlearner: What can 8 learned tokens do for images and videos? In NeurIPS, 2021

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.143450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T16:22:25.123089Z digest=sha256:e54233d18ddd7a1f8eb98f454467c632bdd4b90b2d33c1d90e2b63654e9093ce

Observation 9cd15a53-f043-4a68-9079-f5e6871c87d9 · outbound

This paper cites something something.

Extending Video Masked Autoencoders to 128 frames something something

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.130927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T16:22:25.126787Z digest=sha256:643d09be1e5ba8ad9e32ca0e08b1c7e2fe8820198f8fc09a32b6a02b0f6478b5

Observation 319a40db-cab6-42ab-86cd-b86c83095958 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

Extending Video Masked Autoencoders to 128 frames An image is worth 16x16 words: Transformers for image recognition at scale

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.118986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T16:22:25.130449Z digest=sha256:f1e75592e193a95e602c6c7f552bfb14c2b022e75d86f3d2d0f8cae0bb860885

Observation d87ef826-a40a-4764-958d-63623af7b0cc · outbound

This paper cites Mgmae: Motion guided masking for video masked autoencoding.

Extending Video Masked Autoencoders to 128 frames Mgmae: Motion guided masking for video masked autoencoding

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.105953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T16:22:25.134329Z digest=sha256:f1db7f2e969c76abca24efdddb24eb28796e08f9c4436c133ddb7e81f4ff3836

Observation b5b0e6f3-fc2d-408f-a232-3ca3c19afbf2 · outbound

This paper cites Motion-guided masking for spatiotemporal representation learning.

Extending Video Masked Autoencoders to 128 frames Motion-guided masking for spatiotemporal representation learning

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.091040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T16:22:25.138258Z digest=sha256:f041e7745a9bd2e62ef2d6019ec465e129386aff5aba785c511daf3b343d9516

Observation e32096bb-4f64-4512-ac99-52032f45532d · outbound

This paper cites Video codec design: developing image and video compression systems.

Extending Video Masked Autoencoders to 128 frames Video codec design: developing image and video compression systems

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.077412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T16:22:25.142170Z digest=sha256:84cf1fc5aee97dfd558f8ea7d69e54f76502a835490276e4b2c8f51e51bbaecb

Observation 9a8ae6a4-ee82-4266-bced-f94d9ea68a02 · outbound

This paper cites Raft: Recurrent all-pairs field transforms for optical flow.

Extending Video Masked Autoencoders to 128 frames Raft: Recurrent all-pairs field transforms for optical flow

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.062271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T16:22:25.146065Z digest=sha256:bd5c12ad52f9746faff3cf8a06d5517579dee46ed4ed5f4ac407b9ad67ed5c70

Observation d71d0fc2-7964-48d3-9875-5cb23baf2090 · outbound

This paper cites Videoprism: A foundational visual encoder for video understanding.

Extending Video Masked Autoencoders to 128 frames Videoprism: A foundational visual encoder for video understanding

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.050812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T16:22:25.149577Z digest=sha256:de8a0481cd59cb5be9283ade25c59293bede3f6808fb86aad31081503f911e43

Observation 0c81c503-ca16-426b-8f61-e24d49b7afd8 · outbound

This paper cites Internvideo2: Scaling video foundation models for multimodal video understanding.

Extending Video Masked Autoencoders to 128 frames Internvideo2: Scaling video foundation models for multimodal video understanding

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.037738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T16:22:25.153989Z digest=sha256:5b78f9534354dbcd3b6d08875ec27b98648fe4077e56768525bbee6224c20c3a

Observation 55a81fcf-1c28-4640-b3d8-ce4c2b551f17 · outbound

This paper cites Vivit: A video vision transformer.

Extending Video Masked Autoencoders to 128 frames Vivit: A video vision transformer

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.024691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T16:22:25.157567Z digest=sha256:2ba4e12a4ea90afdbd76721f4308581c9f4b3c2771b9ad3c17abad31d3502c84

Observation 07e3e8ce-9ff2-4a61-9d79-e761d97b7bfb · outbound

This paper cites Finite scalar quantization: VQ-V AE made simple.

Extending Video Masked Autoencoders to 128 frames Finite scalar quantization: VQ-V AE made simple

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.011181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T16:22:25.161177Z digest=sha256:1097d84e08df406d670844bce59e741ef7678a11eec4eea2f1615f0705849cde

Observation 02398460-7a6d-40c7-bf63-b314a15cce3f · outbound

This paper cites A Short Note about Kinetics-600.

Extending Video Masked Autoencoders to 128 frames A Short Note about Kinetics-600

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T16:22:25.164843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:22:25.164843Z digest=sha256:8d0596e2750763c33dac54ec8bac5c093ff6f5d7ac36114d57cec98930942cd0

Observation 95ecca2f-f416-4873-9e90-80143ff10102 · outbound

This paper cites A Short Note on the Kinetics-700 Human Action Dataset.

Extending Video Masked Autoencoders to 128 frames A Short Note on the Kinetics-700 Human Action Dataset

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T16:22:25.168857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:22:25.168857Z digest=sha256:33914c7360e771e188f7f9ac1678715408da5156fef4d99ccbe4052aff444ac1

Observation a196cfee-b1a4-49b0-b4a2-0d09a82b791d · outbound

This paper cites Multiview transformers for video recognition.

Extending Video Masked Autoencoders to 128 frames Multiview transformers for video recognition

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.997209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T16:22:25.172795Z digest=sha256:ed1d1d8d31516385113ff9d814410e1ab1c0fe4b4a7a98221380127637256e47

Observation 33cfc5de-95ae-40d3-8b38-2227648e89d7 · outbound

This paper cites Temporally-Adaptive Models for Efficient Video Understanding.

Extending Video Masked Autoencoders to 128 frames Temporally-Adaptive Models for Efficient Video Understanding

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T16:22:25.176362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:22:25.176362Z digest=sha256:965a3eece1b4540e9d227fd6dc24d4452f8d8b623678b34b0ca3b4cc25c67960

Observation 2f028e33-ae02-4d0b-8d47-ea889ce8eda2 · outbound

This paper cites Training a Large Video Model on a Single Machine in a Day.

Extending Video Masked Autoencoders to 128 frames Training a Large Video Model on a Single Machine in a Day

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T16:22:25.180917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:22:25.180917Z digest=sha256:7319eb5e7b82053fd2dab378c29e480c34fee95ef6329b23ca7a4235e6b06596

Observation 72d5d421-0073-40e7-8db3-f1349385d633 · outbound

This paper cites Imagenet-21k pretraining for the masses.

Extending Video Masked Autoencoders to 128 frames Imagenet-21k pretraining for the masses

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.983302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T16:22:25.185019Z digest=sha256:7656ebe28cfa03c65ca00a11989e740a1ddd77fe705a3650b3a405f706219d19

Observation 45a79a2a-274c-496e-b7cf-87f33caffe1d · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video.

Extending Video Masked Autoencoders to 128 frames Ego4d: Around the world in 3,000 hours of egocentric video

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.972248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T16:22:25.188571Z digest=sha256:ca1496cda9a02c255a2ae59dbec29a60939420948195dce067f0b739ccfbf8fb

Observation bdb7e0cb-169d-4786-bfcf-e6475c709066 · outbound

This paper cites Verbs in action: Improving verb understanding in video-language models.

Extending Video Masked Autoencoders to 128 frames Verbs in action: Improving verb understanding in video-language models

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.961247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T16:22:25.192158Z digest=sha256:b3ca3f88a00bfc3f15c45b8a22f737e1dc832f5e2d7b8eb2fa48188f9b8f8588

Observation ec83fec6-7af5-4875-aefd-eb8f725487be · outbound

This paper cites Slowfast networks for video recognition.

Extending Video Masked Autoencoders to 128 frames Slowfast networks for video recognition

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.949510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T16:22:25.196373Z digest=sha256:b79c77d33cafe0668108c8034f1b8369098af1af55cf231a3f68e4b049a92d4a

Observation 24b0f31e-184e-43ca-9067-4636c6be323f · outbound

This paper cites Interactive prototype learning for egocentric action recognition.

Extending Video Masked Autoencoders to 128 frames Interactive prototype learning for egocentric action recognition

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.939197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T16:22:25.200298Z digest=sha256:a8731c6cdef3b152d18723463350d3831b2186e6c7340ca8d3f8981ece4b5f45

Observation 0be7e6c2-5f50-4e9f-aa24-cd85927732bb · outbound

This paper cites Movinets: Mobile video networks for efficient video recognition.

Extending Video Masked Autoencoders to 128 frames Movinets: Mobile video networks for efficient video recognition

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.928022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T16:22:25.203795Z digest=sha256:6682029f6b62b7c0b54f0efb40ba66bbf0e9a747c1dc4b75865bccb87fe7aa9f

Observation a08bd723-c388-4787-9478-9e4cdde59e2d · outbound

This paper cites Omnivore: A single model for many visual modalities.

Extending Video Masked Autoencoders to 128 frames Omnivore: A single model for many visual modalities

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.916079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T16:22:25.207333Z digest=sha256:c9a6b2f6b5b71531c1c92f35adcccfa2c1341c80f67476b2f92a2b30bc73bd1b

Observation 08528d82-d877-41a5-ba58-e4c47280dda9 · outbound

This paper cites Learning video representations from large language models.

Extending Video Masked Autoencoders to 128 frames Learning video representations from large language models

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.903493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T16:22:25.211110Z digest=sha256:4f9a03ed7d14ff4e2897ca0115c0cf440fb9bc6e666405a1d4e990955b09fe49

Observation 994e1ecd-afde-4267-9eff-b2a21a9e9a74 · outbound

This paper cites Is space-time attention all you need for video understanding? In ICML, 2021.

Extending Video Masked Autoencoders to 128 frames Is space-time attention all you need for video understanding? In ICML, 2021

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.891843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T16:22:25.214507Z digest=sha256:892ce4d8700aea51ff2654948fe365d44971b62a9df8b8d3b5f50e77d4df68c5

Observation 62167356-5e55-4fb7-8620-e87d6e8f161f · outbound

This paper cites Video swin transformer.

Extending Video Masked Autoencoders to 128 frames Video swin transformer

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.879718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T16:22:25.218269Z digest=sha256:c78f2e300e92c71ba0e0cdef17afa6fb418038046d336e1bea8e24c42d00a9ff

Observation be25e975-2432-4536-8ce7-8d7427f0d842 · outbound

This paper cites Can an image classifier suffice for action recognition? In ICLR, 2022.

Extending Video Masked Autoencoders to 128 frames Can an image classifier suffice for action recognition? In ICLR, 2022

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.868902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T16:22:25.221993Z digest=sha256:7f253ccb0d96dea36cff8ddefe5118550f10b484dadb8e15de2be6e3925c7caa

Observation 5efa996f-7043-42bb-bb20-13945d3dc966 · outbound

This paper cites Object-region video transformers.

Extending Video Masked Autoencoders to 128 frames Object-region video transformers

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.857839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T16:22:25.225841Z digest=sha256:45ca3e5a58376d22d8d3d8270c06efaefeb808cd34df34f66700904cb0a344ff

Observation 738d1098-9f5d-4766-8627-53e4d24a5d64 · outbound

This paper cites Aim: Adapting image models for efficient video action recognition.

Extending Video Masked Autoencoders to 128 frames Aim: Adapting image models for efficient video action recognition

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.845804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T16:22:25.229883Z digest=sha256:689052ff9873277e8a49bbb975b27b9bece52fad1a907c71d9f8830e02e4b7d2

Observation 26189b8b-33f3-4750-9990-c658006ca20a · outbound

This paper cites Video-focalnets: Spatio-temporal focal modulation for video action recognition.

Extending Video Masked Autoencoders to 128 frames Video-focalnets: Spatio-temporal focal modulation for video action recognition

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.834222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T16:22:25.233575Z digest=sha256:96f44391821cd825df6346001dbe95b8c2f7444a21b4c2ebab9a98301463ff9a

Observation 6d0e83f9-ab6a-4b0e-910c-dc8b000828e3 · outbound

This paper cites Language model beats diffusion–tokenizer is key to visual generation.

Extending Video Masked Autoencoders to 128 frames Language model beats diffusion–tokenizer is key to visual generation

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.819649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T16:22:25.237643Z digest=sha256:d24113a2f193c7a77f86c980b22b22dff3b76a7df9a3fb1803424641df15b8c6

Observation 35bf2940-3183-4085-8371-02c89b5ee713 · outbound

This paper cites Finegym: A hierarchical video dataset for fine-grained action understanding.

Extending Video Masked Autoencoders to 128 frames Finegym: A hierarchical video dataset for fine-grained action understanding

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.804450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T16:22:25.241462Z digest=sha256:135e79a5141059f8d16a3d1c057c10fa7c5e75067d2a1350c454dd3ab9f57347

Observation c0e06bc4-9991-4702-8521-b65e8aa77aa3 · outbound

This paper cites Learning temporal cues for fine-grained action recognition.

Extending Video Masked Autoencoders to 128 frames Learning temporal cues for fine-grained action recognition

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.790238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T16:22:25.245443Z digest=sha256:bf4aca5cb83dab7210b7fabfa3762be293d9de387f36159e1cc0b5500164725a

Observation b2aa5f83-862d-4df6-b1f7-ebe7564d0c09 · outbound

This paper cites Tsm: Temporal shift module for efficient video understanding.

Extending Video Masked Autoencoders to 128 frames Tsm: Temporal shift module for efficient video understanding

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.775928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T16:22:25.249266Z digest=sha256:8ee7d0fb2f5630fe402f0bdaa6173e9c56844500909c58c3ebced4df891c300f

Observation 33ddac3d-e7d2-4c61-85b0-c7df2c2651a0 · outbound

This paper cites Temporal query networks for fine-grained video understanding.

Extending Video Masked Autoencoders to 128 frames Temporal query networks for fine-grained video understanding

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.762566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T16:22:25.253411Z digest=sha256:6e4fc875b1275d1616aec0197b1fd3e65ac8973d99b19c95159fc33a7bc8215f

Observation d5c37d89-8b6f-4e04-b893-76292cb9fbaf · outbound

This paper cites Combined cnn transformer encoder for enhanced fine-grained human action recognition.

Extending Video Masked Autoencoders to 128 frames Combined cnn transformer encoder for enhanced fine-grained human action recognition

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.751023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T16:22:25.257990Z digest=sha256:a7f980feb064ded5f4ea029fa3d38a5bb145999c17d93a8c8c71a44471735d0e

Observation c1217cf0-e260-4354-a23a-649cf1548550 · outbound

This paper cites Going deeper with image transformers.

Extending Video Masked Autoencoders to 128 frames Going deeper with image transformers

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.738041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T16:22:25.262106Z digest=sha256:9867903583468b18c63ab9c07a6136dd9712f1c7aef0966c028fd8263c8dc8d2

Observation e4c247d4-831e-46c6-a25f-64cfd87e3220 · outbound

This paper cites Scenic: A jax library for computer vision research and beyond.

Extending Video Masked Autoencoders to 128 frames Scenic: A jax library for computer vision research and beyond

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.722729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T16:22:25.266075Z digest=sha256:c7c524dc6721b88dd83a9112436bc75d49974e50918dcd7ad3017e312d396f38

Observation 0a74a736-5133-4c9c-a196-f891b24fb370 · outbound

This paper cites We found this slightly improved accuracy ( +0.5 points on EPIC-Kitchens-100 Verbs and +1.2 points on Diving48).

Extending Video Masked Autoencoders to 128 frames We found this slightly improved accuracy ( +0.5 points on EPIC-Kitchens-100 Verbs and +1.2 points on Diving48)

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.700975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T16:22:25.270584Z digest=sha256:cd2d77b30b2f772408286bbb2d336ebb7d6fbb92df18389eed302bf5798ffc05

Observation e8883ef5-e7c8-4e61-a54c-8143413ce151 · outbound

This paper cites an unresolved cited work.

Extending Video Masked Autoencoders to 128 frames Unresolved cited work

Reference 66

Resolution
unresolved
raw_fallback, observed 2026-08-12T16:22:25.681314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T16:22:25.274933Z digest=sha256:307beb7006e15050470639436839812172962204333a1ae84680c4a11c5e07e1

Observation 7121ac07-0f58-47f2-ad01-88e579276083 · outbound

This paper cites 17 Table 9: Model size vs frames.

Extending Video Masked Autoencoders to 128 frames 17 Table 9: Model size vs frames

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.667471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T16:22:25.279010Z digest=sha256:261dd85aade6f352d37fd68bf807e9a9333cb72c41d7f3636e1ca6ecd84d0a9d

Observation 13270ed7-aa41-42bd-a1da-25f5ef192c76 · outbound

This paper cites Guidelines: • The answer NA means that the abstract and introduction do not include the claims made in the paper.

Extending Video Masked Autoencoders to 128 frames Guidelines: • The answer NA means that the abstract and introduction do not include the claims made in the paper

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.653781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T16:22:25.283416Z digest=sha256:615880fdaeee4fd2f9568ddb627c7eac96682f31966b2057b0e0ab1997173132

Observation fd7f5d79-0999-4578-ba91-897e965c755c · outbound

This paper cites Limitations.

Extending Video Masked Autoencoders to 128 frames Limitations

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.636688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T16:22:25.287958Z digest=sha256:ea7b46ee99f5c61d96f0ce9e504158c61dd48f2473a31e32867d0f9f65c1b07c

Observation 10456fc6-93cc-44fe-982e-3888d20813f1 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include theoretical results.

Extending Video Masked Autoencoders to 128 frames Guidelines: • The answer NA means that the paper does not include theoretical results

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.619738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T16:22:25.292611Z digest=sha256:8b86b19e6348c90f3690087c61164b465f0922cf43a8c85e11f2ad7dfbee43a8

Observation 955057ce-da5a-4eab-b8a1-3bb1e743a1f6 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include experiments.

Extending Video Masked Autoencoders to 128 frames Guidelines: • The answer NA means that the paper does not include experiments

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.605281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T16:22:25.296935Z digest=sha256:7342d73a44f3d33f5e6564c368b47af67e0264bc8ff3879f976c288e2671f4ef

Observation 05b446cf-6eb3-4330-affd-25d7dd01dfe0 · outbound

This paper cites Guidelines: • The answer NA means that paper does not include experiments requiring code.

Extending Video Masked Autoencoders to 128 frames Guidelines: • The answer NA means that paper does not include experiments requiring code

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.588626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T16:22:25.302631Z digest=sha256:70b33d94e6fa09346eb9f6f048b62cda28c855116c58cd82db33936cd0aebef3

Observation 903717b9-95c5-4a0b-ab71-e54f72805d1b · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include experiments.

Extending Video Masked Autoencoders to 128 frames Guidelines: • The answer NA means that the paper does not include experiments

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.572605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T16:22:25.307592Z digest=sha256:578b750397458588f774dd780b010875fd4c20fbb3b833d312b9230eeb784e4c

Observation 7504251a-762c-41d2-ad08-4e797b83ce0a · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include experiments.

Extending Video Masked Autoencoders to 128 frames Guidelines: • The answer NA means that the paper does not include experiments

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.559349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T16:22:25.312618Z digest=sha256:f7809149cb8eb7e293a93ad7684d6a6c1ec59333b888640a446129a5c51ea04c

Observation e409290e-ad8e-4ba9-95c5-8ff826b82976 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include experiments.

Extending Video Masked Autoencoders to 128 frames Guidelines: • The answer NA means that the paper does not include experiments

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.543694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T16:22:25.318955Z digest=sha256:f5374359d439941c39837aee62fbf6d6557f1f1f513dc83afa2f49efa9316a92

Observation 151092d4-fe23-4c42-8e99-94dce5af1702 · outbound

This paper cites Guidelines: • The answer NA means that the authors have not reviewed the NeurIPS Code of Ethics.

Extending Video Masked Autoencoders to 128 frames Guidelines: • The answer NA means that the authors have not reviewed the NeurIPS Code of Ethics

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.528011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T16:22:25.323122Z digest=sha256:5152f1900a5508ee06290c975e065ed8541e5851ad77f8e7a0b12ae45aad7ce1

Observation 27550e91-dc95-4ad8-834c-1d3ca347c615 · outbound

This paper cites • If the authors answer NA or No, they should explain why their work has no societal impact or why the paper does not address societal impact.

Extending Video Masked Autoencoders to 128 frames • If the authors answer NA or No, they should explain why their work has no societal impact or why the paper does not address societal impact

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.510829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T16:22:25.327717Z digest=sha256:a31258f6eb1cb8a19c460a901671c25e9d75b7ed556eddf548ad1822818f8381

Observation 1832e7fb-510a-4c78-80da-e7490525eabf · outbound

This paper cites Guidelines: • The answer NA means that the paper poses no such risks.

Extending Video Masked Autoencoders to 128 frames Guidelines: • The answer NA means that the paper poses no such risks

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.497541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T16:22:25.332877Z digest=sha256:cefd4f769bb58aa86ed616e85450951c90eef5b8f844df87fc1e9c8c1ac81e12

Observation e87e9080-6d21-4d13-8e92-c9dc39889c6e · outbound

This paper cites Guidelines: • The answer NA means that the paper does not use existing assets.

Extending Video Masked Autoencoders to 128 frames Guidelines: • The answer NA means that the paper does not use existing assets

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.484090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T16:22:25.343151Z digest=sha256:b0477248fef03facd4729f44a7552c60f9ebb42e30daaab991bf561fad7e56ed

Observation 0f1ef82c-5084-4726-b5f5-3e7278d9683a · outbound

This paper cites Guidelines: • The answer NA means that the paper does not release new assets.

Extending Video Masked Autoencoders to 128 frames Guidelines: • The answer NA means that the paper does not release new assets

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-12T16:22:25.347041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:22:25.347041Z digest=sha256:c04eaea5ca6aaf2445b9fb2fc173c82e1998fc904f0c3c6aeb250bb511404f57

Observation 32ea79b1-a7bb-4d37-9d42-eee4bc252c05 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects.

Extending Video Masked Autoencoders to 128 frames Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.464389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T16:22:25.351339Z digest=sha256:60006a75751ee4643fb92705a1821362312fcaa249d3ff29d90de88a7e9d0d1b

Observation 1f04bade-1819-48ac-ab45-49e80ce8e8d8 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects.

Extending Video Masked Autoencoders to 128 frames Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.450458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T16:22:25.355326Z digest=sha256:ec5eaab61a11ed4b53a468e78d0c06012c47cb41005bf7fa223cfe1a1777336e

Observation 17e2c672-1e70-46a4-a07b-68f28e108efe · outbound

This paper cites an unresolved cited work.

Extending Video Masked Autoencoders to 128 frames Unresolved cited work

Reference 2023

Resolution
parse uncertain
raw_fallback, observed 2026-08-12T16:22:26.154889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T16:22:25.119455Z digest=sha256:d61a7d19187f5a866024b79548ee828fd4ab2dcae36730fbf48dceaf6cf532b0

Pith citing papers

No inbound Pith citation observations are available.