Pith. sign in

Paper Citation Record · LEDGER

Extending Video Masked Autoencoders to 128 frames

As of 18 August 2026, this Paper Citation Record lists 82 of 82 outbound references and 0 inbound Pith citation observations for arXiv:2411.13683.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.13683 v1

Coverage vector

measured 82 of 82 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T16:22:25.355326Z

measured 82 of 82 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

82 of 82 outbound references displayed

  • verified exact0
  • verified fuzzy71
  • unresolved10
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 84cf1d58-7cf6-4e59-b24a-18365b0499be · outbound

This paper cites Towards long-form video understanding.

Extending Video Masked Autoencoders to 128 frames Towards long-form video understanding

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T16:22:25.014117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:22:25.014117Z digest=sha256:df4b9cc82fdc4a581019a0717f456cc000c94eb7d535897179dcf0b3d87b9db9

Observation 57d954e0-e57e-4b91-834e-9520d8a00794 · outbound

This paper cites Egoschema: A diagnostic benchmark for very long-form video language understanding.

Extending Video Masked Autoencoders to 128 frames Egoschema: A diagnostic benchmark for very long-form video language understanding

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T16:22:25.018939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:22:25.018939Z digest=sha256:e78b7f74fcd78aa4bfd49ad33e99fc1fcf48604401e6f624be7e21db6bd3c463

Observation 1e2090b2-dcc2-4824-91ac-5108d39886e7 · outbound

This paper cites Long movie clip classification with state-space video models.

Extending Video Masked Autoencoders to 128 frames Long movie clip classification with state-space video models

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.442059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T16:22:25.023182Z digest=sha256:9bd6e0bf9655f1800ab2add05e699e02612733f84f5eaf73928cd09920cad11a

Observation 3abf5ab4-b5ec-4dfe-a7e7-e0e7e9f6434f · outbound

This paper cites Selective structured state-spaces for long-form video understanding.

Extending Video Masked Autoencoders to 128 frames Selective structured state-spaces for long-form video understanding

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.428871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T16:22:25.027369Z digest=sha256:8a8cee01337f8229d1b8482b08c95fb4cd98535695bdb56e9f1ce5d869f9d370

Observation e1d2eb4b-a9ef-4679-8eb6-ccfb70df90db · outbound

This paper cites Memory consolidation enables long-context video understanding.

Extending Video Masked Autoencoders to 128 frames Memory consolidation enables long-context video understanding

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.415408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T16:22:25.031425Z digest=sha256:040935741369d7df1aeab92be838f1ec6a1940c4047c1d215a12aba8ff2c6461

Observation fad0160e-a622-46f0-8cbe-1ba5c085f687 · outbound

This paper cites Memvit: Memory-augmented multiscale vision transformer for efficient long-term video recognition.

Extending Video Masked Autoencoders to 128 frames Memvit: Memory-augmented multiscale vision transformer for efficient long-term video recognition

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.401384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T16:22:25.036030Z digest=sha256:f8eb2d77357b8eef2231143121744b671e4f056f32f5cd73a71ea291fe98bd58

Observation 9136e534-9098-4a59-9d9b-2b43c8578db5 · outbound

This paper cites Token turing machines.

Extending Video Masked Autoencoders to 128 frames Token turing machines

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.386697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T16:22:25.040506Z digest=sha256:cc50ad11e1ff65e739879125715e78a754df740b30b56d2c96b8ba3845b2235d

Observation 1ee8d817-2f90-4da5-b5ff-b16c8255302b · outbound

This paper cites Video recap: Recursive captioning of hour-long videos.

Extending Video Masked Autoencoders to 128 frames Video recap: Recursive captioning of hour-long videos

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.374304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T16:22:25.045808Z digest=sha256:a507cdd1c886a5c36423be2674b1a0675bb13c2fbe9008add85a7e652308947f

Observation 15282ec6-ac03-48a8-bb04-1f898e7a3bc4 · outbound

This paper cites A simple recipe for contrastively pre-training video-first encoders beyond 16 frames.

Extending Video Masked Autoencoders to 128 frames A simple recipe for contrastively pre-training video-first encoders beyond 16 frames

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.360794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T16:22:25.050401Z digest=sha256:7301f8495eade25e1b6e40cbae8b50fd962a414e601e3ed6bdb2f57909f3e2a8

Observation 96c04c42-cb02-4ce3-a8e5-aba75807fbeb · outbound

This paper cites A simple llm framework for long-range video question-answering.

Extending Video Masked Autoencoders to 128 frames A simple llm framework for long-range video question-answering

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T16:22:25.054724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:22:25.054724Z digest=sha256:c93345087037783725f166de49c0ef2a251d4e28c7d83ff670223c770decac52

Observation b28fee82-238e-451f-a888-05e92652f15a · outbound

This paper cites Long-form video- language pre-training with multimodal temporal contrastive learning.

Extending Video Masked Autoencoders to 128 frames Long-form video- language pre-training with multimodal temporal contrastive learning

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.341094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T16:22:25.058702Z digest=sha256:182bd5be0841787597b069a5e816ce30c00df30be58fdf21e66bcb6d47e0333c

Observation 4dc4e05a-4e81-45f9-ba36-ef5bdde63ccb · outbound

This paper cites Koala: Key frame-conditioned long video-llm.

Extending Video Masked Autoencoders to 128 frames Koala: Key frame-conditioned long video-llm

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.328284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T16:22:25.062400Z digest=sha256:962d76cc1ae2895ee39730d43001261d4d90ce16dd5b011f6a72bce50569d72d

Observation 9104b8a1-ea81-4e5d-b003-51f7c116eb30 · outbound

This paper cites Masked autoencoders as spatiotemporal learners.

Extending Video Masked Autoencoders to 128 frames Masked autoencoders as spatiotemporal learners

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.316949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T16:22:25.066391Z digest=sha256:8d99f0b8513d92d9e7c6cde2b074ac33ce1ee5443d8e6b19f7ea572262f637cc

Observation 30a10d9c-6a53-456b-9d89-fc6543a31d37 · outbound

This paper cites Videomae v2: Scaling video masked autoencoders with dual masking.

Extending Video Masked Autoencoders to 128 frames Videomae v2: Scaling video masked autoencoders with dual masking

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.304662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T16:22:25.071077Z digest=sha256:529d166a760ac3f27e1706d14dcbb32b88eff296d4b473644fd2637109133aa4

Observation 61aa62bd-3850-4120-ade0-88283e9a3003 · outbound

This paper cites VideoMAE: Masked autoencoders are data-efficient learners for self-supervised video pre-training.

Extending Video Masked Autoencoders to 128 frames VideoMAE: Masked autoencoders are data-efficient learners for self-supervised video pre-training

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.290950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T16:22:25.074927Z digest=sha256:8c45c8788154549c8893003ae8295a0d310dd3b3048abafcc7b8ad684adfb96f

Observation cdd82006-37e8-4192-a1fe-5725f23d39cc · outbound

This paper cites How can objects help action recognition? In CVPR, 2023.

Extending Video Masked Autoencoders to 128 frames How can objects help action recognition? In CVPR, 2023

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.278784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T16:22:25.078492Z digest=sha256:e4e3c0bdb8f277c3ac5a84feccdd10926bc5777bb7e383ee3e09dd06f38c7f1c

Observation 2d832d9b-485e-4aca-8091-c506b8885877 · outbound

This paper cites Everest: Efficient masked video autoencoder by removing redundant spatiotemporal tokens.

Extending Video Masked Autoencoders to 128 frames Everest: Efficient masked video autoencoder by removing redundant spatiotemporal tokens

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.266111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T16:22:25.082930Z digest=sha256:b3b9654973a4aa5715fd69afba48893c49744473b86ada55f3cb1c0d62ffbe4d

Observation 9d1350a1-3813-4f80-bbf4-6a55f4baba20 · outbound

This paper cites Rescaling egocentric vision: Collection, pipeline and challenges for epic-kitchens-100.

Extending Video Masked Autoencoders to 128 frames Rescaling egocentric vision: Collection, pipeline and challenges for epic-kitchens-100

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.253869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T16:22:25.086594Z digest=sha256:46879adcbceace0e7125707363fa0a09c9b7667ada9006f2792bc80b3c7ba5c1

Observation 6c0d1bea-e94f-46af-ab04-b4c051bd18c1 · outbound

This paper cites Resound: Towards action recognition without representation bias.

Extending Video Masked Autoencoders to 128 frames Resound: Towards action recognition without representation bias

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.241287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T16:22:25.090422Z digest=sha256:aa7260c7322b74b89b18f6b6fa3945f3828b703fbd0e959c0d3d455cf68fdec4

Observation 21f1223f-011f-4832-ade6-2ee718f0afb4 · outbound

This paper cites an unresolved cited work.

Extending Video Masked Autoencoders to 128 frames Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-12T16:22:26.229238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T16:22:25.094637Z digest=sha256:e01f4ec2420474f1c3a8996f4dd430a19170bb512308d50be1fc3d33169125ec

Observation 05497a92-60ff-443e-9f04-71f253a54044 · outbound

This paper cites Bevt: Bert pretraining of video transformers.

Extending Video Masked Autoencoders to 128 frames Bevt: Bert pretraining of video transformers

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.218024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T16:22:25.098703Z digest=sha256:cde2c18d97caa9575a7690559fe47be9d428c7fb2757105263f387cb852cca9c

Observation e5e7babb-2187-497a-af9d-76469a1373b0 · outbound

This paper cites Girdhar, A.

Extending Video Masked Autoencoders to 128 frames Girdhar, A

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.205639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T16:22:25.103922Z digest=sha256:c3787758e67c7bb210589e6b00d857fe96e9240f24fcf8ed8e11baa096950d59

Observation 713e89ef-4263-417a-b095-74a19c92e633 · outbound

This paper cites Zero-shot text-to-image generation.

Extending Video Masked Autoencoders to 128 frames Zero-shot text-to-image generation

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.193574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T16:22:25.107960Z digest=sha256:01bfb73e1c571b58e846657e9f2a3f68d83a2fdde5884c71835839485da5ebc7

Observation bc7e4fd2-2979-4311-a58b-b08529e687ad · outbound

This paper cites Masked video distillation: Rethinking masked feature modeling for self-supervised video representation learning.

Extending Video Masked Autoencoders to 128 frames Masked video distillation: Rethinking masked feature modeling for self-supervised video representation learning

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.182211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T16:22:25.111942Z digest=sha256:b4c418b6b368986186e779bb3b2f5d4b506245be84f5cfb1c969bf0f96e5ffd0

Observation e5511110-75f1-465e-bd26-2f2d846c7286 · outbound

This paper cites Magvit: Masked generative video transformer.

Extending Video Masked Autoencoders to 128 frames Magvit: Masked generative video transformer

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.170158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T16:22:25.115504Z digest=sha256:2b904b1945d8c9ae773b74cf71db656082030b10c858c3f0b0e66569e37f58d2

Observation 929db5b8-04b5-4e3d-a6cc-91d48a12431c · outbound

This paper cites Tokenlearner: What can 8 learned tokens do for images and videos? In NeurIPS, 2021.

Extending Video Masked Autoencoders to 128 frames Tokenlearner: What can 8 learned tokens do for images and videos? In NeurIPS, 2021

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.143450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T16:22:25.123089Z digest=sha256:e5c8c3198ec20bc795ad39f5976e25095259caa157b7b6a717a5327a0e16fd4b

Observation 9cd15a53-f043-4a68-9079-f5e6871c87d9 · outbound

This paper cites something something.

Extending Video Masked Autoencoders to 128 frames something something

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.130927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T16:22:25.126787Z digest=sha256:da06a387493c082ac0716c0dbc9f17a967630045ba2ffa81279deb33e928b941

Observation 319a40db-cab6-42ab-86cd-b86c83095958 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

Extending Video Masked Autoencoders to 128 frames An image is worth 16x16 words: Transformers for image recognition at scale

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.118986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T16:22:25.130449Z digest=sha256:2a33e4e217ca004d57afe11b536bcf4e41a178ed8df30c7a565f4c830d739d98

Observation d87ef826-a40a-4764-958d-63623af7b0cc · outbound

This paper cites Mgmae: Motion guided masking for video masked autoencoding.

Extending Video Masked Autoencoders to 128 frames Mgmae: Motion guided masking for video masked autoencoding

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.105953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T16:22:25.134329Z digest=sha256:a97ada0e1784d65ebf3e4405951f7e68d838d21c9e242530c0faba82240bf0ed

Observation b5b0e6f3-fc2d-408f-a232-3ca3c19afbf2 · outbound

This paper cites Motion-guided masking for spatiotemporal representation learning.

Extending Video Masked Autoencoders to 128 frames Motion-guided masking for spatiotemporal representation learning

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.091040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T16:22:25.138258Z digest=sha256:fc6d0de630521dcbeb87f72d22cb9fe8bde20b3a92a5a207f0d315367f594e58

Observation e32096bb-4f64-4512-ac99-52032f45532d · outbound

This paper cites Video codec design: developing image and video compression systems.

Extending Video Masked Autoencoders to 128 frames Video codec design: developing image and video compression systems

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.077412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T16:22:25.142170Z digest=sha256:c0e91a5f24e32d1702ca850ac09e377e662e10b622a23310b08ab70ae5e79d48

Observation 9a8ae6a4-ee82-4266-bced-f94d9ea68a02 · outbound

This paper cites Raft: Recurrent all-pairs field transforms for optical flow.

Extending Video Masked Autoencoders to 128 frames Raft: Recurrent all-pairs field transforms for optical flow

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.062271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T16:22:25.146065Z digest=sha256:f48dfe975265801a9aa53ad8c555368107a0e007c9f4d1eba6221eb66ce9533a

Observation d71d0fc2-7964-48d3-9875-5cb23baf2090 · outbound

This paper cites Videoprism: A foundational visual encoder for video understanding.

Extending Video Masked Autoencoders to 128 frames Videoprism: A foundational visual encoder for video understanding

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.050812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T16:22:25.149577Z digest=sha256:d605f1d0d14508d6a6fa6734900107e76de75fb486b68564bb4a6f0261732ac5

Observation 0c81c503-ca16-426b-8f61-e24d49b7afd8 · outbound

This paper cites Internvideo2: Scaling video foundation models for multimodal video understanding.

Extending Video Masked Autoencoders to 128 frames Internvideo2: Scaling video foundation models for multimodal video understanding

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.037738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T16:22:25.153989Z digest=sha256:8f2d72a204172ad696e3a6a6d6d4ca22e19e2dec44a8f2dc75ddd60313296903

Observation 55a81fcf-1c28-4640-b3d8-ce4c2b551f17 · outbound

This paper cites Vivit: A video vision transformer.

Extending Video Masked Autoencoders to 128 frames Vivit: A video vision transformer

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.024691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T16:22:25.157567Z digest=sha256:621a8641b5133553341984c958e2d02d633257b4f9f8d65c52318051c769225f

Observation 07e3e8ce-9ff2-4a61-9d79-e761d97b7bfb · outbound

This paper cites Finite scalar quantization: VQ-V AE made simple.

Extending Video Masked Autoencoders to 128 frames Finite scalar quantization: VQ-V AE made simple

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:26.011181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T16:22:25.161177Z digest=sha256:dbb6f7cb3d9b2ad189502fa1e7aaf81db0556caf332b1c207f1580c2b3f826da

Observation 02398460-7a6d-40c7-bf63-b314a15cce3f · outbound

This paper cites A Short Note about Kinetics-600.

Extending Video Masked Autoencoders to 128 frames A Short Note about Kinetics-600

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T16:22:25.164843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:22:25.164843Z digest=sha256:8d0596e2750763c33dac54ec8bac5c093ff6f5d7ac36114d57cec98930942cd0

Observation 95ecca2f-f416-4873-9e90-80143ff10102 · outbound

This paper cites A Short Note on the Kinetics-700 Human Action Dataset.

Extending Video Masked Autoencoders to 128 frames A Short Note on the Kinetics-700 Human Action Dataset

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T16:22:25.168857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:22:25.168857Z digest=sha256:33914c7360e771e188f7f9ac1678715408da5156fef4d99ccbe4052aff444ac1

Observation a196cfee-b1a4-49b0-b4a2-0d09a82b791d · outbound

This paper cites Multiview transformers for video recognition.

Extending Video Masked Autoencoders to 128 frames Multiview transformers for video recognition

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.997209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T16:22:25.172795Z digest=sha256:e8a147bce4f3872b5fd039bc710d460be11918f957acbe2372d8552626f9dbe0

Observation 33cfc5de-95ae-40d3-8b38-2227648e89d7 · outbound

This paper cites Temporally-Adaptive Models for Efficient Video Understanding.

Extending Video Masked Autoencoders to 128 frames Temporally-Adaptive Models for Efficient Video Understanding

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T16:22:25.176362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:22:25.176362Z digest=sha256:965a3eece1b4540e9d227fd6dc24d4452f8d8b623678b34b0ca3b4cc25c67960

Observation 2f028e33-ae02-4d0b-8d47-ea889ce8eda2 · outbound

This paper cites Training a Large Video Model on a Single Machine in a Day.

Extending Video Masked Autoencoders to 128 frames Training a Large Video Model on a Single Machine in a Day

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T16:22:25.180917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:22:25.180917Z digest=sha256:7319eb5e7b82053fd2dab378c29e480c34fee95ef6329b23ca7a4235e6b06596

Observation 72d5d421-0073-40e7-8db3-f1349385d633 · outbound

This paper cites Imagenet-21k pretraining for the masses.

Extending Video Masked Autoencoders to 128 frames Imagenet-21k pretraining for the masses

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.983302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T16:22:25.185019Z digest=sha256:7e03a9abd0ec1a4eb37833efa83618c5f46ab4fdad464708bcfdf4deee8541c4

Observation 45a79a2a-274c-496e-b7cf-87f33caffe1d · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video.

Extending Video Masked Autoencoders to 128 frames Ego4d: Around the world in 3,000 hours of egocentric video

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.972248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T16:22:25.188571Z digest=sha256:230840c4aec5781e17d3534cb94c5575a6fc3546c4f0f9bef41986c053ddc24b

Observation bdb7e0cb-169d-4786-bfcf-e6475c709066 · outbound

This paper cites Verbs in action: Improving verb understanding in video-language models.

Extending Video Masked Autoencoders to 128 frames Verbs in action: Improving verb understanding in video-language models

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.961247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T16:22:25.192158Z digest=sha256:11a64bc77a2cbba3317d569beeb41da7bf49f7eee32009de9091deca4249866b

Observation ec83fec6-7af5-4875-aefd-eb8f725487be · outbound

This paper cites Slowfast networks for video recognition.

Extending Video Masked Autoencoders to 128 frames Slowfast networks for video recognition

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.949510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T16:22:25.196373Z digest=sha256:0d92d303a68a6b5980abffa60fcf567c7ecd4f1977c785973786c145c853476e

Observation 24b0f31e-184e-43ca-9067-4636c6be323f · outbound

This paper cites Interactive prototype learning for egocentric action recognition.

Extending Video Masked Autoencoders to 128 frames Interactive prototype learning for egocentric action recognition

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.939197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T16:22:25.200298Z digest=sha256:8a4d539df9a9285831ab9ff0157195f45713a02014232b56b11ad6bb4b4b3188

Observation 0be7e6c2-5f50-4e9f-aa24-cd85927732bb · outbound

This paper cites Movinets: Mobile video networks for efficient video recognition.

Extending Video Masked Autoencoders to 128 frames Movinets: Mobile video networks for efficient video recognition

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.928022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T16:22:25.203795Z digest=sha256:a032d56911af31c951a5898c94053bc6617fc1bece8fb423d6eabbce54c26423

Observation a08bd723-c388-4787-9478-9e4cdde59e2d · outbound

This paper cites Omnivore: A single model for many visual modalities.

Extending Video Masked Autoencoders to 128 frames Omnivore: A single model for many visual modalities

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.916079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T16:22:25.207333Z digest=sha256:0d06ac04318605fc66915618f3c293fac4ca64c4a25d806fc671fe9bcd40e07d

Observation 08528d82-d877-41a5-ba58-e4c47280dda9 · outbound

This paper cites Learning video representations from large language models.

Extending Video Masked Autoencoders to 128 frames Learning video representations from large language models

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.903493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T16:22:25.211110Z digest=sha256:8c46be72030f2aa1db245dd03847f907698ea408f664b7e74982fdbcd04da53f

Observation 994e1ecd-afde-4267-9eff-b2a21a9e9a74 · outbound

This paper cites Is space-time attention all you need for video understanding? In ICML, 2021.

Extending Video Masked Autoencoders to 128 frames Is space-time attention all you need for video understanding? In ICML, 2021

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.891843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T16:22:25.214507Z digest=sha256:a95494226249f8fad2673d8aed31cddf4f138e0c0ca56685dc33011da9632fe5

Observation 62167356-5e55-4fb7-8620-e87d6e8f161f · outbound

This paper cites Video swin transformer.

Extending Video Masked Autoencoders to 128 frames Video swin transformer

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.879718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T16:22:25.218269Z digest=sha256:414ef6e144952745f1741a4ea8eb06df2c817cc03c3100cf431defbe72c81a1a

Observation be25e975-2432-4536-8ce7-8d7427f0d842 · outbound

This paper cites Can an image classifier suffice for action recognition? In ICLR, 2022.

Extending Video Masked Autoencoders to 128 frames Can an image classifier suffice for action recognition? In ICLR, 2022

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.868902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T16:22:25.221993Z digest=sha256:86e8d41aad3067abb9af5be2a13af7bafae67afd2c86bdf92c941a7d3c67591a

Observation 5efa996f-7043-42bb-bb20-13945d3dc966 · outbound

This paper cites Object-region video transformers.

Extending Video Masked Autoencoders to 128 frames Object-region video transformers

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.857839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T16:22:25.225841Z digest=sha256:e1f576fcbcdceb31259dce2f14d75360556f62e77b2b621089ac73ceca722f47

Observation 738d1098-9f5d-4766-8627-53e4d24a5d64 · outbound

This paper cites Aim: Adapting image models for efficient video action recognition.

Extending Video Masked Autoencoders to 128 frames Aim: Adapting image models for efficient video action recognition

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.845804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T16:22:25.229883Z digest=sha256:dbb665931de1b7d2f8be2462d2b3369c7946af153adae12af4e6fc27c9d2d0dd

Observation 26189b8b-33f3-4750-9990-c658006ca20a · outbound

This paper cites Video-focalnets: Spatio-temporal focal modulation for video action recognition.

Extending Video Masked Autoencoders to 128 frames Video-focalnets: Spatio-temporal focal modulation for video action recognition

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.834222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T16:22:25.233575Z digest=sha256:10c5a97398d9d4cd8586aad72690165650add02a056abc72f8a6dc4ad5a97c5a

Observation 6d0e83f9-ab6a-4b0e-910c-dc8b000828e3 · outbound

This paper cites Language model beats diffusion–tokenizer is key to visual generation.

Extending Video Masked Autoencoders to 128 frames Language model beats diffusion–tokenizer is key to visual generation

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.819649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T16:22:25.237643Z digest=sha256:d270843c361a780762b467d3ca1fd661720e07c338f8b8cbed113c0e8d920269

Observation 35bf2940-3183-4085-8371-02c89b5ee713 · outbound

This paper cites Finegym: A hierarchical video dataset for fine-grained action understanding.

Extending Video Masked Autoencoders to 128 frames Finegym: A hierarchical video dataset for fine-grained action understanding

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.804450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T16:22:25.241462Z digest=sha256:6c53b5bd7e2555c20c47fba78d13b9be58b0dc36b2d3d8f5f5ec3e478e66bcad

Observation c0e06bc4-9991-4702-8521-b65e8aa77aa3 · outbound

This paper cites Learning temporal cues for fine-grained action recognition.

Extending Video Masked Autoencoders to 128 frames Learning temporal cues for fine-grained action recognition

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.790238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T16:22:25.245443Z digest=sha256:4bf2fa91d13b480360b56acae57301bda5c0ee55c462e802abc523cbaecd037c

Observation b2aa5f83-862d-4df6-b1f7-ebe7564d0c09 · outbound

This paper cites Tsm: Temporal shift module for efficient video understanding.

Extending Video Masked Autoencoders to 128 frames Tsm: Temporal shift module for efficient video understanding

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.775928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T16:22:25.249266Z digest=sha256:ab668e520ba13353b8518d62052485b2292770b325d873923974beca01fcd59b

Observation 33ddac3d-e7d2-4c61-85b0-c7df2c2651a0 · outbound

This paper cites Temporal query networks for fine-grained video understanding.

Extending Video Masked Autoencoders to 128 frames Temporal query networks for fine-grained video understanding

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.762566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T16:22:25.253411Z digest=sha256:5c19dbfbdc5f2b4cf60362061061a1f643b7ac7861f121503d6613901c0df5b9

Observation d5c37d89-8b6f-4e04-b893-76292cb9fbaf · outbound

This paper cites Combined cnn transformer encoder for enhanced fine-grained human action recognition.

Extending Video Masked Autoencoders to 128 frames Combined cnn transformer encoder for enhanced fine-grained human action recognition

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.751023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T16:22:25.257990Z digest=sha256:28359289fdc79b8f86adcf5417c7afacc7e33abd7454f66928523eef9c3f9e5c

Observation c1217cf0-e260-4354-a23a-649cf1548550 · outbound

This paper cites Going deeper with image transformers.

Extending Video Masked Autoencoders to 128 frames Going deeper with image transformers

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.738041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T16:22:25.262106Z digest=sha256:2cfa68f9396e881ebf30d05749e305a53b288e80fa947f1e0f234cb5e67ca4d7

Observation e4c247d4-831e-46c6-a25f-64cfd87e3220 · outbound

This paper cites Scenic: A jax library for computer vision research and beyond.

Extending Video Masked Autoencoders to 128 frames Scenic: A jax library for computer vision research and beyond

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.722729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T16:22:25.266075Z digest=sha256:ac051d9fb8eb602c0e2f3d6e7fe72c9c7ee7d8cb0517b75642af2784f1371caa

Observation 0a74a736-5133-4c9c-a196-f891b24fb370 · outbound

This paper cites We found this slightly improved accuracy ( +0.5 points on EPIC-Kitchens-100 Verbs and +1.2 points on Diving48).

Extending Video Masked Autoencoders to 128 frames We found this slightly improved accuracy ( +0.5 points on EPIC-Kitchens-100 Verbs and +1.2 points on Diving48)

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.700975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T16:22:25.270584Z digest=sha256:b3fcb3259bf13e065d4f091e45a10731cb89da34ceac0aff7a6cc35072c2ba76

Observation e8883ef5-e7c8-4e61-a54c-8143413ce151 · outbound

This paper cites an unresolved cited work.

Extending Video Masked Autoencoders to 128 frames Unresolved cited work

Reference 66

Resolution
unresolved
raw_fallback, observed 2026-08-12T16:22:25.681314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T16:22:25.274933Z digest=sha256:3c1c582c0ace60f35e6675e912f6b8779901164fc6981d4af8257c04e48afd71

Observation 7121ac07-0f58-47f2-ad01-88e579276083 · outbound

This paper cites 17 Table 9: Model size vs frames.

Extending Video Masked Autoencoders to 128 frames 17 Table 9: Model size vs frames

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.667471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T16:22:25.279010Z digest=sha256:de6d8e332b5b2f072649a8df01a98956587e12706222ca85af26cd081d8970a9

Observation 13270ed7-aa41-42bd-a1da-25f5ef192c76 · outbound

This paper cites Guidelines: • The answer NA means that the abstract and introduction do not include the claims made in the paper.

Extending Video Masked Autoencoders to 128 frames Guidelines: • The answer NA means that the abstract and introduction do not include the claims made in the paper

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.653781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T16:22:25.283416Z digest=sha256:7133341c8c4b8dc35ef58510cb831239d1f8a00d47554e5cc8e89566e549c90f

Observation fd7f5d79-0999-4578-ba91-897e965c755c · outbound

This paper cites Limitations.

Extending Video Masked Autoencoders to 128 frames Limitations

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.636688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T16:22:25.287958Z digest=sha256:a42042b35d7d284986b876affded301476bfe98bc581d1b48e39098d30c1d494

Observation 10456fc6-93cc-44fe-982e-3888d20813f1 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include theoretical results.

Extending Video Masked Autoencoders to 128 frames Guidelines: • The answer NA means that the paper does not include theoretical results

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.619738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T16:22:25.292611Z digest=sha256:6eb68c0080083867b8e1be3cc8a327e75471e87f4de2c50d1ec64f4643b33c6e

Observation 955057ce-da5a-4eab-b8a1-3bb1e743a1f6 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include experiments.

Extending Video Masked Autoencoders to 128 frames Guidelines: • The answer NA means that the paper does not include experiments

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.605281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T16:22:25.296935Z digest=sha256:c43274665d8c757b2123af1da631345b2a7b3c4afb3fd44667bf49396b36e6a6

Observation 05b446cf-6eb3-4330-affd-25d7dd01dfe0 · outbound

This paper cites Guidelines: • The answer NA means that paper does not include experiments requiring code.

Extending Video Masked Autoencoders to 128 frames Guidelines: • The answer NA means that paper does not include experiments requiring code

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.588626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T16:22:25.302631Z digest=sha256:21567431616a7da90140fa365c6f7c0a6d4d34a8a84441b7a5b05955ff9e17d0

Observation 903717b9-95c5-4a0b-ab71-e54f72805d1b · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include experiments.

Extending Video Masked Autoencoders to 128 frames Guidelines: • The answer NA means that the paper does not include experiments

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.572605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T16:22:25.307592Z digest=sha256:9e04fceb2376de7a013c4c70afe4febab202e8b8f07d974fcb22c763a691222e

Observation 7504251a-762c-41d2-ad08-4e797b83ce0a · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include experiments.

Extending Video Masked Autoencoders to 128 frames Guidelines: • The answer NA means that the paper does not include experiments

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.559349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T16:22:25.312618Z digest=sha256:d35cb5a2f099cf5d661a0263978d2a2feff45b911a2db500bb309b2726aff174

Observation e409290e-ad8e-4ba9-95c5-8ff826b82976 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include experiments.

Extending Video Masked Autoencoders to 128 frames Guidelines: • The answer NA means that the paper does not include experiments

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.543694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T16:22:25.318955Z digest=sha256:944d305b2eec588cc11aa127f5463e223c47c299eb310a3945cdbf2498ce9fbd

Observation 151092d4-fe23-4c42-8e99-94dce5af1702 · outbound

This paper cites Guidelines: • The answer NA means that the authors have not reviewed the NeurIPS Code of Ethics.

Extending Video Masked Autoencoders to 128 frames Guidelines: • The answer NA means that the authors have not reviewed the NeurIPS Code of Ethics

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.528011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T16:22:25.323122Z digest=sha256:7fd3fffb7e5ec2ed68b984796485fa28801495a1ff8e136310a1b59920faad40

Observation 27550e91-dc95-4ad8-834c-1d3ca347c615 · outbound

This paper cites • If the authors answer NA or No, they should explain why their work has no societal impact or why the paper does not address societal impact.

Extending Video Masked Autoencoders to 128 frames • If the authors answer NA or No, they should explain why their work has no societal impact or why the paper does not address societal impact

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.510829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T16:22:25.327717Z digest=sha256:0a688c4a6b7a6553d39ed0e1130599625b23e46abe519deab05cdd6fdba943a3

Observation 1832e7fb-510a-4c78-80da-e7490525eabf · outbound

This paper cites Guidelines: • The answer NA means that the paper poses no such risks.

Extending Video Masked Autoencoders to 128 frames Guidelines: • The answer NA means that the paper poses no such risks

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.497541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T16:22:25.332877Z digest=sha256:127e6df145794bdcb7f465b5fd9d040c409a0cc8bc13bce0818e0f37d6048bfb

Observation e87e9080-6d21-4d13-8e92-c9dc39889c6e · outbound

This paper cites Guidelines: • The answer NA means that the paper does not use existing assets.

Extending Video Masked Autoencoders to 128 frames Guidelines: • The answer NA means that the paper does not use existing assets

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.484090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T16:22:25.343151Z digest=sha256:9cf2d0f821caf4a892d017718dabc8bec0f01ab4154073e73f66f1ad01eab0ec

Observation 0f1ef82c-5084-4726-b5f5-3e7278d9683a · outbound

This paper cites Guidelines: • The answer NA means that the paper does not release new assets.

Extending Video Masked Autoencoders to 128 frames Guidelines: • The answer NA means that the paper does not release new assets

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-12T16:22:25.347041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:22:25.347041Z digest=sha256:c04eaea5ca6aaf2445b9fb2fc173c82e1998fc904f0c3c6aeb250bb511404f57

Observation 32ea79b1-a7bb-4d37-9d42-eee4bc252c05 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects.

Extending Video Masked Autoencoders to 128 frames Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.464389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T16:22:25.351339Z digest=sha256:b13386f0e1b0618f2b8d9330031b3a369e8b3267d6a5cf1c7586f3512a820017

Observation 1f04bade-1819-48ac-ab45-49e80ce8e8d8 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects.

Extending Video Masked Autoencoders to 128 frames Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:22:25.450458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T16:22:25.355326Z digest=sha256:fc734acf35a9c56f772da3fb3946e78093164da64f36a9f9476e1be69d3a5c88

Observation 17e2c672-1e70-46a4-a07b-68f28e108efe · outbound

This paper cites an unresolved cited work.

Extending Video Masked Autoencoders to 128 frames Unresolved cited work

Reference 2023

Resolution
parse uncertain
raw_fallback, observed 2026-08-12T16:22:26.154889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T16:22:25.119455Z digest=sha256:d06df2b262aa3e1949d362aa49dd9323340eabb41cb8594a42b86137e6eab02f

Pith citing papers

No inbound Pith citation observations are available.