Pith. sign in

Paper Citation Record · LEDGER

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders

As of 9 August 2026, this Paper Citation Record lists 100 of 111 outbound references and 1 inbound Pith citation observation for arXiv:2502.07811.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.07811 v1

Coverage vector

measured 100 of 111 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T19:17:39.829689Z

measured 101 of 101 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:09:18.061419Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T18:16:30.548575Z

Reference resolution

100 of 111 outbound references displayed

  • verified exact5
  • verified fuzzy54
  • unresolved41
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation aa31f2e1-7314-4142-8a0e-8c9e1ca4047e · outbound

This paper cites Self-supervised multimodal versatile networks.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Self-supervised multimodal versatile networks

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:39.436477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:39.436477Z digest=sha256:38d06f76b5d11e791f14bd8ba58cfc4ffefc0115410ed6981aa53ab529294b71

Observation 5c6b4a85-fdfa-49f7-8089-cc2eca4b9ba8 · outbound

This paper cites Self-supervised learning by cross-modal audio-video clustering.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Self-supervised learning by cross-modal audio-video clustering

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:39.441116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:39.441116Z digest=sha256:b51b60fc75574469c871bf5f57da814442fbbd485ba7e39a1d1305a68176fbdf

Observation 514f8041-bb2a-4464-bd57-0a3a6b5da576 · outbound

This paper cites Look, listen and learn.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Look, listen and learn

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:39.445250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:39.445250Z digest=sha256:428c2d81f0961d80c415e4de8a70a72e97403d729e7a88c1057ca9ec554e32b8

Observation a8586362-2712-4b57-8c07-76dd79501518 · outbound

This paper cites Objects that sound.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Objects that sound

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:39.449872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:39.449872Z digest=sha256:65f2267c11915ece07d03ec62b934a5cea8aa180084e850c26dc5af5840d1d7a

Observation 5f683016-7c5c-4ee6-9ad5-a37dcad2619e · outbound

This paper cites Vivit: A video vision transformer.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Vivit: A video vision transformer

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:39.453863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:39.453863Z digest=sha256:58ac75f5837049c68f85f6f7553be54e681066fddb16d666cc71a7eb6c3bfba8

Observation 417c24bb-c924-400d-b216-af913765b5b8 · outbound

This paper cites Adamae: Adaptive masking for efficient spatiotempo- ral learning with masked autoencoders.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Adamae: Adaptive masking for efficient spatiotempo- ral learning with masked autoencoders

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:39.457906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:39.457906Z digest=sha256:3341736fad25c6e3a635ef97d80ce8aa673d68eeaa58726e735658886348e943

Observation e49641c3-d57b-4f0a-9d1b-5c40625657aa · outbound

This paper cites BEit: BERT pre-training of image transformers.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders BEit: BERT pre-training of image transformers

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:39.462043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:39.462043Z digest=sha256:fb8765c90c8a46e2ca4ec143818b5789f0029e50d20a328995b490e2c448bc41

Observation 6611a668-d9f4-4c85-a89b-53f609f485a1 · outbound

This paper cites Speednet: Learning the speediness in videos.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Speednet: Learning the speediness in videos

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:39.466647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:39.466647Z digest=sha256:cc35811821ba3f20955af4b0a3002171daf04873b9c891349d7e5c538cf7cf24

Observation 0b2949be-ce36-4de2-b3ea-b842c2f58b5e · outbound

This paper cites Is space-time attention all you need for video understanding? In Proceedings of the 38th International Conference on Ma- chine Learning (ICML), pages 813–824.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Is space-time attention all you need for video understanding? In Proceedings of the 38th International Conference on Ma- chine Learning (ICML), pages 813–824

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:39.470804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:39.470804Z digest=sha256:67386be2321a3cef8c8ea0d28883801d2be6619d247e96183d10bc2f2eaca31b

Observation 6668cefe-edd9-4bd6-9124-7a21f1422748 · outbound

This paper cites Emerging properties in self-supervised vision transformers.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Emerging properties in self-supervised vision transformers

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:39.475016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:39.475016Z digest=sha256:9551f33080aae1ddd42005a5b6e9424843c4baca467484149523e134f9527d67

Observation 65788d6c-4189-41e8-8422-35b2a69464e0 · outbound

This paper cites Learning aligned cross-modal representations from weakly aligned data.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Learning aligned cross-modal representations from weakly aligned data

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:39.479038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:39.479038Z digest=sha256:374c7e86356341aaaaf0d3b7ae400da38d063e5afabdfc1e0db9379ac249de1b

Observation 17edde5c-77e6-4dd2-ba52-2736aa6f6c61 · outbound

This paper cites Generative pre- training from pixels.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Generative pre- training from pixels

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:39.483141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:39.483141Z digest=sha256:d5a357caacf874bca8b8720268821ceaf8979dffce94d867d3b95e1bc20ffee4

Observation 5539f8f5-1b13-4a00-9020-c74c0b3e8240 · outbound

This paper cites Rspnet: Relative speed perception for unsupervised video representation learning.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Rspnet: Relative speed perception for unsupervised video representation learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:39.487073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:39.487073Z digest=sha256:f03c34ea72f4b601229e73de6fa6ccb81f71fd613cb04a9d5067e44e7a31ecdb

Observation 75f840e0-2139-4217-8264-8cf28e8a5d9e · outbound

This paper cites A simple framework for contrastive learning of visual representations.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders A simple framework for contrastive learning of visual representations

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:39.490967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:39.490967Z digest=sha256:d42a99b40dc507976c1d26ac2a2768b779e7fa623c000b01a59ea2b1982809e4

Observation f3e82289-834d-4ddd-ae23-8b5524d2c663 · outbound

This paper cites Electra: Pre-training text encoders as dis- criminators rather than generators.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Electra: Pre-training text encoders as dis- criminators rather than generators

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:39.495458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:39.495458Z digest=sha256:0bc28e2f8d608e9dbc9b3004b68140fd0a0324007dfb812e1c63be78b65bd148

Observation 959bdc18-f39f-4383-bf50-293dafaa4dd7 · outbound

This paper cites Randaugment: Practical automated data augmen- tation with a reduced search space.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Randaugment: Practical automated data augmen- tation with a reduced search space

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:39.499741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:39.499741Z digest=sha256:008ad0d16512c92acb03458001649599aa424453ea27ce70bc892357b73c111e

Observation 155e3fce-1ef4-475f-8c7a-dd0a7ab896b7 · outbound

This paper cites Unifying video self-supervised learning across families of tasks: A survey.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Unifying video self-supervised learning across families of tasks: A survey

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:39.503662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:39.503662Z digest=sha256:2ed4b1650e68382ca586be081ce48f26b28b30670f2777917b77c98bffd1b50c

Observation ea42b726-5ceb-4e61-8f21-4b2dec2cd59f · outbound

This paper cites Imagenet: A large-scale hierarchical im- age database.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Imagenet: A large-scale hierarchical im- age database

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:39.507721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:39.507721Z digest=sha256:54921ffda1964f456248853ca21b5ca69766c69072b66f857237ed0bf9f22c87

Observation 312f0708-8db9-4eb9-857c-05931df10a21 · outbound

This paper cites Virtex: Learning visual representations from textual annotations.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Virtex: Learning visual representations from textual annotations

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:39.512352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:39.512352Z digest=sha256:15e34efc3482297a75880f6d17fc10768e4a4cdd117d5d675cccbbc67d2f8e28

Observation 139a9f0e-f182-4047-a18f-4ec727dfb35e · outbound

This paper cites 9 Vi2clr: Video and image for visual contrastive learning of representation.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders 9 Vi2clr: Video and image for visual contrastive learning of representation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:39.516832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:39.516832Z digest=sha256:fddd3f429d37f5236f0307ff4c241bbe97ba0b80e63015594da2a0d7dbb3634a

Observation 43e118a0-44e5-4587-aef1-6d255efc6c90 · outbound

This paper cites Taming transformers for high-resolution image synthesis.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Taming transformers for high-resolution image synthesis

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:39.520528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:39.520528Z digest=sha256:62eb5dbeff03596807091c9059a6fe1e2d11bc1f3bf1be2cf738536eb3898129

Observation 20cab6dd-82ee-4f9d-9672-fc366b2f2c78 · outbound

This paper cites Multiscale vision transformers.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Multiscale vision transformers

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:39.524439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:39.524439Z digest=sha256:c24bb331818b48c297d6e81edd733052ce74473aefd74dbab581c3157967f919

Observation 96353754-9cec-44fc-a127-806e7f99f547 · outbound

This paper cites Slowfast networks for video recognition.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Slowfast networks for video recognition

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:39.528415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:39.528415Z digest=sha256:648c685fa25f96a119f35776f4ea52f8bf0e43f4a5a723de87f8dc97ac848f60

Observation 0bcce743-4f81-45bd-bdd4-f3b969a60443 · outbound

This paper cites A large-scale study on unsuper- vised spatiotemporal representation learning.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders A large-scale study on unsuper- vised spatiotemporal representation learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:39.532554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:39.532554Z digest=sha256:27ada2c8fd2b7e3fdf532a79dd189c6cf2b70c9336277889bce9a876f2792141

Observation 901d0767-fcc8-42e2-ac84-f82b43be35f0 · outbound

This paper cites Masked autoencoders as spatiotemporal learners.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Masked autoencoders as spatiotemporal learners

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:39.537655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:39.537655Z digest=sha256:b16e43a4745be52d7cf4e6e74fa4229a0a5bab259f1fe736ee67623cb29d0fab

Observation 63f6a578-731e-49e2-8427-51835d2d17e4 · outbound

This paper cites Mcmae: Masked convolution meets masked autoencoders.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Mcmae: Masked convolution meets masked autoencoders

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:39.542371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:39.542371Z digest=sha256:869ed9aeaf4cffab690f5bf61d437bb3e56f27cd462f5ccd039ca9a12dce7c1a

Observation 8466d3cb-f8e2-40a2-b371-f9372e09c879 · outbound

This paper cites Revitalizing cnn attention via transformers in self-supervised visual representation learn- ing.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Revitalizing cnn attention via transformers in self-supervised visual representation learn- ing

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:39.546400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:39.546400Z digest=sha256:dc1964613cab51b17eb346325c5868dd4093e48f4eb8feb9423f8b02925c2293

Observation 902a23cd-ffc7-42aa-83fd-437ffbfaaf93 · outbound

This paper cites Omnimae: Single model masked pretraining on images and videos.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Omnimae: Single model masked pretraining on images and videos

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:39.549998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:39.549998Z digest=sha256:ae046dca1c7dd4f5390bf6c721e2da96845890feb2d14b0315986108eae2b6b5

Observation cd0451dc-47d5-41b0-b9eb-34aed2a78ebc · outbound

This paper cites Improving image-sentence embeddings using large weakly annotated photo collec- tions.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Improving image-sentence embeddings using large weakly annotated photo collec- tions

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:39.553347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:39.553347Z digest=sha256:a6dcf9f5b3de720122f704fe274a48b5d2c557210ca2af01575b32537ca98c47

Observation 00ed9763-81a6-48b8-a699-eb1a39e57c45 · outbound

This paper cites Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:39.557394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:39.557394Z digest=sha256:836a0233ddddda79ec3b296afec86704d1f564e7f1a88f7459d7add7db78c119

Observation 226ba7d6-3b09-4de9-ad54-5722c4e210c3 · outbound

This paper cites The” something something” video database for learning and evaluating visual common sense.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders The” something something” video database for learning and evaluating visual common sense

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:39.561828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:39.561828Z digest=sha256:ea70eac662dc770d46ca6c0169500a6c3d79547b7a84b82fe928b167635ab7bd

Observation a1541263-211e-48db-aaec-9b861248f306 · outbound

This paper cites Bootstrap your own latent-a new approach to self-supervised learning.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Bootstrap your own latent-a new approach to self-supervised learning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:39.564924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:39.564924Z digest=sha256:eaf29237e805006f3f6ad1a5ab533e7cb14aba0baca37eee32ea4239a7a7ee42

Observation 191c8262-b90c-4e24-bd6d-88c7ebdf13a6 · outbound

This paper cites How Effective are Self-Supervised Models for Contact Identification in Videos.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders How Effective are Self-Supervised Models for Contact Identification in Videos

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-08-08T19:17:40.109081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:17:39.568572Z digest=sha256:e6eba31f8e714c58df478d64cf2f7e11a9c3622fed331ed937348d738170d7a1

Observation f230e108-8784-43cc-9750-f5125c4111f8 · outbound

This paper cites MaskViT: Masked Visual Pre-Training for Video Prediction.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders MaskViT: Masked Visual Pre-Training for Video Prediction

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:39.573335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:39.573335Z digest=sha256:2334cde0df7932adeb49d0333880bc9b60b40837288647e15fe35bfb3ea0ab50

Observation 43e74f4b-1dca-4739-924b-b2392e12b043 · outbound

This paper cites Memory- augmented dense predictive coding for video representa- tion learning.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Memory- augmented dense predictive coding for video representa- tion learning

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.823654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:17:39.577840Z digest=sha256:a123b123531d3ce047abbb1d66f7b3e7aec0bd5756c40e724711e30b1f8b08fd

Observation 3d180f1e-242b-40fb-ac7a-1ba0012db08a · outbound

This paper cites Self- supervised co-training for video representation learning.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Self- supervised co-training for video representation learning

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.813186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:17:39.581200Z digest=sha256:d929ca2a0802b2cd3a0f283f981a2f45e3df59c1e8d79049a872b724747f3127

Observation 85459afd-c541-4341-82f4-2a7c78a58e0a · outbound

This paper cites Momentum contrast for unsupervised visual rep- resentation learning.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Momentum contrast for unsupervised visual rep- resentation learning

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.802759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:17:39.584842Z digest=sha256:4a66b10499c4c8a54bcefc43c96957e13a73273eaaa20b254f5d8f5075fca338

Observation 04165458-3c61-40e2-a83c-6295736d3104 · outbound

This paper cites Masked autoencoders are scal- able vision learners.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Masked autoencoders are scal- able vision learners

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.791735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:17:39.588393Z digest=sha256:72cc46d13d99123746b6e4e922c11fe205e40ceeb80ce8b3da570061f900e2e7

Observation 0de94c77-96c4-45f6-9dbf-983a67480395 · outbound

This paper cites Human gaze control during real-world scene perception.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Human gaze control during real-world scene perception

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.780607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:17:39.591639Z digest=sha256:cc512bb1ff1d175ebbdfd3c57c0f2c7a81baa0b86a5c82e82cdeb20d947bfac0

Observation 68eaaf99-ec76-43fc-a872-bc7721e897d7 · outbound

This paper cites ViC-MAE: Self-Supervised Representation Learning from Images and Video with Contrastive Masked Autoencoders.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders ViC-MAE: Self-Supervised Representation Learning from Images and Video with Contrastive Masked Autoencoders

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-08-08T19:17:40.087326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:17:39.594919Z digest=sha256:1d362bcd637c6c6a8265fb8bbe4b7eec24f3190ee448d5b546d75a626631cc24

Observation f3fe09b1-b083-491d-bedb-dc09a8ea5e3a · outbound

This paper cites Augment your batch: Improv- ing generalization through instance repetition.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Augment your batch: Improv- ing generalization through instance repetition

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.770627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:17:39.598583Z digest=sha256:6c9b3851d2ebe62bf1ff5ce47dad791c77beee0ab4945fdb0a7c562ab5247bc6

Observation b906e515-298f-40f0-bbf5-b0ce88e3aab6 · outbound

This paper cites Contrast and order representa- tions for video self-supervised learning.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Contrast and order representa- tions for video self-supervised learning

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.761185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:17:39.601608Z digest=sha256:796a122aa94201026cf5650fa72738d89ef0adffcf457f58cd29cca3e88e4145

Observation 6e6c79f1-2abf-4621-b7f9-5d30f07241d9 · outbound

This paper cites Mgmae: Motion guided masking for video masked autoencoding.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Mgmae: Motion guided masking for video masked autoencoding

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.750253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:17:39.605271Z digest=sha256:ced50557d0be1878d07ce3f5008a66cf29ac44f6a591d92329b560da4e6d1759

Observation fe7d86f1-dd41-43b5-a621-217655a7b909 · outbound

This paper cites Self-supervised video representation learn- ing by context and motion decoupling.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Self-supervised video representation learn- ing by context and motion decoupling

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.739266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:17:39.608840Z digest=sha256:9161241bf6fdded1856e4bbf81be0746ce0a6d08829692f2a1ecb6773c839dec

Observation 738c14af-7d8d-4dc7-b1ec-4b44aeab3eed · outbound

This paper cites TAda! Temporally-Adaptive Convolutions for Video Understanding.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders TAda! Temporally-Adaptive Convolutions for Video Understanding

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:39.612598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:39.612598Z digest=sha256:8f86a0194386ebb9a1edde6af11f578a29bfcebf81f0a857b3ca3f9f576ee286

Observation 871e1068-622d-4070-998b-01da2f9af763 · outbound

This paper cites Self-Supervised Spatiotemporal Feature Learning via Video Rotation Prediction.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Self-Supervised Spatiotemporal Feature Learning via Video Rotation Prediction

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:39.616909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:39.616909Z digest=sha256:bc3e06e33a96543e42210b2bbb7cb7becb198c66f33e6c012b32625bb35d81d7

Observation 4cc14dd9-12dc-4812-83e2-50e587072d44 · outbound

This paper cites Hard negative mixing for contrastive learning.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Hard negative mixing for contrastive learning

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.729598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:17:39.621618Z digest=sha256:1b2418ee308bf303990b31c7219473aac72c3c609c5b301db78d10564a9feb61

Observation d283f6db-8757-4960-b160-66b335ad927e · outbound

This paper cites Deep visual-semantic alignments for generating image descriptions.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Deep visual-semantic alignments for generating image descriptions

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.719841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:17:39.625420Z digest=sha256:6d859c9199bbe9ee8d1154e56dce61d18b17a4ddcb92ebeae218a83911eb711e

Observation c873d323-5d1e-45bd-805a-8dcc16c1fcc4 · outbound

This paper cites The Kinetics Human Action Video Dataset.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders The Kinetics Human Action Video Dataset

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:39.629499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:39.629499Z digest=sha256:bbdaa821625fcbbca4d090e306fc940f07f2f25e5e68aca4b7dd3ebf9ac5e889

Observation 5baf0a28-b25a-4b8c-9b71-3cfbecb9a128 · outbound

This paper cites Fusion attention for action recognition: Integrating sparse-dense and global at- tention for video action recognition.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Fusion attention for action recognition: Integrating sparse-dense and global at- tention for video action recognition

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.711257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:17:39.634738Z digest=sha256:a9d771d2b4acd59a70302e16f3c2db298e623d527e1c38a8c9122de6e2cb0a57

Observation 9813b497-3a07-4597-a37d-0bf3bf129886 · outbound

This paper cites Co- operative learning of audio and video models from self- supervised synchronization.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Co- operative learning of audio and video models from self- supervised synchronization

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.701974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:17:39.641055Z digest=sha256:3c3c8541a6672600268c476b29694c5e1a739727b039592226f37a53a859b7e6

Observation 500f507f-515d-4222-b5a5-7434d5fd9663 · outbound

This paper cites Hmdb: a large video database for human motion recognition.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Hmdb: a large video database for human motion recognition

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.692379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:17:39.645261Z digest=sha256:8915c941530c74694ea2e506215b7cd16a92e825527e7f0003f90db9bca984fd

Observation 7e6406a4-b4dd-4641-9d63-875b5c9f1d29 · outbound

This paper cites A large-scale analysis on self- supervised video representation learning.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders A large-scale analysis on self- supervised video representation learning

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.682078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:17:39.649317Z digest=sha256:89f6a10e8b207bf4aa17cdab0718931ca3591724fee3723910e18a2e49f2866b

Observation 0799014a-b89c-478a-a0c5-1b1021f610dc · outbound

This paper cites Semmae: Semantic-guided masking for learning masked autoencoders.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Semmae: Semantic-guided masking for learning masked autoencoders

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.671868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:17:39.653578Z digest=sha256:bd4fe7d9261a71bdfd85992edd55c3c22995b7a17369196bfb28d623d437bb93

Observation 4bd49512-7c46-4fb3-b6b2-8f2c829f6e61 · outbound

This paper cites Improve unsuper- vised pretraining for few-label transfer.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Improve unsuper- vised pretraining for few-label transfer

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.661450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:17:39.657779Z digest=sha256:841c09f1fef50f5b7e1515d10de089f71e416d1543297537eabfa6efc004a38a

Observation b6fa720b-5e50-44cb-8783-ce9ae101ed0f · outbound

This paper cites Learning Spatiotemporal Features via Video and Text Pair Discrimination.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Learning Spatiotemporal Features via Video and Text Pair Discrimination

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:39.665769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:39.665769Z digest=sha256:88fb59e41bb2ae158e54824d4c1c409b639747a0ed453fb405008c52af609007

Observation dd473744-3a29-415c-9c46-0d7152670e70 · outbound

This paper cites Tsm: Temporal shift module for efficient video understanding.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Tsm: Temporal shift module for efficient video understanding

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.650601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:17:39.670766Z digest=sha256:f16eac32a34f6f3d4bbadbe970df02d795f3c9b84815a9876179ebf0a4c21cb5

Observation b0003748-d48d-4714-b611-efbcedd98d0b · outbound

This paper cites Self-supervised video-based action recognition with distur- bances.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Self-supervised video-based action recognition with distur- bances

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.640547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:17:39.674852Z digest=sha256:8a91e3a5a3daaa0743f2ce4c6475d8f57d1b10863cd2ad778124c2111ac1f5be

Observation 5d93b97f-445d-4dba-83ad-edd6c0a371c7 · outbound

This paper cites CrossVideo: Self-supervised Cross-modal Contrastive Learning for Point Cloud Video Understanding.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders CrossVideo: Self-supervised Cross-modal Contrastive Learning for Point Cloud Video Understanding

Reference 59

Resolution
verified exact
local_arxiv, observed 2026-08-08T19:17:40.026614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:17:39.679785Z digest=sha256:2ef4b22df1382f82f9746e66327baae448333804ee4f4bd0d6e44e10302315ba

Observation f63ef84d-b5e8-45f5-84ff-585c00b310ff · outbound

This paper cites Teinet: Towards an efficient architecture for video recognition.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Teinet: Towards an efficient architecture for video recognition

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.629344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:17:39.683826Z digest=sha256:561b51c9444eca073485b39ad6813d8f5381b88f4b0bc13d53b42a34b9a05ea7

Observation 792fa445-64f4-4b00-8df2-15369d8e8ddd · outbound

This paper cites Tam: Temporal adaptive module for video recog- nition.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Tam: Temporal adaptive module for video recog- nition

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.618146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:17:39.687932Z digest=sha256:1189b3e0a17fdd477c959382f1ff05dfad30487afee62ae9305641f787c17377

Observation 03ae3348-07ab-4823-9671-0e5c13c61487 · outbound

This paper cites Video swin transformer.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Video swin transformer

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.606934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:17:39.692108Z digest=sha256:c93a3aa4b9725b18dc0990c2b6460aec88c32bb31a3c417408a680e2c8a4f9c0

Observation 4ad99798-c765-4d09-ba86-c350fc00c566 · outbound

This paper cites SGDR: Stochastic Gradient Descent with Warm Restarts.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders SGDR: Stochastic Gradient Descent with Warm Restarts

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:39.696217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:39.696217Z digest=sha256:65447a6f9709f071e58ab6dd4ffdf1646a36e3bad20608a4b4bebfdfe78f49ab

Observation 9d1f7e83-dfba-417d-abda-a809af85e874 · outbound

This paper cites Decoupled weight de- cay regularization.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Decoupled weight de- cay regularization

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.597155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:17:39.700346Z digest=sha256:0ce5b1330621178734517a6b7f016ec7b581ca52a98cef6cd1b6c7abfb16b992

Observation fbe91960-dfe9-40cf-9114-eba79172bcb1 · outbound

This paper cites CMAE-V: Contrastive Masked Autoencoders for Video Action Recognition.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders CMAE-V: Contrastive Masked Autoencoders for Video Action Recognition

Reference 65

Resolution
verified exact
local_arxiv, observed 2026-08-08T19:17:39.991655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:17:39.703294Z digest=sha256:5a4e7e9448eb7c6ba67b2f8ae4672eab94600883aad3ece8f023de78b0f7e14f

Observation e280b4a2-9b2c-4658-ac55-94cd84302cda · outbound

This paper cites 12-in-1: Multi-task vision and language representation learning.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders 12-in-1: Multi-task vision and language representation learning

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.586283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:17:39.707292Z digest=sha256:216154d2cc0b7800b6b88a456a3c34eaf46eb977d3a0d7c8e733b23d028d8bfe

Observation c2c981d9-c97d-41a6-96a6-d97c93d2512d · outbound

This paper cites End-to-end learning of visual representations from uncurated instruc- tional videos.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders End-to-end learning of visual representations from uncurated instruc- tional videos

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.575610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:17:39.711733Z digest=sha256:808de7eb56beababc3ca82003f45a93e23cbfa9592f9032cc26d8434de315e97

Observation 084a12c4-3429-471b-bfa8-7986637c7087 · outbound

This paper cites Self-supervised learning of pretext-invariant representations.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Self-supervised learning of pretext-invariant representations

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.563057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:17:39.715441Z digest=sha256:f94d6752fe167efe6cf183314cdc9775f38e0121729ee80c01fdc76b576c9c0e

Observation 13389bba-7981-4c78-ab3a-5b9f71333ee5 · outbound

This paper cites Shuffle and learn: unsupervised learning using temporal order verification.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Shuffle and learn: unsupervised learning using temporal order verification

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.551556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:17:39.719271Z digest=sha256:a8386321fd68a1f9b19585c9043244766c5c2b7ee96a1fccda46e0899430e333

Observation 72dca34f-316f-4a03-b152-e32f1f507d0d · outbound

This paper cites Ro- bust audio-visual instance discrimination.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Ro- bust audio-visual instance discrimination

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.539755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:17:39.722912Z digest=sha256:9f2085316e8a5811a2cf5fbe3c7da855091e5ad710428c324ddb63c8a5d6be9f

Observation 8b33199a-aaee-47d7-a634-907428509832 · outbound

This paper cites Audio-visual instance discrimination with cross-modal agreement.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Audio-visual instance discrimination with cross-modal agreement

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.527661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:17:39.726729Z digest=sha256:d9c6fbcd81ec3c50ee55faf239ddfa83c6ade5c7073cf4a4a06be9571a38fa35

Observation 7c5b6ad5-ee64-4192-b19c-21661f1d02e8 · outbound

This paper cites Audio-visual scene analysis with self-supervised multisensory features.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Audio-visual scene analysis with self-supervised multisensory features

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:39.730563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:39.730563Z digest=sha256:68c43ec58b10af6c868398a5814416adf675afac18e0e9d41dddba10ee704121

Observation c739f4a5-045d-40b1-9468-201f92560579 · outbound

This paper cites Videomoco: Contrastive video representation learning with temporally adversarial examples.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Videomoco: Contrastive video representation learning with temporally adversarial examples

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.505976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:17:39.733882Z digest=sha256:667a4d6374de21438cac33a221df052a0980d6057ce969378946a0a079271b83

Observation a7a40403-6a6d-4c2e-8fdf-d059a87425f1 · outbound

This paper cites Multi-modal self-supervision from generalized data transformations.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Multi-modal self-supervision from generalized data transformations

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.496205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:17:39.737476Z digest=sha256:fee0d61147f0d0eaf866c64460bb326e53b7ebc41b2dde79e9cab52f9f9abe76

Observation 7cb4e5ca-9546-4928-ba59-a484f56495c9 · outbound

This paper cites Keeping your eye on the ball: Trajectory attention in video transformers.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Keeping your eye on the ball: Trajectory attention in video transformers

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.484561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:17:39.742292Z digest=sha256:6525321c753e92e5a7e8e8e3c9df061a7d284d3d56190f41ba22f1013219235e

Observation 59e4f32a-76b0-4639-861e-d44248bd2633 · outbound

This paper cites Evolving losses for unsupervised video representation learning.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Evolving losses for unsupervised video representation learning

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.473778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:17:39.745476Z digest=sha256:fec48cfcac5ad72f876fefa11fa9ced3ec424675d908b1e8670308b6b86e63a4

Observation b3be13be-da0b-4586-bcdd-c7d238515bd0 · outbound

This paper cites Spa- tiotemporal contrastive video representation learning.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Spa- tiotemporal contrastive video representation learning

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.459220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:17:39.749310Z digest=sha256:d322ad12cd8cf4a6563a01dc932b43437bd1f394afd0fd2f66cf00fdc428a88c

Observation 2e6d248a-812e-4f05-b38a-3b4fec0a1a50 · outbound

This paper cites Learning from untrimmed videos: Self-supervised video representation learning with hierarchical consistency.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Learning from untrimmed videos: Self-supervised video representation learning with hierarchical consistency

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.445236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:17:39.752423Z digest=sha256:77715064b539621899e1249f0ddb754f221057b87a38a8458599970a823e3e8c

Observation d8fa6935-4bef-4384-aef3-57926f397d74 · outbound

This paper cites Mar: Masked autoencoders for efficient action recognition.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Mar: Masked autoencoders for efficient action recognition

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.432594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:17:39.755665Z digest=sha256:628b7b277ce5ab4a7d5178dfdbfe258b046ea2c23486e4e0ee795f24e1ef3852

Observation 05e5e519-21f4-40e6-a1a8-789046110996 · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Learn- ing transferable visual models from natural language super- vision

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.420141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:17:39.759382Z digest=sha256:171ab559b619215a4b7f7916829704a135af237a46759726d47938e44e1df467

Observation adf4fcdb-1454-4f16-acf8-4b9dca135e39 · outbound

This paper cites Self-supervised video transformer.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Self-supervised video transformer

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.405711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:17:39.762718Z digest=sha256:66d3058c6642164eab5bf0ecc888106ae94110ae8e747726618b6208131f06cb

Observation f272e378-e5b7-42cc-ba3c-20e64d8078f4 · outbound

This paper cites The development of visual attention and the brain.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders The development of visual attention and the brain

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.394787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:17:39.766117Z digest=sha256:527b3189f9594d0f0144fb4b037ca872542310aa8e04fb620a51cae6a5c8a97e

Observation ac81765d-19a8-4a6f-82ed-ebc0f7a27dd9 · outbound

This paper cites Im- agenet large scale visual recognition challenge.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Im- agenet large scale visual recognition challenge

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.382306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:17:39.769664Z digest=sha256:dceec3e0bce29a055b58d0af9946e1b76b5e96e8dbd76271bab32478b16471dd

Observation aacb1bdb-4c33-43d1-8c50-b1e8c18e2730 · outbound

This paper cites Learning visual representations with caption annotations.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Learning visual representations with caption annotations

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.370700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:17:39.773677Z digest=sha256:6f4dc6d17fb7a3026f6cda416460f3da039872b858be848f87a8964dca511143

Observation 57cc4b17-a994-4024-b402-77c003c9f91a · outbound

This paper cites Two-stream con- volutional networks for action recognition in videos.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Two-stream con- volutional networks for action recognition in videos

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.359348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:17:39.777062Z digest=sha256:e4da61b9bda7ed84a187b0404ac5dd6a2a457ff3733387c1b917b2efc34b5c16

Observation c2bb9469-7334-4d1b-bdfe-745d0885e127 · outbound

This paper cites UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:39.780236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:39.780236Z digest=sha256:57ec9a5a14176e182ee8d0080b3cc83a8f039ecfffadba1e0340c534fa79fddd

Observation 0cf22c64-cdd8-4834-9d4f-d56f4bd303b9 · outbound

This paper cites Masked motion encoding for self-supervised video representation learning.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Masked motion encoding for self-supervised video representation learning

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.347994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:17:39.783726Z digest=sha256:1b52cc52bb763394cee1379df91dee054d7f189f215e1bc39c4c67af0add9d2d

Observation 45d830d2-4e75-4bdc-a2a2-8feb0d6a47a2 · outbound

This paper cites Rethinking the inception architecture for computer vision.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Rethinking the inception architecture for computer vision

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.335542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:17:39.787937Z digest=sha256:c580b6c7babc6e7b6b1e10d30a452d14ca856c2e8caa66fa291efedb79c8874e

Observation 090a7cd6-6cc6-48d3-9c96-edda12c8ba7e · outbound

This paper cites VIMPAC: Video Pre-Training via Masked Token Prediction and Contrastive Learning.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders VIMPAC: Video Pre-Training via Masked Token Prediction and Contrastive Learning

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:39.791309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:39.791309Z digest=sha256:8db3d4289645a76e8429a4189523b3914cefc1e952c08ee50c652d7f558ed74b

Observation 33997c7c-e833-473b-8077-43da076bacfd · outbound

This paper cites Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.324599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:17:39.794882Z digest=sha256:ab19a27555320f4039d3d781eeafcadac3b124f2273a2c164000e22a8df3039f

Observation ec8c52c4-5395-4d2d-80de-931e41f183ef · outbound

This paper cites Learning spatiotemporal fea- tures with 3d convolutional networks.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Learning spatiotemporal fea- tures with 3d convolutional networks

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.313412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:17:39.798379Z digest=sha256:4dae0c7e1cdae35e486c4469740618927a5f001cc31d1a8d3bc5e90b516d7480

Observation 40feb0f8-378a-4761-a06b-4345b2024495 · outbound

This paper cites Video classification with channel-separated convolutional networks.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Video classification with channel-separated convolutional networks

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.303081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:17:39.801539Z digest=sha256:b106a9ef957873df5a466105475a3c8f298579cb25b7a634a7decab69df78acb

Observation 1a1311ce-37c9-4453-ac70-7e0d9043c430 · outbound

This paper cites Spa- tiotemporal integration and object perception in infancy: Perceiving unity versus form.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Spa- tiotemporal integration and object perception in infancy: Perceiving unity versus form

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.292395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:17:39.804641Z digest=sha256:01c0c0c054b0635dc2385ba5edd8ee56a4dc2c045ff157e70643e1a866af87f9

Observation a09874af-8021-4a54-9bc8-254dfdd7e0a9 · outbound

This paper cites Generating videos with scene dynamics.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Generating videos with scene dynamics

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.281255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:17:39.807748Z digest=sha256:a545e7c288a1d13fa0bfa9edecab32a9ef691c20487777d2e6edcfd0d6148b03

Observation 4c8b68c1-5415-4cc2-ad1b-714acea31bb0 · outbound

This paper cites Unsupervised visual rep- resentation learning by tracking patches in video.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Unsupervised visual rep- resentation learning by tracking patches in video

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.270594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:17:39.810778Z digest=sha256:8966b42b338faf4315584a63007812f729aa929c2bd9ee491781fe8b22fb2335

Observation 1dc366f0-d074-4d3d-8f04-608407d27283 · outbound

This paper cites Self- supervised video representation learning by pace predic- tion.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Self- supervised video representation learning by pace predic- tion

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.259895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:17:39.814884Z digest=sha256:58201c0d1ad597f327223c57c74ccf4ad23eed3b415f1b7499a4b399966910eb

Observation ccc4a9ca-1992-4714-8caf-4a8ddcac3375 · outbound

This paper cites Self-supervised Temporal Discriminative Learning for Video Representation Learning.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Self-supervised Temporal Discriminative Learning for Video Representation Learning

Reference 97

Resolution
verified exact
local_arxiv, observed 2026-08-08T19:17:39.950848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:17:39.818340Z digest=sha256:aa02e0375efc34d59e64996b3d1e973f3d9831273514fea8d715a252748cf41d

Observation d6d43169-c1ea-4ac8-9025-d1a943d08a96 · outbound

This paper cites Temporal segment networks for action recognition in videos.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Temporal segment networks for action recognition in videos

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.249145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:17:39.822225Z digest=sha256:f9511b4c3ac5662ddc428ebe5968d940e5e0303db93014ae2ac660926da3c80f

Observation 3aa3cc15-d187-430b-8e3b-d61f85b8e5b8 · outbound

This paper cites Tdn: Temporal difference networks for efficient action recog- nition.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Tdn: Temporal difference networks for efficient action recog- nition

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.239461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:17:39.825930Z digest=sha256:eacf722f2e91029895a677a3fbc5d25d331a5540f07fe503a48aaf0b5735fc0e

Observation 6e06deb0-1d76-4d4b-9eae-ce9c0bfb6918 · outbound

This paper cites Videomae v2: Scaling video masked autoencoders with dual masking.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Videomae v2: Scaling video masked autoencoders with dual masking

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.229038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:17:39.829689Z digest=sha256:810bf0c62bb3ec3d845235e7f6abdc738b45ee970a2c0474c7675e2ce22d6645

Pith citing papers

Observation ff0a1ede-3856-4a74-ad1e-e3544cd0a8b4 · inbound

Infinite Video Understanding cites this paper.

Infinite Video Understanding CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:10:16.731929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:09:18.061419Z digest=sha256:6e94cd5b7e7e8f37b642528a9f57abe50f8c46b2b9260d9306523e6f95240311