Pith. sign in

Paper Citation Record · LEDGER

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders

As of 14 August 2026, this Paper Citation Record lists 100 of 111 outbound references and 1 inbound Pith citation observation for arXiv:2502.07811.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.07811 v1

Coverage vector

measured 100 of 111 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T19:17:39.829689Z

measured 101 of 101 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:09:18.061419Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-10T05:30:23.456663Z

Reference resolution

100 of 111 outbound references displayed

  • verified exact5
  • verified fuzzy54
  • unresolved41
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1
pith, observed 2026-08-10T05:30:23.456663Z

Outbound references

Observation aa31f2e1-7314-4142-8a0e-8c9e1ca4047e · outbound

This paper cites Self-supervised multimodal versatile networks.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Self-supervised multimodal versatile networks

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:39.436477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:39.436477Z digest=sha256:38d06f76b5d11e791f14bd8ba58cfc4ffefc0115410ed6981aa53ab529294b71

Observation 5c6b4a85-fdfa-49f7-8089-cc2eca4b9ba8 · outbound

This paper cites Self-supervised learning by cross-modal audio-video clustering.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Self-supervised learning by cross-modal audio-video clustering

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:39.441116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:39.441116Z digest=sha256:b51b60fc75574469c871bf5f57da814442fbbd485ba7e39a1d1305a68176fbdf

Observation 514f8041-bb2a-4464-bd57-0a3a6b5da576 · outbound

This paper cites Look, listen and learn.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Look, listen and learn

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:39.445250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:39.445250Z digest=sha256:428c2d81f0961d80c415e4de8a70a72e97403d729e7a88c1057ca9ec554e32b8

Observation a8586362-2712-4b57-8c07-76dd79501518 · outbound

This paper cites Objects that sound.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Objects that sound

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:39.449872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:39.449872Z digest=sha256:65f2267c11915ece07d03ec62b934a5cea8aa180084e850c26dc5af5840d1d7a

Observation 5f683016-7c5c-4ee6-9ad5-a37dcad2619e · outbound

This paper cites Vivit: A video vision transformer.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Vivit: A video vision transformer

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:39.453863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:39.453863Z digest=sha256:58ac75f5837049c68f85f6f7553be54e681066fddb16d666cc71a7eb6c3bfba8

Observation 417c24bb-c924-400d-b216-af913765b5b8 · outbound

This paper cites Adamae: Adaptive masking for efficient spatiotempo- ral learning with masked autoencoders.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Adamae: Adaptive masking for efficient spatiotempo- ral learning with masked autoencoders

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:39.457906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:39.457906Z digest=sha256:3341736fad25c6e3a635ef97d80ce8aa673d68eeaa58726e735658886348e943

Observation e49641c3-d57b-4f0a-9d1b-5c40625657aa · outbound

This paper cites BEit: BERT pre-training of image transformers.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders BEit: BERT pre-training of image transformers

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:39.462043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:39.462043Z digest=sha256:fb8765c90c8a46e2ca4ec143818b5789f0029e50d20a328995b490e2c448bc41

Observation 6611a668-d9f4-4c85-a89b-53f609f485a1 · outbound

This paper cites Speednet: Learning the speediness in videos.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Speednet: Learning the speediness in videos

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:39.466647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:39.466647Z digest=sha256:cc35811821ba3f20955af4b0a3002171daf04873b9c891349d7e5c538cf7cf24

Observation 0b2949be-ce36-4de2-b3ea-b842c2f58b5e · outbound

This paper cites Is space-time attention all you need for video understanding? In Proceedings of the 38th International Conference on Ma- chine Learning (ICML), pages 813–824.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Is space-time attention all you need for video understanding? In Proceedings of the 38th International Conference on Ma- chine Learning (ICML), pages 813–824

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:39.470804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:39.470804Z digest=sha256:67386be2321a3cef8c8ea0d28883801d2be6619d247e96183d10bc2f2eaca31b

Observation 6668cefe-edd9-4bd6-9124-7a21f1422748 · outbound

This paper cites Emerging properties in self-supervised vision transformers.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Emerging properties in self-supervised vision transformers

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:39.475016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:39.475016Z digest=sha256:9551f33080aae1ddd42005a5b6e9424843c4baca467484149523e134f9527d67

Observation 65788d6c-4189-41e8-8422-35b2a69464e0 · outbound

This paper cites Learning aligned cross-modal representations from weakly aligned data.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Learning aligned cross-modal representations from weakly aligned data

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:39.479038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:39.479038Z digest=sha256:374c7e86356341aaaaf0d3b7ae400da38d063e5afabdfc1e0db9379ac249de1b

Observation 17edde5c-77e6-4dd2-ba52-2736aa6f6c61 · outbound

This paper cites Generative pre- training from pixels.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Generative pre- training from pixels

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:39.483141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:39.483141Z digest=sha256:d5a357caacf874bca8b8720268821ceaf8979dffce94d867d3b95e1bc20ffee4

Observation 5539f8f5-1b13-4a00-9020-c74c0b3e8240 · outbound

This paper cites Rspnet: Relative speed perception for unsupervised video representation learning.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Rspnet: Relative speed perception for unsupervised video representation learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:39.487073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:39.487073Z digest=sha256:f03c34ea72f4b601229e73de6fa6ccb81f71fd613cb04a9d5067e44e7a31ecdb

Observation 75f840e0-2139-4217-8264-8cf28e8a5d9e · outbound

This paper cites A simple framework for contrastive learning of visual representations.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders A simple framework for contrastive learning of visual representations

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:39.490967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:39.490967Z digest=sha256:d42a99b40dc507976c1d26ac2a2768b779e7fa623c000b01a59ea2b1982809e4

Observation f3e82289-834d-4ddd-ae23-8b5524d2c663 · outbound

This paper cites Electra: Pre-training text encoders as dis- criminators rather than generators.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Electra: Pre-training text encoders as dis- criminators rather than generators

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:39.495458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:39.495458Z digest=sha256:0bc28e2f8d608e9dbc9b3004b68140fd0a0324007dfb812e1c63be78b65bd148

Observation 959bdc18-f39f-4383-bf50-293dafaa4dd7 · outbound

This paper cites Randaugment: Practical automated data augmen- tation with a reduced search space.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Randaugment: Practical automated data augmen- tation with a reduced search space

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:39.499741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:39.499741Z digest=sha256:008ad0d16512c92acb03458001649599aa424453ea27ce70bc892357b73c111e

Observation 155e3fce-1ef4-475f-8c7a-dd0a7ab896b7 · outbound

This paper cites Unifying video self-supervised learning across families of tasks: A survey.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Unifying video self-supervised learning across families of tasks: A survey

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:39.503662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:39.503662Z digest=sha256:2ed4b1650e68382ca586be081ce48f26b28b30670f2777917b77c98bffd1b50c

Observation ea42b726-5ceb-4e61-8f21-4b2dec2cd59f · outbound

This paper cites Imagenet: A large-scale hierarchical im- age database.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Imagenet: A large-scale hierarchical im- age database

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:39.507721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:39.507721Z digest=sha256:54921ffda1964f456248853ca21b5ca69766c69072b66f857237ed0bf9f22c87

Observation 312f0708-8db9-4eb9-857c-05931df10a21 · outbound

This paper cites Virtex: Learning visual representations from textual annotations.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Virtex: Learning visual representations from textual annotations

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:39.512352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:39.512352Z digest=sha256:15e34efc3482297a75880f6d17fc10768e4a4cdd117d5d675cccbbc67d2f8e28

Observation 139a9f0e-f182-4047-a18f-4ec727dfb35e · outbound

This paper cites 9 Vi2clr: Video and image for visual contrastive learning of representation.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders 9 Vi2clr: Video and image for visual contrastive learning of representation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:39.516832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:39.516832Z digest=sha256:fddd3f429d37f5236f0307ff4c241bbe97ba0b80e63015594da2a0d7dbb3634a

Observation 43e118a0-44e5-4587-aef1-6d255efc6c90 · outbound

This paper cites Taming transformers for high-resolution image synthesis.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Taming transformers for high-resolution image synthesis

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:39.520528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:39.520528Z digest=sha256:62eb5dbeff03596807091c9059a6fe1e2d11bc1f3bf1be2cf738536eb3898129

Observation 20cab6dd-82ee-4f9d-9672-fc366b2f2c78 · outbound

This paper cites Multiscale vision transformers.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Multiscale vision transformers

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:39.524439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:39.524439Z digest=sha256:c24bb331818b48c297d6e81edd733052ce74473aefd74dbab581c3157967f919

Observation 96353754-9cec-44fc-a127-806e7f99f547 · outbound

This paper cites Slowfast networks for video recognition.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Slowfast networks for video recognition

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:39.528415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:39.528415Z digest=sha256:648c685fa25f96a119f35776f4ea52f8bf0e43f4a5a723de87f8dc97ac848f60

Observation 0bcce743-4f81-45bd-bdd4-f3b969a60443 · outbound

This paper cites A large-scale study on unsuper- vised spatiotemporal representation learning.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders A large-scale study on unsuper- vised spatiotemporal representation learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:39.532554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:39.532554Z digest=sha256:27ada2c8fd2b7e3fdf532a79dd189c6cf2b70c9336277889bce9a876f2792141

Observation 901d0767-fcc8-42e2-ac84-f82b43be35f0 · outbound

This paper cites Masked autoencoders as spatiotemporal learners.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Masked autoencoders as spatiotemporal learners

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:39.537655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:39.537655Z digest=sha256:b16e43a4745be52d7cf4e6e74fa4229a0a5bab259f1fe736ee67623cb29d0fab

Observation 63f6a578-731e-49e2-8427-51835d2d17e4 · outbound

This paper cites Mcmae: Masked convolution meets masked autoencoders.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Mcmae: Masked convolution meets masked autoencoders

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:39.542371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:39.542371Z digest=sha256:869ed9aeaf4cffab690f5bf61d437bb3e56f27cd462f5ccd039ca9a12dce7c1a

Observation 8466d3cb-f8e2-40a2-b371-f9372e09c879 · outbound

This paper cites Revitalizing cnn attention via transformers in self-supervised visual representation learn- ing.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Revitalizing cnn attention via transformers in self-supervised visual representation learn- ing

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:39.546400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:39.546400Z digest=sha256:dc1964613cab51b17eb346325c5868dd4093e48f4eb8feb9423f8b02925c2293

Observation 902a23cd-ffc7-42aa-83fd-437ffbfaaf93 · outbound

This paper cites Omnimae: Single model masked pretraining on images and videos.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Omnimae: Single model masked pretraining on images and videos

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:39.549998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:39.549998Z digest=sha256:ae046dca1c7dd4f5390bf6c721e2da96845890feb2d14b0315986108eae2b6b5

Observation cd0451dc-47d5-41b0-b9eb-34aed2a78ebc · outbound

This paper cites Improving image-sentence embeddings using large weakly annotated photo collec- tions.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Improving image-sentence embeddings using large weakly annotated photo collec- tions

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:39.553347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:39.553347Z digest=sha256:a6dcf9f5b3de720122f704fe274a48b5d2c557210ca2af01575b32537ca98c47

Observation 00ed9763-81a6-48b8-a699-eb1a39e57c45 · outbound

This paper cites Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:39.557394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:39.557394Z digest=sha256:836a0233ddddda79ec3b296afec86704d1f564e7f1a88f7459d7add7db78c119

Observation 226ba7d6-3b09-4de9-ad54-5722c4e210c3 · outbound

This paper cites The” something something” video database for learning and evaluating visual common sense.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders The” something something” video database for learning and evaluating visual common sense

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:39.561828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:39.561828Z digest=sha256:ea70eac662dc770d46ca6c0169500a6c3d79547b7a84b82fe928b167635ab7bd

Observation a1541263-211e-48db-aaec-9b861248f306 · outbound

This paper cites Bootstrap your own latent-a new approach to self-supervised learning.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Bootstrap your own latent-a new approach to self-supervised learning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:39.564924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:39.564924Z digest=sha256:eaf29237e805006f3f6ad1a5ab533e7cb14aba0baca37eee32ea4239a7a7ee42

Observation 191c8262-b90c-4e24-bd6d-88c7ebdf13a6 · outbound

This paper cites How Effective are Self-Supervised Models for Contact Identification in Videos.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders How Effective are Self-Supervised Models for Contact Identification in Videos

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-08-08T19:17:40.109081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:17:39.568572Z digest=sha256:cac0267f9b7f21e5ae4790c46c4ad1af4cf7a8ca089816dca0d3e75caa8b16e8

Observation f230e108-8784-43cc-9750-f5125c4111f8 · outbound

This paper cites MaskViT: Masked Visual Pre-Training for Video Prediction.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders MaskViT: Masked Visual Pre-Training for Video Prediction

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:39.573335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:39.573335Z digest=sha256:246afcfeebd1d143d4207277811d4ddbb940644902e17b746035ac5a6398a635

Observation 43e74f4b-1dca-4739-924b-b2392e12b043 · outbound

This paper cites Memory- augmented dense predictive coding for video representa- tion learning.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Memory- augmented dense predictive coding for video representa- tion learning

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.823654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:17:39.577840Z digest=sha256:f1bb4237bda4fc2ffe353885b416719bd4943ad3824d98fc8b8c01c58a23f4a9

Observation 3d180f1e-242b-40fb-ac7a-1ba0012db08a · outbound

This paper cites Self- supervised co-training for video representation learning.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Self- supervised co-training for video representation learning

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.813186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:17:39.581200Z digest=sha256:9c6f95215a3375897d6098bb68358209ad9ee32e16b60211b48416dc54fca5af

Observation 85459afd-c541-4341-82f4-2a7c78a58e0a · outbound

This paper cites Momentum contrast for unsupervised visual rep- resentation learning.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Momentum contrast for unsupervised visual rep- resentation learning

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.802759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:17:39.584842Z digest=sha256:7d92f096571558ca1ae076f5fbe4d8978369ad99c0482405cf46d1bfab0b7e1d

Observation 04165458-3c61-40e2-a83c-6295736d3104 · outbound

This paper cites Masked autoencoders are scal- able vision learners.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Masked autoencoders are scal- able vision learners

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.791735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:17:39.588393Z digest=sha256:33a371782488a84e1585c4e02eca0e87266c4fd73ad9ca4df914eb6aed35acec

Observation 0de94c77-96c4-45f6-9dbf-983a67480395 · outbound

This paper cites Human gaze control during real-world scene perception.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Human gaze control during real-world scene perception

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.780607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:17:39.591639Z digest=sha256:bc16ee60f7bd7426c56ed7053be55e8ff7739fba9208c01affea9af973b49a1f

Observation 68eaaf99-ec76-43fc-a872-bc7721e897d7 · outbound

This paper cites ViC-MAE: Self-Supervised Representation Learning from Images and Video with Contrastive Masked Autoencoders.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders ViC-MAE: Self-Supervised Representation Learning from Images and Video with Contrastive Masked Autoencoders

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-08-08T19:17:40.087326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:17:39.594919Z digest=sha256:fbc96f3aa2e5de83b49f69fddc7571c767e40fb6d223f818bc820fcc2e97a206

Observation f3fe09b1-b083-491d-bedb-dc09a8ea5e3a · outbound

This paper cites Augment your batch: Improv- ing generalization through instance repetition.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Augment your batch: Improv- ing generalization through instance repetition

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.770627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:17:39.598583Z digest=sha256:f0386bc3aaa77dc015c7ec471fc4c90bf715bc9933365662acdeaf91e5e383c5

Observation b906e515-298f-40f0-bbf5-b0ce88e3aab6 · outbound

This paper cites Contrast and order representa- tions for video self-supervised learning.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Contrast and order representa- tions for video self-supervised learning

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.761185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:17:39.601608Z digest=sha256:d0ae87464e313692367ffd9398e2cf90e5325050a440b6f974f2e8c4c4f9163b

Observation 6e6c79f1-2abf-4621-b7f9-5d30f07241d9 · outbound

This paper cites Mgmae: Motion guided masking for video masked autoencoding.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Mgmae: Motion guided masking for video masked autoencoding

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.750253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:17:39.605271Z digest=sha256:034cb1e8d0f507c80a4997c5682dcc5c3f37bb7695c5655fbdba48ef9c14243c

Observation fe7d86f1-dd41-43b5-a621-217655a7b909 · outbound

This paper cites Self-supervised video representation learn- ing by context and motion decoupling.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Self-supervised video representation learn- ing by context and motion decoupling

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.739266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:17:39.608840Z digest=sha256:414fa7bba3754069176fd2137ea1119dd591f13e6af4e3cf14e9b926b6145b74

Observation 738c14af-7d8d-4dc7-b1ec-4b44aeab3eed · outbound

This paper cites TAda! Temporally-Adaptive Convolutions for Video Understanding.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders TAda! Temporally-Adaptive Convolutions for Video Understanding

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:39.612598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:39.612598Z digest=sha256:8f86a0194386ebb9a1edde6af11f578a29bfcebf81f0a857b3ca3f9f576ee286

Observation 871e1068-622d-4070-998b-01da2f9af763 · outbound

This paper cites Self-Supervised Spatiotemporal Feature Learning via Video Rotation Prediction.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Self-Supervised Spatiotemporal Feature Learning via Video Rotation Prediction

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:39.616909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:39.616909Z digest=sha256:bc3e06e33a96543e42210b2bbb7cb7becb198c66f33e6c012b32625bb35d81d7

Observation 4cc14dd9-12dc-4812-83e2-50e587072d44 · outbound

This paper cites Hard negative mixing for contrastive learning.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Hard negative mixing for contrastive learning

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.729598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:17:39.621618Z digest=sha256:b1eaba12fe1bce81f61e06af184aad1bc2b50e0a681b0bc0e23c8a4664b6bc85

Observation d283f6db-8757-4960-b160-66b335ad927e · outbound

This paper cites Deep visual-semantic alignments for generating image descriptions.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Deep visual-semantic alignments for generating image descriptions

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.719841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:17:39.625420Z digest=sha256:ae7f63e682c0ce535510005d7991013a8b41b1b87e92a22af2518b77941f0b43

Observation c873d323-5d1e-45bd-805a-8dcc16c1fcc4 · outbound

This paper cites The Kinetics Human Action Video Dataset.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders The Kinetics Human Action Video Dataset

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:39.629499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:39.629499Z digest=sha256:bbdaa821625fcbbca4d090e306fc940f07f2f25e5e68aca4b7dd3ebf9ac5e889

Observation 5baf0a28-b25a-4b8c-9b71-3cfbecb9a128 · outbound

This paper cites Fusion attention for action recognition: Integrating sparse-dense and global at- tention for video action recognition.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Fusion attention for action recognition: Integrating sparse-dense and global at- tention for video action recognition

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.711257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:17:39.634738Z digest=sha256:caffa7d4b017b6b01af11935a7e68314531933bad3038f6da8891babd9e90c66

Observation 9813b497-3a07-4597-a37d-0bf3bf129886 · outbound

This paper cites Co- operative learning of audio and video models from self- supervised synchronization.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Co- operative learning of audio and video models from self- supervised synchronization

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.701974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:17:39.641055Z digest=sha256:7cf288ef0ccfb4caae4e44ad339b9cc7c38c977af9421c734f78d26eb9c947bd

Observation 500f507f-515d-4222-b5a5-7434d5fd9663 · outbound

This paper cites Hmdb: a large video database for human motion recognition.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Hmdb: a large video database for human motion recognition

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.692379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:17:39.645261Z digest=sha256:bccbde3394c87d74cde5811438815ffa366cc8b6bbe711fb78e805f580d93022

Observation 7e6406a4-b4dd-4641-9d63-875b5c9f1d29 · outbound

This paper cites A large-scale analysis on self- supervised video representation learning.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders A large-scale analysis on self- supervised video representation learning

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.682078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:17:39.649317Z digest=sha256:4a09ff7c3edf6ad97007f808f6fd0bdce110d9c9f09ec0a16c06883e2d668aef

Observation 0799014a-b89c-478a-a0c5-1b1021f610dc · outbound

This paper cites Semmae: Semantic-guided masking for learning masked autoencoders.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Semmae: Semantic-guided masking for learning masked autoencoders

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.671868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:17:39.653578Z digest=sha256:8667edb142c6337cd8fb06c3961123ecfef983f584148c6b537ce1e9f0d86515

Observation 4bd49512-7c46-4fb3-b6b2-8f2c829f6e61 · outbound

This paper cites Improve unsuper- vised pretraining for few-label transfer.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Improve unsuper- vised pretraining for few-label transfer

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.661450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:17:39.657779Z digest=sha256:8ab2fa789ed8b1495c444cd761db09d759b231547482843016fd39f5166dbbd4

Observation b6fa720b-5e50-44cb-8783-ce9ae101ed0f · outbound

This paper cites Learning Spatiotemporal Features via Video and Text Pair Discrimination.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Learning Spatiotemporal Features via Video and Text Pair Discrimination

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:39.665769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:39.665769Z digest=sha256:009de3bcc329dbeb8358ce32ffd303309e688cea481ada36a3383c2672b9a753

Observation dd473744-3a29-415c-9c46-0d7152670e70 · outbound

This paper cites Tsm: Temporal shift module for efficient video understanding.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Tsm: Temporal shift module for efficient video understanding

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.650601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:17:39.670766Z digest=sha256:9d147da9a6b3ca633f0ed83f388bb28f814226564912418f99d37b959e82807e

Observation b0003748-d48d-4714-b611-efbcedd98d0b · outbound

This paper cites Self-supervised video-based action recognition with distur- bances.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Self-supervised video-based action recognition with distur- bances

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.640547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:17:39.674852Z digest=sha256:b10bb5198a3b6298bb47c9edf3ce9ab33c2222914d6fd439ac82854cbbe287d7

Observation 5d93b97f-445d-4dba-83ad-edd6c0a371c7 · outbound

This paper cites CrossVideo: Self-supervised Cross-modal Contrastive Learning for Point Cloud Video Understanding.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders CrossVideo: Self-supervised Cross-modal Contrastive Learning for Point Cloud Video Understanding

Reference 59

Resolution
verified exact
local_arxiv, observed 2026-08-08T19:17:40.026614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:17:39.679785Z digest=sha256:eda272fc513ff6d4eb17fb4092f8044ccba9a3680d9ba2a0838a241cc7373f49

Observation f63ef84d-b5e8-45f5-84ff-585c00b310ff · outbound

This paper cites Teinet: Towards an efficient architecture for video recognition.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Teinet: Towards an efficient architecture for video recognition

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.629344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:17:39.683826Z digest=sha256:bca58f2558dc62eff5fc1ec044c2a7649dbd375d46433195032423abf4f8e70b

Observation 792fa445-64f4-4b00-8df2-15369d8e8ddd · outbound

This paper cites Tam: Temporal adaptive module for video recog- nition.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Tam: Temporal adaptive module for video recog- nition

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.618146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:17:39.687932Z digest=sha256:1b105854c4c1371417704a0d3d6ddf07c7de2e14ed9be6f691ac2a45835ccf38

Observation 03ae3348-07ab-4823-9671-0e5c13c61487 · outbound

This paper cites Video swin transformer.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Video swin transformer

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.606934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:17:39.692108Z digest=sha256:a4c9e5c87cc454dbcf47c0f6a446b7a2820713987f9a23c6d7a0dfac2a558fee

Observation 4ad99798-c765-4d09-ba86-c350fc00c566 · outbound

This paper cites SGDR: Stochastic Gradient Descent with Warm Restarts.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders SGDR: Stochastic Gradient Descent with Warm Restarts

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:39.696217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:39.696217Z digest=sha256:65447a6f9709f071e58ab6dd4ffdf1646a36e3bad20608a4b4bebfdfe78f49ab

Observation 9d1f7e83-dfba-417d-abda-a809af85e874 · outbound

This paper cites Decoupled weight de- cay regularization.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Decoupled weight de- cay regularization

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.597155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:17:39.700346Z digest=sha256:a2560e18256e21e5f84dbeb2234a809c8d83bf23d618b1697f5f48a2d568d6fb

Observation fbe91960-dfe9-40cf-9114-eba79172bcb1 · outbound

This paper cites CMAE-V: Contrastive Masked Autoencoders for Video Action Recognition.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders CMAE-V: Contrastive Masked Autoencoders for Video Action Recognition

Reference 65

Resolution
verified exact
local_arxiv, observed 2026-08-08T19:17:39.991655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:17:39.703294Z digest=sha256:9737e2936463c149072173fd5f9310e854536637a45c65066b86c5262d26cf1f

Observation e280b4a2-9b2c-4658-ac55-94cd84302cda · outbound

This paper cites 12-in-1: Multi-task vision and language representation learning.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders 12-in-1: Multi-task vision and language representation learning

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.586283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:17:39.707292Z digest=sha256:773cabd6a112e1f5cf82d00eb82d905e5eeef3a21dd522266f060f9fc9293a46

Observation c2c981d9-c97d-41a6-96a6-d97c93d2512d · outbound

This paper cites End-to-end learning of visual representations from uncurated instruc- tional videos.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders End-to-end learning of visual representations from uncurated instruc- tional videos

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.575610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:17:39.711733Z digest=sha256:c22ac7cd95a681d2eb0a759896cddc8b6b1a61eb4269430fbc0d1adb8cd69f1d

Observation 084a12c4-3429-471b-bfa8-7986637c7087 · outbound

This paper cites Self-supervised learning of pretext-invariant representations.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Self-supervised learning of pretext-invariant representations

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.563057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:17:39.715441Z digest=sha256:d55e4c287330c74ab75c59c6701d92c2a9c57f2ed96ba9af74a38d7c5366d3e8

Observation 13389bba-7981-4c78-ab3a-5b9f71333ee5 · outbound

This paper cites Shuffle and learn: unsupervised learning using temporal order verification.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Shuffle and learn: unsupervised learning using temporal order verification

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.551556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:17:39.719271Z digest=sha256:b9ef30a0e757312f4fbaaf65f90f35a4fc997ca84ee82f56c651d9c2195176b5

Observation 72dca34f-316f-4a03-b152-e32f1f507d0d · outbound

This paper cites Ro- bust audio-visual instance discrimination.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Ro- bust audio-visual instance discrimination

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.539755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:17:39.722912Z digest=sha256:def883c1647d277da991ddf7260ca17bf1fc22344f73783ce63b3dd8687b56e1

Observation 8b33199a-aaee-47d7-a634-907428509832 · outbound

This paper cites Audio-visual instance discrimination with cross-modal agreement.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Audio-visual instance discrimination with cross-modal agreement

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.527661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:17:39.726729Z digest=sha256:3d20196ae5cd1f8f7549720ba184c5b3b64ce6d38375750834c8c70718324110

Observation 7c5b6ad5-ee64-4192-b19c-21661f1d02e8 · outbound

This paper cites Audio-visual scene analysis with self-supervised multisensory features.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Audio-visual scene analysis with self-supervised multisensory features

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:39.730563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:39.730563Z digest=sha256:68c43ec58b10af6c868398a5814416adf675afac18e0e9d41dddba10ee704121

Observation c739f4a5-045d-40b1-9468-201f92560579 · outbound

This paper cites Videomoco: Contrastive video representation learning with temporally adversarial examples.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Videomoco: Contrastive video representation learning with temporally adversarial examples

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.505976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:17:39.733882Z digest=sha256:26f201ef8fe77afe60f1c3f21df09f7a0505c7d933924e7a64ad0692b02ad066

Observation a7a40403-6a6d-4c2e-8fdf-d059a87425f1 · outbound

This paper cites Multi-modal self-supervision from generalized data transformations.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Multi-modal self-supervision from generalized data transformations

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.496205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:17:39.737476Z digest=sha256:ad1d33c99da407f8d76e586fecafba0b26199089322a8436c75715ed9d05fa30

Observation 7cb4e5ca-9546-4928-ba59-a484f56495c9 · outbound

This paper cites Keeping your eye on the ball: Trajectory attention in video transformers.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Keeping your eye on the ball: Trajectory attention in video transformers

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.484561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:17:39.742292Z digest=sha256:df731dde946cbeff647a5b9118de5d924154590739af6935834f16f43b8124e2

Observation 59e4f32a-76b0-4639-861e-d44248bd2633 · outbound

This paper cites Evolving losses for unsupervised video representation learning.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Evolving losses for unsupervised video representation learning

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.473778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:17:39.745476Z digest=sha256:0bdc688b3817b85995c2d8bd7938157b51f7dac25633f1f0681f493eb5bc28cf

Observation b3be13be-da0b-4586-bcdd-c7d238515bd0 · outbound

This paper cites Spa- tiotemporal contrastive video representation learning.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Spa- tiotemporal contrastive video representation learning

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.459220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:17:39.749310Z digest=sha256:2029bc306ecd2bdd92d70e5e0bb49f63c38b36964be249fb335e2523e42b9c50

Observation 2e6d248a-812e-4f05-b38a-3b4fec0a1a50 · outbound

This paper cites Learning from untrimmed videos: Self-supervised video representation learning with hierarchical consistency.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Learning from untrimmed videos: Self-supervised video representation learning with hierarchical consistency

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.445236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:17:39.752423Z digest=sha256:d778de4fbe1e46d651a91a0a4c510703f90aaec29dc43b848f32fc1d369c537c

Observation d8fa6935-4bef-4384-aef3-57926f397d74 · outbound

This paper cites Mar: Masked autoencoders for efficient action recognition.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Mar: Masked autoencoders for efficient action recognition

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.432594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:17:39.755665Z digest=sha256:d13384d34c6555a63adf0e6d1f59c73af0168eae1b85fe458321261a46d6d403

Observation 05e5e519-21f4-40e6-a1a8-789046110996 · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Learn- ing transferable visual models from natural language super- vision

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.420141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:17:39.759382Z digest=sha256:34c1dfe33b4919515884600f4e6b8b1ebfe42b0afb05b5fe1edf0923e9cf8b99

Observation adf4fcdb-1454-4f16-acf8-4b9dca135e39 · outbound

This paper cites Self-supervised video transformer.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Self-supervised video transformer

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.405711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:17:39.762718Z digest=sha256:a8947965170ef4b28010ab1efb1339bfbb7a768fb3ce55469447f5a89818a9c5

Observation f272e378-e5b7-42cc-ba3c-20e64d8078f4 · outbound

This paper cites The development of visual attention and the brain.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders The development of visual attention and the brain

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.394787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:17:39.766117Z digest=sha256:5a0e90ec242cd4c44c6876994da62fb2267233f1a1170187d80b0f3eee9673d6

Observation ac81765d-19a8-4a6f-82ed-ebc0f7a27dd9 · outbound

This paper cites Im- agenet large scale visual recognition challenge.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Im- agenet large scale visual recognition challenge

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.382306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:17:39.769664Z digest=sha256:eb4a0a317629f44e2e43dac771d1ac2a290f6541cc699d325e96028c594c8fbd

Observation aacb1bdb-4c33-43d1-8c50-b1e8c18e2730 · outbound

This paper cites Learning visual representations with caption annotations.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Learning visual representations with caption annotations

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.370700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:17:39.773677Z digest=sha256:2250efe79f4462f52062c65415d9dea31ffaa0bfe9669f0b9588b36fbc663414

Observation 57cc4b17-a994-4024-b402-77c003c9f91a · outbound

This paper cites Two-stream con- volutional networks for action recognition in videos.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Two-stream con- volutional networks for action recognition in videos

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.359348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:17:39.777062Z digest=sha256:2c6ff51efb67b127145efc219c475845008db0e66813b7bb8fd49391fb5afe2a

Observation c2bb9469-7334-4d1b-bdfe-745d0885e127 · outbound

This paper cites UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:39.780236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:39.780236Z digest=sha256:041727505b498b20af1289d2537580881f35c83925505113f010f40ba9ab18db

Observation 0cf22c64-cdd8-4834-9d4f-d56f4bd303b9 · outbound

This paper cites Masked motion encoding for self-supervised video representation learning.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Masked motion encoding for self-supervised video representation learning

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.347994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:17:39.783726Z digest=sha256:88fbc688337c9e84a1882f8a44c055f8528f94fca8199023514587bae9653f30

Observation 45d830d2-4e75-4bdc-a2a2-8feb0d6a47a2 · outbound

This paper cites Rethinking the inception architecture for computer vision.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Rethinking the inception architecture for computer vision

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.335542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:17:39.787937Z digest=sha256:53e3fe254ffd42f2a65a2caab1d79e25c83be3f61adcd287680eaf1c90e8bf0f

Observation 090a7cd6-6cc6-48d3-9c96-edda12c8ba7e · outbound

This paper cites VIMPAC: Video Pre-Training via Masked Token Prediction and Contrastive Learning.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders VIMPAC: Video Pre-Training via Masked Token Prediction and Contrastive Learning

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:39.791309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:39.791309Z digest=sha256:a3e9c7a50d1009a2a758061f567bcdee556c7b15f54c661bd660790bc3ade9bd

Observation 33997c7c-e833-473b-8077-43da076bacfd · outbound

This paper cites Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.324599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:17:39.794882Z digest=sha256:79b12632cf79b4a176c2c76ac6d85cdfda15dc6fc33b4e4a19583543cc491949

Observation ec8c52c4-5395-4d2d-80de-931e41f183ef · outbound

This paper cites Learning spatiotemporal fea- tures with 3d convolutional networks.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Learning spatiotemporal fea- tures with 3d convolutional networks

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.313412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:17:39.798379Z digest=sha256:48bcce2d59712e34029a3779d91759814f0e641e5ef745dc376a1ddb9fabf922

Observation 40feb0f8-378a-4761-a06b-4345b2024495 · outbound

This paper cites Video classification with channel-separated convolutional networks.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Video classification with channel-separated convolutional networks

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.303081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:17:39.801539Z digest=sha256:1feedc75256a5f9884142e870d991421b8293cc104db45be84ff8ae595eed98f

Observation 1a1311ce-37c9-4453-ac70-7e0d9043c430 · outbound

This paper cites Spa- tiotemporal integration and object perception in infancy: Perceiving unity versus form.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Spa- tiotemporal integration and object perception in infancy: Perceiving unity versus form

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.292395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:17:39.804641Z digest=sha256:318773f1dac8835c84e9cf0000c518aee6ab8414cf451110509b70c6a84c7483

Observation a09874af-8021-4a54-9bc8-254dfdd7e0a9 · outbound

This paper cites Generating videos with scene dynamics.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Generating videos with scene dynamics

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.281255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:17:39.807748Z digest=sha256:551276e9b5ab1e36d19d8008de48284305135e0ef3aa482b1119d4317b8b264a

Observation 4c8b68c1-5415-4cc2-ad1b-714acea31bb0 · outbound

This paper cites Unsupervised visual rep- resentation learning by tracking patches in video.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Unsupervised visual rep- resentation learning by tracking patches in video

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.270594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:17:39.810778Z digest=sha256:4e9a993c00efe01613d7337c5fb5c8df636e909c44f916601da35fd771e9d908

Observation 1dc366f0-d074-4d3d-8f04-608407d27283 · outbound

This paper cites Self- supervised video representation learning by pace predic- tion.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Self- supervised video representation learning by pace predic- tion

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.259895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:17:39.814884Z digest=sha256:25afe5f4cb4ed8f2c0f2d8504305c826a4da69db5a43026bdaa60aa1fc3b6f51

Observation ccc4a9ca-1992-4714-8caf-4a8ddcac3375 · outbound

This paper cites Self-supervised Temporal Discriminative Learning for Video Representation Learning.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Self-supervised Temporal Discriminative Learning for Video Representation Learning

Reference 97

Resolution
verified exact
local_arxiv, observed 2026-08-08T19:17:39.950848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:17:39.818340Z digest=sha256:03e8db1225337e476c81e9b0a4cadb4ba3d0395b4c37523756894d27b3152112

Observation d6d43169-c1ea-4ac8-9025-d1a943d08a96 · outbound

This paper cites Temporal segment networks for action recognition in videos.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Temporal segment networks for action recognition in videos

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.249145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:17:39.822225Z digest=sha256:39bab5a610e5efe16ec9ae455af90dedda634a69e4694f86562728742c50e45d

Observation 3aa3cc15-d187-430b-8e3b-d61f85b8e5b8 · outbound

This paper cites Tdn: Temporal difference networks for efficient action recog- nition.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Tdn: Temporal difference networks for efficient action recog- nition

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.239461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:17:39.825930Z digest=sha256:13363c77651b4347207bcf7950d279bbfb5cb3da4e3d88a35c148e27f3f68af8

Observation 6e06deb0-1d76-4d4b-9eae-ce9c0bfb6918 · outbound

This paper cites Videomae v2: Scaling video masked autoencoders with dual masking.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Videomae v2: Scaling video masked autoencoders with dual masking

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:40.229038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:17:39.829689Z digest=sha256:cb996feaa28e32102defdb3da46b5c23ae97a0fea0ec871f51b62ec96ca0cd8a

Pith citing papers

Observation ff0a1ede-3856-4a74-ad1e-e3544cd0a8b4 · inbound

Infinite Video Understanding cites this paper.

Infinite Video Understanding CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:10:16.731929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T18:09:18.061419Z digest=sha256:5f6d1a3fe5951844c4743cb783ad6f4562a8f28d619c9b29e0d42f93b6f8368f