Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-08T19:17:39.829689Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 100 of 111 outbound references and 1 inbound Pith citation observation for arXiv:2502.07811.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-08T19:17:39.829689Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T18:09:18.061419Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-06T18:16:30.548575Z
100 of 111 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation aa31f2e1-7314-4142-8a0e-8c9e1ca4047e · outbound
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Self-supervised multimodal versatile networks
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c6b4a85-fdfa-49f7-8089-cc2eca4b9ba8 · outbound
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Self-supervised learning by cross-modal audio-video clustering
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 514f8041-bb2a-4464-bd57-0a3a6b5da576 · outbound
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Look, listen and learn
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8586362-2712-4b57-8c07-76dd79501518 · outbound
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Objects that sound
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f683016-7c5c-4ee6-9ad5-a37dcad2619e · outbound
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Vivit: A video vision transformer
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 417c24bb-c924-400d-b216-af913765b5b8 · outbound
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Adamae: Adaptive masking for efficient spatiotempo- ral learning with masked autoencoders
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e49641c3-d57b-4f0a-9d1b-5c40625657aa · outbound
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders BEit: BERT pre-training of image transformers
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6611a668-d9f4-4c85-a89b-53f609f485a1 · outbound
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Speednet: Learning the speediness in videos
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b2949be-ce36-4de2-b3ea-b842c2f58b5e · outbound
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Is space-time attention all you need for video understanding? In Proceedings of the 38th International Conference on Ma- chine Learning (ICML), pages 813–824
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6668cefe-edd9-4bd6-9124-7a21f1422748 · outbound
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Emerging properties in self-supervised vision transformers
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65788d6c-4189-41e8-8422-35b2a69464e0 · outbound
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Learning aligned cross-modal representations from weakly aligned data
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17edde5c-77e6-4dd2-ba52-2736aa6f6c61 · outbound
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Generative pre- training from pixels
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5539f8f5-1b13-4a00-9020-c74c0b3e8240 · outbound
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Rspnet: Relative speed perception for unsupervised video representation learning
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75f840e0-2139-4217-8264-8cf28e8a5d9e · outbound
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders A simple framework for contrastive learning of visual representations
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3e82289-834d-4ddd-ae23-8b5524d2c663 · outbound
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Electra: Pre-training text encoders as dis- criminators rather than generators
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 959bdc18-f39f-4383-bf50-293dafaa4dd7 · outbound
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Randaugment: Practical automated data augmen- tation with a reduced search space
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 155e3fce-1ef4-475f-8c7a-dd0a7ab896b7 · outbound
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Unifying video self-supervised learning across families of tasks: A survey
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea42b726-5ceb-4e61-8f21-4b2dec2cd59f · outbound
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Imagenet: A large-scale hierarchical im- age database
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 312f0708-8db9-4eb9-857c-05931df10a21 · outbound
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Virtex: Learning visual representations from textual annotations
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 139a9f0e-f182-4047-a18f-4ec727dfb35e · outbound
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders 9 Vi2clr: Video and image for visual contrastive learning of representation
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43e118a0-44e5-4587-aef1-6d255efc6c90 · outbound
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Taming transformers for high-resolution image synthesis
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20cab6dd-82ee-4f9d-9672-fc366b2f2c78 · outbound
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Multiscale vision transformers
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96353754-9cec-44fc-a127-806e7f99f547 · outbound
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Slowfast networks for video recognition
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0bcce743-4f81-45bd-bdd4-f3b969a60443 · outbound
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders A large-scale study on unsuper- vised spatiotemporal representation learning
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 901d0767-fcc8-42e2-ac84-f82b43be35f0 · outbound
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Masked autoencoders as spatiotemporal learners
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63f6a578-731e-49e2-8427-51835d2d17e4 · outbound
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Mcmae: Masked convolution meets masked autoencoders
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8466d3cb-f8e2-40a2-b371-f9372e09c879 · outbound
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Revitalizing cnn attention via transformers in self-supervised visual representation learn- ing
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 902a23cd-ffc7-42aa-83fd-437ffbfaaf93 · outbound
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Omnimae: Single model masked pretraining on images and videos
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd0451dc-47d5-41b0-b9eb-34aed2a78ebc · outbound
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Improving image-sentence embeddings using large weakly annotated photo collec- tions
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00ed9763-81a6-48b8-a699-eb1a39e57c45 · outbound
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 226ba7d6-3b09-4de9-ad54-5722c4e210c3 · outbound
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders The” something something” video database for learning and evaluating visual common sense
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1541263-211e-48db-aaec-9b861248f306 · outbound
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Bootstrap your own latent-a new approach to self-supervised learning
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 191c8262-b90c-4e24-bd6d-88c7ebdf13a6 · outbound
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders How Effective are Self-Supervised Models for Contact Identification in Videos
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f230e108-8784-43cc-9750-f5125c4111f8 · outbound
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders MaskViT: Masked Visual Pre-Training for Video Prediction
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43e74f4b-1dca-4739-924b-b2392e12b043 · outbound
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Memory- augmented dense predictive coding for video representa- tion learning
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 3d180f1e-242b-40fb-ac7a-1ba0012db08a · outbound
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Self- supervised co-training for video representation learning
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 85459afd-c541-4341-82f4-2a7c78a58e0a · outbound
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Momentum contrast for unsupervised visual rep- resentation learning
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 04165458-3c61-40e2-a83c-6295736d3104 · outbound
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Masked autoencoders are scal- able vision learners
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 0de94c77-96c4-45f6-9dbf-983a67480395 · outbound
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Human gaze control during real-world scene perception
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 68eaaf99-ec76-43fc-a872-bc7721e897d7 · outbound
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders ViC-MAE: Self-Supervised Representation Learning from Images and Video with Contrastive Masked Autoencoders
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f3fe09b1-b083-491d-bedb-dc09a8ea5e3a · outbound
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Augment your batch: Improv- ing generalization through instance repetition
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b906e515-298f-40f0-bbf5-b0ce88e3aab6 · outbound
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Contrast and order representa- tions for video self-supervised learning
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 6e6c79f1-2abf-4621-b7f9-5d30f07241d9 · outbound
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Mgmae: Motion guided masking for video masked autoencoding
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation fe7d86f1-dd41-43b5-a621-217655a7b909 · outbound
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Self-supervised video representation learn- ing by context and motion decoupling
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 738c14af-7d8d-4dc7-b1ec-4b44aeab3eed · outbound
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders TAda! Temporally-Adaptive Convolutions for Video Understanding
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 871e1068-622d-4070-998b-01da2f9af763 · outbound
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Self-Supervised Spatiotemporal Feature Learning via Video Rotation Prediction
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4cc14dd9-12dc-4812-83e2-50e587072d44 · outbound
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Hard negative mixing for contrastive learning
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d283f6db-8757-4960-b160-66b335ad927e · outbound
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Deep visual-semantic alignments for generating image descriptions
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c873d323-5d1e-45bd-805a-8dcc16c1fcc4 · outbound
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders The Kinetics Human Action Video Dataset
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5baf0a28-b25a-4b8c-9b71-3cfbecb9a128 · outbound
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Fusion attention for action recognition: Integrating sparse-dense and global at- tention for video action recognition
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 9813b497-3a07-4597-a37d-0bf3bf129886 · outbound
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Co- operative learning of audio and video models from self- supervised synchronization
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 500f507f-515d-4222-b5a5-7434d5fd9663 · outbound
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Hmdb: a large video database for human motion recognition
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7e6406a4-b4dd-4641-9d63-875b5c9f1d29 · outbound
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders A large-scale analysis on self- supervised video representation learning
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 0799014a-b89c-478a-a0c5-1b1021f610dc · outbound
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Semmae: Semantic-guided masking for learning masked autoencoders
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 4bd49512-7c46-4fb3-b6b2-8f2c829f6e61 · outbound
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Improve unsuper- vised pretraining for few-label transfer
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b6fa720b-5e50-44cb-8783-ce9ae101ed0f · outbound
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Learning Spatiotemporal Features via Video and Text Pair Discrimination
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd473744-3a29-415c-9c46-0d7152670e70 · outbound
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Tsm: Temporal shift module for efficient video understanding
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b0003748-d48d-4714-b611-efbcedd98d0b · outbound
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Self-supervised video-based action recognition with distur- bances
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 5d93b97f-445d-4dba-83ad-edd6c0a371c7 · outbound
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders CrossVideo: Self-supervised Cross-modal Contrastive Learning for Point Cloud Video Understanding
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f63ef84d-b5e8-45f5-84ff-585c00b310ff · outbound
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Teinet: Towards an efficient architecture for video recognition
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 792fa445-64f4-4b00-8df2-15369d8e8ddd · outbound
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Tam: Temporal adaptive module for video recog- nition
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 03ae3348-07ab-4823-9671-0e5c13c61487 · outbound
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Video swin transformer
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 4ad99798-c765-4d09-ba86-c350fc00c566 · outbound
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders SGDR: Stochastic Gradient Descent with Warm Restarts
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d1f7e83-dfba-417d-abda-a809af85e874 · outbound
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Decoupled weight de- cay regularization
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation fbe91960-dfe9-40cf-9114-eba79172bcb1 · outbound
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders CMAE-V: Contrastive Masked Autoencoders for Video Action Recognition
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation e280b4a2-9b2c-4658-ac55-94cd84302cda · outbound
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders 12-in-1: Multi-task vision and language representation learning
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c2c981d9-c97d-41a6-96a6-d97c93d2512d · outbound
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders End-to-end learning of visual representations from uncurated instruc- tional videos
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 084a12c4-3429-471b-bfa8-7986637c7087 · outbound
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Self-supervised learning of pretext-invariant representations
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 13389bba-7981-4c78-ab3a-5b9f71333ee5 · outbound
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Shuffle and learn: unsupervised learning using temporal order verification
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 72dca34f-316f-4a03-b152-e32f1f507d0d · outbound
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Ro- bust audio-visual instance discrimination
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8b33199a-aaee-47d7-a634-907428509832 · outbound
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Audio-visual instance discrimination with cross-modal agreement
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7c5b6ad5-ee64-4192-b19c-21661f1d02e8 · outbound
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Audio-visual scene analysis with self-supervised multisensory features
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c739f4a5-045d-40b1-9468-201f92560579 · outbound
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Videomoco: Contrastive video representation learning with temporally adversarial examples
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a7a40403-6a6d-4c2e-8fdf-d059a87425f1 · outbound
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Multi-modal self-supervision from generalized data transformations
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7cb4e5ca-9546-4928-ba59-a484f56495c9 · outbound
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Keeping your eye on the ball: Trajectory attention in video transformers
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 59e4f32a-76b0-4639-861e-d44248bd2633 · outbound
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Evolving losses for unsupervised video representation learning
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b3be13be-da0b-4586-bcdd-c7d238515bd0 · outbound
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Spa- tiotemporal contrastive video representation learning
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 2e6d248a-812e-4f05-b38a-3b4fec0a1a50 · outbound
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Learning from untrimmed videos: Self-supervised video representation learning with hierarchical consistency
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d8fa6935-4bef-4384-aef3-57926f397d74 · outbound
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Mar: Masked autoencoders for efficient action recognition
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 05e5e519-21f4-40e6-a1a8-789046110996 · outbound
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Learn- ing transferable visual models from natural language super- vision
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation adf4fcdb-1454-4f16-acf8-4b9dca135e39 · outbound
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Self-supervised video transformer
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f272e378-e5b7-42cc-ba3c-20e64d8078f4 · outbound
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders The development of visual attention and the brain
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ac81765d-19a8-4a6f-82ed-ebc0f7a27dd9 · outbound
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Im- agenet large scale visual recognition challenge
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation aacb1bdb-4c33-43d1-8c50-b1e8c18e2730 · outbound
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Learning visual representations with caption annotations
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 57cc4b17-a994-4024-b402-77c003c9f91a · outbound
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Two-stream con- volutional networks for action recognition in videos
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c2bb9469-7334-4d1b-bdfe-745d0885e127 · outbound
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0cf22c64-cdd8-4834-9d4f-d56f4bd303b9 · outbound
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Masked motion encoding for self-supervised video representation learning
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 45d830d2-4e75-4bdc-a2a2-8feb0d6a47a2 · outbound
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Rethinking the inception architecture for computer vision
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 090a7cd6-6cc6-48d3-9c96-edda12c8ba7e · outbound
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders VIMPAC: Video Pre-Training via Masked Token Prediction and Contrastive Learning
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33997c7c-e833-473b-8077-43da076bacfd · outbound
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ec8c52c4-5395-4d2d-80de-931e41f183ef · outbound
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Learning spatiotemporal fea- tures with 3d convolutional networks
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 40feb0f8-378a-4761-a06b-4345b2024495 · outbound
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Video classification with channel-separated convolutional networks
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 1a1311ce-37c9-4453-ac70-7e0d9043c430 · outbound
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Spa- tiotemporal integration and object perception in infancy: Perceiving unity versus form
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a09874af-8021-4a54-9bc8-254dfdd7e0a9 · outbound
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Generating videos with scene dynamics
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 4c8b68c1-5415-4cc2-ad1b-714acea31bb0 · outbound
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Unsupervised visual rep- resentation learning by tracking patches in video
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 1dc366f0-d074-4d3d-8f04-608407d27283 · outbound
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Self- supervised video representation learning by pace predic- tion
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ccc4a9ca-1992-4714-8caf-4a8ddcac3375 · outbound
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Self-supervised Temporal Discriminative Learning for Video Representation Learning
Reference 97
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d6d43169-c1ea-4ac8-9025-d1a943d08a96 · outbound
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Temporal segment networks for action recognition in videos
Reference 98
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 3aa3cc15-d187-430b-8e3b-d61f85b8e5b8 · outbound
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Tdn: Temporal difference networks for efficient action recog- nition
Reference 99
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 6e06deb0-1d76-4d4b-9eae-ce9c0bfb6918 · outbound
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Videomae v2: Scaling video masked autoencoders with dual masking
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ff0a1ede-3856-4a74-ad1e-e3544cd0a8b4 · inbound
Infinite Video Understanding CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.