Pith. sign in

Paper Citation Record · LEDGER

VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 25 inbound Pith citation observations for arXiv:2203.12602.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2203.12602 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 25 of 25 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:46:14.151441Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

436
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 1a18dcdb-9497-4d96-8ac2-20f2d22e7a76 · inbound

LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment cites this paper.

LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training

Reference 124

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T03:27:59.064233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-17T03:27:58.952076Z digest=sha256:15336acf02ec6646876ba2030cecd8468792b914004da045ef8c40a5144921f6

Observation 1ee0adec-e107-440f-9b09-0ad25428836c · inbound

Fine-Tuning Video Transformers for Word-Level Bangla Sign Language: A Comparative Analysis for Classification Tasks cites this paper.

Fine-Tuning Video Transformers for Word-Level Bangla Sign Language: A Comparative Analysis for Classification Tasks VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T10:46:14.151441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:46:14.151441Z digest=sha256:4efe9946b2b3f5cc3438b237c7ae56a2d70adbc49caefb560233a9cb60b0f8d8

Observation 3a3ce278-b861-4ab0-a19b-46b89a94b4c6 · inbound

FRAME: Pre-Training Video Feature Representations via Anticipation and Memory cites this paper.

FRAME: Pre-Training Video Feature Representations via Anticipation and Memory VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T10:26:59.459159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:26:59.459159Z digest=sha256:8b697d5d9953e46d5fa64f55d5ebd13682d58b17cb6d743a48641dc21151d01b

Observation 596e9c2a-34a5-4024-9e7c-e77d82852611 · inbound

One Video to Steal Them All: 3D-Printing IP Theft through Optical Side-Channels cites this paper.

One Video to Steal Them All: 3D-Printing IP Theft through Optical Side-Channels VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T22:24:19.049850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:24:19.049850Z digest=sha256:ac4ed04ab92fc7fec01b2ba35aa497c2db46f063b48990e41ed8a9ed0ce9c6e0

Observation af449d3e-434b-4f99-ada6-63a5585c1387 · inbound

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition cites this paper.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:57.583666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:57.583666Z digest=sha256:0494c45426d224734008294bf8ead0e4c5d60d475aebed7716c45b3d15d075dc

Observation 1343f094-5c33-4b3f-b3ad-46901b5913de · inbound

MVP: Winning Solution to SMP Challenge 2025 Video Track cites this paper.

MVP: Winning Solution to SMP Challenge 2025 Video Track VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T21:06:56.193790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:06:56.193790Z digest=sha256:057e1724c9bd58601c49ae728ce41787499ff876002b8e0023eee267e0fa8d54

Observation 13d5428e-6ba5-44f3-9677-ad9af44217c8 · inbound

Multimodal Framework for Explainable Autonomous Driving: Integrating Video, Sensor, and Textual Data for Enhanced Decision-Making and Transparency cites this paper.

Multimodal Framework for Explainable Autonomous Driving: Integrating Video, Sensor, and Textual Data for Enhanced Decision-Making and Transparency VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T18:33:25.677404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:33:25.677404Z digest=sha256:6fea8c91464695127d6d05bdefe483485b5da8dfd807f5d398b35f962d1f6028

Observation 8627080a-8146-4fd5-893f-d0b40c9a5e1c · inbound

Infinite Video Understanding cites this paper.

Infinite Video Understanding VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:13.574181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:13.574181Z digest=sha256:a5c7ea3ab6d004ee82c33d775dea402e57035a20747042da53ef1cb9b2b72b4c

Observation 097ec171-d4f5-48ea-8624-e42e5edbdbd3 · inbound

Interpreting Video Representations with Spatio-Temporal Sparse Autoencoders cites this paper.

Interpreting Video Representations with Spatio-Temporal Sparse Autoencoders VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-13T12:01:02.997589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T12:01:02.997589Z digest=sha256:bb6d81392bf1b98809d08a5dd83796b42ceaf82cad3abfe04da766c0f9b50679

Observation bd06e627-9cbf-48b4-9d99-0592fa36bc05 · inbound

Insights from Visual Cognition: Understanding Human Action Dynamics with Overall Glance and Refined Gaze Transformer cites this paper.

Insights from Visual Cognition: Understanding Human Action Dynamics with Overall Glance and Refined Gaze Transformer VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:01.661100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T18:07:41.426142Z digest=sha256:20d0e6044fb797c1394e6d4fb5e3de491e160b12b633081aa18f9b0e04d20cac

Observation 32eda57c-720b-401e-8905-3813f7ac41bd · inbound

Zero-shot World Models Are Developmentally Efficient Learners cites this paper.

Zero-shot World Models Are Developmentally Efficient Learners VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:21:00.069998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T15:33:39.342672Z digest=sha256:5da5a66b77ba885c47de6fc768b8f69749455b2f37ef267194f9f5926b67729f

Observation 436e807f-cc52-4e35-b968-5a6601d722d9 · inbound

Beyond Independent Frames: Latent Attention Masked Autoencoders for Multi-View Echocardiography cites this paper.

Beyond Independent Frames: Latent Attention Masked Autoencoders for Multi-View Echocardiography VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:00:22.231240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T11:56:05.271784Z digest=sha256:8312a085e52ac03c81ef16a20d672e1f4e1ab245662505188c63d4cd0dc0b5cf

Observation 10c72199-24da-4342-ad3b-d63142d15e99 · inbound

Mask World Model: Predicting What Matters for Robust Robot Policy Learning cites this paper.

Mask World Model: Predicting What Matters for Robust Robot Policy Learning VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T13:11:05.988492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T02:14:17.676675Z digest=sha256:ec551bec94883e27bdafccc9f221c8c681f0955404c031f160fa06c95bfbf0bc

Observation 5c24090c-e261-4807-9828-8c5940807490 · inbound

SignMAE: Segmentation-Driven Self-Supervised Learning for Sign Language Recognition cites this paper.

SignMAE: Segmentation-Driven Self-Supervised Learning for Sign Language Recognition VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-09T05:55:30.819738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T19:23:07.258575Z digest=sha256:57cd8456c07806984a443bcc13218e2cc1f6e48a69aca2ae4497bca14117734f

Observation 012078d3-2e08-4cc7-9d3d-bd8b845b2204 · inbound

SpecSem-Net: Integrating Spectral and Semantic Features for Robust AI-generated Video Detection cites this paper.

SpecSem-Net: Integrating Spectral and Semantic Features for Robust AI-generated Video Detection VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-20T14:23:21.618120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T14:19:36.263592Z digest=sha256:9ecc01a2cd05a803a29afc3cc5662c01725dc70d5440d0b71ce180854401a5b0

Observation eb425c87-a1ff-4375-8e13-c999bb6f4426 · inbound

Prognostic Value of Lung Ultrasound Biomarkers for Readmission Risk in Congestive Heart Failure: A Pilot Data-Driven Analysis cites this paper.

Prognostic Value of Lung Ultrasound Biomarkers for Readmission Risk in Congestive Heart Failure: A Pilot Data-Driven Analysis VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T15:43:26.997735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-20T15:41:31.181025Z digest=sha256:41619f421a8155e579f0000e7bb0b6bf3965d493b10d7c5b566d51556d1980b1

Observation 666c8eb3-efaf-45fa-bfc3-08b6c9fdab57 · inbound

FAST-ME: Foundation-aware Adaptive Stopping for Motion Estimation for Efficient IoT Video Analysis cites this paper.

FAST-ME: Foundation-aware Adaptive Stopping for Motion Estimation for Efficient IoT Video Analysis VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-25T04:50:21.230946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T04:47:22.381247Z digest=sha256:f9db92a01e0ca368789cc77e31ccf56825bc5c25f8fd5bbc752bbad9c1085e25

Observation a4990119-6ac8-4626-9ad9-6aedb9c67083 · inbound

EVA-Net: Subject-Independent EEG Motor Decoding with Video-Derived Motor Priors cites this paper.

EVA-Net: Subject-Independent EEG Motor Decoding with Video-Derived Motor Priors VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T22:46:20.318495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T14:59:24.109366Z digest=sha256:d663525e8a78b166d4acd94d559a597bcdff447de3f019befad0ee532b044441

Observation 60ed711d-f744-49c6-8541-5677007360d6 · inbound

EVA-Net: Subject-Independent EEG Motor Decoding with Video-Derived Motor Priors cites this paper.

EVA-Net: Subject-Independent EEG Motor Decoding with Video-Derived Motor Priors VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-12T15:23:58.657504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T15:23:58.657504Z digest=sha256:e2495bbc3c5686d9e22744a99a89c6c10e6be2d17ab7e83b04ae0a80d58edf1d

Observation ec84fac8-62f7-4d77-82c0-93596628fa38 · inbound

VTI-CoT: Visual-Textual Interleaved Chain of Thought for Video Reasoning cites this paper.

VTI-CoT: Visual-Textual Interleaved Chain of Thought for Video Reasoning VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:46:56.824828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T01:52:44.785582Z digest=sha256:bda8e7e741152d75d7bf3f3a253da525632cca55d7371d79f9dcb87f083e01c3

Observation d398b846-83d2-478c-a491-b75ce976e56f · inbound

BioVid: Autoregressive Video Generation with Biological Behavior Semantic Comprehension cites this paper.

BioVid: Autoregressive Video Generation with Biological Behavior Semantic Comprehension VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-06-27T18:41:07.978312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T18:35:14.331659Z digest=sha256:5777bc256e42f149d1905b4ee02a7f0783ba3a08ff0df7926fb8797e5799b612

Observation b91c577c-86eb-4d8e-87eb-fcc1e100c55a · inbound

Do Video Foundation Models Understand Intuitive Physics? A Layerwise Probing Analysis cites this paper.

Do Video Foundation Models Understand Intuitive Physics? A Layerwise Probing Analysis VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T00:07:28.693423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T17:25:15.493925Z digest=sha256:5d67681c953f9882d97a9225bff6bb01f25f8efb2cc6240651af4455777c56f8

Observation 12dc5c0b-df9a-48d7-a577-d48b2aa1b896 · inbound

Deep Learning-Based Sign Language Recognition from Videos and Cross-Lingual Translation to Indian Vernaculars cites this paper.

Deep Learning-Based Sign Language Recognition from Videos and Cross-Lingual Translation to Indian Vernaculars VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:49:42.467011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T10:55:03.728012Z digest=sha256:49086e1be80df015ff2530ba72d26285d385b1d1d9aee51ccc43c10aa3272216

Observation ab485a0f-7d57-4e89-84c5-b056706ae7a3 · inbound

Empirical Evaluation of Multi-Modal Touch Detection in Over-the-Shoulder Video Surveillance cites this paper.

Empirical Evaluation of Multi-Modal Touch Detection in Over-the-Shoulder Video Surveillance VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-06-30T07:04:20.959297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T07:02:23.993806Z digest=sha256:c91682fd68b08ab426c08ab3ee372ce3459b1c2e8d30bc56802427c05f2a9bad

Observation beb3c69b-1797-4c66-853a-d40f3b182606 · inbound

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models cites this paper.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.632640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.632640Z digest=sha256:ed43d2488ad6a8e7f5c6ee96d1a27574c1bf5e002fd3824d61b64f6f6af75833