Pith. sign in

REVIEW 10 cited by

VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2203.12602 v3 pith:4SUUPXFF submitted 2022-03-23 cs.CV

classification cs.CV
keywords videovideomaepre-trainingdatadatasetsextraimportantmasking
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Pre-training video transformers on extra large-scale datasets is generally required to achieve premier performance on relatively small datasets. In this paper, we show that video masked autoencoders (VideoMAE) are data-efficient learners for self-supervised video pre-training (SSVP). We are inspired by the recent ImageMAE and propose customized video tube masking with an extremely high ratio. This simple design makes video reconstruction a more challenging self-supervision task, thus encouraging extracting more effective video representations during this pre-training process. We obtain three important findings on SSVP: (1) An extremely high proportion of masking ratio (i.e., 90% to 95%) still yields favorable performance of VideoMAE. The temporally redundant video content enables a higher masking ratio than that of images. (2) VideoMAE achieves impressive results on very small datasets (i.e., around 3k-4k videos) without using any extra data. (3) VideoMAE shows that data quality is more important than data quantity for SSVP. Domain shift between pre-training and target datasets is an important issue. Notably, our VideoMAE with the vanilla ViT can achieve 87.4% on Kinetics-400, 75.4% on Something-Something V2, 91.3% on UCF101, and 62.6% on HMDB51, without using any extra data. Code is available at https://github.com/MCG-NJU/VideoMAE.

Discussion (0). Sign in to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 436 citations worldwide. Full citation record

  1. One Video to Steal Them All: 3D-Printing IP Theft through Optical Side-Channels

    cs.CR 2025-06 conditional novelty 7.0 of 10

    An optical side-channel on 3D printing can recover printable G-code instructions, with functional counterfeit objects demonstrated on a key and a gear.

  2. CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models

    cs.CV 2026-08 conditional novelty 6.0 of 10

    CRAFT recursively merges video tokens with training-free similarity selection plus learnable gated fusion, retaining ~97% of average accuracy at 8x compression across six benchmarks.

  3. Interpreting Video Representations with Spatio-Temporal Sparse Autoencoders

    cs.CV 2026-04 conditional novelty 6.0 of 10

    Spatio-temporal contrastive SAEs recover temporal coherence lost by hard TopK, improve action probes by +3.9% and retrieval by up to 2.8× R@1, and expose a monosemanticity metric artifact.

  4. FRAME: Pre-Training Video Feature Representations via Anticipation and Memory

    cs.CV 2025-06 conditional novelty 6.0 of 10

    FRAME distills DINO and CLIP features into a compact video encoder with a memory module and future-frame prediction, outperforming image-based and self-supervised video baselines on dense video tasks.

  5. EVA-Net: Subject-Independent EEG Motor Decoding with Video-Derived Motor Priors

    cs.AI 2026-06 unverdicted novelty 5.0 of 10

    EVA-Net improves subject-independent EEG motor decoding by using video action priors via cross-modal contrastive alignment and knowledge distillation, reporting an 8.66% LOSO accuracy gain on EEGMMI.

  6. MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition

    cs.CV 2025-06 conditional novelty 5.0 of 10

    MoMa adapts frozen CLIP to video by injecting Mamba-computed scale and bias into each layer, improving accuracy and efficiency on multiple action recognition benchmarks.

  7. Fine-Tuning Video Transformers for Word-Level Bangla Sign Language: A Comparative Analysis for Classification Tasks

    cs.CV 2025-06 conditional novelty 5.0 of 10

    Off-the-shelf video transformers (VideoMAE, ViViT, TimeSformer) fine-tuned on Bangla sign language videos reach 95.5% top-1 accuracy on BdSLW60 and 81.04% on the BdSLW401 front subset.

  8. Infinite Video Understanding

    cs.CV 2025-07 conditional novelty 3.0 of 10

    The paper argues that video understanding research should aim at processing streams of arbitrary, unbounded duration and outlines the challenges, directions, and metrics needed.

  9. Multimodal Framework for Explainable Autonomous Driving: Integrating Video, Sensor, and Textual Data for Enhanced Decision-Making and Transparency

    cs.MM 2025-07 reject novelty 3.0 of 10

    A three-branch model (VideoMAE, a two-layer sensor MLP, BERT, fused into a BART-based explainer) reports 92.5% action accuracy and a 0.75 BLEU-4 on nuScenes, with simulated attention maps and no released code.

  10. MVP: Winning Solution to SMP Challenge 2025 Video Track

    cs.CV 2025-07 conditional novelty 3.0 of 10

    MVP, a pipeline using XCLIP video features, user metadata, and a CatBoost regressor, won the SMP Challenge 2025 Video Track with a MAPE of 0.1754.

Pith tools