Pith. sign in

REVIEW 6 cited by

A Survey on Masked Autoencoder for Self-supervised Learning in Vision and Beyond

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2208.00173 v1 pith:K44GRPE2 submitted 2022-07-30 cs.CV cs.AIcs.LG

classification cs.CVcs.AIcs.LG
keywords maskedvisionautoencoderautoencoderslearningbertbeyondgenerative
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Masked autoencoders are scalable vision learners, as the title of MAE \cite{he2022masked}, which suggests that self-supervised learning (SSL) in vision might undertake a similar trajectory as in NLP. Specifically, generative pretext tasks with the masked prediction (e.g., BERT) have become a de facto standard SSL practice in NLP. By contrast, early attempts at generative methods in vision have been buried by their discriminative counterparts (like contrastive learning); however, the success of mask image modeling has revived the masking autoencoder (often termed denoising autoencoder in the past). As a milestone to bridge the gap with BERT in NLP, masked autoencoder has attracted unprecedented attention for SSL in vision and beyond. This work conducts a comprehensive survey of masked autoencoders to shed insight on a promising direction of SSL. As the first to review SSL with masked autoencoders, this work focuses on its application in vision by discussing its historical developments, recent progress, and implications for diverse applications.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. GViT: Representing Images as Gaussians for Visual Recognition

    cs.CV 2025-06 conditional novelty 7.0 of 10

    Images encoded as a few hundred learnable 2D Gaussians, steered by classifier gradients, support a ViT that reaches 76.9% top-1 on ImageNet-1k, close to patch-based ViTs.

  2. NAE: Normalizing AutoEncoder

    cs.LG 2026-08 conditional novelty 6.0 of 10

    A conditional surrogate loss that always picks the gradient estimate aligned with the reconstruction loss improves flow autoencoder training and reaches state-of-the-art generative performance on molecules, tabular da...

  3. Pre-Training Graph Contrastive Masked Autoencoders are Strong Distillers for EEG

    cs.LG 2024-11 conditional novelty 6.0 of 10

    EEG-DisGCMAE combines graph contrastive and masked autoencoder pre-training with a graph topology distillation loss to improve low-density EEG classification using high-density and unlabeled data.

  4. MM-Prompt: Cross-Modal Prompt Tuning for Continual Visual Question Answering

    cs.CV 2025-05 conditional novelty 5.0 of 10

    MM-Prompt couples the visual and language prompt paths in continual VQA, and reports higher average accuracy and lower forgetting than existing prompt-based methods.

  5. Cluster Specific Representation Learning

    cs.LG 2024-12 conditional novelty 4.0 of 10

    A meta-algorithm that learns cluster-specific embedding functions jointly with cluster assignments improves clustering and denoising over standard representation learning baselines.

  6. KDC-MAE: Knowledge Distilled Contrastive Mask Auto-Encoder

    cs.CV 2024-11 conditional novelty 4.0 of 10

    KDC-MAE pretrains an audio-video transformer with two complementary masks plus KL self-distillation and reports small, partly inconsistent accuracy gains over CAV-MAE.

Pith tools