REVIEW 6 cited by
A Survey on Masked Autoencoder for Self-supervised Learning in Vision and Beyond
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Masked autoencoders are scalable vision learners, as the title of MAE \cite{he2022masked}, which suggests that self-supervised learning (SSL) in vision might undertake a similar trajectory as in NLP. Specifically, generative pretext tasks with the masked prediction (e.g., BERT) have become a de facto standard SSL practice in NLP. By contrast, early attempts at generative methods in vision have been buried by their discriminative counterparts (like contrastive learning); however, the success of mask image modeling has revived the masking autoencoder (often termed denoising autoencoder in the past). As a milestone to bridge the gap with BERT in NLP, masked autoencoder has attracted unprecedented attention for SSL in vision and beyond. This work conducts a comprehensive survey of masked autoencoders to shed insight on a promising direction of SSL. As the first to review SSL with masked autoencoders, this work focuses on its application in vision by discussing its historical developments, recent progress, and implications for diverse applications.
Forward citations
Cited by 6 Pith papers
-
GViT: Representing Images as Gaussians for Visual Recognition
Images encoded as a few hundred learnable 2D Gaussians, steered by classifier gradients, support a ViT that reaches 76.9% top-1 on ImageNet-1k, close to patch-based ViTs.
-
NAE: Normalizing AutoEncoder
A conditional surrogate loss that always picks the gradient estimate aligned with the reconstruction loss improves flow autoencoder training and reaches state-of-the-art generative performance on molecules, tabular da...
-
Pre-Training Graph Contrastive Masked Autoencoders are Strong Distillers for EEG
EEG-DisGCMAE combines graph contrastive and masked autoencoder pre-training with a graph topology distillation loss to improve low-density EEG classification using high-density and unlabeled data.
-
MM-Prompt: Cross-Modal Prompt Tuning for Continual Visual Question Answering
MM-Prompt couples the visual and language prompt paths in continual VQA, and reports higher average accuracy and lower forgetting than existing prompt-based methods.
-
Cluster Specific Representation Learning
A meta-algorithm that learns cluster-specific embedding functions jointly with cluster assignments improves clustering and denoising over standard representation learning baselines.
-
KDC-MAE: Knowledge Distilled Contrastive Mask Auto-Encoder
KDC-MAE pretrains an audio-video transformer with two complementary masks plus KL self-distillation and reports small, partly inconsistent accuracy gains over CAV-MAE.
Discussion (0). Continue with ORCID to comment.