REVIEW 3 cited by
Denoising Masked AutoEncoders Help Robust Classification
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
In this paper, we propose a new self-supervised method, which is called Denoising Masked AutoEncoders (DMAE), for learning certified robust classifiers of images. In DMAE, we corrupt each image by adding Gaussian noises to each pixel value and randomly masking several patches. A Transformer-based encoder-decoder model is then trained to reconstruct the original image from the corrupted one. In this learning paradigm, the encoder will learn to capture relevant semantics for the downstream tasks, which is also robust to Gaussian additive noises. We show that the pre-trained encoder can naturally be used as the base classifier in Gaussian smoothed models, where we can analytically compute the certified radius for any data point. Although the proposed method is simple, it yields significant performance improvement in downstream classification tasks. We show that the DMAE ViT-Base model, which just uses 1/10 parameters of the model developed in recent work arXiv:2206.10550, achieves competitive or better certified accuracy in various settings. The DMAE ViT-Large model significantly surpasses all previous results, establishing a new state-of-the-art on ImageNet dataset. We further demonstrate that the pre-trained model has good transferability to the CIFAR-10 dataset, suggesting its wide adaptability. Models and code are available at https://github.com/quanlin-wu/dmae.
Forward citations
Cited by 3 Pith papers
-
Video-GPT via Next Clip Diffusion
A 3.8B transformer pretrained on 70M unlabeled videos with next clip diffusion (autoregressive clean-history conditioning plus in-clip denoising) reports state-of-the-art Physics-IQ and Kinetics-600 video prediction.
-
MINR: Implicit Neural Representations with Masked Image Modelling
A hybrid of implicit neural representations and masked image modeling, called MINR, reconstructs masked image patches better than MAE in the reported in-domain and out-of-distribution tests with fewer parameters.
-
Block-based Symmetric Pruning and Fusion for Efficient Vision Transformers
A block-based attention pruning and fusion method for ViTs that reports large accuracy gains at reduced FLOPs, but the gain is mostly from the chunk-attention backbone and the core symmetry claim is false.
Discussion (0). Continue with ORCID to comment.